Return
Advanced machine learning for subject-oriented inappropriate content classification: A topic modeling approach
DOI:10.1016/j.eij.2025.100882.png)
Abstract
En 中文
The rapid escalation of inappropriate online content calls for sophisticated and highly accurate methods for content classification and filtering. Previous approaches primarily focused on classifying web pages based on their textual and visual contents, ignoring the subject-oriented aspect, leading to an inaccurate classification of inappropriate content topics. This study proposes a novel Subject-Oriented Filtering (SOF) integration and formalization that couples dynamic URL whitelists/blacklists with HTML topic vectors fed directly to discriminative classifiers for accurate inappropriate-content classification. By exploiting the semantic richness of HTML structure and inputting the topic vectors as features for advanced machine learning classifiers, this methodology noticeably increases the accuracy of webpage filtering and classification. This study performed extensive experiments, which show that SOF achieves an accuracy exceeding 94%, substantially outperforming conventional methods. The methodological innovation of this study establishes a new state-of-the-art baseline in subject-oriented web content classification, representing significant progress over previous studies and contributing to safer online environments.
Keywords:
Topic modeling
Subject-oriented filtering
Content classification
HTML web filtering
Inappropriate content
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
4.3
Papers:
770
Citations:
1.4K

