返回
Learning to crawl: Comparing classification schemes
DOI:10.1145/1095872.1095875.png)
摘要
En 中文
Topical crawling is a young and creative area of research that holds the promise of benefiting from several sophisticated data mining techniques. The use of classification algorithms to guide topical crawlers has been sporadically suggested in the literature. No systematic study, however, has been done on their relative merits. Using the lessons learned from our previous crawler evaluation studies, we experiment with multiple versions of different classification schemes. The crawling process is modeled as a parallel best-first search over a graph defined by the Web. The classifiers provide heuristics to the crawler thus biasing it towards certain portions of the Web graph. Our results show that Naive Bayes is a weak choice for guiding a topical crawler when compared with Support Vector Machine or Neural Network. Further, the weak performance of Naive Bayes can be partly explained by extreme skewness of posterior probabilities generated by it. We also observe that despite similar performances, different topical crawlers cover subspaces on the Web with low overlap.
Keyword:
topical crawlers
focused crawlers
classifiers
machine learning
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
9.1
论文数:
1.2K
被引数:
4.7K
机构
暂无机构信息
引用论文
30 YEARS OF ADAPTIVE NEURAL NETWORKS - PERCEPTRON, MADALINE, AND BACKPROPAGATION
PROCEEDINGS OF THE IEEE
IF25.9
Adaptive retrieval agents: Internalizing local context and scaling up to the Web
MACHINE LEARNING
IF2.9
Cross-linkable Polymer Matrix for Enhanced Thermal Stability of Succinonitrile-based Polymer Electrolyte in Lithium Rechargeable Batteries可交联聚合物基质,用于增强锂可充电电池中基于丁二腈的聚合物电解质的热稳定性

