返回
A new adaptive sampling algorithm for big data classification
DOI:10.1016/j.jocs.2022.101653.png)
摘要
En 中文
The exponential growth of the quantity of data that circulates on the web led to the emergence of the big data phenomenon. This fact is a natural consequence of the proliferation of social media, mobile devices, the abundance of free online storage, and new technologies like the internet of things. Subsequently, big data has created several challenges to the computer science community, among which the large size of data is the most challenging. Traditional machine learning algorithms used mostly for insight extraction find themselves inadequate, even on high-performance computer architectures. For instance, big data analytics algorithms can overcome the size issue by either: (1) adapting the existing machine learning techniques to the scale of the big data; or, (2) by sampling big datasets, choosing randomly much smaller subsets of the data population, to meet what current algorithms can handle. In the present work, we aim to proceed through the second alternative to address the size challenge in the big data context. We propose intelligent sampling techniques based on Scalable Simple Random Sampling (ScaSRS) and Subsampled Double Bootstrap (SDB). Test results carried out on public generic datasets show that our proposal is able to address the size dimension efficiently. The proposed algorithms were evaluative on a classification task where the obtained results provided significant improvement compared to the state-of-the-art.
Keyword:
Big data
Data classification
Sampling methods
Subsampled Double Bootstrap
Naive Bayes classifier
期刊
IF:
18.3
论文数:
3.1K
被引数:
4.0K
机构
引用论文
A survey towards an integration of big data analytics to big insights for value-creation关于将大数据分析与大洞察相结合以创造价值的调查
A comprehensive survey on support vector machine classification: Applications, challenges and trends支持向量机分类综述: 应用、挑战与趋势
NEUROCOMPUTING
IF6.5

