返回
Tension in big data using machine learning: Analysis and applications
DOI:10.1016/j.techfore.2020.120175.png)
摘要
En 中文
The access of machine learning techniques in popular programming languages and the exponentially expanding big data from social media, news, surveys, and markets provide exciting challenges and invaluable opportunities for organizations and individuals to explore implicit information for decision making. Nevertheless, the users of machine learning usually find that these sophisticated techniques could incur a high level of tensions caused by the selection of the appropriate size of the training data set among other factors. In this paper, we provide a systematic way of resolving such tensions by examining practical examples of predicting popularity and sentiment of posts on Twitter and Facebook, blogs on Mashable, news on Google and Yahoo, the US house survey, and Bitcoin prices. Interesting results show that for the case of big data, using around 20% of the full sample often leads to a better prediction accuracy than opting for the full sample. Our conclusion is found to be consistent across a series of experiments. The managerial implication is that using more is not necessarily the best and users need to be cautious about such an important sensitivity as the simplistic approach may easily lead to inferior solutions with potentially detrimental consequences.
Keyword:
Big data
Machine learning
Data size
Prediction accuracy
Social media
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
13.3
论文数:
7.8K
被引数:
6.2W
机构
引用论文
Meta-learning for evolutionary parameter optimization of classifiers用于分类器进化参数优化的元学习
MACHINE LEARNING
IF2.9
Instantaneous vehicle fuel consumption estimation using smartphones and recurrent neural networks基于智能手机和递归神经网络的瞬时汽车油耗估算

