返回
Evaluation Methods for Statistically Dependent Text
DOI:10.1162/COLI_a_00230.png)
摘要
En 中文
In recent years, many studies have been published on data collected from social media, especially microblogs such as Twitter. However, rather few of these studies have considered evaluation methodologies that take into account the statistically dependent nature of such data, which breaks the theoretical conditions for using cross-validation. Despite concerns raised in the past about using cross-validation for data of similar characteristics, such as time series, some of these studies evaluate their work using standard k-fold cross-validation. Through experiments on Twitter data collected during a two-year period that includes disastrous events, we show that by ignoring the statistical dependence of the text messages published in social media, standard cross-validation can result in misleading conclusions in a machine learning task. We explore alternative evaluation methods that explicitly deal with statistical dependence in text. Our work also raises concerns for any other data for which similar conditions might hold.
Keyword:
CROSS-VALIDATION
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
5.3
论文数:
837
被引数:
2.7K
机构
引用论文
Sustainability in the prospective scenarios methods: A case study of scenarios for biodiesel industry in Brazil, for 2030
Futures
IF0
Dispersion and Beating of Bacterial Cellulose and their Influence on Paper Properties
BioResources
IF0

