返回
Time Series Impact Through Topic Modeling
DOI:10.1109/ACCESS.2022.3202960.png)
摘要
En 中文
A time-series of numerical data and a sequence of time-ordered documents are often correlated. This paper aims at modeling the impact that the underlying themes discussed in the text data have on the time series. To do so, we introduce an original topic model, Time Series Impact Through Topic Modeling (TSITM), that includes contextual data by coupling Latent Dirichlet Allocation (LDA) with linear regression, using an elastic net prior to set to zero the impact of uncorrelated topics. The resulting topics act as explanatory variables for the regression of the numerical time series, which allows us to understand the time series movements based on the events described on the text data. We have tested our model on two datasets: first, we used political news to explain the US president's disapproval ratings; then, we considered a corpus of economic news to explain the financial returns of 4 different multinational corporations. Our experiments show that an appropriate selection of hyperparameters (via repeated random subsampling validation and Bayesian optimization) leads to significant correlations: both an intrinsic baseline and state of the art methods were significantly outperformed by TSITM in MSE, MAE and out-of-sample R-2, according to our hypothesis tests. We believe that this framework can be useful in the context of reputational risk management.
Keyword:
Expectation-maximization algorithms
natural language processing
regression analysis
text mining
text analysis
time series analysis
期刊
IF:
3.6
论文数:
9.8W
被引数:
29.4W
机构
引用论文
On exploring the impact of users' bullish-bearish tendencies in online community on the stock market
Comparative analysis of fat and muscle proteins in fenofibratefed type II diabetic OLETF rats: the fenofibrate-dependent expression of PEBP or C11orf59 protein
BMB Reports
IF0
Financial Latent Dirichlet Allocation (FinLDA): Feature Extraction in Text and Data Mining for Financial Time Series Prediction金融潜在狄利克雷分配 (FinLDA): 用于金融时间序列预测的文本和数据挖掘中的特征提取
IEEE ACCESS
IF3.6

