返回
A Model for Enhancing Unstructured Big Data Warehouse Execution Time
DOI:10.3390/bdcc8020017.png)
摘要
En 中文
Traditional data warehouses (DWs) have played a key role in business intelligence and decision support systems. However, the rapid growth of the data generated by the current applications requires new data warehousing systems. In big data, it is important to adapt the existing warehouse systems to overcome new issues and limitations. The main drawbacks of traditional Extract-Transform-Load (ETL) are that a huge amount of data cannot be processed over ETL and that the execution time is very high when the data are unstructured. This paper focuses on a new model consisting of four layers: Extract-Clean-Load-Transform (ECLT), designed for processing unstructured big data, with specific emphasis on text. The model aims to reduce execution time through experimental procedures. ECLT is applied and tested using Spark, which is a framework employed in Python. Finally, this paper compares the execution time of ECLT with different models by applying two datasets. Experimental results showed that for a data size of 1 TB, the execution time of ECLT is 41.8 s. When the data size increases to 1 million articles, the execution time is 119.6 s. These findings demonstrate that ECLT outperforms ETL, ELT, DELT, ELTL, and ELTA in terms of execution time.
Keyword:
big data
unstructured data warehouse
ELT
ETL
期刊
B
IF:
4.4
论文数:
1.3K
被引数:
2.4K
机构
引用论文
京津冀城市群水资源开发利用的时空特征与政策启示The spatial-temporal characteristics and policy implications of water resources development and utilization in the Beijing-Tianjin-Hebei urban agglomeration
地理科学进展
IF0
PSMD11 modulates circadian clock function through PER and CRY nuclear translocationPSMD11通过PER和CRY核转位调节生物钟功能
PLOS ONE
IF0
Evaluating partitioning and bucketing strategies for Hive-based Big Data Warehousing systems
JOURNAL OF BIG DATA
IF6.4

