返回
Preprocessing framework for scholarly big data management
DOI:10.1007/s11042-022-13513-8.png)
摘要
En 中文
Big data technologies have found applications in disparate domains. One of the largest sources of textual big data is scientific documents and papers. Scholarly big data has been used in numerous ways to develop innovative applications such as collaborator discovery, expert finding and research management systems. With the evolution of machine and deep learning techniques, the efficacy of such applications has risen manifold. However, the biggest challenge in the development of deep learning models for scholarly applications in cloud-based environment is the under-utilization of resources because of the excessive time required for textual preprocessing. This paper presents a preprocessing pipeline that uses Spark for data ingestion and Spark ML for performing preprocessing tasks. The proposed approach is evaluated with the help of a case study, which uses LSTM-based text summarization to generate title or summaries from abstracts of scholarly articles. Results indicate a substantial reduction in ingestion, preprocessing and cumulative time for the proposed approach, which shall manifest reduction in development time and costs as well.
Keyword:
Deep learning applications
Preprocessing pipeline
Scholarly big data
Scholarly data applications
Spark ML
期刊
IF:
3
论文数:
1.9W
被引数:
3.2W
机构
暂无机构信息
引用论文
Capacity optimization of hybrid renewable energy system considering part-load ratio and resource endowment考虑部分负荷率和资源禀赋的混合可再生能源系统容量优化
Enhancement of pyramid solar distiller performance using reflectors, cooling cycle, and dangled cords of wicks
Desalination
IF0
Myocardial gene delivery using molecular cardiac surgery with recombinant adeno-associated virus vectors in vivo
Gene Therapy
IF0

