Return
Preprocessing framework for scholarly big data management
DOI:10.1007/s11042-022-13513-8.png)
Abstract
En 中文
Big data technologies have found applications in disparate domains. One of the largest sources of textual big data is scientific documents and papers. Scholarly big data has been used in numerous ways to develop innovative applications such as collaborator discovery, expert finding and research management systems. With the evolution of machine and deep learning techniques, the efficacy of such applications has risen manifold. However, the biggest challenge in the development of deep learning models for scholarly applications in cloud-based environment is the under-utilization of resources because of the excessive time required for textual preprocessing. This paper presents a preprocessing pipeline that uses Spark for data ingestion and Spark ML for performing preprocessing tasks. The proposed approach is evaluated with the help of a case study, which uses LSTM-based text summarization to generate title or summaries from abstracts of scholarly articles. Results indicate a substantial reduction in ingestion, preprocessing and cumulative time for the proposed approach, which shall manifest reduction in development time and costs as well.
Keywords:
Deep learning applications
Preprocessing pipeline
Scholarly big data
Scholarly data applications
Spark ML
Journal
IF:
3
Papers:
2.0W
Citations:
3.2W
Organization
No organization information available
Cited Papers
Enhancement of pyramid solar distiller performance using reflectors, cooling cycle, and dangled cords of wicks
Desalination
IF0
Myocardial gene delivery using molecular cardiac surgery with recombinant adeno-associated virus vectors in vivo
Gene Therapy
IF0

