Return
Optimizing Representation for Abstractive Multidocument Summarization Based on Adversarial Learning Strategy
DOI:10.1109/TCDS.2025.3563357.png)
Abstract
En 中文
The Abstractive multi-document summarization (MDS) is a crucial technique in cognitive computing, enabling the efficient synthesis of a documents cluster into a concise and complete summary. Despite recent advances, existing approaches still face challenges in representation learning when processing large-scale documents clusters: 1) incomplete semantic learning caused by documents truncation or exclusion; 2) the incorporation of noise, such as irrelevant or redundant information from documents; and 3) the potential omission of critical content due to partial coverage of documents. These limitations collectively undermine the semantic integrity and conciseness of the generated summaries. To address these issues, we propose TALER, a two-stage representation architecture enhanced by adversarial learning for abstractive MDS, which reformulates the MDS task as a single-document optimization problem. In Stage I, TALER focuses on enhancing single-document representations by maximizing semantic learning from each document in the cluster and employing the adversarial learning to suppress the introduction of documents noise. In Stage II, TALER conducts multidocument semantic fusion and summary generation by aggregating the learned document embeddings based on Stage I into a cluster-level representation through a pooling mechanism, followed by a self-attention module to capture salient content and produce the final summary. Experimental results on the Multi-News, DUC04, and Multi-XScience datasets demonstrate that TALER consistently outperforms existing baseline models across multiple evaluation metrics.
Keywords:
Semantics
Training
Noise
Adversarial machine learning
Decoding
Data mining
Transformers
Representation learning
Fans
Encoding
Abstractive summarization
adversarial learning
multidocument summarization (MDS)
natural language processing (NLP)
Journal
IF:
4.9
Papers:
1.0K
Citations:
3.5K

