返回
Improving BERT With Self-Supervised Attention
DOI:10.1109/ACCESS.2021.3122273.png)
摘要
En 中文
One of the most popular paradigms of applying large pre-trained NLP models such as BERT is to fine-tune it on a smaller dataset. However, one challenge remains as the fine-tuned model often overfits on smaller datasets. A symptom of this phenomenon is that irrelevant or misleading words in the sentence, which are easy to understand for human beings, can substantially degrade the performance of these fine-tuned BERT models. In this paper, we propose a novel technique, called Self-Supervised Attention (SSA) to help facilitate this generalization challenge. Specifically, SSA automatically generates weak, token-level attention labels iteratively by probing the fine-tuned model from the previous iteration. We investigate two different ways of integrating SSA into BERT and propose a hybrid approach to combine their benefits. Empirically, through a variety of public datasets, we illustrate significant performance improvement using our SSA-enhanced BERT model.
Keyword:
Task analysis
Bit error rate
Predictive models
Data models
Training
Training data
Licenses
Natural language processing
attention model
text classification
BERT
pre-trained model
期刊
IF:
3.6
论文数:
9.8W
被引数:
29.4W
机构
引用论文
Facile synthesis of tremelliform Co0.85Se nanosheets: An efficient catalyst for the decomposition of hydrazine hydratetremelliform Co0.85Se纳米片的简易合成: 水合肼分解的有效催化剂
Oxidative stress in liver of grass carp Ctenopharyngodon idella naturally infected with Saprolegnia parasitica and its influence on disease pathogenesis天然感染水蚤的草鱼肝脏氧化应激及其对疾病发病机制的影响
Determining the three-dimensional geometry of a dike swarm and its impact on later rift geometry using seismic reflection data
Geology
IF0

