返回
FeaMix: Feature Mix With Memory Batch Based on Self-Consistency Learning for Code Generation and Code Translation
DOI:10.1109/TETCI.2024.3395531.png)
摘要
En 中文
Data augmentation algorithms, such as back translation, have shown to be effective in various deep-learning tasks. Despite their remarkable success, there has been a hurdle to applying data augmentation algorithms to code-related tasks since code consists of discrete tokens with uniqueness and certainty. In this work, we propose FeaMix, a novel yet simple data augmentation approach designed for the feature mix with memory batch based on self-consistency learning. FeaMix has a couple of uniqueness. First, it specially selects the samples to be mixed by memory batch to guarantee that the generated features are in the same spatial distribution as the mixed features. Second, it extends the self-consistency learning technique to optimize the language model for code-related tasks. With extensive experiments, we empirically validate that our method outperforms several baseline models and traditional data augmentation methods on code generation and code translation. It is noteworthy that we achieve state-of-the-art results in the CoNaLa and CodeTrans benchmarks, with a significant improvement of 1.9% in the Exact Match accuracy metric for code translation tasks.
Keyword:
Code generation
code translation
data augmentation
memory batch
self-consistency learning
期刊
I
IF:
6.5
论文数:
1.4K
被引数:
4.5K
机构
引用论文
Evaluation of awareness about pharmacovigilance and adverse drug reaction monitoring in resident doctors of a tertiary care teaching hospital对三级护理教学医院住院医师关于药物警戒和不良反应监测的认知评估

