1
Return

From Robustness to Improved Generalization and Calibration in Pre-trained Language Models

delete2025-03-19
delete0
delete
OA
AI
J
Josip Jukić *
J
Jan Šnajder
DOI:10.1162/tacl_a_00739delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Enforcing representation smoothness in pre-trained language models (PLMs) through Jacobian and Hessian regularization provides an effective approach for enhancing both robustness and generalization. Although such regularization methods have proven effective in computer vision, their application in natural language processing, where PLM inputs are derived from a discrete domain, poses unique challenges. We introduce JACHESS, a regularization approach for PLMs that minimizes the norms of the Jacobian and Hessian matrices in intermediate representations, using embeddings as substitutes for discrete token inputs. JACHESS supports dual-mode regularization, alternating between fine-tuning with labeled data and regularization with unlabeled data. We evaluate JACHESS on the GLUE benchmark and demonstrate that it consistently and significantly improves in-distribution generalization and enhances performance under domain shift. Across diverse PLMs, JACHESS outperforms comparable representation-based regularization methods and unregularized fine-tuning, while also improving model calibration. Our findings, coupled with a computationally efficient estimator for the Jacobian and Hessian norms, position JACHESS as a robust and widely applicable solution for enhancing PLM performance.

Journal

T
Transactions of the Association for Computational Linguistics
IF:
6.9
Papers:
486
Citations:
5.7K

Organization

U
Univ Zagreb
Scholars:
898
Papers: 417
Citations: 132
Cited Papers

Cited Papers

Citing Papers

Citing Papers