arrow
Return

Hyperbolic Pre-Trained Language Model

delete2024-01-01
delete0
PRE
AI
W
Weize Chen
X
Xu Han
Y
Yankai Lin
K
Kaichen He
R
Ruobing Xie
J
Jie Zhou
Z
Zhiyuan Liu *
M
Maosong Sun
DOI:10.1109/TASLP.2024.3407575delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In recent years, we have witnessed significant improvements in pre-trained language models (PLM) brought about by the scaling of parameter sizes and data amounts. However, this also brings high computational and storage costs. In this paper, we present a new direction to improve PLMs without scaling parameters and data: adopting a geometric feature space that is more suitable for encoding the intrinsic structured features of text. Although text is generally considered unstructured data, it possesses rich intrinsic structured features that signify syntactic and semantic relationships. Leveraging these structured features is vital for text understanding. Given that structured features are better encoded in hyperbolic spaces than in the Euclidean spaces used by conventional PLMs, we propose that PLMs should operate entirely within hyperbolic spaces. Our experiments demonstrate the superiority of hyperbolic PLMs over Euclidean PLMs across a wide variety of tasks, using the same parameter and data settings. This suggests that altering the geometry of model representation is a promising direction for model enhancement.
Keywords:
Pre-trained language model
fine-tuning
hyperbolic geometry

Journal

I
IEEE-ACM Transactions on Audio Speech and Language Processing
IF:
5.1
Papers:
2.6K
Citations:
1.1W

Organization

T
tsinghua university
Scholars:
11.7W
Papers: 9.9W
Citations: 137
R
Renmin University of China
Scholars:
8.1K
Papers: 7.7K
Citations: 1.1W
T
Tencent
Scholars:
1.1K
Papers: 891
Citations: 5
researcher View more organizations