arrow
Return

Learning Discriminative Syntax-Semantic Patterns With Transformer-Based Contrastive Learning

delete2026-01-01
delete0
PRE
AI
K
Kiran Mayee Adavala *
O
Om Adavala
DOI:10.54364/AAIML.2026.61273delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
This paper proposes a dual-encoder framework that learns syntax-and semantics-aligned sentence embeddings using contrastive learning. The proposed Syntax-Semantic Contrastive Pretraining (SSCP) model employs two transformer encoders to separately model syntactic structure and semantic content, which are aligned in a shared embedding space via a symmetric contrastive objective. Across standard benchmarks, SSCP consistently outperforms strong baselines such as BERT, SimCSE, and SyntaxBERT. In particular, SSCP improves Spearman correlation on STS-B by +1.4 points over SimCSE, increases PAWS paraphrase accuracy by +2.5 points, and achieves 96.1% accuracy on TREC question classification, exceeding existing syntax-aware models. Probing experiments further show gains of up to +4-5 points on syntactic structure prediction tasks, confirming that SSCP preserves grammatical information while maintaining semantic robustness. These results demonstrate that explicitly aligning syntactic and semantic views yields representations that are more discriminative, interpretable, and robust than single-view or syntax-augmented approaches, positioning SSCP as a principled multi-view pretraining strategy for structure-aware language understanding.
Keywords:
Transformer models
Contrastive learning
Syntax-Semantic representation
Sentence embeddings
Pattern recognition
Representation learning
Natural language processing
Multi-View learning
Machine learning

Journal

A
Advances in Artificial Intelligence and Machine Learning
IF:
0.5
Papers:
41
Citations:
0

Organization

Kakatiya University cover
Kakatiya University
Scholars:
596
Papers: 401
Citations: 469
I
indian institute of technology system (iit system)
Scholars:
9.4W
Papers: 9.9W
Citations: 93