arrow
Return

A Bond-Level Sequence Framework for Molecular Representation Learning with Structural Constraints

delete2026-06-05
delete0
delete
OA
AI
H
Haoran Fan
H
Haoqiang Qi
X
Xin Huang
D
Dongyang Zhu
N
Na Wang *
T
Ting Wang *
郝红勋 cover
郝红勋 (Hongxun Hao) *
DOI:10.3390/molecules31111972delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Molecular property prediction is a fundamental task in drug discovery and materials design. While graph neural networks (GNNs) and SMILES-based Transformers have made significant strides, the former are often limited by local message-passing bottlenecks such as over-squashing, while the latter frequently lack explicit topological constraints and suffer from severe vocabulary imbalance. In this work, we revisit the granularity of molecular modeling and propose a representation learning framework built upon bond-level sequences. Our framework models molecules as sequences of directed bond tokens and introduces a structure-aware hybrid attention mechanism. By imposing hard topological constraints on a subset of attention heads to reinforce local connectivity while preserving global receptive fields in the remaining heads, the design is intended to separate short-range chemical bonding from long-range contextual dependencies. For pre-training, we implemented a multi-scale consistency learning paradigm, which utilizes an atom-centric group masking strategy to induce a hierarchical loss of local structural information and employs contrastive and triplet losses to ensure identity consistency across varying scales of structural degradation. Furthermore, by incorporating macro-scale physicochemical descriptors (e.g., LogP, TPSA) as global anchors, we examined how the inclusion of global attribute bias can provide weak physicochemical priors during pre-training, while its effect during downstream fine-tuning remains task-dependent. Experimental results demonstrate that our lightweight model, with approximately 3.5 million parameters, exhibits a dataset-dependent performance profile across MoleculeNet benchmarks and shows promising behavior on selected topology-sensitive tasks, particularly MUV. Ablation studies further analyze the contribution of bond-level connectivity, the stage-dependent dynamics of global attribute bias, structured masking, and pre-training configurations. Ultimately, this work provides an alternative representation design for molecular modeling, offering a parameter-efficient option for future molecular learning systems alongside traditional SMILES-based and graph-based formulations.
Keywords:
molecular property prediction
transformer
bond-level representation
self-supervised learning
structural degeneracy

Journal

Molecules cover
Molecules
IF:
4.6
Papers:
6.4W
Citations:
23.7W

Organization

T
tianjin university
Scholars:
7.9W
Papers: 5.7W
Citations: 88