返回
Quantifying Pairwise Similarity for Complex Polymers
DOI:10.1021/acs.macromol.3c00761.png)
摘要
En 中文
Defining the similarity between chemical entities is an essential task in polymer informatics, enabling ranking, clustering, and classification. Despite its importance, the pairwise chemical similarity of polymers remains an open problem. Here, a similarity function for polymers with well-defined backbones is designed based on polymers' stochastic graph representations generated from canonical BigSMILES, a structurally based line notation for describing macromolecules. The stochastic graph representations are separated into three parts: repeat units, end groups, and polymer topology. The earth mover's distance is utilized to calculate the similarity of the repeat units and end groups, while the graph edit distance is used to calculate the similarity of the topology. These three values can be linearly or nonlinearly combined to yield an overall pairwise chemical similarity score for polymers that is largely consistent with the chemical intuition of expert users and is adjustable based on the relative importance of different chemical features for a given similarity problem. This method gives a reliable solution to quantitatively calculate the pairwise chemical similarity score for polymers and represents a vital step toward building search engines and quantitative design tools for polymer data.
Keyword:
CHEMICAL LANGUAGE
ALGORITHM
SMILES
期刊
IF:
5.2
论文数:
3.7W
被引数:
9.4W
机构
引用论文
SMILES, a chemical language and information system. 1. Introduction to methodology and encoding rules微笑,一种化学语言和信息系统。1.介绍方法和编码规则

