返回
PI1M: A Benchmark Database for Polymer Informatics
DOI:10.1021/acs.jcim.0c00726.png)
摘要
En 中文
Open-source data on large scale are the cornerstones for data-driven research, but they are not readily available for polymers. In this work, we build a benchmark database, called PI1M (referring to similar to 1 million polymers for polymer informatics), to provide data resources that can be used for machine learning research in polymer informatics. A generative model is trained on similar to 12 000 polymers manually collected from the largest existing polymer database PolyInfo, and then the model is used to generate similar to 1 million polymers. A new representation for polymers, polymer embedding (PE), is introduced, which is then used to perform several polymer informatics regression tasks for density, glass transition temperature, melting temperature, and dielectric constants. By comparing the PE trained by the PolyInfo data and that by the PI1M data, we conclude that the PI1M database covers similar chemical space as PolyInfo, but significantly populate regions where PolyInfo data are sparse. We believe that PI1M will serve as a good benchmark database for future research in polymer informatics.
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
5.3
论文数:
9.1K
被引数:
4.0W
机构
引用论文
Predicting Materials Properties with Little Data Using Shotgun Transfer Learning使用shot弹枪迁移学习在少量数据的情况下预测材料特性
ACS CENTRAL SCIENCE
IF10.4
SMILES, a chemical language and information system. 1. Introduction to methodology and encoding rules微笑,一种化学语言和信息系统。1.介绍方法和编码规则
Automatic Chemical Design Using a Data-Driven Continuous Representation of Molecules使用数据驱动的分子连续表示的自动化学设计
ACS CENTRAL SCIENCE
IF10.4
Polymer Genome: A Data-Powered Polymer Informatics Platform for Property Predictions聚合物基因组: 用于属性预测的数据驱动的聚合物信息学平台

