1
Return

Challenges and Vision for Standardization of Biopolymer Data Sets for Machine Learning

delete2026-06-10
delete0
delete
OA
AI
J
Jessica Lalonde *
D
Defne Circi
B
Babetta L. Marrone
S
Stefan Zauscher *
L
L. Catherine Brinson
DOI:10.1021/acs.biomac.6c00211delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Machine learning (ML) is transforming materials research, yet potential for biopolymer discovery remains constrained by fragmented data and nonstandardized reporting. Biopolymers differ significantly from synthetic polymers, requiring specialized approaches to represent their biosynthetic origins, hierarchical structures, and application-specific metrics. In this Perspective, we identify three core challenges limiting biopolymer representation: information encoding, data quality, and data sharing. We describe the most pressing issues and propose commensurate approaches to address each key challenge. Recommendations include the design and adoption of biopolymer-specific fingerprinting and representation frameworks, development of hybrid human-large language model (LLM) data extraction strategies, and expanding Findable, Accessible, Interoperable, Reusable (FAIR)-compliant repositories. We propose a robust foundation to define interoperable, high-quality data sets that capture the full context of biopolymer materials. Standardized metadata, shared ontologies, and community-driven infrastructure would enable scalable, reproducible workflows and accelerate the ML-driven development of biopolymers.
Keywords:
Biomaterials
Biopolymers
Machine learning
Plastics
Polymers

Journal

Biomacromolecules cover
Biomacromolecules
IF:
5.4
Papers:
1.2W
Citations:
4.1W

Organization

L
los alamos national laboratory
Scholars:
641
Papers: 238
Citations: 0
D
duke university
Scholars:
7.3K
Papers: 2.9K
Citations: 2
Cited Papers

Cited Papers

Citing Papers

Citing Papers