arrow
Return

Understanding generative AI output with embedding models

delete2025-11-26
delete0
PRE
AI
M
Max Vargas
R
Reilly Cannon
A
Andrew G. Engel
A
Anand D. Sarwate *
T
Tony Chiang *
DOI:10.1126/sciadv.adx4082delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Constructing high-quality features is critical to any quantitative data analysis. While feature engineering was historically addressed by carefully handcrafting data representations on the basis of domain expertise, deep neural networks (DNNs) now offer a radically different approach. DNNs implicitly engineer features by transforming their input data into hidden feature vectors called embeddings. For embedding vectors produced by foundation models—which are trained to be useful across many contexts—we demonstrate that simple and well-studied dimensionality-reduction techniques such as principal components analysis uncover inherent heterogeneity in input data concordant with human-understandable explanations. Of the many applications for this framework, we find empirical evidence that there is intrinsic separability between real samples and those generated by artificial intelligence.

Journal

Science Advances cover
Science Advances
IF:
12.5
Papers:
2.0W
Citations:
18.1W

Organization

P
Pacific Northwest National Laboratory
Scholars:
9.0K
Papers: 6.3K
Citations: 14
R
Rutgers University
Scholars:
1.3K
Papers: 683
Citations: 4.9W