arrow
Return

Efficient Similarity-Based Datasheet Retrieval and Analysis Using Retrieval-Augmented Generation for Electronic Component Selection

delete2026-05-30
delete0
delete
OA
AI
D
Dan Curavale
G
Georgian Nicolae
A
Alexandru Caranica
H
Horia Cucu *
C
Corneliu Burileanu
V
Valentina Davidoiu
A
Andi Buzo
G
Georg Pelz
DOI:10.3390/electronics15112301delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Component obsolescence and supply-chain disruptions increasingly force engineers to spend significant time manually searching and comparing PDF datasheets to identify compatible replacement parts. We propose an AI-powered datasheet assistant based on a Retrieval-Augmented Generation (RAG) pipeline that automatically processes datasheets to accelerate component identification and matching. The core contribution is a summary-driven retrieval mechanism: a Large Language Model (LLM) generates a structured semantic summary of an input datasheet, and the vector embedding of this summary is used to retrieve semantically similar components from a reference database. The system also supports natural language question answering and structured component comparison. Its architecture separates scalable text-only reference indexing from more expensive query-time summarization and reranking. Validation includes a controlled synthetic benchmark and a pilot-scale real-world evaluation on 18 publicly listed microcontroller datasheets grouped into six engineering families. The synthetic benchmark is used to assess pipeline behavior under controlled conditions, while the real-world evaluation measures performance on heterogeneous manufacturer datasheets. In the real-world evaluation, structured summaries generated with Claude Sonnet 4.5 combined with cross-encoder reranking achieved a 72.2% Family Retrieval Rate at k = 1 (13/18; Wilson 95% CI: 49.1–87.5%). Additional experiments with local LLM summaries indicate that retrieval performance depends strongly on summary quality and model capability, with lightweight local summarizers producing lower first-candidate retrieval performance in this setup. The analysis further reports confidence intervals, no-summary baselines, chunking sensitivity, and an Image Reference Rate metric used as a lexical reference proxy rather than a direct measure of visual grounding.
Keywords:
retrieval-augmented generation (RAG)
datasheet analysis
component similarity search
large language models (LLMs)
multimodal document understanding
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Electronics cover
Electronics
IF:
2.6
Papers:
9.6K
Citations:
4.7W

Organization

I
infineon technologies
Scholars:
778
Papers: 499
Citations: 0
P
politehnica university of bucharest
Scholars:
55
Papers: 34
Citations: 0