arrow
Return

Vision Transformer Representations for Efficient Content-Based Image Retrieval

delete2026-01-01
delete0
PRE
AI
S
Stanisław Łażewski *
B
Bogusław Cyganek
DOI:10.1007/978-3-032-03708-4_11delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Content-based image retrieval (CBIR) is one of the basic tasks of computer vision. Numerous studies have been conducted, leading to many groundbreaking methods based on deep neural networks and even more recently on vision transformers (ViT). In this article, we propose a new CBIR method based on the original self-distilled with no labels semantic features (DINO), obtained using ViT, and then additionally compressed using the principal and neighbourhood component analysis. We show highly accurate results on non trivial datasets such as Caltech-256, as well as on histopathological scans such as Kather and BreaKHis. Our method freely compares with the best CBIR approaches while having very compact image representations.
Keywords:
content-based image retrieval
ViT
DINO
histopathology
component analysis
colorectal cancer
breast cancer
BreaKHis

Journal

A
ARTIFICIAL INTELLIGENCE AND SOFT COMPUTING, ICAISC 2025, PT II
IF:
0
Papers:
30
Citations:
0

Organization

A
agh university of krakow
Scholars:
1.5K
Papers: 693
Citations: 0