1
Return

Photon: Efficient Prefix-Conditioned Image Captioning with Lightweight Transformer Decoding

delete2026-07-15
delete0
PRE
AI
L
Lakshmi Ganapathi Kodi *
D
Dharmendra Chauhan
D
Dhanush Reddy Gangireddy
S
Sonam Kumari Chaudhary
K
Kalidas Yeturu *
N
Nikhil Srivastava
DOI:10.1007/s10994-026-07072-4delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Image captioning systems based on large vision–language models often rely on computationally intensive cross-attention architectures, limiting their suitability for real-time and resource-constrained deployment. In this work, we propose Photon, a lightweight prefix-conditioned multimodal captioning framework that combines a frozen MobileCLIP vision encoder with a compact decoder-only Transformer incorporating Rotary Positional Embeddings, RMSNorm, and SwiGLU activations. Visual information is injected through a small set of learned prefix tokens, enabling efficient multimodal conditioning without region-based processing or heavy cross-modal attention. On the MS-COCO Karpathy split, Photon achieves a CIDEr score of 108.59 while requiring only 12.41 M trainable parameters and 3.72 GFLOPs for image to caption generation. The model demonstrates competitive semantic alignment performance and improves inference efficiency, achieving 2.41 $$\times $$ GPU and 8.28 $$\times $$ CPU speed-ups compared to larger pretrained baselines. Zero-shot evaluation on Flickr30K, NoCaps and TextCaps further indicates consistent cross-dataset generalization across lexical, consensus-based, and embedding-based metrics. Batch-scaling analysis reveals near-linear throughput growth up to batch size 512, highlighting effective parallelization of the decoder. These results suggest that prefix-based multimodal conditioning with a modern lightweight decoder provides a favorable balance between caption quality and computational efficiency, making it suitable for practical deployment scenarios.
Keywords:
Image captioning
Multimodal learning
Prefix conditioning
Vision–language models
Transformer decoding

Journal

Machine Learning cover
Machine Learning
IF:
2.9
Papers:
2.6K
Citations:
3.4W

Organization

I
Indian Institute of Technology Tirupati
Scholars:
136
Papers: 61
Citations: 610
N
National Institute of Technology
Scholars:
1.5K
Papers: 752
Citations: 2.1K
Cited Papers

Cited Papers

Citing Papers

Citing Papers