Return
Photon: Efficient Prefix-Conditioned Image Captioning with Lightweight Transformer Decoding
L
D
D
S
K
N
DOI:10.1007/s10994-026-07072-4.png)
Abstract
En 中文
Image captioning systems based on large vision–language models often rely on computationally intensive cross-attention architectures, limiting their suitability for real-time and resource-constrained deployment. In this work, we propose Photon, a lightweight prefix-conditioned multimodal captioning framework that combines a frozen MobileCLIP vision encoder with a compact decoder-only Transformer incorporating Rotary Positional Embeddings, RMSNorm, and SwiGLU activations. Visual information is injected through a small set of learned prefix tokens, enabling efficient multimodal conditioning without region-based processing or heavy cross-modal attention. On the MS-COCO Karpathy split, Photon achieves a CIDEr score of 108.59 while requiring only 12.41 M trainable parameters and 3.72 GFLOPs for image to caption generation. The model demonstrates competitive semantic alignment performance and improves inference efficiency, achieving 2.41 $$\times $$ GPU and 8.28 $$\times $$ CPU speed-ups compared to larger pretrained baselines. Zero-shot evaluation on Flickr30K, NoCaps and TextCaps further indicates consistent cross-dataset generalization across lexical, consensus-based, and embedding-based metrics. Batch-scaling analysis reveals near-linear throughput growth up to batch size 512, highlighting effective parallelization of the decoder. These results suggest that prefix-based multimodal conditioning with a modern lightweight decoder provides a favorable balance between caption quality and computational efficiency, making it suitable for practical deployment scenarios.
Keywords:
Image captioning
Multimodal learning
Prefix conditioning
Vision–language models
Transformer decoding
Journal
IF:
2.9
Papers:
2.6K
Citations:
3.4W
