arrow
Return

A modular image manipulation localization framework using a dual-stream classifier and conditional diffusion models

delete2026-09-24
delete0
PRE
AI
M
Mohammad Zohaib Hamdule *
V
Venkatanareshbabu Kuppili
DOI:10.1007/s11042-026-21940-0delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Image manipulation localization (IML), the task of accurately identifying the manipulated regions within an image, remains a significant challenge. Conventional deep learning methods often struggle with generalization and adapting to new manipulation schemes. To address these limitations, a novel two-stage IML system is proposed. The first stage employs a Dual-Stream Manipulation Classifier, fusing features from the standard RGB domain with low-level forensic noise artifacts extracted via Steganalysis Rich Model (SRM) filters using a ResNet like 4-stage backbone, to accurately classify the input image into one of four common manipulation types: Copy-Move, Splicing, Removal (Inpainting), or Enhancement. In the second stage, the classified image is directed to a specialized Conditional Diffusion Model (CDM), which treats the localization task as an image-to-mask generation problem. These CDMs are conditioned on the input image, and are trained with a modified loss function incorporating both Mean Squared Error (MSE) and Intersection over Union (IoU) to better handle smaller manipulation masks and encourage better spatial structure. Evaluated on the DF2023 dataset, the manipulation classifier achieves an accuracy of 89% and localization CDMs achieve average IoU of 0.70 and F1-Score of 0.77. The system demonstrates competitive performance on benchmark datasets, confirming the efficacy of a modular, generative approach for tackling the complexity of image forgeries.
Keywords:
Image Manipulation Localization
Conditional Diffusion
Dual-Stream Classifier
Image Forensics

Journal

Multimedia Tools and Applications cover
Multimedia Tools and Applications
IF:
3
Papers:
2.0W
Citations:
3.2W

Organization

No organization information available
Cited Papers

Cited Papers

Photorealistic Text-To-Image Diffusion Models with Deep Language Understanding
err2022-01-01
err0
PREAI
errChan,William; Denton,Emily; Fleet,David; Ghasemipour,Kamyar; Lopes,Raphael Gontijo; Ho,Jonathan; Ayan,Burcu Karagol; Li,Lala; Norouzi,Mohammad; Saharia,Chitwan; Salimans,Tim; Saxena,Saurabh; Whang,Jay
errShare
errSave
errShare
errSave
The spread of true and false news online
errSCIENCE
IF45.8
err2018-03-09
err4.1K
errOAAI
errVosoughi, Soroush; Roy, Deb; Aral, Sinan
errShare
errSave
researcher View more