Return
GSID: Guided single-image generation with diffusion models
DOI:10.1016/j.jvcir.2026.104921.png)
Abstract
En 中文
To address the prevalent challenges of mode collapse and the fidelity-diversity trade-off in single-image generation, we introduce GSID, a multi-scale framework for structurally-guided synthesis via diffusion models. GSID utilizes a modified ConvNeXt Block to enhance structural regularization, effectively capturing global context while preventing overfitting in a data-scarce environment. We also introduce a High-Frequency Extractor Module (H-FEM) to improve the preservation of fine-grained details, achieving explicit structure-texture decoupling. Furthermore, we formulate the Dynamic Mask Strategy as a reactive closed-loop inference mechanism. Instead of relying on manually predefined static perturbation parameters, GSID periodically measures perceptual deviation during sampling, computes an error signal against a dynamic fidelity threshold, and adaptively adjusts the mask strength to regulate the generation trajectory. This feedback-regulated inference process enhances image diversity and controls structural variation without requiring additional training or per-instance manual parameter tuning. Evaluated on various natural and medical image datasets, GSID generates diverse, high-quality samples, demonstrating superior structural consistency and perceptual quality compared to existing state-of-the-art methods.
Keywords:
Single-image generation
Diffusion models
Structure-texture decoupling
Mask Strategy
Journal
IF:
3.1
Papers:
540
Citations:
5.6K
Organization
No organization information available
Cited Papers
No cited papers available

