Return
FoodDiff: A Collaborative Relationship Perception Framework for Food Image Synthesis Using Diffusion Models
DOI:10.1109/tmm.2026.3654435.png)
Abstract
En 中文
Food image generation is a typical application of text-to-image (T2I) models. The core difference between food image synthesis and other T2I tasks is that there exist complex collaborative relationships among ingredients, cooking actions, and food images, which determine the appearance of dishes. However, existing food image generation models generally ignore or fail to sufficiently utilize such collaborative relationships, which hinders the model from precisely perceiving the shapes and details of food. Furthermore, the pre-training distribution of T2I models is usually noisy and differs from the user-preferred food feature distributions, resulting in deviations from human aesthetics. To address the above issues, we propose FoodDiff, a collaborative relationship-aware diffusion model for food image generation, which consists of three key components: (1) To perceive collaborative relationships, we propose a collaborative relation module to extract these relations and inject them into the image generation process. (2) To sufficiently interact with the relationships between recipe semantics and food representations, we propose a recipe fusion fine-tuning module to precisely fuse recipe semantics with visual features and fine-tune the pre-trained model. (3) To make the pre-training feature distribution conform to human preference, we introduce an image reward feedback mechanism to optimize the aesthetics of food images. In addition, we propose a high-quality food dataset named Food-Aesthetic with exquisite plates and elaborate annotations. Extensive experiments and human evaluations show that FoodDiff has superior image aesthetics and semantic consistency.
Keywords:
Collaboration
Image synthesis
Diffusion models
Visualization
Semantics
Text to image
Noise reduction
Feature extraction
Telecommunications
Generators
Text-to-image generation
aesthetic food images
diffusion model
image reward feedback
Journal
IF:
9.7
Papers:
4.5K
Citations:
2.4W
Organization
Cited Papers
No cited papers available

