Text-to-image diffusion models can generate high-quality images from individual prompts, yet it remains unclear how reliably they can fuse two distinct concepts into a single coherent image in a strictly zero-shot setting. Inspired by human combinational creativity, we operationalize visual concept blending here as conditioning a diffusion model on a pair of text prompts and ask whether the resulting image preserves recognizable traits of both concepts while remaining visually coherent. We study four inference-time blending strategies that intervene at different stages of the diffusion process, including embedding-space interpolation, mid-denoising prompt switching, timestep-wise prompt alternation, and layer-wise conditioning within the denoising network. We evaluate these strategies across diverse blend categories, spanning object–object blends, compound concepts, style–content combinations, and architectural landmarks, and we complement the analysis with a user study involving 100 participants. Results indicate that no single strategy dominates across all categories: timestep-scheduling approaches are often preferred for producing recognisable and well-integrated hybrids, while embedding- and layer-based interventions exhibit characteristic strengths and failure modes in specific settings. Finally, we analyse practical control factors such as prompt ordering, blend ratio, and random seed, showing how they can substantially alter outcomes and providing concrete levers for steering blends more predictably in creative applications.

Olearo, L., Longari, G., Raganato, A., Peñaloza, R., Melzi, S. (2026). Blending concepts with text-to-image diffusion models. DISCOVER ARTIFICIAL INTELLIGENCE, 6(1) [10.1007/s44163-026-01134-1].

Blending concepts with text-to-image diffusion models

Olearo L.
;
Raganato A.;Peñaloza R.;Melzi S.
2026

Abstract

Text-to-image diffusion models can generate high-quality images from individual prompts, yet it remains unclear how reliably they can fuse two distinct concepts into a single coherent image in a strictly zero-shot setting. Inspired by human combinational creativity, we operationalize visual concept blending here as conditioning a diffusion model on a pair of text prompts and ask whether the resulting image preserves recognizable traits of both concepts while remaining visually coherent. We study four inference-time blending strategies that intervene at different stages of the diffusion process, including embedding-space interpolation, mid-denoising prompt switching, timestep-wise prompt alternation, and layer-wise conditioning within the denoising network. We evaluate these strategies across diverse blend categories, spanning object–object blends, compound concepts, style–content combinations, and architectural landmarks, and we complement the analysis with a user study involving 100 participants. Results indicate that no single strategy dominates across all categories: timestep-scheduling approaches are often preferred for producing recognisable and well-integrated hybrids, while embedding- and layer-based interventions exhibit characteristic strengths and failure modes in specific settings. Finally, we analyse practical control factors such as prompt ordering, blend ratio, and random seed, showing how they can substantially alter outcomes and providing concrete levers for steering blends more predictably in creative applications.
Articolo in rivista - Articolo scientifico
diffusion, concept blending
English
13-apr-2026
2026
6
1
487
open
Olearo, L., Longari, G., Raganato, A., Peñaloza, R., Melzi, S. (2026). Blending concepts with text-to-image diffusion models. DISCOVER ARTIFICIAL INTELLIGENCE, 6(1) [10.1007/s44163-026-01134-1].
File in questo prodotto:
File Dimensione Formato  
Olearo-2026-Discover Artificial Intelligence-preprint.pdf

accesso aperto

Tipologia di allegato: Submitted Version (Pre-print)
Licenza: Creative Commons
Dimensione 6.11 MB
Formato Adobe PDF
6.11 MB Adobe PDF Visualizza/Apri
Olearo-2026-Discover Artificial Intelligence-VoR.pdf

accesso aperto

Tipologia di allegato: Publisher’s Version (Version of Record, VoR)
Licenza: Creative Commons
Dimensione 6.25 MB
Formato Adobe PDF
6.25 MB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/10281/618081
Citazioni
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
Social impact