Text → Video
Connect the prompt to the visual subject it describes.
1Tel Aviv University 2Simon Fraser University
These baseline outputs show the requested sound attaching to the nearby distractor. Turn on sound to hear the leakage.
Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same mechanism can introduce subtle and systematic semantic leakage. We study these models by probing and analyzing the “attention triangle,” comprising the three cross-attention edges connecting the text, audio, and video streams, and examine how semantic information is routed across modalities during generation. Our analysis reveals that routing along the audio-video edge is bidirectional: audio can influence video generation, while video can influence audio generation. This edge is shaped by biases encoded in the model's parameters and emerges as a major contributor to leakage: when prompts are in tension with learned priors, cross-modal interactions may override the intended conditioning and reroute semantics toward visually canonical but incorrect outcomes. These effects suggest that semantic artifacts arise not merely from attention spreading beyond its intended target, but from structured, bias-driven interactions along specific pathways. Building on this perspective, we extract attention-derived signals that expose how semantics are distributed and grounded across modalities, and use them as a diagnostic tool to both analyze and deliberately incur leakage under controlled conditions. This enables us to probe the internal dynamics of cross-modal routing and isolate the role of individual interactions. We further leverage these signals to guide inference-time interventions that encourage more consistent cross-modal alignment. Extensive experiments support our analysis and demonstrate improved semantic grounding while preserving generation quality.
Three modalities describe one event. In LTX-2, the audio-video edge emerges as a major contributor to source leakage; full steering coordinates it with both text-conditioned edges.
Connect the prompt to the visual subject it describes.
Connect the instruction to the intended sound or voice.
Ground sound in the correct source, place, and time.
Turn on sound, then use “Play together” to restart each matched pair.
Static locked-off backyard close-up. A black dog stands beside a fluffy white cat, both centered on the lawn under afternoon sun. The black dog stands perfectly still with its mouth closed and ears relaxed. The white cat is the only sound source, opening its mouth wide. It produces loud guttural dog barks, woofs, and growls. The black dog does not move its lips.
Static locked-off desert cartoon close-up at sunset. A silent cowboy stands beside a tall cactus with a friendly face, both centered near a wooden fence. The cowboy keeps his mouth closed. The cactus is the only speaker, its carved mouth moving with clear timing. In a dusty cowboy drawl with a warm western twang, the cactus says: “Howdy partner” Wind whistles faintly through the desert.
Static locked-off cozy bedroom cartoon shot. A furry monster sits beside a small teddy bear on a blanket, both facing the camera. The monster stays quiet, hands folded. The teddy bear is the only sound source, opening its stitched mouth. It produces a deep rumbling monster growl, scratchy and cavernous with a low vibrating finish. The tone is funny rather than scary.
Static locked-off animated kitchen cave hybrid shot. A dragon crouches beside a round blue teapot with a face, both centered under warm light. The dragon stays silent. The teapot is the only sound source, lid-mouth moving in sync. It releases a deep dragon roar mixed with smoky breath and a low fiery rumble. Steam curls from the spout without whistling.
If you find this work useful, please cite it using the BibTeX entry below.