The Attention Triangle
in Audio-Video Models

Sagi Polaczek1, Noa Kraicer1, Gal Metzer1, Zhuo Ning2, Ali Mahdavi-Amiri2, Raja Giryes1, Daniel Cohen-Or1

1Tel Aviv University    2Simon Fraser University

The Attention Triangle in Audio-Video Models. Top-left: baseline T2AV generation exhibits strong leakage; the pirate visually inherits parrot-like attributes and becomes the speaking source. Top-right: Ours-Text restores the pirate's appearance, but speech remains incorrectly localized to the pirate. Bottom-left: Ours-AV correctly localizes speech to the parrot, but appearance leakage persists in the pirate. Bottom-right: Ours-Full jointly restores correct appearance and localizes speech to the intended source. The prompt is shown below the figure.
Prompt

Static locked-off medium close-up, no camera movement. A weathered pirate with a black tricorn hat and thick grey beard stands on a wooden dock at sunset. A bright green-and-red macaw parrot perches on his right shoulder, facing forward. Golden hour light reflects off calm harbor water behind them. The pirate stands motionless with arms crossed, staring stoically past the camera. The parrot shifts its weight slightly, ruffling its feathers, head tilted toward the lens. The parrot is the only character that speaks. The macaw opens its beak and begins articulating words clearly synced to speech. In a shrill, scratchy, high-pitched voice, the parrot says: “We are lost at sea.” The pirate does not move or speak. The only audio is the parrot's voice and lapping waves.

Where the sound leaks

These baseline outputs show the requested sound attaching to the nearby distractor. Turn on sound to hear the leakage.

Baseline leakageThe cowboy speaks instead.
Baseline leakageThe monster growls instead.
Baseline leakageThe dragon roars instead.
Baseline leakageThe dog barks instead.

Abstract

Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same mechanism can introduce subtle and systematic semantic leakage. We study these models by probing and analyzing the “attention triangle,” comprising the three cross-attention edges connecting the text, audio, and video streams, and examine how semantic information is routed across modalities during generation. Our analysis reveals that routing along the audio-video edge is bidirectional: audio can influence video generation, while video can influence audio generation. This edge is shaped by biases encoded in the model's parameters and emerges as a major contributor to leakage: when prompts are in tension with learned priors, cross-modal interactions may override the intended conditioning and reroute semantics toward visually canonical but incorrect outcomes. These effects suggest that semantic artifacts arise not merely from attention spreading beyond its intended target, but from structured, bias-driven interactions along specific pathways. Building on this perspective, we extract attention-derived signals that expose how semantics are distributed and grounded across modalities, and use them as a diagnostic tool to both analyze and deliberately incur leakage under controlled conditions. This enables us to probe the internal dynamics of cross-modal routing and isolate the role of individual interactions. We further leverage these signals to guide inference-time interventions that encourage more consistent cross-modal alignment. Extensive experiments support our analysis and demonstrate improved semantic grounding while preserving generation quality.

The Attention Triangle

Three modalities describe one event. In LTX-2, the audio-video edge emerges as a major contributor to source leakage; full steering coordinates it with both text-conditioned edges.

Attention seeded on the parrot spreads across audio tokens and returns to the video stream concentrated on the pirate.
Audio-mediated video-to-video leakage. Starting from the prompt-designated parrot (left), the video → audio hop spreads attention across audio tokens (middle). The following audio → video hop moves the attention onto the pirate (right), exposing an indirect path for source-attribution leakage.

Baseline audio → video attention

Baseline speech-active audio attention concentrated across the pirate.

Steered audio → video attention

Steered speech-active audio attention concentrated on the parrot.
Speech-active audio Audio waveform with speech between approximately 1.9 and 3.5 seconds.
Steering source attribution. In the baseline, attention from speech-active audio concentrates on the pirate. Full triangle steering redirects it to the prompt-designated parrot.

Text → Video

Connect the prompt to the visual subject it describes.

Text → Audio

Connect the instruction to the intended sound or voice.

Audio ↔ Video

Ground sound in the correct source, place, and time.

Baseline vs. Full Triangle Steering

Turn on sound, then use “Play together” to restart each matched pair.

The cat barks. The dog stays silent.

Baseline
Full triangle steering
Prompt

Static locked-off backyard close-up. A black dog stands beside a fluffy white cat, both centered on the lawn under afternoon sun. The black dog stands perfectly still with its mouth closed and ears relaxed. The white cat is the only sound source, opening its mouth wide. It produces loud guttural dog barks, woofs, and growls. The black dog does not move its lips.

The cactus says “Howdy, partner.”

Baseline
Full triangle steering
Prompt

Static locked-off desert cartoon close-up at sunset. A silent cowboy stands beside a tall cactus with a friendly face, both centered near a wooden fence. The cowboy keeps his mouth closed. The cactus is the only speaker, its carved mouth moving with clear timing. In a dusty cowboy drawl with a warm western twang, the cactus says: “Howdy partner” Wind whistles faintly through the desert.

The teddy bear growls. The monster stays quiet.

Baseline
Full triangle steering
Prompt

Static locked-off cozy bedroom cartoon shot. A furry monster sits beside a small teddy bear on a blanket, both facing the camera. The monster stays quiet, hands folded. The teddy bear is the only sound source, opening its stitched mouth. It produces a deep rumbling monster growl, scratchy and cavernous with a low vibrating finish. The tone is funny rather than scary.

The teapot roars. The dragon stays silent.

Baseline
Full triangle steering
Prompt

Static locked-off animated kitchen cave hybrid shot. A dragon crouches beside a round blue teapot with a face, both centered under warm light. The dragon stays silent. The teapot is the only sound source, lid-mouth moving in sync. It releases a deep dragon roar mixed with smoky breath and a low fiery rumble. Steam curls from the spout without whistling.

Citation

If you find this work useful, please cite it using the BibTeX entry below.