The image is enough to guess a usually correct sound
On a congruent audiovisual clip, a visual guess and genuine auditory recognition produce the same answer, so conventional accuracy cannot distinguish them.
A diagnostic benchmark that locks the video and changes only the auditory facts. Its matched Core tests agreement, absence, contradiction, and coexistence; CLASH keeps an old fact visible while speech supplies a mutually exclusive new one.
Visual objects and typical sounds are highly correlated in natural videos. When shown a piano, car, or dog, a model can ignore the audio track and still generate “piano music,” “engine noise,” or “barking” from visual priors. The answer looks plausible, but it does not prove that the model listened.
On a congruent audiovisual clip, a visual guess and genuine auditory recognition produce the same answer, so conventional accuracy cannot distinguish them.
If the same piano video actually contains birdsong, answering “piano” reveals visual override. If piano and bird coexist but the model reports only piano, it reveals auditory omission.
OST-DiagBench turns every source clip into a paired experiment. Video content and visual cues are locked; only the presence of the complete CLEAN track and the addition of a donor sound are controlled. Changes in model output can therefore be attributed to changes in auditory evidence.
Vision and audio agree: can the model report the original source?
Source and donor coexist: can the model capture both?
The sound is absent: does the image still induce the model to report the source?
The donor replaces the original track: does the model trust hearing or vision?
Construction follows one rule: within a group, change only preregistered variables. MIX retains the complete CLEAN waveform and overlays the donor; it requires no source separation and makes no claim of having an isolated source stem.
Select source clips with visible events and reliable acoustic labels.
Assign near / far events in a balanced way from a non-overlapping donor bank.
Generate CLEAN, MUTE, matched SWAP, and matched MIX.
Apply the same attenuation to the entire group to avoid condition-specific distortion.
Use multilabel validation; reject the entire group if any member fails.
Share donor identity, a natural 2–4 s segment, onset/offset, 50–100 ms fades, gain, codec, and mastering. Donor-parameter differences are excluded from the comparison.
MIX uses overlap-window DLR: δ = Ldonor − Lsource. SWAP has no source, so it uses donor-to-background SNR.
CLEAN/SWAP/MIX-AONLY retain final audio that is byte-identical to the corresponding AV condition. A donor recognized in A-ONLY but missed in AV is direct diagnostic evidence of cross-modal suppression.
Core fixes the onset at the middle. A separate panel uses the same donor segment and gain while varying only early / middle / late position to analyze position, recency, and retention.
FAR-SWAP crosses broad acoustic classes; NEAR-SWAP changes the fine-grained label within the same superclass. This distinguishes “not heard,” “heard only at a coarse level,” and “heard but overridden by vision.”
Check the locked-video signature, source/donor recognizability, DLR/SNR within ±1 dB, true peak ≤ −1 dBTP, A-ONLY hashes, and editing artifacts.
The table reports only the final Challenge benchmark. Constructed counts records generated from 1,068 frozen source clips or balanced subsets; Accepted counts records that passed audio, video, and group-consistency validation and entered the report. Native and neutral-video A-ONLY are separate input routes and are counted separately, but they share the final audio of the corresponding AV condition.
| Suite / operation | What changes | Constructed records | Accepted records |
|---|---|---|---|
| Factorial Core · locked video | |||
| CLEAN | Retain the complete source-containing CLEAN track | 1,068 | 510 |
| MUTE | Remove the complete CLEAN track and replace it with matched low-level background | 1,068 | 510 |
| SWAP | Background + donor, creating a replacement contradiction | 1,068 | 510 |
| MIX | Complete CLEAN track + donor, creating auditory coexistence | 1,068 | 510 |
| Calibration Controls · native audio-only + neutral-video surrogate | |||
| CLEAN-AONLY | Identical CLEAN audio through two non-semantic video routes | 2,136 | 1,020 |
| SWAP-AONLY | Identical SWAP audio through two non-semantic video routes | 2,136 | 1,020 |
| MIX-AONLY | Identical MIX audio through two non-semantic video routes | 2,136 | 1,020 |
| Evidence Strength · independently balanced panel | |||
| MIX DLR −10 / 0 / +10 | Change only the donor-to-source loudness ratio | 288 | 157 |
| Matched SWAP gain/SNR | Match MIX donor gain; report donor/background SNR | 288 | 157 |
| Temporal Position · independently balanced panel | |||
| MIX early / middle / late | Use the same donor and gain; change only the onset | 288 | 158 |
| SWAP early / middle / late | Match the temporal position used in MIX | 288 | 158 |
| Semantic Granularity · fine-label auditable subset | |||
| FAR-SWAP + AONLY | Donor from a different broad acoustic class; pair AV with native A-ONLY | 96 | 86 |
| NEAR-SWAP + AONLY | Same superclass but a different fine-grained category; pair AV with native A-ONLY | 96 | 92 |
| All released event-level Challenge records | 12,024 | 5,908 | |
| Fact-level CLASH · separate source-utterance path | |||
| CLASH A+V + A-ONLY | Keep the old visible fact; replace aligned speech with a mutually exclusive new fact | — | 510 items |
All four players come from the same accepted source group, ave-0a7683f182e2. Expected facts come from the manifest; model text is the actual output of Qwen2.5-Omni-7B offline.
The players use a uniform neutral-video carrier; their audio is byte-identical to the corresponding AV condition. Below are real outputs from the same model on the neutral-video route.
Matched contrast: The same MIX audio yields only “guitar” with the original guitar video, but “guitar + dog” with neutral video. This is exactly the phenomenon that the Cross-Modal Suppression Gap is designed to quantify; aggregate metrics still report native and neutral routes separately.
The three MIX videos use the same source, donor, segment, middle onset, and mastering; only overlap-window DLR changes. Real responses progress from hearing only the flute, to hearing both, to the donor overwhelming the source.
All three videos use the same guitar + dog pair and 0 dB DLR. Real Qwen3 offline outputs show that early and middle donors are omitted, while the late donor is retained in the final answer.
The original vehicle-horn track is removed from both SWAP videos below. A-ONLY recognizes the donor, but when paired with the original video, Qwen2.5 instead reports a visually compatible train horn in both cases.
OST-DiagBench contains two frozen stress regimes. The event-level branch admits 1,068 source clips and reports the complete Core on 510 groups valid in CLEAN, MUTE, SWAP, and MIX. The separate CLASH branch retains 510 manually reviewed fact conflicts from 349 source videos.
The 1,068 frozen source clips are stratified across four acoustic–visual categories.
Validity gates, difficulty selection, and audio-quality gates are separate stages; 510 is the shared denominator for the complete Core.
A unified judge maps each response to one of four states and interprets it according to which sounds are actually present in that condition. This separates hallucination, donor omission, source omission, and joint recognition.
MIX never enters VHR. In MIX, the source is genuinely present, so reporting only the source is donor omission rather than hallucination. SWAP AFR and VHR are not complements either: a model may report both the donor and the visually implied source in the same response.
Eleven reported routes across two frozen 510-item stress regimes: the matched event-level Core and the separately reviewed CLASH fact set.
Naming the right sound on CLEAN cannot separate hearing from guessing, so we added two vision-only models that physically cannot read the audio track. They fix the zero point of the scale. Measured against it, every open omni-modal model we tested — across four families, three scales, and two generations — sits closer to that zero point than to OST.
Auditory evidence use on OST-DiagBench. (a) CLEAN source recall is plotted against MUTE hallucination rate. The diagonal marks audio-blind behavior, the dashed contour marks d′ = 1.4, circles denote offline models, and diamonds denote naive streaming. The bracket shows the 11.6-point CLEAN gap between Qwen3-Omni and Qwen3-VL. (b) Under SWAP contradiction and MIX coexistence, capped outlines show donor recall from audio alone while filled bars show what remains after video is restored. (c) The same paired comparison measures whether models follow the new spoken fact when CLASH keeps the stale fact visible. Panels (b) and (c) report baselines.
d′ = z(CLEAN-SR) − z(MUTE-VHR), treating CLEAN as signal-present and MUTE as signal-absent. A model that cannot read the audio track scores exactly 0, because for it naming the true sound and hallucinating an absent one are the same event.
| Model | d′ ↑ | CLEAN-SR ↑ | MUTE-VHR ↓ | SWAP-AFR ↑ | MIX-JR ↑ | CLASH-Acc ↑ |
|---|---|---|---|---|---|---|
| Audio-blind reference · cannot read the audio track | ||||||
| Qwen2.5-VL-7B | 0.00 | 73.14% | 73.14% | 0.39% | 0.39% | 0.00% |
| Qwen3-VL-30B-A3B | 0.00 | 80.98% | 80.98% | 0.20% | 0.00% | 0.00% |
| Open omni-modal · offline | ||||||
| Qwen2.5-Omni-3B | 0.71 | 89.41% | 70.59% | 50.98% | 19.02% | 2.75% |
| Qwen2.5-Omni-7B | 0.59 | 81.57% | 61.96% | 60.00% | 19.41% | 2.75% |
| MiniCPM-o 2.6 | 0.88 | 68.24% | 34.31% | 17.25% | 0.78% | 2.75% |
| Baichuan-Omni-1.5 | 1.27 | 62.94% | 17.25% | 34.90% | 2.35% | 2.75% |
| Qwen3-Omni-30B-A3B | 0.81 | 92.55% | 73.73% | 35.29% | 4.51% | 1.18% |
| Open omni-modal · naive streaming thinking | ||||||
| Qwen2.5-Omni-7B | 0.94 | 62.55% | 26.67% | 28.43% | 1.76% | 0.59% |
| MiniCPM-o 2.6 | 1.38 | 79.02% | 28.24% | 25.10% | 3.33% | 4.90% |
| Qwen3-Omni-30B-A3B | 0.53 | 92.55% | 81.96% | 18.04% | 1.57% | 0.00% |
| OST (ours, streaming) | 2.95 | 95.49% | 10.39% | 81.76% | 51.37% | 46.47% |
The audio pathway buys agreement, not listening. Qwen3-Omni-30B-A3B reaches 92.55% on CLEAN against 80.98% for Qwen3-VL-30B-A3B, its audio-blind twin of the same generation and scale. Under agreement the entire audio pathway is therefore worth 11.6 points — small enough that a visual prior is a workable substitute for hearing, which is exactly what a CLEAN-only metric rewards.
The same waveform is played twice: once through the native audio-only route with no video, and once with the locked video restored. The first is what the model can hear; the second is what it still reports. Their ratio is how much of its own hearing survives the image.
| Model (offline) | SWAP heard alone | with video | CMSG | retained | MIX heard alone | with video | CMSG | retained |
|---|---|---|---|---|---|---|---|---|
| Qwen2.5-Omni-3B | 66.86% | 50.98% | 15.88 | 0.76 | 41.57% | 22.35% | 19.22 | 0.54 |
| Qwen2.5-Omni-7B | 77.84% | 60.00% | 17.84 | 0.77 | 43.33% | 21.57% | 21.76 | 0.50 |
| MiniCPM-o 2.6 | 45.69% | 17.25% | 28.43 | 0.38 | 24.71% | 4.12% | 20.59 | 0.17 |
| Baichuan-Omni-1.5 | 56.67% | 34.90% | 21.76 | 0.62 | 22.94% | 7.25% | 15.69 | 0.32 |
| Qwen3-Omni-30B-A3B | 79.80% | 35.29% | 44.51 | 0.44 | 32.16% | 6.08% | 26.08 | 0.19 |
These four runs completed but measure the protocol rather than auditory ability. We report them so the exclusions are auditable rather than silent.
</think> never appears, and output length is exactly 64 tokens on all 3,570 samples. MIX-JR = 0.00% is therefore not evidence that the donor was inaudible; the model never emitted an answer.JUDGE_V2.0.0 is an English lexical judge, so mentions almost never fire: mean events per condition ≈ 0.1 despite ~40 mean tokens. CLEAN-SR 5.69% is a language mismatch under the locked prompt and judge, not a measurement of hearing.The event-level results use the same 510 matched Core groups with identical condition waveforms, PROMPT_V2, decoding contract, and unified judge. CLASH is a separate 510-item fact-conflict set with its own neutral cloze and deterministic fact matcher.
How to read: Higher is better for CLEAN-SR, SWAP-AFR, and MIX-DR/SR/JR; lower is better for MUTE-VHR and SWAP-VHR. The players above illustrate single-sample failure modes only; all percentages are aggregated over the complete 510-group Core.
The same response reports both the true source and the added donor. The best baseline reaches 19.41%, an absolute improvement of 31.96 points.
CLEAN-SR and MIX-SR (%, higher is better)
MUTE-VHR and SWAP-VHR (%, lower is better)
SWAP-AFR and MIX-DR (%, higher is better)
MIX-JR: source and donor reported together (%, higher is better)
Interpretation: Baselines appear to recognize the visible source reliably on CLEAN/MIX-SR, yet commonly omit the concurrent donor. OST improves both donor recall and joint recall.
CLASH-Acc (%, N=510). A-ONLY and A+V share the identical edited waveform; A+V additionally exposes the conflicting old fact in visible text.
| Method | A-ONLY ↑ | A+V / CLASH-Acc ↑ | Suppression gap ↓ |
|---|---|---|---|
| Qwen2.5-Omni-7B · naive streaming | 95.69% | 0.59% | 95.10 |
| MiniCPM-o 2.6 · naive streaming | 80.78% | 4.90% | 75.88 |
| Qwen3-Omni-30B-A3B · naive streaming | 94.12% | 0.00% | 94.12 |
| OST (ours, native streaming) | 100.00% | 46.47% | 53.53 |
CLASH reading. All three naive-streaming models recognize the edited speech in isolation but almost always revert to the visible old fact when video is restored. OST narrows this same-item suppression gap to 53.53 points, although the visible fact still wins often.
Event-level Core rows use the same 510 groups, audio routing, PROMPT_V2, greedy decoding, and unified event judge. CLASH uses a separate frozen 510-item set, the same neutral cloze for every route, and a deterministic matcher for the mutually exclusive spoken and visible facts. Naive streaming thinking receives nominal 1 s chunks and answers only at the end; it is not native online generation.
The final event-level comparison retains seven complementary Core metrics, each computed from the same 510 matched source groups. Together they cover agreement, absence, contradiction, and coexistence. CLASH is reported separately above because it uses a distinct fact-level construction and judge.
| Method | CLEAN-SR ↑ | MUTE-VHR ↓ | SWAP-AFR ↑ | SWAP-VHR ↓ | MIX-DR ↑ | MIX-SR ↑ | MIX-JR ↑ |
|---|---|---|---|---|---|---|---|
| Audio-blind reference · every condition collapses onto the CLEAN value | |||||||
| Qwen2.5-VL-7B (Offline) | 73.14% | 73.14% | 0.39% | 73.14% | 0.39% | 73.14% | 0.39% |
| Qwen2.5-VL-7B (Naive streaming thinking) | 64.12% | 64.12% | 0.00% | 64.12% | 0.00% | 64.12% | 0.00% |
| Qwen3-VL-30B-A3B (Offline) | 80.98% | 80.98% | 0.20% | 80.98% | 0.20% | 80.98% | 0.00% |
| Qwen3-VL-30B-A3B (Naive streaming thinking) | 77.45% | 77.45% | 0.00% | 77.45% | 0.00% | 77.45% | 0.00% |
| Open omni-modal | |||||||
| Qwen2.5-Omni-3B (Offline) | 89.41% | 70.59% | 50.98% | 51.57% | 22.35% | 88.43% | 19.02% |
| Qwen2.5-Omni-3B (Naive streaming thinking) | 84.31% | 72.35% | 25.69% | 55.29% | 9.80% | 83.14% | 6.08% |
| Qwen2.5-Omni-7B (Offline) | 81.57% | 61.96% | 60.00% | 55.88% | 21.57% | 85.10% | 19.41% |
| Qwen2.5-Omni-7B (Naive streaming thinking) | 62.55% | 26.67% | 28.43% | 20.00% | 4.31% | 61.76% | 1.76% |
| MiniCPM-o 2.6 (Offline) | 68.24% | 34.31% | 17.25% | 35.88% | 4.12% | 66.27% | 0.78% |
| MiniCPM-o 2.6 (Naive streaming thinking) | 79.02% | 28.24% | 25.10% | 31.18% | 6.27% | 74.31% | 3.33% |
| Baichuan-Omni-1.5 (Offline) | 62.94% | 17.25% | 34.90% | 26.67% | 7.25% | 63.33% | 2.35% |
| Qwen3-Omni-30B-A3B (Offline) | 92.55% | 73.73% | 35.29% | 40.20% | 6.08% | 90.59% | 4.51% |
| Qwen3-Omni-30B-A3B (Naive streaming thinking) | 92.55% | 81.96% | 18.04% | 67.84% | 1.76% | 92.94% | 1.57% |
| OST (ours, streaming) | 95.49% | 10.39% | 81.76% | 4.71% | 54.31% | 96.67% | 51.37% |
OST exact counts (N=510): CLEAN-SR 487/510 · MUTE-VHR 53/510 · SWAP-AFR 417/510 · SWAP-VHR 24/510 · MIX-DR 277/510 · MIX-SR 493/510 · MIX-JR 262/510. CLASH-Acc is 237/510 on its separate fact-level set. CLASH A-ONLY is 510/510. SWAP-AFR and SWAP-VHR are not complements; MIX uses only DR/SR/JR and is excluded from VHR.