Back to OST
Auditory evidence under visual bias

OST-DiagBenchWhat did the model actually hear?

A diagnostic benchmark that locks the video and changes only the auditory facts. Its matched Core tests agreement, absence, contradiction, and coexistence; CLASH keeps an old fact visible while speech supplies a mutually exclusive new one.

1,068event-level source clips
510matched Core groups
510reviewed CLASH items
349CLASH source videos
Visible scene · fixed in every conditionPerson playing piano
Video locked
CLEANpiano
MUTEroom tone
SWAPbird
MIXpiano + bird
CLASHtext: old · speech: new
The video never changes; the correct answer can only come from the audio difference.

Why do conventional AV benchmarks overestimate auditory ability?

Visual objects and typical sounds are highly correlated in natural videos. When shown a piano, car, or dog, a model can ignore the audio track and still generate “piano music,” “engine noise,” or “barking” from visual priors. The answer looks plausible, but it does not prove that the model listened.

The shortcut

The image is enough to guess a usually correct sound

On a congruent audiovisual clip, a visual guess and genuine auditory recognition produce the same answer, so conventional accuracy cannot distinguish them.

VisibleModel answer
piano → “I hear piano music.” ✓
The diagnostic counterfactual

Keep the image plausible while changing the audio facts

If the same piano video actually contains birdsong, answering “piano” reveals visual override. If piano and bird coexist but the model reports only piano, it reveals auditory omission.

Actually audibleModel answer
bird → “I hear piano.” ✕
Being correct on CLEAN does not imply that the model used audio.Genuine auditory ability must persist when visual priors fail, sounds compete, or the evidence weakens.

Locked video, change only the audio facts.

OST-DiagBench turns every source clip into a paired experiment. Video content and visual cues are locked; only the presence of the complete CLEAN track and the addition of a donor sound are controlled. Changes in model output can therefore be attributed to changes in auditory evidence.

Construct four acoustic realities for the same video and observe when the model listens, misses a sound, or substitutes vision for hearing.
01Lock the video to eliminate visual-content differences
02Control the source / donor composition
03Compare responses within the same source group
Donor absent
Donor present
Source-containing CLEAN track kept
S=1 · D=0

CLEAN

Vision and audio agree: can the model report the original source?

S=1 · D=1

MIX

Source and donor coexist: can the model capture both?

Source-containing CLEAN track removed
S=0 · D=0

MUTE

The sound is absent: does the image still induce the model to report the source?

S=0 · D=1

SWAP

The donor replaces the original track: does the model trust hearing or vision?

CLEAN · agreementBasic recognition ability, but not sufficient on its own to prove audio use.
MUTE · absenceMeasures visually induced false sound reports.
SWAP · contradictionMeasures whether a counter-visual donor overrides the visual script.
MIX · coexistenceMeasures donor omission and joint recall under concurrent sources.
CLASH · subtitle–speech fact conflict A visible subtitle retains the old fact while aligned speech states a locally plausible, mutually exclusive new fact. The correct answer is the spoken value; an identical-waveform A-ONLY route controls audibility.

From real videos to interpretable intervention groups

Construction follows one rule: within a group, change only preregistered variables. MIX retains the complete CLEAN waveform and overlays the donor; it requires no source separation and makes no claim of having an isolated source stem.

1
Admit source

Select source clips with visible events and reliable acoustic labels.

2
Assign donor

Assign near / far events in a balanced way from a non-overlapping donor bank.

3
Synthesize group

Generate CLEAN, MUTE, matched SWAP, and matched MIX.

4
Common master

Apply the same attenuation to the entire group to avoid condition-specific distortion.

5
Verify together

Use multilabel validation; reject the entire group if any member fails.

Matched SWAP / MIX

Share donor identity, a natural 2–4 s segment, onset/offset, 50–100 ms fades, gain, codec, and mastering. Donor-parameter differences are excluded from the comparison.

dB

Perceptual loudness control

MIX uses overlap-window DLR: δ = Ldonor − Lsource. SWAP has no source, so it uses donor-to-background SNR.

−10 dB
masked
0 dB
balanced
+10 dB
salient
A

Matched A-ONLY calibration

CLEAN/SWAP/MIX-AONLY retain final audio that is byte-identical to the corresponding AV condition. A donor recognized in A-ONLY but missed in AV is direct diagnostic evidence of cross-modal suppression.

Temporal position

Core fixes the onset at the middle. A separate panel uses the same donor segment and gain while varying only early / middle / late position to analyze position, recency, and retention.

Δ

Semantic granularity

FAR-SWAP crosses broad acoustic classes; NEAR-SWAP changes the fine-grained label within the same superclass. This distinguishes “not heard,” “heard only at a coarse level,” and “heard but overridden by vision.”

Artifact-aware acceptance

Check the locked-video signature, source/donor recognizability, DLR/SNR within ±1 dB, true peak ≤ −1 dBTP, A-ONLY hashes, and editing artifacts.

Evidence strengthDLR / SNR curves
Temporal positionearly · middle · late
Semantic proximityNEAR · FAR
CLASH follows a separate fact-level path. It starts from a real source-aligned utterance, leaves the old fact visible, replaces the spoken value with a mutually exclusive alternative, and retains an item only after ASR integrity checks and manual review of the question, edit, visible text, and conflict difficulty.

Which operations did we generate, and how much data does each contain?

The table reports only the final Challenge benchmark. Constructed counts records generated from 1,068 frozen source clips or balanced subsets; Accepted counts records that passed audio, video, and group-consistency validation and entered the report. Native and neutral-video A-ONLY are separate input routes and are counted separately, but they share the final audio of the corresponding AV condition.

Suite / operationWhat changesConstructed recordsAccepted records
Factorial Core · locked video
CLEANRetain the complete source-containing CLEAN track1,068510
MUTERemove the complete CLEAN track and replace it with matched low-level background1,068510
SWAPBackground + donor, creating a replacement contradiction1,068510
MIXComplete CLEAN track + donor, creating auditory coexistence1,068510
Calibration Controls · native audio-only + neutral-video surrogate
CLEAN-AONLYIdentical CLEAN audio through two non-semantic video routes2,1361,020
SWAP-AONLYIdentical SWAP audio through two non-semantic video routes2,1361,020
MIX-AONLYIdentical MIX audio through two non-semantic video routes2,1361,020
Evidence Strength · independently balanced panel
MIX DLR −10 / 0 / +10Change only the donor-to-source loudness ratio288157
Matched SWAP gain/SNRMatch MIX donor gain; report donor/background SNR288157
Temporal Position · independently balanced panel
MIX early / middle / lateUse the same donor and gain; change only the onset288158
SWAP early / middle / lateMatch the temporal position used in MIX288158
Semantic Granularity · fine-label auditable subset
FAR-SWAP + AONLYDonor from a different broad acoustic class; pair AV with native A-ONLY9686
NEAR-SWAP + AONLYSame superclass but a different fine-grained category; pair AV with native A-ONLY9692
All released event-level Challenge records12,0245,908
Fact-level CLASH · separate source-utterance path
CLASH A+V + A-ONLYKeep the old visible fact; replace aligned speech with a mutually exclusive new fact510 items
A record is not a unique clip. One source generates multiple paired conditions, A-ONLY routes, and panel variants.
Group-level acceptance. If any Core member fails, the CLEAN/MUTE/SWAP/MIX group for that source is removed in full, preventing condition-specific attrition.
Two reporting scopes. Seven event metrics use the same 510 Core groups; CLASH-Acc uses a separate 510-item manually reviewed fact set.

One guitar video, four audio realities

All four players come from the same accepted source group, ave-0a7683f182e2. Expected facts come from the manifest; model text is the actual output of Qwen2.5-Omni-7B offline.

CLEAN1,068 → 510
Expected auditory factsmusical instrument / guitar
Actual model answer“The audio contains the sound of a guitar being played, with a focus on the plucking of the strings.”
✓ source detectedave-0a7683f182e2__FC__CLEAN__AV
MUTE1,068 → 510
Expected auditory factsno salient source or donor; matched background only
Actual model answer“The sound of a guitar being played is heard.”
✕ visual hallucinationave-0a7683f182e2__FC__MUTE__AV
SWAP1,068 → 510
Expected auditory factsdog barking; visual musical source is absent
Actual model answer“The audio contains the sound of a dog barking in the background.”
✓ donor detectedave-0a7683f182e2__FC__SWAP__AV
MIX1,068 → 510
Expected auditory factsmusical instrument / guitar + dog barking
Actual model answer“The audio contains a guitar being played, with a focus on the strings being plucked.”
! donor omittedave-0a7683f182e2__FC__MIX__AV
!This matched group demonstrates the central value of the benchmark.CLEAN appears to show that the model can listen, but MUTE reveals visual auditory hallucination and MIX reveals omission of the concurrent dog. CLEAN accuracy alone cannot expose either failure.

The donor becomes audible again when semantic video is removed

The players use a uniform neutral-video carrier; their audio is byte-identical to the corresponding AV condition. Below are real outputs from the same model on the neutral-video route.

CLEAN-AONLY2,136 → 1,020
Expectedmusical instrument / guitar
Neutral-route answer“A lively instrumental piece ... featuring a guitar and a bass guitar.”
✓ source detectedave-0a7683f182e2__FC__CLEAN__AONLY_NEUTRAL
SWAP-AONLY2,136 → 1,020
Expecteddog barking
Neutral-route answer“The audio contains a dog barking and a sound effect.”
✓ donor detectedave-0a7683f182e2__FC__SWAP__AONLY_NEUTRAL
MIX-AONLY2,136 → 1,020
Expectedmusical instrument / guitar + dog barking
Neutral-route answer“A lively music track with a guitar and a dog barking in the background.”
✓ both detectedave-0a7683f182e2__FC__MIX__AONLY_NEUTRAL

Matched contrast: The same MIX audio yields only “guitar” with the original guitar video, but “guitar + dog” with neutral video. This is exactly the phenomenon that the Cross-Modal Suppression Gap is designed to quantify; aggregate metrics still report native and neutral routes separately.

Change only donor loudness: flute music + frog

The three MIX videos use the same source, donor, segment, middle onset, and mastering; only overlap-window DLR changes. Real responses progress from hearing only the flute, to hearing both, to the donor overwhelming the source.

MIX · −10 dB96 → 23 valid
Expectedflute music + frog (donor masked)
Qwen2.5 offline“There is a sound of a flute being played.”
! frog omittedmusic-avqa-r-3daaf1254d42__DLRM10__MIX__AV
MIX · 0 dB96 → 51 valid
Expectedflute music + frog (balanced)
Qwen2.5 offline“There is a sound of a frog croaking and a woman playing a flute.”
✓ both detectedmusic-avqa-r-3daaf1254d42__DLRP0__MIX__AV
MIX · +10 dB96 → 83 valid
Expectedflute music + frog (donor salient)
Qwen2.5 offline“The audio contains a sound of a frog croaking.”
! flute omittedmusic-avqa-r-3daaf1254d42__DLRP10__MIX__AV
Matched SWAP288 → 157 valid
Expectedfrog only; visible flute source is absent
Qwen2.5 offline“The audio contains the sound of a frog croaking and a stream flowing.”
✓ donor detectedmusic-avqa-r-3daaf1254d42__DLRP0__SWAP__AV

Change only when the dog appears: early / middle / late

All three videos use the same guitar + dog pair and 0 dB DLR. Real Qwen3 offline outputs show that early and middle donors are omitted, while the late donor is retained in the final answer.

MIX · EARLY96 → 54 valid
ExpectedDOG EARLY · guitar + dog
Qwen3 offline excerpt“The audio clip features a solo acoustic guitar performance ...”
! dog omittedave-0a7683f182e2__TPEARLY__MIX__AV
MIX · MIDDLE96 → 50 valid
ExpectedDOG MIDDLE · guitar + dog
Qwen3 offline excerpt“The audio clip features a solo acoustic guitar performance ...”
! dog omittedave-0a7683f182e2__TPMIDDLE__MIX__AV
MIX · LATE96 → 54 valid
ExpectedDOG LATE · guitar + dog
Qwen3 offline excerpt“... guitar ... accompanied by the distinct sound of a dog barking in the background.”
✓ both detectedave-0a7683f182e2__TPLATE__MIX__AV
SWAP · MIDDLE96 → 50 valid
ExpectedDOG MIDDLE · dog only; guitar absent
Qwen2.5 offline“The audio contains the sound of a dog barking in the background.”
✓ donor detectedave-0a7683f182e2__TPMIDDLE__SWAP__AV

FAR and NEAR donors with the same visible vehicle-horn source

The original vehicle-horn track is removed from both SWAP videos below. A-ONLY recognizes the donor, but when paired with the original video, Qwen2.5 instead reports a visually compatible train horn in both cases.

FAR-SWAPAV 48 → 43 valid
Expectedclock alarm only; visible vehicle horn is absent
AV answer“The audio contains the sound of a train horn blowing.”
Matched A-ONLY answer“... a device beeping ...”
✕ semantic overrideave-5380b0f69bf4__SEMFAR__SWAP__AV
NEAR-SWAPAV 48 → 46 valid
Expectedairplane only; visible vehicle horn is absent
AV answer“... a train horn blowing and ... a train moving.”
Matched A-ONLY answer“... a jet engine ... a propeller plane ...”
✕ semantic overrideave-5380b0f69bf4__SEMNEAR__SWAP__AV

Dataset scale, composition, and valid samples

OST-DiagBench contains two frozen stress regimes. The event-level branch admits 1,068 source clips and reports the complete Core on 510 groups valid in CLEAN, MUTE, SWAP, and MIX. The separate CLASH branch retains 510 manually reviewed fact conflicts from 349 source videos.

1,068frozen Challenge source clips
510accepted matched Core groups
510reviewed CLASH fact conflicts
349source videos in the CLASH branch

Challenge source composition

The 1,068 frozen source clips are stratified across four acoustic–visual categories.

156299499114
Animal · 14.6%
Instrument · 28.0%
Vehicle · 46.7%
Voice · 10.7%

From candidate pool to headline comparison

Validity gates, difficulty selection, and audio-quality gates are separate stages; 510 is the shared denominator for the complete Core.

1,414screened candidates
1,068admitted source clips
914automated-valid matched groups
510reviewed matched Core groups
12,024 constructed recordsCore, controls and extension panels
5,908 event records914 matched event-level groups
240 human-audit clips480 double-annotation assignments
0 donor overlapdonor recordings are disjoint from evaluation sources
Challenge selection disclosure. For the event-level branch, validity is decided by source identifiability, visual affordance, and intervention integrity before a Qwen2.5-Omni-7B difficulty pre-filter and decisive human confirmation pass. Because that model contributed to admission, its own event-level row is biased downward. CLASH follows a separate quality, integrity, multi-checkpoint difficulty, and manual-review funnel. Both branches are frozen stress sets, so their rates diagnose retained hard cases rather than performance on unscreened natural video.

Beyond accuracy: record the joint source–donor state

A unified judge maps each response to one of four states and interprets it according to which sounds are actually present in that condition. This separates hallucination, donor omission, source omission, and joint recognition.

P00source 0 · donor 0
Neither is reported
P01source 0 · donor 1
Donor only
P10source 1 · donor 0
Source only
P11source 1 · donor 1
Both are reported
CLEANSource Recall · original-sound recognitionhigher ↑
MUTEVHR · report a visually implied source that is absentlower ↓
SWAPAFR · donor recall / VHR · visual hallucinationjoint view
MIXDR · donor / SR · source / JR · bothhigher ↑
A-ONLYCMSG · audio-only donor recall − AV donor recalllower ↓
CLASHAccuracy for the new spoken fact over the visible old facthigher ↑

MIX never enters VHR. In MIX, the source is genuinely present, so reporting only the source is donor omission rather than hallucination. SWAP AFR and VHR are not complements either: a model may report both the donor and the visually implied source in the same response.

Does any open omni-modal model actually listen?

Eleven reported routes across two frozen 510-item stress regimes: the matched event-level Core and the separately reviewed CLASH fact set.

Naming the right sound on CLEAN cannot separate hearing from guessing, so we added two vision-only models that physically cannot read the audio track. They fix the zero point of the scale. Measured against it, every open omni-modal model we tested — across four families, three scales, and two generations — sits closer to that zero point than to OST.

d′ < 1.4Sensitivity of every open omni-modal model, offline or streaming
11.6 ptsWhat the entire audio pathway is worth when sound and image agree
0.77 → 0.44Audible evidence retained across the Qwen generation, while audibility itself does not move
2.95OST sensitivity · 2.1× the best baseline row
Three-panel OST-DiagBench analysis: the CLEAN versus MUTE detection plane; audio-only versus audio-visual donor recall under SWAP and MIX; and audio-only versus audio-visual recognition of the new spoken fact under CLASH.

Auditory evidence use on OST-DiagBench. (a) CLEAN source recall is plotted against MUTE hallucination rate. The diagonal marks audio-blind behavior, the dashed contour marks d′ = 1.4, circles denote offline models, and diamonds denote naive streaming. The bracket shows the 11.6-point CLEAN gap between Qwen3-Omni and Qwen3-VL. (b) Under SWAP contradiction and MIX coexistence, capped outlines show donor recall from audio alone while filled bars show what remains after video is restored. (c) The same paired comparison measures whether models follow the new spoken fact when CLASH keeps the stale fact visible. Panels (b) and (c) report baselines.

Sensitivity against the audio-blind zero point

d′ = z(CLEAN-SR) − z(MUTE-VHR), treating CLEAN as signal-present and MUTE as signal-absent. A model that cannot read the audio track scores exactly 0, because for it naming the true sound and hallucinating an absent one are the same event.

Modeld′ ↑CLEAN-SR ↑MUTE-VHR ↓SWAP-AFR ↑MIX-JR ↑CLASH-Acc ↑
Audio-blind reference · cannot read the audio track
Qwen2.5-VL-7B0.0073.14%73.14%0.39%0.39%0.00%
Qwen3-VL-30B-A3B0.0080.98%80.98%0.20%0.00%0.00%
Open omni-modal · offline
Qwen2.5-Omni-3B0.7189.41%70.59%50.98%19.02%2.75%
Qwen2.5-Omni-7B0.5981.57%61.96%60.00%19.41%2.75%
MiniCPM-o 2.60.8868.24%34.31%17.25%0.78%2.75%
Baichuan-Omni-1.51.2762.94%17.25%34.90%2.35%2.75%
Qwen3-Omni-30B-A3B0.8192.55%73.73%35.29%4.51%1.18%
Open omni-modal · naive streaming thinking
Qwen2.5-Omni-7B0.9462.55%26.67%28.43%1.76%0.59%
MiniCPM-o 2.61.3879.02%28.24%25.10%3.33%4.90%
Qwen3-Omni-30B-A3B0.5392.55%81.96%18.04%1.57%0.00%
OST (ours, streaming)2.9595.49%10.39%81.76%51.37%46.47%

The audio pathway buys agreement, not listening. Qwen3-Omni-30B-A3B reaches 92.55% on CLEAN against 80.98% for Qwen3-VL-30B-A3B, its audio-blind twin of the same generation and scale. Under agreement the entire audio pathway is therefore worth 11.6 points — small enough that a visual prior is a workable substitute for hearing, which is exactly what a CLEAN-only metric rewards.

The bottleneck is not the audio encoder

The same waveform is played twice: once through the native audio-only route with no video, and once with the locked video restored. The first is what the model can hear; the second is what it still reports. Their ratio is how much of its own hearing survives the image.

Model (offline)SWAP heard alonewith videoCMSGretainedMIX heard alonewith videoCMSGretained
Qwen2.5-Omni-3B66.86%50.98%15.880.7641.57%22.35%19.220.54
Qwen2.5-Omni-7B77.84%60.00%17.840.7743.33%21.57%21.760.50
MiniCPM-o 2.645.69%17.25%28.430.3824.71%4.12%20.590.17
Baichuan-Omni-1.556.67%34.90%21.760.6222.94%7.25%15.690.32
Qwen3-Omni-30B-A3B79.80%35.29%44.510.4432.16%6.08%26.080.19
Audibility holds; entitlement collapsesFrom Qwen2.5-Omni-7B to Qwen3-Omni-30B-A3B the donor stays just as audible when only the waveform is played (77.84% → 79.80%), yet attribution with the video present falls by nearly half (60.00% → 35.29%) and suppression more than doubles (17.84 → 44.51). What changed is not the audio encoder but which channel is entitled to settle the question once the image is there.
Not a matter of scaleWithin one recipe, retention barely moves from Qwen2.5-Omni-3B to 7B (0.76 and 0.77). It only falls to 0.44 across the generation. The regression tracks the training recipe, not parameter count — which is why more scale alone will not fix it.
Not particular to one familyMiniCPM-o 2.6 fails differently: its audibility ceiling is already low (45.69%) and retention is low as well (0.38). Baichuan-Omni-1.5 keeps 0.62. Low retention appears in every family we could run, so it is a property of how omni models are trained rather than of one lab’s recipe.
Restraint alone does not repair itBaichuan-Omni-1.5 is by far the quietest model here (MUTE-VHR 17.25% against 61.96–73.73% for the Qwen offline rows) and by that route the most sensitive baseline (d′ = 1.27). It still discards 38% of what it demonstrably hears. Speaking less lowers hallucination without making the model listen.
Streaming pushes models toward the blind lineUnder naive streaming thinking, Qwen3-Omni-30B-A3B reports an absent sound 81.96% of the time, statistically indistinguishable from the 77.45% of its audio-blind twin, while its hit rate does not fall at all (CLEAN-SR is still 92.55%). What degrades under streaming is not hearing but staying silent when there is nothing to report.
The collapse is not “it says less”Qwen3-Omni-30B-A3B writes nearly twice as much as Qwen2.5-Omni-7B (61.9 vs 33.2 mean tokens) and reports more events on MUTE (1.376 vs 0.929 per answer), yet fewer on SWAP (1.239 vs 1.706). It is more talkative overall and specifically quieter about the donor, so verbosity cannot explain the drop.

Systems excluded from the headline, and why

These four runs completed but measure the protocol rather than auditory ability. We report them so the exclusions are auditable rather than silent.

Qwen3-Omni-30B-A3B-ThinkingUnder the locked 64-token budget every single answer is an unfinished scratchpad — </think> never appears, and output length is exactly 64 tokens on all 3,570 samples. MIX-JR = 0.00% is therefore not evidence that the donor was inaudible; the model never emitted an answer.
VITA-1.5Answers arrive as Chinese sound captions while JUDGE_V2.0.0 is an English lexical judge, so mentions almost never fire: mean events per condition ≈ 0.1 despite ~40 mean tokens. CLEAN-SR 5.69% is a language mismatch under the locked prompt and judge, not a measurement of hearing.
Qwen3-Omni-30B-A3B-CaptionerRun on native A-ONLY only, to test whether a captioning objective raises the audibility ceiling. It does not: SWAP donor recall is 18.04% against 79.80% for Instruct on the same waveforms. It spends its 64 tokens on ambient and music texture and under-names the donor.
Baichuan-Omni-1.5 · naive streamingIncomplete at 2,683 / 3,570. The 887 missing samples are all long locked-video clips that exhaust an H200 during a single packed prefill. Judging the remainder as N = 510 would bias CLEAN and MUTE, so only the offline row is reported.

OST follows auditory evidence across absence, contradiction, coexistence, and fact conflict

The event-level results use the same 510 matched Core groups with identical condition waveforms, PROMPT_V2, decoding contract, and unified judge. CLASH is a separate 510-item fact-conflict set with its own neutral cloze and deterministic fact matcher.

How to read: Higher is better for CLEAN-SR, SWAP-AFR, and MIX-DR/SR/JR; lower is better for MUTE-VHR and SWAP-VHR. The players above illustrate single-sample failure modes only; all percentages are aggregated over the complete 510-group Core.

51.37%

OST MIX joint recall

The same response reports both the true source and the added donor. The best baseline reaches 19.41%, an absolute improvement of 31.96 points.

2.95 d′auditory sensitivity, 2.1× the best baseline row
81.76%SWAP-AFR when donor audio contradicts the locked video
51.37%MIX-JR for reporting source and donor together
46.47%CLASH-Acc for following the new spoken fact

① Source recognition stays high

CLEAN-SR and MIX-SR (%, higher is better)

CLEAN-SRMIX-SR
Q2.5 · Offline
CLEAN
81.57
↳ MIX source
85.10
Q2.5 · Naive streaming thinking
CLEAN
62.55
↳ MIX source
61.76
Q3 · Offline
CLEAN
92.55
↳ MIX source
90.59
Q3 · Naive streaming thinking
CLEAN
92.55
↳ MIX source
92.94
OST
CLEAN
95.49
↳ MIX source
96.67

② Visual hallucination under absence and conflict

MUTE-VHR and SWAP-VHR (%, lower is better)

MUTE-VHRSWAP-VHR
Q2.5 · Offline
MUTE
61.96
↳ SWAP
55.88
Q2.5 · Naive streaming thinking
MUTE
26.67
↳ SWAP
20.00
Q3 · Offline
MUTE
73.73
↳ SWAP
40.20
Q3 · Naive streaming thinking
MUTE
81.96
↳ SWAP
67.84
OST
MUTE
10.39
↳ SWAP
4.71

③ Following the actual donor

SWAP-AFR and MIX-DR (%, higher is better)

SWAP-AFRMIX-DR
Q2.5 · Offline
SWAP
60.00
↳ MIX donor
21.57
Q2.5 · Naive streaming thinking
SWAP
28.43
↳ MIX donor
4.31
Q3 · Offline
SWAP
35.29
↳ MIX donor
6.08
Q3 · Naive streaming thinking
SWAP
18.04
↳ MIX donor
1.76
OST
SWAP
81.76
↳ MIX donor
54.31

④ Joint recognition under coexistence

MIX-JR: source and donor reported together (%, higher is better)

Q2.5 · Offline
19.41
Q2.5 · Naive streaming thinking
1.76
Q3 · Offline
4.51
Q3 · Naive streaming thinking
1.57
OST
51.37

Interpretation: Baselines appear to recognize the visible source reliably on CLEAN/MIX-SR, yet commonly omit the concurrent donor. OST improves both donor recall and joint recall.

⑤ Following speech when visible text disagrees

CLASH-Acc (%, N=510). A-ONLY and A+V share the identical edited waveform; A+V additionally exposes the conflicting old fact in visible text.

MethodA-ONLY ↑A+V / CLASH-Acc ↑Suppression gap ↓
Qwen2.5-Omni-7B · naive streaming95.69%0.59%95.10
MiniCPM-o 2.6 · naive streaming80.78%4.90%75.88
Qwen3-Omni-30B-A3B · naive streaming94.12%0.00%94.12
OST (ours, native streaming)100.00%46.47%53.53

CLASH reading. All three naive-streaming models recognize the edited speech in isolation but almost always revert to the visible old fact when video is restored. OST narrows this same-item suppression gap to 53.53 points, although the visible fact still wins often.

Lower hallucination is not achieved through silenceOST reaches 95.49% / 96.67% CLEAN-SR / MIX-SR while reducing MUTE/SWAP VHR to 10.39% / 4.71%.
Replacement and coexistence both improveOST reaches 81.76% SWAP-AFR, 54.31% MIX-DR, and 51.37% MIX-JR; the lower MIX values reflect genuine source competition.
Results are limited to hard casesChallenge is a reference-model-specific hard-case set. Point differences support stress diagnosis and should not be extrapolated to overall performance on natural data.

Unified evaluation contract and reproducible artifacts

Event-level Core rows use the same 510 groups, audio routing, PROMPT_V2, greedy decoding, and unified event judge. CLASH uses a separate frozen 510-item set, the same neutral cloze for every route, and a deterministic matcher for the mutually exclusive spoken and visible facts. Naive streaming thinking receives nominal 1 s chunks and answers only at the end; it is not native online generation.

Prompt“Describe all salient sounds you hear in this clip.”
JudgeJUDGE_V2.0.0 · assertion / negation / uncertainty aware
Statistical unitpaired source-cluster bootstrap; crossed source/donor clustering for reused donors
Predictionsoutputs/v2/inference/{qwen25,qwen25_omni3b,qwen3,qwen3_thinking,qwen3_captioner,qwen25_vl7b,qwen3_vl30b,minicpm_o26,baichuan_omni15,vita15}/{offline,chunked}/
Metricsseven event-level Core endpoints · d′ and CMSG · CLASH-Acc on a separate N=510 fact set
Human audit240 clips · 480 double-annotation assignments
Consistency contract. All seven event metrics use the same Core denominator. SWAP and MIX exactly match donor asset, segment, onset/offset, fades, gain, codec, and mastering. MIX is excluded from VHR, and A-ONLY retains audio identical to the linked AV condition. In CLASH, A+V and A-ONLY likewise share the identical edited waveform; only the conflicting visible old fact is added.

All reported metrics in one table

The final event-level comparison retains seven complementary Core metrics, each computed from the same 510 matched source groups. Together they cover agreement, absence, contradiction, and coexistence. CLASH is reported separately above because it uses a distinct fact-level construction and judge.

MethodCLEAN-SR ↑MUTE-VHR ↓SWAP-AFR ↑SWAP-VHR ↓MIX-DR ↑MIX-SR ↑MIX-JR ↑
Audio-blind reference · every condition collapses onto the CLEAN value
Qwen2.5-VL-7B (Offline)73.14%73.14%0.39%73.14%0.39%73.14%0.39%
Qwen2.5-VL-7B (Naive streaming thinking)64.12%64.12%0.00%64.12%0.00%64.12%0.00%
Qwen3-VL-30B-A3B (Offline)80.98%80.98%0.20%80.98%0.20%80.98%0.00%
Qwen3-VL-30B-A3B (Naive streaming thinking)77.45%77.45%0.00%77.45%0.00%77.45%0.00%
Open omni-modal
Qwen2.5-Omni-3B (Offline)89.41%70.59%50.98%51.57%22.35%88.43%19.02%
Qwen2.5-Omni-3B (Naive streaming thinking)84.31%72.35%25.69%55.29%9.80%83.14%6.08%
Qwen2.5-Omni-7B (Offline)81.57%61.96%60.00%55.88%21.57%85.10%19.41%
Qwen2.5-Omni-7B (Naive streaming thinking)62.55%26.67%28.43%20.00%4.31%61.76%1.76%
MiniCPM-o 2.6 (Offline)68.24%34.31%17.25%35.88%4.12%66.27%0.78%
MiniCPM-o 2.6 (Naive streaming thinking)79.02%28.24%25.10%31.18%6.27%74.31%3.33%
Baichuan-Omni-1.5 (Offline)62.94%17.25%34.90%26.67%7.25%63.33%2.35%
Qwen3-Omni-30B-A3B (Offline)92.55%73.73%35.29%40.20%6.08%90.59%4.51%
Qwen3-Omni-30B-A3B (Naive streaming thinking)92.55%81.96%18.04%67.84%1.76%92.94%1.57%
OST (ours, streaming)95.49%10.39%81.76%4.71%54.31%96.67%51.37%

OST exact counts (N=510): CLEAN-SR 487/510 · MUTE-VHR 53/510 · SWAP-AFR 417/510 · SWAP-VHR 24/510 · MIX-DR 277/510 · MIX-SR 493/510 · MIX-JR 262/510. CLASH-Acc is 237/510 on its separate fact-level set. CLASH A-ONLY is 510/510. SWAP-AFR and SWAP-VHR are not complements; MIX uses only DR/SR/JR and is excluded from VHR.