A guess enters memory as if observed.
Free-form think-while-watching merges what the model saw with what it merely expected. The error can survive compression and compound across the stream.



Future-testable claims for deciding what and when to answer from a causal audio-visual stream.
Visual evidence is dense and often locally decidable. Auditory evidence is sparse, temporally extended, and vulnerable to eviction. Under bounded memory, this asymmetry creates premature cross-modal commitment: a plausible visual explanation becomes a working fact before audio can settle it.
Free-form think-while-watching merges what the model saw with what it merely expected. The error can survive compression and compound across the stream.
Each unresolved interpretation names a verifying modality, an evidence interval, and the states that depend on it. The answer gate waits while an answer-critical review remains scheduled.
OST runs one closed loop at every decision step. New evidence first checks earlier assumptions; corrections propagate through the exact reasoning they influenced; only then can the model decide whether the stream is sufficient.
Keep visual and audio native tokens in separately budgeted buffers.
Record the prediction, verifying modality, future interval, and support scope.
Check the claim against realized evidence from its specified modality.
Reduce refuted influence through recorded dependencies, then rewrite the state.
Wait while any answer-critical claim remains unsettled.

The six-field Omni-State makes omission visible. Silence is a valid fact; uncertainty is legal; fabricated audio is not required. Claims preserve their modality and provenance even after history is folded into summaries.
OST-DiagBench diagnoses whether an omni-modal model actually uses auditory evidence. Its 510 matched Core groups keep video fixed while audio is retained, muted, swapped, or mixed. A separate set of 510 reviewed CLASH examples leaves an old fact visible in subtitles while speech states a mutually exclusive new fact.
With a frozen Qwen3-Omni-30B-A3B-Instruct backbone and lightweight adaptation, OST outperforms the strongest open baselines on five streaming and audio-visual benchmarks by more than 10% relative on average.
Best online result across real-time, recall, and proactive questions.
SOVBench-O average under causal, prefix-only inference.
Respond/wait trigger quality under native streaming.
Reporting note. SOVBench values reproduce Table 1 of the current arXiv manuscript; diagnostic values reproduce Table 2 on 510 matched Core groups and 510 reviewed CLASH examples.
The latest manuscript connects a failure mode, a mechanism, and an evaluation: premature commitment becomes measurable, future evidence can revise memory, and the effect is tested across both standard benchmarks and controlled audio interventions.
OST identifies how an early visual interpretation can become a working fact and persist in streaming memory after later audio contradicts it.
Future-testable claims connect a modality-specific review to dependency-aware retraction, corrected state decoding, and answer timing.
OST leads open systems across five streaming and audio-visual benchmarks and exposes audio use under agreement, absence, contradiction, coexistence, and fact conflict.
Streaming omni-modal models must decide what and when to answer from the video chunks and synchronized audio observed so far. Visual cues often support an interpretation before an utterance or sound event is complete. If that interpretation enters memory as a fact, later reasoning can keep relaying it even after audio contradicts it. We call this failure premature cross-modal commitment. We propose Omni-Streaming Thinking (OST), which generates structured outputs that include evidence observed so far, forecasts of future evidence, and claims based on this evidence. Each claim is initially marked as pending and linked to a future verification interval. Audio and visual evidence are stored separately, and OST checks a claim against the evidence from the specified modality at the end of the verification interval. When contradictory evidence is detected, a refutation process reduces the influence of the claim and its dependent states, and then guides a state update using the new evidence. An answer gate decides whether the answer-critical claims meet the conditions for giving a response. Using a frozen Qwen3-Omni-30B-A3B-Instruct backbone with lightweight adaptation, OST outperforms the strongest open baselines on five streaming and audio-visual benchmarks by more than 10% relative on average. We also introduce OST-DiagBench, which holds video fixed and edits audio to test agreement, absence, contradiction, coexistence, and subtitle-speech conflict. OST reaches d′ = 2.95, compared with at most 1.38 for open baselines, while reducing vision-induced auditory hallucinations.
@article{du2026omni,
title={Omni-Streaming Thinking},
author={Enjun Du and Siyi Liu and Ziyu Zheng and Jingyu Li and Yiwen Guo and Yongqi Zhang and Difan Zou},
year={2026}
}