Omni-Streaming Thinking
Sep 14, 2026·
,,,,,,·
0 min read
Enjun Du
Siyi Liu
Ziyu Zheng
Jingyu Li
Yiwen Guo*
Yongqi Zhang
Difan Zou*
Overview of Omni-Streaming Thinking.Abstract
Streaming omni-modal models must decide what and when to answer from the video chunks and synchronized audio observed so far. Visual cues can support an interpretation before an utterance or sound event is complete, allowing a premature assumption to persist in memory even after later audio contradicts it. Omni-Streaming Thinking (OST) generates structured outputs containing observed evidence, forecasts of future evidence, and pending claims tied to modality-specific verification intervals. When later evidence refutes a claim, OST reduces its influence through recorded dependencies, updates the reasoning state, and gates the final answer until answer-critical reviews are settled. OST also introduces OST-DiagBench, a controlled benchmark for testing agreement, absence, contradiction, coexistence, and subtitle-speech conflict while holding video fixed and intervening on audio.
Type