Multimodal Models · Long Context · Retrieval-Augmented Generation
X-Stream vs V-RAGBench: streaming video understanding, measured honestly
X-Stream tests watching concurrent live streams (best model 49.60% vs 91.84% human); V-RAGBench/CARVE tests retrieval over 11-99h egocentric video (Recall@5 0.603 vs 0.510).
Why long and streaming video break the old benchmark playbook
Short-clip video QA was a solved-looking game: take 30 seconds of footage, brute-force every frame into the context window, and let a strong MLLM answer. That protocol dies the moment the video is an hour long, or live and unrewindable, or arriving as several feeds at once. The honest question in 2026 is no longer “higher accuracy on dataset X”. It is now where the evidence lives in a stream you cannot fully re-read, and how a system behaves when its context budget is smaller than the video.
The watch-remember-reason survey gives this page its map. It argues the field should stop indexing papers by dataset and start indexing them by what the model actually has to do: watch (acquire evidence from the stream), remember (hold context across minutes or hours), and reason (produce a grounded answer). It organizes 100+ methods under that skeleton and names streaming egocentric understanding as the open frontier, which is exactly the regime the two benchmark papers below stress.
Two of those abilities now have benchmarks honest enough to trust. Both exist because their authors found existing evaluations were lying, in opposite directions: one let models look capable without ever watching more than one stream, the other let models score well without retrieving anything at all.
X-Stream: can any model watch several live feeds at once?
X-Stream is the first benchmark for concurrent multi-stream understanding: not one clip, but several live video feeds the model must watch and reason across simultaneously: multi-window desktop footage, multi-view scenes, multi-device camera streams. It contributes 4,220 curated QA pairs over 932 videos across 11 subtasks, and the conceptual lens of treating an MLLM as a multiplexer that must interleave and route between concurrent signals.
The headline result is a clear no. The strongest model, Gemini 3 Pro, reaches 49.60% overall against a 91.84% human baseline, a 42-point gap that no amount of generation scaling has closed. GPT-5 manages 27.78% and GPT-4o 22.46%, and every open-source model lands in the 6-34% band. Purpose-built streaming architectures do worst of all: Dispider scores 15.44%, VideoLLM-online-8B 8.48%, MMDuet2 6.79%. Models engineered for single-stream online video are not merely failing at concurrency; they are failing harder than general frontier models.
The subtask breakdown localizes the failure. Even Gemini 3 Pro drops to 41.79% on causal reasoning and 44.18% on decision-making, versus 66.72% on the easier visual grounding. Proactive ability (reacting at the right moment, unprompted, across streams) collapses below 21% for most models. This is the capability a surveillance operator, a live-sports analyst, or a multi-camera robot actually needs, and it is precisely what single-stream benchmarks never measured.
V-RAGBench and CARVE: when the video cannot fit, retrieval is the only route
The complementary regime is the video you could never watch in full. V-RAGBench and CARVE start from an uncomfortable finding: in widely used video QA benchmarks, more than half of the queries can be answered without the video at all, from language priors, world knowledge, or static cues. Those benchmarks quietly reward retrieval systems that never retrieve.
V-RAGBench is built to close that hole. It draws 216 egocentric videos (174 from Ego4D averaging 86 minutes, 42 from EgoLife averaging 379 minutes; durations from 11 to 99 hours), then passes candidate queries through five sequential filters: semantic-similarity removal, answerability verification, shortcut-bias elimination, empirical answerability confirmation, and an evidence-uniqueness check. What survives is 2,100 labeled query/evidence/answer triplets (1,800 train / 300 test), each tied to a known evidence chunk, so retrieval and generation are scored separately: Recall@5 and nDCG@5 for retrieval, LLM-judged pass rate for generation.
CARVE, the method that ships with it, retrieves over long video by running four retrievers in parallel (one per modality-granularity configuration) and assigning each chunk the config that found it best. The numbers:
- Recall@5: 0.603 versus 0.510 for VideoRAG-A, the strongest of eight baselines.
- nDCG@5: 0.433 versus 0.340 for the best baseline on that metric.
- Generation pass rate: 0.357 with Qwen3-VL-8B (0.315 for the best baseline), 0.367 vs 0.317 with Qwen3-VL-32B, and 0.320 vs 0.307 with Gemma-4-26B; the retrieval gain carries through three different generators.
The ablations explain where the gain comes from, and they are the part worth copying. The best single retrieval config reaches only 0.507; the best two-way partial combination reaches 0.567; all four configs with per-chunk selection reach 0.603. So the finding is not “more retrievers” but the best config is a property of the chunk’s content, not of the query; fixing one config per query, the standard practice, leaves recall on the table. Reranking each chunk only under its own config is load-bearing: random config assignment or joint concatenation both drop Recall@5 to 0.513. And this hand-built per-chunk rule outperforms trained routing (0.357 pass rate against a trained non-LLM router at 0.329 and a LoRA-tuned LLM router at 0.310) without any training at all.
The two failure modes are complementary, and that is the real story
Read side by side, the two papers describe the same wall from opposite sides. X-Stream shows that when you give frontier models the whole video and simply ask them to watch more of it at once, they still fail: 49.60% against 91.84% human, with the proactive, multi-stream regime below 21% for most. V-RAGBench shows that when you give up on watching everything and retrieve instead, the retrieval itself is fragile: even the best pipeline misses the right evidence within the top 5 about 40% of the time, on footage a model was never going to fit in context anyway.
Neither number is directly comparable to the other: different tasks, different harnesses, different metrics, and V-RAGBench’s deliberately leak-filtered design makes its absolute scores harsher than legacy video-QA leaderboards. The honest cross-paper claim is narrower and more useful: watch-everything and retrieve-what-matters are both still losing, in measurable, localized ways, and a system design that assumes either one is solved will ship with a known failure surface. That is also exactly the watch-remember-reason decomposition’s point: the two benchmarks attack the watch and remember pillars separately, and neither touches the reason pillar’s faithfulness problem.
Key numbers
| Axis | X-Stream | V-RAGBench / CARVE | Comparable? |
|---|---|---|---|
| What it tests | Concurrent multi-stream watching + reasoning | Retrieval over 11-99h egocentric video | Different tasks |
| Scale | 4,220 QA pairs, 932 videos, 11 subtasks | 2,100 triplets, 216 videos (1,800 train / 300 test) | / |
| Best result | Gemini 3 Pro 49.60% overall | CARVE Recall@5 0.603 | Not directly comparable |
| Reference point | Human baseline 91.84% (42.2-point gap) | Best baseline VideoRAG-A 0.510 (0.093 gap) | / |
| Proactive / hardest axis | Proactive under 21% for most; causal 41.79%, decision 44.18% (Gemini 3 Pro) | Generation pass rate 0.357 (Qwen3-VL-8B) | / |
| Purpose-built systems | Dispider 15.44%, VideoLLM-online-8B 8.48%, MMDuet2 6.79% | Best single config 0.507; random assignment 0.513 | / |
Ablations with the same harness inside each paper:
| Comparison | Score | Setting |
|---|---|---|
| CARVE all four configs, per-chunk | 0.603 Recall@5 | V-RAGBench test |
| CARVE best two-config partial combo | 0.567 Recall@5 | Same harness |
| CARVE best single config | 0.507 Recall@5 | Same harness |
| Random config / joint concat | 0.513 Recall@5 | Same harness |
| CARVE rule vs trained non-LLM router vs LoRA LLM router | 0.357 vs 0.329 vs 0.310 pass rate | Qwen3-VL-8B generator |
When to use which
Benchmark a frontier MLLM on multi-stream readiness → X-Stream. If your product involves more than one concurrent feed (multi-camera, multi-window, multi-device) single-stream video QA numbers tell you almost nothing. X-Stream’s 11 subtasks and its proactive-ability slice are the closest published measure of “can this model be a live operator.”
Build QA over hour-scale egocentric footage → a CARVE-style per-chunk retrieval pipeline. The recipe that survives its ablations: multiple modality-granularity configs in parallel, assign each chunk the config that retrieved it best, rerank only under that config. Expect the honest operating point to be near 0.6 Recall@5, not 0.9; plan your UX for missed evidence.
Evaluate any retrieval pipeline → leak-filter first, like V-RAGBench. If more than half of your eval queries are answerable without the video, your leaderboard is measuring language priors. The five-filter construction (shortcut-bias elimination plus evidence-uniqueness checks) is the transferable part, more than CARVE itself.
Navigate the field → the watch-remember-reason survey. It is the only one of the three with no new numbers, and it does not pretend otherwise; use it to place method papers under the right pillar before trusting their benchmark claims.
Limits and open questions
Both benchmarks are photographs of mid-2026 models, and X-Stream’s spread (Gemini 3 Pro at 49.60%, GPT-5 at 27.78%) will compress as frontier models refresh. V-RAGBench’s own limits are structural: evidence lives in non-overlapping 2-minute chunks, so queries whose evidence straddles a boundary or spans long ranges are out of scope, and knowledge-graph or metadata retrieval is set aside entirely. CARVE runs four retrievers plus a reranker per query, and the paper frames quality gains without a latency or cost budget, so the tradeoff against a cheaper fixed config is unquantified. Both papers also lean on specific judges (GPT-5.2-chat for filtering, LLM pass-rate judging), and how much the numbers depend on those exact judges is not stressed.
The open question neither paper answers: whether watch-everything and retrieve-what-matters compose. Nobody has run a frontier model with CARVE-style retrieval over a multi-stream X-Stream-style workload, and the watch-remember-reason survey explicitly flags streaming egocentric understanding as the unsolved regime. A system that retrieves across concurrent live feeds under a latency budget, with proactive triggers, is the benchmark that would actually settle the question.
FAQ
Which is harder, X-Stream or V-RAGBench?
They stress different things and their scores are not directly comparable. X-Stream’s headline gap is between models and humans on the same 4,220 QA pairs: the best model, Gemini 3 Pro, scores 49.60% against a 91.84% human baseline. V-RAGBench’s headline gap is between retrieval systems: CARVE reaches 0.603 Recall@5 against 0.510 for the best baseline on 2,100 leak-filtered triplets. X-Stream measures whether models can watch concurrent streams at all; V-RAGBench measures whether retrieval finds evidence in video too long to watch.
Can MLLMs watch multiple video streams at the same time?
Barely, as of the X-Stream evaluation. The best model, Gemini 3 Pro, reaches 49.60% overall on 4,220 multi-stream QA pairs while human annotators reach 91.84%, and proactive ability (reacting unprompted at the right moment) stays below 21% for most models. Purpose-built streaming architectures score worst: Dispider 15.44%, VideoLLM-online-8B 8.48%, MMDuet2 6.79%.
What is CARVE’s 0.603 Recall@5 measured against?
It is measured on V-RAGBench, a leak-filtered benchmark of 2,100 query/evidence/answer triplets over 216 egocentric videos lasting 11 to 99 hours. The 0.603 Recall@5 for CARVE compares against 0.510 for VideoRAG-A, the strongest of eight baselines, with each query’s evidence chunk labeled so retrieval is scored independently of generation.
Why does V-RAGBench filter out half of typical video QA queries?
Because the CARVE/V-RAGBench authors found that in widely used video QA benchmarks, more than half of the queries can be answered without the video at all, from language priors, world knowledge, or static cues. Such queries inflate generation accuracy while hiding retrieval failure, so V-RAGBench applies five filters, including shortcut-bias elimination and an evidence-uniqueness check, before keeping its 2,100 triplets.
Where does the watch-remember-reason survey fit next to these two benchmarks?
It is the map, not a measurement. The survey organizes 100+ long-video methods into three abilities (watching, remembering, and reasoning) and positions streaming egocentric understanding as the open frontier. X-Stream and V-RAGBench each benchmark one pillar (watching under concurrency, remembering via retrieval), and the survey explains why the third pillar, faithful reasoning over what was actually perceived, remains the least measurable of the three.