CoFiE: Coarse-to-Fine Evidence Selection for Efficient Streaming Video Understanding
CoFiE removes redundant frames before expensive visual encoding, then refines the retained candidates with query-specific evidence during LLM prefill.

Abstract
Streaming video understanding requires Vision Language Models (VLLMs) to process growing video streams and answer user questions under tight latency constraints. Existing methods improve efficiency through token pruning and memory-bank schemes, but mainly reduce visual tokens after visual encoding—after the expensive frame encoding cost has already been incurred.
We propose CoFiE, a Coarse-to-Fine Evidence Selection framework that decouples selection into a coarse, query-agnostic filtering stage before the vision encoder and a fine, query-specific refinement stage during LLM prefill. Experiments show that CoFiE establishes a new state-of-the-art accuracy–efficiency trade-off across multiple video understanding benchmarks, improving over prior methods by up to 3.15% while accelerating end-to-end inference by up to 2.54×.
Why coarse-to-fine?
Most acceleration methods prune visual tokens only after every frame has already passed through the vision encoder. That reduces part of the LLM workload, but cannot recover the dominant front-end computation. CoFiE aligns each selection decision with the information available—and the computation that can still be avoided—at that stage.
Novelty-Guided Frame Filtering
Uses lightweight temporal novelty to discard redundant raw frames while retaining a high-recall candidate set.
Query-agnostic · encoder cost avoidedQuery-Specific Evidence Refinement
Ranks encoded candidates with text-to-visual attention and keeps the evidence most relevant to the user query.
Query-aware · irrelevant context removed
Accuracy meets efficiency
CoFiE does not treat compression as a fixed trade-off. Removing redundant frames can reduce computation and interference from irrelevant context at the same time. Across the evaluated settings, the method maintains strong accuracy even under aggressive frame removal.


Filtering frames before visual encoding produces real end-to-end savings; refining evidence during prefill preserves the visual cues needed for accurate answers.
Full experimental results
Results across streaming, real-time perception, offline long-video understanding, and component-level ablations.
StreamingBench
| Method | OP | CR | CS | ATP | EU | TR | PR | SU | ACP | CT | Acc. |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Proprietary VLLMs | |||||||||||
| Gemini 1.5 Pro | 79.0 | 80.5 | 83.5 | 79.7 | 80.0 | 84.7 | 77.8 | 64.2 | 72.0 | 48.7 | 75.7 |
| GPT-4o | 77.1 | 80.5 | 83.9 | 76.5 | 70.2 | 83.8 | 66.7 | 62.2 | 69.1 | 49.2 | 73.3 |
| Claude 3.5 Sonnet | 80.49 | 77.34 | 82.02 | 81.73 | 72.33 | 75.39 | 61.11 | 61.79 | 69.32 | 43.09 | 72.44 |
| Open-source offline video VLLMs | |||||||||||
| MiniCPM-V 2.6 | 71.93 | 71.09 | 77.92 | 75.82 | 64.60 | 65.73 | 70.37 | 56.10 | 62.32 | 53.37 | 67.44 |
| InternVL-V2 | 68.12 | 60.94 | 69.40 | 77.12 | 67.70 | 62.93 | 59.26 | 53.25 | 54.96 | 56.48 | 63.72 |
| VILA-1.5 | 53.68 | 49.22 | 70.98 | 56.86 | 53.42 | 53.89 | 54.63 | 48.78 | 50.14 | 17.62 | 52.32 |
| Video-LLaMA2 | 55.86 | 55.47 | 57.41 | 58.17 | 52.80 | 43.61 | 39.81 | 42.68 | 45.61 | 35.23 | 49.52 |
| LLaVA-OneVision | 80.38 | 74.22 | 76.03 | 80.72 | 72.67 | 71.65 | 67.59 | 65.45 | 65.72 | 45.08 | 71.12 |
| Qwen2-VL-7B | 75.2 | 82.81 | 73.19 | 77.45 | 68.32 | 71.03 | 72.22 | 61.19 | 61.47 | 46.11 | 69.04 |
| Qwen2.5-VL-7B | 78.32 | 80.47 | 78.86 | 80.45 | 76.73 | 78.50 | 79.63 | 63.41 | 66.19 | 53.19 | 73.68 |
| Qwen3-VL-8B | 81.84 | 80.47 | 82.97 | 83.01 | 75.47 | 82.87 | 81.48 | 65.04 | 67.05 | 53.19 | 75.88 |
| Qwen3-VL-8B · drop 50% | 81.57 | 81.25 | 81.70 | 83.97 | 76.73 | 82.24 | 80.56 | 65.85 | 69.03 | 54.79 | 76.28 |
| Qwen3-VL-8B · drop 80% | 79.40 | 79.69 | 79.50 | 82.37 | 71.70 | 76.64 | 81.48 | 65.04 | 63.64 | 54.79 | 73.56 |
| Open-source online video VLLMs | |||||||||||
| Dispider | 74.92 | 75.53 | 74.10 | 73.08 | 74.44 | 59.92 | 76.14 | 62.91 | 62.16 | 45.80 | 67.63 |
| Flash-VStream | 25.89 | 43.57 | 24.91 | 23.87 | 27.33 | 13.08 | 18.52 | 25.20 | 23.87 | 48.70 | 23.23 |
| ViSpeak | 79.80 | 88.30 | 83.30 | 81.10 | 76.40 | 75.10 | 70.40 | 65.90 | 77.30 | 34.20 | 74.40 |
| StreamForest | 83.11 | 82.81 | 82.65 | 84.26 | 77.50 | 78.19 | 76.85 | 69.11 | 75.64 | 54.40 | 77.26 |
| TimeChat-Online | 80.22 | 82.03 | 79.50 | 83.33 | 76.10 | 78.50 | 78.70 | 64.63 | 69.60 | 57.98 | 75.36 |
| TimeChat-Online† | 81.84 | 80.47 | 82.97 | 83.01 | 75.47 | 82.87 | 81.48 | 65.04 | 67.05 | 53.19 | 75.88 |
| TimeChat-Online† · drop 50% | 80.23 | 82.11 | 78.80 | 81.87 | 78.30 | 75.54 | 75.47 | 63.39 | 63.96 | 54.44 | 72.53 |
| TimeChat-Online† · drop 80% | 79.46 | 78.05 | 79.60 | 81.35 | 76.42 | 70.82 | 77.36 | 61.61 | 61.26 | 55.56 | 71.14 |
| Ours | |||||||||||
| CoFiE | 81.03 | 84.38 | 90.54 | 83.92 | 74.68 | 81.31 | 88.89 | 67.07 | 73.30 | 58.51 | 78.58 |
| CoFiE · drop 50% | 82.93 | 84.38 | 89.91 | 82.96 | 75.32 | 84.11 | 87.96 | 67.48 | 72.44 | 57.45 | 78.86 |
| CoFiE · drop 80% | 82.93 | 83.59 | 88.96 | 83.28 | 74.68 | 79.44 | 87.96 | 65.85 | 72.44 | 55.85 | 77.82 |
OP: Object Perception · CR: Causal Reasoning · CS: Clip Summarization · ATP: Attribute Perception · EU: Event Understanding · TR: Text-Rich Understanding · PR: Prospective Reasoning · SU: Spatial Understanding · ACP: Action Perception · CT: Counting. † Improved reimplementation.
OvO-Bench · Real-Time Visual Perception
| Method | OCR | ACR | ATR | STU | FPD | OJR | Avg. |
|---|---|---|---|---|---|---|---|
| Proprietary VLLMs | |||||||
| Gemini 1.5 Pro | 87.30 | 67.00 | 80.20 | 54.50 | 68.30 | 67.40 | 70.80 |
| GPT-4o | 69.10 | 65.10 | 65.50 | 50.00 | 68.30 | 63.70 | 63.60 |
| Open-source offline video VLLMs | |||||||
| Qwen2-VL-7B | 69.13 | 53.21 | 63.79 | 50.56 | 66.34 | 60.87 | 60.65 |
| LLaVA-NeXT-Video-7B | 69.80 | 59.60 | 66.40 | 50.60 | 72.30 | 61.40 | 63.30 |
| LLaVA-OneVision-7B | 67.10 | 58.70 | 69.80 | 49.40 | 71.30 | 60.30 | 62.80 |
| InternVL-V2-8B | 68.50 | 58.70 | 69.00 | 44.90 | 67.30 | 56.00 | 60.70 |
| LongVU-7B | 55.70 | 49.50 | 59.50 | 48.30 | 68.30 | 63.00 | 57.40 |
| Qwen3-VL-8B | 77.85 | 59.63 | 73.28 | 53.37 | 69.00 | 57.07 | 65.03 |
| Qwen3-VL-8B · drop 50% | 79.19 | 61.47 | 71.55 | 55.62 | 68.00 | 57.61 | 65.57 |
| Qwen3-VL-8B · drop 80% | 75.17 | 56.88 | 71.55 | 50.00 | 67.00 | 59.78 | 63.40 |
| Open-source online video VLLMs | |||||||
| Dispider | 57.72 | 49.54 | 62.07 | 44.94 | 61.39 | 51.63 | 54.55 |
| TimeChat-Online | 75.20 | 46.80 | 70.70 | 47.80 | 69.30 | 61.40 | 61.90 |
| Flash-VStream | 25.50 | 32.10 | 29.30 | 33.70 | 29.70 | 28.80 | 29.90 |
| StreamForest | 68.46 | 53.21 | 71.55 | 47.75 | 65.35 | 60.87 | 61.20 |
| TimeChat-Online† | 77.85 | 59.63 | 73.28 | 53.37 | 69.00 | 57.07 | 65.03 |
| TimeChat-Online† · drop 50% | 78.52 | 58.72 | 73.28 | 51.12 | 66.00 | 59.24 | 64.48 |
| TimeChat-Online† · drop 80% | 72.48 | 55.05 | 68.97 | 48.88 | 66.00 | 53.26 | 60.77 |
| Ours | |||||||
| CoFiE | 79.87 | 58.72 | 78.45 | 56.18 | 75.00 | 60.33 | 68.09 |
| CoFiE · drop 50% | 80.54 | 59.63 | 76.72 | 56.74 | 74.00 | 64.67 | 68.72 |
| CoFiE · drop 80% | 76.51 | 52.29 | 75.86 | 53.37 | 76.00 | 61.96 | 66.00 |
OCR: Optical Character Recognition · ACR: Action Recognition · ATR: Attribute Recognition · STU: Spatial Understanding · FPD: Future Prediction · OJR: Object Recognition. † Improved reimplementation.
Offline long-video benchmarks
| Method | MLVU | LVB | MVB | Video-MME Long | Video-MME All |
|---|---|---|---|---|---|
| Qwen2-VL-7B | — | — | 67.0 | — | 63.3 |
| Qwen2.5-VL-7B | — | — | — | 50.4 | 63.2 |
| Qwen3-VL-8B | 76.94 | 63.73 | 66.88 | 65.1 | 71.7 |
| Dispider | 61.7 | — | — | — | 57.2 |
| TimeChat-Online | 62.6 | 55.4 | — | 48.4 | 62.4 |
| Flash-VStream | 66.3 | 42.0 | 65.4 | — | — |
| StreamForest | 70.0 | — | 70.2 | — | 61.9 |
| CoFiE | 77.79 | 66.26 | 71.12 | 66.4 | 72.1 |
0% drop. LVB and MVB denote LongVideoBench and MVBench; — indicates unreported results.
NGFF and QER components
| Variant | Drop (%) | Avg. | ΔAvg. |
|---|---|---|---|
| Qwen3-VL-8B | 0.0 | 65.03 | — |
| NGFF-only | 20.0 | 65.57 | +0.54 |
| QER-only | 37.5 | 65.03 | +0.00 |
| NGFF20% + QER37.5% | 50.0 | 66.85 | +1.82 |
| NGFF-only | 50.0 | 63.40 | −1.63 |
| QER-only | 60.0 | 65.03 | +0.00 |
| NGFF50% + QER60% | 80.0 | 64.95 | −0.08 |
| NGFF-only | 60.0 | 61.96 | −3.07 |
| QER-only | 75.0 | 65.03 | +0.00 |
| NGFF60% + QER75% | 90.0 | 62.99 | −2.04 |
OvO-Bench Real-Time Visual Perception. Two-stage selection improves the 50% drop setting by +1.82 points.
Lightweight filtering signals
| Filtering signal | Time (s) | Accuracy |
|---|---|---|
| Uniform Sampling · drop 0% | 319 | 65.03 |
| Low-Res Frame Difference | 292 | 61.01 |
| dHash | 298 | 61.73 |
| Downsampled SSIM | 297 | 62.54 |
| Optical Flow | 321 | 62.39 |
| Scene Change Detection + H-Diff | 294 | 61.01 |
| Histogram Difference | 293 | 63.58 |
Alternative signals use a 20% drop ratio; Histogram Difference is CoFiE’s default NGFF signal.
Method at a glance
Given a video prefix and user query , CoFiE first produces a compact candidate set before visual encoding. After those candidates interact with the query during multimodal prefill, the framework refines them into a smaller evidence set for answer generation:
The design is lightweight and compatible with modern video-capable VLLMs. Our implementation uses Qwen3-VL-8B-Instruct as the backbone.
Citation
@inproceedings{jiang2026cofie,
title = {CoFiE: Coarse-to-Fine Evidence Selection for Efficient Streaming Video Understanding},
author = {Jiang, Jing and Ling, Yiran and Li, Ruonan and Stamoulis, Dimitrios and Liu, Jie},
booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing},
year = {2026}
}Acknowledgments
This work was supported by the National Natural Science Foundation of China (grant No. 62350710797) and the NSFC Excellent Young Scientists Fund Program (Overseas).
