CoFiE: Coarse-to-Fine Evidence Selection for Efficient Streaming Video Understanding
Published in Conference on Empirical Methods in Natural Language Processing (EMNLP), 2026

CoFiE is a coarse-to-fine evidence selection framework for efficient streaming video understanding. It first uses lightweight, query-agnostic novelty signals to remove redundant frames before the vision encoder, avoiding unnecessary visual encoding computation.
During multimodal prefill, CoFiE then uses text-to-video attention to retain frames most relevant to the user query. Across StreamingBench and OvO-Bench, the framework establishes a strong accuracy-efficiency trade-off, preserving robust performance under aggressive frame filtering while improving end-to-end inference latency by up to 2.54×.
Accepted at the Conference on Empirical Methods in Natural Language Processing (EMNLP), 2026.
Recommended citation: Jing Jiang*, Yiran Ling*, Ruonan Li, Dimitrios Stamoulis, and Jie Liu. CoFiE: Coarse-to-Fine Evidence Selection for Efficient Streaming Video Understanding. Conference on Empirical Methods in Natural Language Processing (EMNLP), 2026. (* Equal contribution.)
