CoFiE: Coarse-to-Fine Evidence Selection for Efficient Streaming Video Understanding

Published in Conference on Empirical Methods in Natural Language Processing (EMNLP), 2026

CoFiE overview

CoFiE is a coarse-to-fine evidence selection framework for efficient streaming video understanding. It first uses lightweight, query-agnostic novelty signals to remove redundant frames before the vision encoder, avoiding unnecessary visual encoding computation.

During multimodal prefill, CoFiE then uses text-to-video attention to retain frames most relevant to the user query. Across StreamingBench and OvO-Bench, the framework establishes a strong accuracy-efficiency trade-off, preserving robust performance under aggressive frame filtering while improving end-to-end inference latency by up to 2.54×.

Accepted at the Conference on Empirical Methods in Natural Language Processing (EMNLP), 2026.

Project Page

Recommended citation: Jing Jiang*, Yiran Ling*, Ruonan Li, Dimitrios Stamoulis, and Jie Liu. CoFiE: Coarse-to-Fine Evidence Selection for Efficient Streaming Video Understanding. Conference on Empirical Methods in Natural Language Processing (EMNLP), 2026. (* Equal contribution.)