EviDETR: Preserving Query-Relevant Temporal Evidence for Moment Retrieval and Highlight Detection
Beijing Normal-Hong Kong Baptist University · Zhuhai, China
* Equal contribution.
TL;DR
EviDETR keeps query-relevant evidence alive across the whole pipeline: reweight it before decoding (SFR), refine it per query with sparse Top-2 expert routing (TTop2MoE), then carry span-level retrieval evidence into clip-level saliency (MR2HD).
QVHighlights validation split, CLIP+SlowFast features, mean ± std over three seeds.
— view full size
Abstract
Joint video moment retrieval and highlight detection requires identifying query-relevant temporal segments while estimating clip-level saliency, yet DETR-style pipelines do not explicitly preserve query-relevant evidence throughout encoding, decoding, and cross-task prediction. We propose EviDETR, an evidence-preserving framework with three components. Semantic-aware Feature Reweighting (SFR) enhances query-relevant clip representations through saliency estimation and cross-modal interaction. A Temporal Top-2 Mixture-of-Experts (TTop2MoE) decoder performs query-adaptive refinement via sparse expert routing. MR-to-HD (MR2HD) fusion transfers span-level retrieval evidence to clip-level highlight prediction through confidence-weighted multi-scale aggregation. Using CLIP+SlowFast features, EviDETR achieves 69.29 R1@0.5, 54.77 R1@0.7, and 48.41 Avg. mAP for moment retrieval on QVHighlights, together with 41.83 HD-mAP and 68.33 HIT@1. Strong results on TACoS and Charades-STA further demonstrate cross-dataset transferability.
Visualization
— view full size
BibTeX
If you use EviDETR in academic work, please cite the paper.
@misc{sun2026evidetr, title = {EviDETR: Preserving Query-Relevant Temporal Evidence for Moment Retrieval and Highlight Detection}, author = {Sun, Haoran and Li, Yufan and Zhang, Qichen and Zhao, Haoran and Wang, Shuqi}, year = {2026}, note = {Preprint} }