EviDETR

EviDETR: Preserving Query-Relevant Temporal Evidence for Moment Retrieval and Highlight Detection

Haoran Sun*  ·  Yufan Li*  ·  Qichen Zhang*  ·  Haoran Zhao  ·  Shuqi Wang

Beijing Normal-Hong Kong Baptist University · Zhuhai, China

* Equal contribution.

TL;DR

EviDETR keeps query-relevant evidence alive across the whole pipeline: reweight it before decoding (SFR), refine it per query with sparse Top-2 expert routing (TTop2MoE), then carry span-level retrieval evidence into clip-level saliency (MR2HD).

R1@0.5 69.29±0.52
Moment retrieval
R1@0.7 54.77±0.53
Moment retrieval
Avg. mAP 48.41±0.09
Moment retrieval
HD-mAP 41.83±0.39
Highlight detection
HIT@1 68.33±1.26
Highlight detection

QVHighlights validation split, CLIP+SlowFast features, mean ± std over three seeds.

EviDETR architecture. A video encoder and a text encoder feed an SFR Encoder that contains a LocalSaliencyHead, a Reweighting Gate, a V2TExtractor and a Convolutional Fuser built from a T2TEncoder, a conv block and an encoder. Its output goes to a TTop2MoE Decoder with three Temporal MoE layers; the first layer expands into a Router and a Top2-Experts Fusion module. The decoder output goes to a prediction head holding an HD head with MR2HD fusion, which produces the saliency score, and an MR head, which produces logits and spans. — view full size
Overview of EviDETR. The SFR encoder reweights clips by query relevance before decoding. The TTop2MoE decoder then refines each temporal query over three MoE layers, routing it to two of eight experts, and the prediction head fuses span evidence into the clip-level saliency score.

Abstract

Joint video moment retrieval and highlight detection requires identifying query-relevant temporal segments while estimating clip-level saliency, yet DETR-style pipelines do not explicitly preserve query-relevant evidence throughout encoding, decoding, and cross-task prediction. We propose EviDETR, an evidence-preserving framework with three components. Semantic-aware Feature Reweighting (SFR) enhances query-relevant clip representations through saliency estimation and cross-modal interaction. A Temporal Top-2 Mixture-of-Experts (TTop2MoE) decoder performs query-adaptive refinement via sparse expert routing. MR-to-HD (MR2HD) fusion transfers span-level retrieval evidence to clip-level highlight prediction through confidence-weighted multi-scale aggregation. Using CLIP+SlowFast features, EviDETR achieves 69.29 R1@0.5, 54.77 R1@0.7, and 48.41 Avg. mAP for moment retrieval on QVHighlights, together with 41.83 HD-mAP and 68.33 HIT@1. Strong results on TACoS and Charades-STA further demonstrate cross-dataset transferability.

Visualization

Three stacked panels for one query. The top panel is a filmstrip of twelve sampled video frames. The middle panel plots predicted temporal spans against the ground-truth span on a 0 to 150 second axis at three confidence levels. The bottom panel plots normalized ground-truth and predicted saliency curves together with ground-truth clip saliency bars. — view full size
One query, end to end. Top: twelve frames sampled from the video. Middle: the retrieved span against the annotated moment and a baseline, over a 90–130 s window. Bottom: normalized saliency curves for the ground truth, EviDETR and the baseline, with per-clip ground-truth saliency on the 0–12 scale. EviDETR's span overlaps the annotated moment more closely than the baseline's, and its saliency curve rises across that segment.

BibTeX

If you use EviDETR in academic work, please cite the paper.

evidetr.bib
@misc{sun2026evidetr,
          title  = {EviDETR: Preserving Query-Relevant Temporal Evidence for Moment Retrieval and Highlight Detection},
          author = {Sun, Haoran and Li, Yufan and Zhang, Qichen and Zhao, Haoran and Wang, Shuqi},
          year   = {2026},
          note   = {Preprint}
        }