Back to papers
July 27, 2026cs.CVcs.AI

The Visual Bottleneck: Sparse-Frame Adaptation of MLLMs for Joint Spatial-Temporal Video Grounding

Categories

cs.CV, cs.AI