Signals 4 · free daily AI digest

The Shape of Time: Video-Token Contrast for Temporal Understanding in VideoLMs

Paper recorded by Signals 4 on 2026-09-03 in cs.CV. Abstract reproduced from arXiv; link to the original below.

Published 2026-09-03 on arXiv · recorded by Signals 4 on 2026-09-04

Category: cs.CV · 计算机视觉 · first seen 2026-09-04

Abstract

Seeing frames in order does not mean representing time. Modern VideoLMs receive ordered video streams, yet their main supervision acts on generated text rather than video-token representations where event dynamics should first emerge. This mismatch allows models to learn temporal answers from shortcuts such as objects, scenes, and language priors, without requiring internal video representations t

Read on arXiv →

#151 most recent of 237 cs.CV papers we have recorded · ↑ newer: BooM-VVT: Boosting Mask-Free Video Virtual Try-On with Image-Level Pse · ↓ older: Adaptive Vision-Language Grasping via Composable Foundation Priors and
Cite this page: The Shape of Time: Video-Token Contrast for Temporal Understanding in VideoLMs: the #151 most recent of 237 cs.CV papers we have recorded (as of 2026-09-03). Source: Signals 4 (Signals API) — https://data.jiangzhang.ca/signals4/t/papers/the-shape-of-time-video-token-contrast-for-temporal-understanding-in-videolms.html
Free to quote with attribution to “Signals 4 (Signals API)”. Machine-readable: papers.json
Related: More cs.CV papers · arXiv signals · All papers · Today in AI
Get 4 AI signals a day by email — free.
Subscribe free → See all plans →
Get 4 AI signals a day by email — free
All models · All repos · By company · Daily editions