Signals 4 · free daily AI digest

What, Where, and How: Probing Spatiotemporal Representations in Video Foundation Models

Paper recorded by Signals 4 on 2026-09-01 in cs.CV. Abstract reproduced from arXiv; link to the original below.

Published 2026-09-01 on arXiv · recorded by Signals 4 on 2026-09-02

Category: cs.CV · 计算机视觉 · first seen 2026-09-02

Abstract

Self-supervised video foundation models learn rich spatiotemporal representations, yet it remains unclear what visual concepts these representations encode, where they emerge across transformer layers, and how they are geometrically organized. In this work, we tackle these three questions through a systematic layer-wise analysis of V-JEPA 2 and VideoMAE-v2. We leverage lightweight probes trained t

Read on arXiv →

#180 most recent of 237 cs.CV papers we have recorded · ↑ newer: SpatialGuard: Harness-Guided Verifiable Spatial Reasoning for Text-to- · ↓ older: Revisiting Cross-View Completion: Self-Supervised Pre-Training via Rec
Cite this page: What, Where, and How: Probing Spatiotemporal Representations in Video Foundation Models: the #180 most recent of 237 cs.CV papers we have recorded (as of 2026-09-01). Source: Signals 4 (Signals API) — https://data.jiangzhang.ca/signals4/t/papers/what-where-and-how-probing-spatiotemporal-representations-in-video-foundation-mo.html
Free to quote with attribution to “Signals 4 (Signals API)”. Machine-readable: papers.json
Related: More cs.CV papers · arXiv signals · All papers · Today in AI
Get 4 AI signals a day by email — free.
Subscribe free → See all plans →
Get 4 AI signals a day by email — free
All models · All repos · By company · Daily editions