237 papers recorded in this category. Newest first.
- Can 4D Foundation Models Remember?
Perceiving and remembering the visual world is fundamental to navigating and interacting with our environment. Current 4D foundation models, such as camera-controllable video model…
2026-09-17
- SplashSplat: Reconstructing Splashing Liquids from Real-World Multi-View Videos
A splash lives for a fraction of a second: sheets tear into ligaments and droplets, appearance is view-dependent and nearly textureless, and little persists long enough to track. R…
2026-09-17
- FlowSGS: Improving Flow Matching Priors for Inverse Imaging with Stochastic Interpolants
Flow matching has emerged as the state-of-the-art generative model and has been used for plug-and-play (PnP) priors to solve inverse problems in computational imaging. However, exi…
2026-09-17
- Should This Case Be Adapted? Prediction Fragmentation Controls Test-Time Adaptation
Episodic test-time adaptation resets a frozen segmenter to source weights $M_0$ on each case and adapts for a fixed step count. A fixed horizon conflates a cohort-level question, h…
2026-09-17
- Towards Scaling Marine Perception with Synthetic Data
Scalable machine learning in challenging underwater environments is strongly limited by the lack of labeled real-world training data. This data is often expensive and laborious to …
2026-09-17
- FunArt: Decoding Functional Structure and Articulation from Generative 3D Latents
To operate effectively in human environments, robots must identify articulated objects, segment their movable and interactive parts, and estimate their kinematic models. Existing a…
2026-09-17
- Learning Foresight without Explicit Trajectories for 3D Diffusion Policies
3D diffusion policies are strong at generating geometrically grounded actions from current observations, but successful manipulation requires not only knowing what motion is feasib…
2026-09-17
- Earth Surface Immune System for Rapid Monitoring of Unknown Anomalies
Earth surface anomalies, driven by escalating climate change, and expanding human activities, are increasing in both frequency and diversity, yet their limited historical data and …
2026-09-17
- DexTouch-WM: Learning Action-Conditioned Tactile World Models from Human Touch for Dexterous Robot Manipulation
Learning predictive models of contact-rich dexterous manipulation requires dense tactile interaction, but such data are costly to scale on real robots and remain tied to embodiment…
2026-09-17
- PROVIA: Procedure State Tracking for Online Mistake Detection in Egocentric Videos
An assistant watching egocentric video should notice a mistake from past frames alone, before the next step begins, and keep working once the person recovers. A mistake changes the…
2026-09-17
- Refinement Is Inherently Editable: Training-Free Prompt-to-Prompt Image Editing with Generative Refinement Network
Text-guided image editing must introduce the requested changes while preserving unrelated source content. Diffusion-based editors rely on spatial controls whose inaccuracies can le…
2026-09-17
- PhGS: Post-Hoc Pruning and Refinement of Single-View Feed-Forward 3D Gaussian Reconstructions
Recent single-view feed-forward 3D Gaussian Splatting (3DGS) generation predicts a fixed number of Gaussians per camera ray, introducing severe spatial redundancy. Most existing co…
2026-09-17
- INSPECT: Learning Robot View Selection from Assistant Use
Robots inspecting an assembly must determine which parts are present and whether they are correctly installed. During egocentric assembly assistance, head motion and workpiece hand…
2026-09-17
- RawSLAM: Online HDR Gaussian SLAM from Linear Radiance
Current dense visual SLAM systems rely almost exclusively on 8-bit tonemapped Low Dynamic Range (LDR) inputs, limiting their robustness in extreme lighting where shadows and highli…
2026-09-17
- CoRef-GS: Cooperative Referring Gaussian Splatting for Multi-Agent Scene Understanding
Referring scene understanding for embodied robots requires grounding object- and relation-centric language queries from a designated viewpoint. While a local semantic Gaussian map …
2026-09-17
- PointZero: 3D Point Track Completion for Learning Transferable 3D Dynamics
World models endow perceptual systems with the ability to predict how scenes evolve under interaction. They are most beneficial when trained on diverse volumes of data, to instill …
2026-09-16
- In-Context Robot Learning with VLM Agents
Enabling robots to adapt to unfamiliar environments as readily as humans remains a moonshot goal of embodied AI. No finite collection of demonstrations can cover every task and sit…
2026-09-16
- Adaptive Convolutional Sparse Coding via Information Bottleneck for Robust Visual Signal Representation
Visual signals require compact yet sufficient representations for robust downstream prediction. Convolutional sparse coding (CSC) provides an explicit mechanism for suppressing red…
2026-09-16
- Track, Articulate, Act: Generating Articulation from Casual Human Videos
Human videos contain rich causal evidence for robot manipulation: they reveal how hand motion induces object motion and produces task-relevant changes in object state. In this work…
2026-09-16
- PhysVGGT: Feed-Forward Dense Physical Property Estimation from A Single Image
Physical properties, such as friction, hardness, stiffness, and density, govern how robots should grasp, manipulate and interact with objects, yet estimating these properties from …
2026-09-16
- NormLift: From Lifted Features To Semantic Reliability In 3D Gaussian Splatting
Training-free weighted aggregation is widely used to lift 2D semantic features onto 3D Gaussians for open-vocabulary scene understanding, yet its theoretical role remains insuffici…
2026-09-16
- Copy What Is Seen, Generate What Is Not: Training-Free Anomaly-Aware Video Restoration
A surveillance system that detects an anomaly often has to repair the footage as well, yet the two tasks are studied in isolation: training-free anomaly detectors stop at a score o…
2026-09-16
- DISTA-Net++: Rethinking Infrared Small Target Unmixing Beyond Sub-Pixel Separation
Long-range infrared imaging frequently confronts dense target clusters whose diffraction-limited signatures merge into a single indistinguishable blob, concealing the number, sub-p…
2026-09-16
- Toward Markerless Video-based Tremor Analysis: Objective Quantification of Pathological Tremor in Mouse Preclinical Models
Tremor is a movement disorder characterized by involuntary, rhythmic oscillations of body parts and is a hallmark of several neurological conditions, including Parkinson's disease …
2026-09-16
- Geometry beneath the Waves: Dense Priors for Sparse-View Underwater 3D Gaussian Splatting
Underwater 3D reconstruction supports applications ranging from marine ecosystem monitoring and subsea inspection to underwater archaeology, education, and immersive visualisation.…
2026-09-16
- Mask IPL: Noise-Free Intrinsic Position Learning via Computation Graph Clipping for Event-Based Spike-Driven Tracking
Spiking Neural Networks (SNNs) match the event-driven nature of event cameras and naturally extract spatiotemporal features. These properties have motivated a series of recent stud…
2026-09-16
- RankGround: Efficient High-Resolution GUI Grounding via Lightweight Reranker-Guided Crop Selection
Graphical User Interface (GUI) grounding is a fundamental perception task for multimodal agents, enabling them to interpret natural language instructions and interact with digital …
2026-09-16
- Generalist-Specialist Mixture-of-Experts for Rare Pathology Detection in Multimodal Imaging
AI models for multimodal medical imaging must balance modality-specific specialization with cross-modal shared representations, a trade-off that pure Mixture-of-Experts (MoE) archi…
2026-09-16
- Video-Based Markerless Motion Capture for Clinical and Rehabilitation Biomechanics: A PRISMA-ScR Scoping Review of Validated Architectures, Clinical Readiness, and Emerging Methods
Background.. Video-based markerless motion capture promises movement analysis without the cost, space and skin-marker constraints of optoelectronic systems, with particular potenti…
2026-09-16
- ORCA: Occlusion-Aware Refinement and Completion for Novel View Synthesis
Novel-view synthesis from a single image is a fundamentally ambiguous problem. As the camera moves away from the input viewpoint, previously hidden regions become visible, exposing…
2026-09-15
- BrainFocus: EEG-Guided ROI Selection for Efficient Vision-Language Models
Vision-language models (VLMs) achieve strong visual question answering (VQA) performance, but processing large cluttered images is computationally expensive when only a small regio…
2026-09-15
- SlotDiT: Object-Centric Representations for Diffusion Transformers
Text-conditioned latent diffusion models perform strongly in video generation and are promising backbones for robotic applications. However, existing approaches rely on pixel-level…
2026-09-15
- SSC-Priors: Exploring Semantic and Visibility Priors to Boost Lidar Semantic Scene Completion
This paper investigates easy strategies to boost the performance of existing networks for lidar semantic scene completion (SSC) without requiring complex architectural redesigns. T…
2026-09-15
- PanoGS-SLAM: Panoramic 3D Gaussian Splatting SLAM
Real-time dense SLAM is a core capability for robotics applications that require robust localization and high- quality mapping in dynamic or fast-changing environments. Recent 3D G…
2026-09-15
- Optical-Flow Wingbeat Counting in MuJoCo: A Comparison of Convolutional, Spiking, and Attention-Based Temporal Models
Visual monitoring of flapping-wing vehicles requires distinguishing individual wingbeats from motion strength and average frequency. This paper presents a controlled MuJoCo evaluat…
2026-09-15
- Semantic-Spatial Agreement Verification for Mitigating Object Hallucination in Multimodal Large Language Models
Multimodal large language models generate natural-language responses from visual inputs, yet may mention objects absent from an image. In medication assistance, accessible percepti…
2026-09-15
- Exploring 2D backbone effects for indoor semantic occupancy prediction
Semantic occupancy prediction gives an embodied agent a voxel-level account of where space is free, occupied, and semantically meaningful. In RGB-D pipelines such as EmbodiedScan, …
2026-09-15
- Video-HolmesV2: Can MLLMs Reason with Spatio-Temporal Audio-Visual Evidence in Long Videos?
Multimodal Large Language Models have demonstrated impressive video understanding, yet their ability to reason over long-form narratives is often masked by visual-centric evaluatio…
2026-09-15
- DecoGS: Adaptive Static-Dynamic Decoupling of 3D Gaussians for Free-Viewpoint Video Streaming
Streaming 3D reconstruction demands both speed and temporal fidelity, goals that existing methods undermine by updating every Gaussian every frame, even in static regions. We prese…
2026-09-15
- FROD: Feature Matching Residual Denoising Oracle Bone Decipher
Oracle bone script (OBS), one of the earliest Chinese writing systems, plays an important role in the study of Chinese etymology. Traditional decipherment relies heavily on domain …
2026-09-15
- InfoTaxa: Information-Calibrated Label-Free Clustering for Fine-Grained Visual Taxonomy
Label-free clustering of frozen pretrained visual embeddings offers a scalable route to biodiversity monitoring, but image-only fine-grained taxonomy exhibits a consistent coarse-t…
2026-09-15
- Probe-VAD: Ordinal Likelihood Probing for Training-Free Video Anomaly Detection
Video anomaly detection (VAD) aims to localize anomalous events in untrimmed videos. Vision-language models (VLMs) provide rich visual understanding for training-free VAD, but exis…
2026-09-15
- EventEgoHands++: Event-based Egocentric 3D Hand Mesh Reconstruction with Real Dataset
3D hand mesh reconstruction is a challenging yet essential task for downstream applications, including human-robot interaction and AR/VR. Although conventional cameras have been wi…
2026-09-15
- Multimodal Cultural Heritage Architectural Style Classification for Residential Buildings in the UAE Based on CLIP Embeddings and SVM
The analysis and classification of cultural heritage architectural styles remain challenging due to the complexity of visual images of buildings, which are highly relied on in trad…
2026-09-15
- LynnReal-Omni: Native multi-modal Video Generation for Agentic Visual Workflows
Video diffusion models are stochastic and hard to control: precise content often requires repeated sampling without guaranteed success, and long-horizon scenes drift in appearance,…
2026-09-14
- TRACE: Two-Stage Detector-Response Estimation With Angular Cosine Expansion for Ring Artifact Correction in Photon-Counting CT
Detector response nonuniformity introduces systematic projection errors and ring artifacts in photon-counting detector computed tomography (PCD-CT). In measured PCD-CT data, residu…
2026-09-14
- Integrating Multi-view Multi-light Surface Reconstruction into Cultural Heritage Workflows
Cultural heritage documentation increasingly relies on image-based 3D surface reconstruction, with photogrammetry software making such workflows accessible to archaeologists, conse…
2026-09-14
- VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention
Diffusion Transformers deliver state-of-the-art video generation, but their long spatiotemporal sequences make attention the dominant deployment cost, and a deployable low-bit kern…
2026-09-14
- SURE-Map: Self-Correcting Streaming Geometric Foundation Model
Streaming geometric foundation models are emerging as a compelling alternative to SLAM systems. Yet this streaming nature introduces a fundamental issue: each prediction is made fr…
2026-09-14
- TopoRig: Topology-Agnostic Facial Rigging via Multi-Source Supervision
Automatic facial rigging across heterogeneous mesh topologies remains challenging because high-quality expression supervision is often tied to canonical templates, while deformatio…
2026-09-14
- Transforming harmonic coefficients for 3D splat compression
We address the problem of color attribute compression for 3D splats. We show that all images generated by 3D splats are linear in the coefficients for each color channel, each sphe…
2026-09-14
- Bench2Dex: Benchmarking Visuo-Tactile Bimanual Dexterous Manipulation Across Dexterous Hands
Tactile sensing provides contact information that can be difficult to infer from vision alone, but tactile hardware for dexterous hands has not converged to a common design. Dexter…
2026-09-14
- Predicting build orientation for SLM dental parts: a comparison of rotation representations and direct vector regression
Build orientation for selective laser melting (SLM) manufacturing of dental parts is usually chosen manually by technicians. We treat orientation prediction as supervised machine l…
2026-09-14
- Can a Neural Encoding Model Replicate an fMRI Visualization Study?
Most knowledge of graphical perception comes from behavioral studies. Understanding from a neural perspective is much more limited due in part to neuroimaging studies' expensivenes…
2026-09-14
- V-ICAL Bench: Evaluating Video In-Context Learning for Multimodal Agents in Interactive Environments
While In-Context Learning (ICL) enables models to adapt from exemplars without parameter updates, multimodal ICL remains largely underexplored, particularly regarding video demonst…
2026-09-14
- MambaMPD: A Mamba-Driven Segmentation Framework for Marine Pollution Detection from Remote Sensing Imagery
Accurate marine pollution detection (MPD) is essential for protecting coastal ecosystems and marine biodiversity. Vision Mamba models have shown promise in remote-sensing semantic …
2026-09-14
- Don't Send What You Don't Need: Question-Guided Token Pruning as a Privacy Defense for Vision-Language Models
Visual Question Answering (VQA) with Vision-Language Models (VLMs) is increasingly used in privacy-sensitive and bandwidth-constrained settings. Federated Learning (FL), Split Lear…
2026-09-14
- Benchmarking Intra-Patient 3D Deformable Multimodal Image Registration
Multimodal image registration is a key component of many clinical workflows, yet it remains challenging because corresponding anatomical structures often exhibit substantially diff…
2026-09-14
- Circuit-MLLM: Topological Logic-Guided Latent-Space Visual Reasoning for Circuit Schematic Understanding
Through pre-training on extensive text and image datasets, current multi-modal large language models (MLLMs) achieve strong performance on general tasks. However, circuit schematic…
2026-09-14
- Kaininja: Extending Native 3D Generators to the Part Level
Native 3D generators turn one image into a single mesh. TRELLIS.2 and its peers deliver high-fidelity non-watertight geometry with materials, but the output is one fused object, wh…
2026-09-14
- From Model Patterns to Abstract Semantics in Compositional Zero-Shot Learning
Compositional Zero Shot Learning aims to recognize unseen compositions by recombining learned primitives. Recent methods rely on vision language models and attempt to explicitly mo…
2026-09-14
- SNAP3D: Physically Grounded 3D Parts for Assembly from a Single Image
Part-aware 3D asset generation enables applications such as editing, articulation, simulation, and fabrication, yet existing methods can generate visually complete individual parts…
2026-09-11
- Recurrent Dynamic Range Extension
We present an approach to progressively extend the highlights of an image. Instead of reconstructing the full dynamic range of a complex scene directly, we learn a simpler task fir…
2026-09-11
- SV-Cine: Diagnosis-Conditioned Segmentation of Single Ventricle Physiology via Generative Data Augmentation
Single Ventricle Physiology (SVP) is a rare subtype of congenital heart disease characterized by the presence of a single functional cardiac ventricle with atypical anatomic config…
2026-09-11
- Investigating Temporal Motion Features for Pose-to-Text Indian Sign Language Translation
We investigate the effect of pretrained T5 model scale and explicit motion features on pose-to-text Indian Sign Language Translation (SLT) for the WSLP 2026 Shared Task. Pose seque…
2026-09-11
- Generative Retrieval for Unsupervised Text-Based Person Search
Text-based person search (TBPS) aims to retrieve images of a target person from a large image gallery based on a given natural language description. Most existing methods rely on s…
2026-09-11
- Fast and Faithful: Principled Conditional Flow Matching for Inverse Problems
Flow matching approaches to imaging inverse problems commonly incorporate measurements in two ways. Conditioning-based approaches supply measurement-derived information as a networ…
2026-09-11
- DementiaCare-Bench: A Modality-Validated Video Benchmark
Dementia affects an estimated 57 million people worldwide, and for most families the hardest part of care is not memory loss but the behavioral and psychological symptoms of dement…
2026-09-11
- Input Resolution Matters: Real-Time Object Detection Latency
We model total latency as the convolution of preprocessing, inference, and postprocessing distributions under a simplifying independence approximation, with selected stage paramete…
2026-09-11
- PA-CDM: Position-Aware Character Detection Matching for Evaluating Handwritten Mathematical Expression Recognition
Handwritten mathematical expression recognition (HMER) is conventionally scored by exact-match rates and string-similarity metrics that are blind to where an error occurs: two pred…
2026-09-11
- Parallel Training Using a CNN-DNN Architecture for Accelerated Development of Diagnostic Models
Artificial intelligence has shown promise in assisting radiologists in imaging-based diagnosis across a wide range of diseases. Efficient training of large deep learning models is …
2026-09-11
- UniPart: Towards Zero-shot Language-Grounded 3D Part Segmentation for Embodied Interaction
Fine-grained robotic manipulation depends on understanding parts, not only whole objects. Existing 3D foundation models tend to be either generalized but object-aware, or part-awar…
2026-09-11
- Beyond Accuracy: Uncertainty-Guided Boundary Refinement for Reliable Biomedical Image Segmentation
Accurate biomedical image segmentation requires not only high global overlap but also reliable delineation of clinically meaningful boundaries. In blood-smear microscopy, cytoplasm…
2026-09-11
- Learning Sign Language Recognition under Label Noise: A Study of Noise-Robust Losses for Isolated and Continuous Settings
In sign language recognition, the isolated (ISLR) classification loss treats a single label as ground truth, as does the frame-level auxiliary classifier over pseudo-labels we add …
2026-09-11
- VideoTok4D: A 4D-Aware Video Tokenizer for Compact World Representation
Video tokenizers have emerged as a cornerstone of modern video modeling, underpinning progress in compression, reconstruction and generation by mapping high-dimensional visual sign…
2026-09-11
- A Multi-Vehicle Dataset with Camera, LiDAR, and Radar Sensors and Scanned 3D Models for Custom Auto-Annotation using RTK-GNSS
Datasets are a crucial element in the development of perception algorithms. They relate sensor measurement data to annotated reference information and allow for the deduction of se…
2026-09-11
- 3D CT-to-PET Translation via Latent Brownian Bridge Diffusion
Computed tomography (CT) and positron emission tomography (PET) provide complementary anatomical and functional information for cancer diagnosis and treatment planning. However, th…
2026-09-11
- MGAvatar: Mesh-Bound Gaussians for Head Avatar Geometry and Appearance Modeling
Accurate head modeling requires a stable yet expressive geometric representation. Existing Gaussian-based head avatars commonly rely on parametric templates (e.g., FLAME) for Gauss…
2026-09-11
- SenseNova-U1.5: Towards Native Unified Visual Intelligence
We launch SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VAE-free architecture. …
2026-09-10
- Caption-once, Frames-on-Demand: Visual-Need Routing for Budget-Aware Agentic Long Video Understanding
Long-video understanding on edge devices must reason over hours of content under tight compute and bandwidth budgets. Subsampling visual tokens loses temporal structure, while text…
2026-09-10