The latest new AI papers on arXiv we picked up from arxiv. Updated as new data arrives; the free daily digest sends one signal per core source.
- Sep 10, 2026 arxiv
SenseNova-U1.5: Towards Native Unified Visual IntelligenceWe launch SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VAE-free architecture.
- Sep 10, 2026 arxiv
GPU-CFR: 80x Faster Counterfactual Regret Minimization by Compiling the Game to Static Dataflow and CUDA Graph ReplayCounterfactual regret minimization (CFR) is one of the few large numerical workloads that still runs faster on CPUs than on GPUs. Each iteration sweeps a game tree with up to billi
- Sep 10, 2026 arxiv
General Quantification of Covariate and Concept ShiftsGeneralization under distribution shift remains a core challenge in modern machine learning, yet existing learning bound theory is limited to narrow, idealized settings and is non-
- Sep 10, 2026 arxiv
Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated DataAs the supply of human-written text is exhausted, it has become standard practice to repeat language model training data. Prior work has studied data repetition for densely activat
- Sep 10, 2026 arxiv
Can Edge-Deployable Vision-Language Models Identify Species?Camera traps often run in the field on edge hardware with limited or no connectivity, making small, locally-deployable vision-language models (VLMs) -- not frontier-scale ones -- t
- Sep 10, 2026 arxiv
Generative Marketing Mix Modeling: A Causal Inference Framework Linking GEO and GEM to Business ImpactGenerative artificial intelligence changes how firms reach customers, but standard marketing data do not record how often users see and notice a firm's name in generated answers. W
- Sep 10, 2026 arxiv
Distance generalization in transformers: why bother with positional encoding?Out-of-distribution length generalization, namely to extrapolate a task from short to longer context, has been studied intensively for transformers. Here we focus on distance gener
- Sep 10, 2026 arxiv
Artificial Id: Drive and Persistent Alignment in Agentic AIAgentic AI is moving from bounded task execution toward systems that retain consequential state, continue operating and adapt across task boundaries. That shift creates a control p
- Sep 10, 2026 arxiv
From Protocols to Evidence: Bounded Claims for AI in Service of the Common GoodArtificial Intelligence does more than create a governance problem. It can also reveal where institutions have already failed to provide responsiveness, belonging, care, and accoun
- Sep 10, 2026 arxiv
TART: A Modular Tool for Technique-Aware Audio-to-Tablature Guitar TranscriptionAutomatic Music Transcription (AMT) for guitar remains limited by three challenges: existing systems often fail to capture expressive techniques such as slides, bends, and percussi
- Sep 10, 2026 arxiv
MindTopo: Can Foundation Models Reason in Topological Space?Spatial reasoning depends not only on metric properties such as distance, angle, and shape, but also on topological relations that remain invariant under continuous deformation. Co
- Sep 10, 2026 arxiv
Caption-once, Frames-on-Demand: Visual-Need Routing for Budget-Aware Agentic Long Video UnderstandingLong-video understanding on edge devices must reason over hours of content under tight compute and bandwidth budgets. Subsampling visual tokens loses temporal structure, while text
- Sep 10, 2026 arxiv
CausalArena: Benchmarking Causal Discovery in the Foundation Model EraCausal discovery aims to uncover causal structures from data and is fundamental to scientific reasoning and intervention-based decision making. Its evaluation relies heavily on str
- Sep 10, 2026 arxiv
3D Point Splatting for mmWave Radar Novel View SynthesisSolving novel view synthesis (NVS) for millimeter-wave (mmWave) radar requires a renderer that is physically faithful, complex-valued, and multi-viewpoint-tractable. No prior metho
- Sep 10, 2026 arxiv
Nuha-Speech: Building General-Purpose Arabic Speech-LLMsAs Speech Large Language Models (speech-LLMs) become increasingly multilingual, Arabic remains significantly underrepresented, highlighting the need for dedicated infrastructure to
- Sep 10, 2026 arxiv
Guided Super-Resolution of Digital Elevation Models with Diffusion-Based Image GeneratorsHigh-resolution digital surface models (DSMs) play an important role in urban analysis, 3D building reconstruction, and infrastructure monitoring, yet their availability remains li
- Sep 10, 2026 arxiv
CoRA-NAS: Coarse Ranking and Anchor-Residual Refinement for Neural Architecture SearchZero-cost proxies rank architectures cheaply, but their reliability varies across search spaces. We introduce CoRA-NAS (COarse Ranking + Anchor-residual), a two-stage framework com
- Sep 10, 2026 arxiv
Domain-Specific Hallucination Detection in Large Language ModelsLarge language models generate fluent text that can contain unfaithful claims -- a phenomenon known as hallucination. We present a multi-signal detection pipeline combining fine-tu
- Sep 10, 2026 arxiv
Biology-in-the-loop: Amortized Adaptive Hit Discovery in CRISPR ScreensMany biological discovery problems require experiments to be selected sequentially under constrained budgets. CRISPR screening is a prominent example, as exhaustive perturbation te
- Sep 10, 2026 arxiv
On the Regularization Landscape for the Linear Recommendation ModelsRecently, a wide range of recommendation algorithms inspired by deep learning techniques have emerged as the performance leaders on several standard recommendation benchmarks. Whil
- Sep 10, 2026 arxiv
The Last AI Built by Humans: Toward Genuine Recursive Self-ImprovementRecursive self-improvement (RSI) enables AI systems to turn experience and feedback into persistent changes that improve both their capabilities and the process of future improveme
- Sep 10, 2026 arxiv
Evaluating Time-Series Foundation Models and Multimodal Dietary Context for CGM ForecastingContinuous glucose monitoring (CGM) provides high-frequency measurements of glucose dynamics and enables short-term glucose forecasting for diabetes management. Although time-serie
- Sep 10, 2026 arxiv
Augustinian BabyLM: What Ostensive Definition Can and Cannot Teach a Small Language ModelA language model normally begins training with random word embeddings: whatever 'banana' means must be learned from training corpora. I implement St. Augustine's picture of word le
- Sep 10, 2026 arxiv
AdamX: Cosine similarity meets gradient descentWe introduce AdamX, a first-order optimizer that incorporates cosine similarity as an adaptive mechanism for controlling update magnitudes. The proposed method is scalable, model-a
- Sep 10, 2026 arxiv
Epistemic orientation predicts legislative effectiveness among members of the US CongressTruth and evidence-based communication provide important foundations for democratic governance, accountability, and collective decision-making. Prior work shows that evidence-orien
- Sep 10, 2026 arxiv
RetroThinker: Enabling Retrospective Thinking in Speech LLMsSpeech large language models (SpeechLLMs) offer reduced latency and retain paralinguistic nuances that are typically lost in cascaded automatic speech recognition (ASR) and text-ba
- Sep 10, 2026 arxiv
Explainability Assistant: A Conversational XAI Interface for Interpreting Energy Consumption ModelsEnergy consumption forecasting relies on increasingly complex machine learning (ML) models, such as Genetic Programming-based symbolic regressors, whose predictions can be difficul
- Sep 10, 2026 arxiv
From Parameters to Answers: How LLMs Retrieve and Use Their Internal KnowledgeHow does a language model's dependence on query-routing information and target knowledge change as it answers a question? We study this question through layerwise interventions on
- Sep 10, 2026 arxiv
IndicTriMix: Developing Language Identification Datasets and Models for Tri-Language Code-MixingLanguage identification in code-mixed text, largely observed in social media, is highly essential when users frequently switch between multiple languages within a single utterance.
- Sep 10, 2026 arxiv
Model-Aware Schedules Improve Generation via Fiberwise Optimal TransportDiffusion and flow-matching schedules control the signal and noise coefficients that mix data and noise along affine probability paths. Minimizing a kinetic action defined on coeff
- Sep 10, 2026 arxiv
Target leakage, not model class, explains reported accuracy in survey-based cardiovascular screening: a leakage-tiered audit of glass-box and tabular foundation modelsCardiovascular screening models trained on national health surveys routinely report areas under the receiver operating characteristic curve (AUROC) near 0.89. We asked whether that
- Sep 10, 2026 arxiv
Near-Optimal Reinforcement Learning with Multi-Step Transition LookaheadWe study reinforcement learning (RL) with transition look-ahead, where the agent may observe which states would be visited upon playing any sequence of $\ell$ actions before decidi
- Sep 10, 2026 arxiv
Understanding Operator Attitudes Toward AI-Supported Decision Making in Maritime OperationsMaritime Autonomous Surface Ships (MASS) and AI- supported decision assistants are expected to transform maritime operations, but their safe integration depends on how maritime pro
- Sep 10, 2026 arxiv
Logit Refiner: Improving Visual Autoregressive Models via Intra-Scale Dependency ModelingVisual Autoregressive Models (VAR) generate images through next-scale prediction, producing all tokens within each scale in parallel. We show that this parallel decoding constitute
- Sep 10, 2026 arxiv
Thinking with Looped FlowsHumans and machines often solve harder problems by spending more time on computation. In deep learning, looped models implement this idea during inference by recurrently updating a
- Sep 10, 2026 arxiv
SpecGuard: Inference-Time Backdoor Detection For FreeLarge language models are often fine-tuned, shared, or downloaded from third parties, so a deployed model may carry a hidden backdoor that behaves normally on benign inputs but swi
- Sep 10, 2026 arxiv
Dynamic language model representations for multi-objective reaction optimisationOptimising chemical reactions across multiple objectives, such as yield, selectivity, and safety, is central to chemical synthesis, and model-driven approaches depend critically on
- Sep 10, 2026 arxiv
Beyond Word Error Rate: A Switch Aware Evaluation of ASR and Audio Language Models on English Yoruba Code-Switched SpeechAutomatic speech recognition (ASR) systems and audio language models (audio LMs) now report low error rates on monolingual benchmarks, but their behavior on code switched speech in
- Sep 10, 2026 arxiv
Predicting Privacy Leakage from Weight Spectral DensityMembership inference attacks (MIAs) are widely used to audit the privacy disclosure risk of machine learning models, however current state-of-the-art attacks require training compu
- Sep 10, 2026 arxiv
Differentially Private EEG Feature Anonymization: A Privacy-Utility Case Study in Clinical NeurophysiologyClinical electroencephalography (EEG) data are valuable for healthcare research and for developing artificial intelligence (AI)-based clinical decision-support systems, but EEG rec
- Sep 10, 2026 arxiv
Whisper-Based Speech Transcription from Videos Across Multiple Languages for Cross-Cultural UnderstandingCross-cultural understanding has become increasingly important in today's highly connected, cross-national world. The success of LLM-based technologies is now driving the developme
- Sep 10, 2026 arxiv
The widening evaluation gap in medical large language model research 2023 to 2026Large language models are superseded every few quarters; clinical evidence takes years. We asked whether medical research is keeping pace with the systems it evaluates. PubMed retu
- Sep 10, 2026 arxiv
Recognizing Is Not Reversing: A Controlled Inversion Test of Fact-Preserving News FramingLarge language models (LLMs) are increasingly used to analyze and rewrite news, yet current framing studies mainly evaluate generation, detection, or whether rewritten text appears
- Sep 10, 2026 arxiv
A Unified Per-Token Gating Family for On-Policy Distillation: FKL/RKL Mixing with Multi-Channel and Bias CoefficientsPer-token gating of forward/reverse KL losses has become a standard technique for on-policy knowledge distillation (OPD), but existing methods such as EOPD (Jin et al., 2026) and T
- Sep 10, 2026 arxiv
Component-Aware Differential Privacy for Federated Multilingual Speech-LLMsPer-layer differential privacy (DP) clipping improves gradient fidelity in federated learning by allocating per-matrix clipping budgets proportional to parameter count. We show tha
- Sep 10, 2026 arxiv
RAG-Safety-Bench: Reliable Evaluation of Retrieval-Augmented LLM SafetyAllowing large language models (LLMs) to retrieve information from a set of trusted documents can increase reliability and reduce hallucination. However, recent work has demonstrat
- Sep 10, 2026 arxiv
Revisiting Avatar-As-Image: High-Fidelity Registration is All You NeedThe representation of 3D clothed humans as standardized 2D UV texture and displacement maps over an underlying body model has long been studied. This compact representation is enti
- Sep 10, 2026 arxiv
MC-DeTra: Motion-Consistent Joint Object Detection and Socially-Aware Trajectory Forecasting in Bird's-Eye-View ImagesUnified models for object detection and trajectory forecasting aim to merge perception and prediction for autonomous driving, refining actor trajectories directly over shared bird'
- Sep 10, 2026 arxiv
Language-Augmented Semantic Priors for B-Spline Surface FittingThe use of B-splines and Non-Uniform Rational B-Splines surfaces constitutes the mathematical foundation of contemporary computer-aided design (CAD) systems. Despite long-term prog
- Sep 10, 2026 arxiv
Spectral Adapters for Segment Anything Model-based Segmentation of Colorectal Liver Metastases in Computed TomographyAccurate segmentation of colorectal liver metastases (CRLM) in contrast-enhanced computed tomography (CT) is important for response assessment, surgical planning, and follow-up. We
- Sep 10, 2026 arxiv
Single-Stream Multi-Feature Fusion with Temporal Robustness for Gait Emotion Recognition3D skeleton-based gait emotion recognition faces high annotation costs, data scarcity, and poor generalization on heterogeneous data. This paper proposes SV-GCN, a single-stream mu
- Sep 10, 2026 arxiv
Multimodal Taxonomic Conditioning for Generative Plankton ImageryAutomated plankton imaging produces severely long-tailed datasets, where the rare taxa of greatest ecological interest have too few images to train or evaluate classifiers reliably
- Sep 10, 2026 arxiv
Self-Supervised Cardiac Phase Detection via Single-Parameter Latent OrbitsAccurate identification of end-diastole (ED) and end-systole (ES) in echocardiography underpins the quantification of ventricular function, yet manual selection of these key frames
- Sep 10, 2026 arxiv
Vidu S2: Real-Time Interactive, Editable, and Spatial Video GenerationWe present Vidu S2, which comprises Vidu S2-Avatar, a real-time interactive digital-character model, and Vidu S2-Editing, a real-time video editing model. Moreover, we explore the
- Sep 10, 2026 arxiv
LangStreet: Persistent Language Fields for Anchor-Decoded Street GaussiansLanguage Gaussian fields implicitly assume that the primitive carrying semantics remains identifiable across views. This assumption breaks in scalable anchor-decoded representation
- Sep 10, 2026 arxiv
MMGait: Benchmarking and Unifying Gait Recognition across Heterogeneous ModalitiesGait recognition is commonly studied using RGB videos or their derived silhouettes and poses. Yet human walking produces heterogeneous photometric, geometric, and motion cues that
- Sep 10, 2026 arxiv
OmniKVQuant: KV Cache Quantization for Omni-LLMsAs Omni-modal large language models (Omni-LLMs) take in audio, video and text together, their KV cache memory cost grows. KV cache quantization is the de facto approach in text-onl
- Sep 10, 2026 arxiv
Learn the Solid, Not the File: Canonical Inputs for Neural Networks on CAD Boundary RepresentationsBoundary representation (B-rep) is the standard format used by modern CAD systems for parametric 3D models. It turns out, the exact same solid can be represented by different B-rep
- Sep 10, 2026 arxiv
A Comparative Evaluation of Pre-trained Convolutional Neural Networks for Melanoma DetectionEarly diagnosis of melanoma is critical for improving patient survival rates. However, accurately distinguishing melanoma from other skin lesions remains a significant clinical cha