Paper recorded by Signals 4 on 2026-09-24 in cs.AI. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-24 on arXiv · recorded by Signals 4 on 2026-09-25
Category: cs.AI · 人工智能 · first seen 2026-09-25
Vision-language models (VLMs) have become a common foundation for autonomous mobile GUI agents, but most existing systems rely on the VLM for both planning and action grounding at nearly every interaction step, leading to substantial latency and model-serving cost. We introduce Jev-Mobile, which shifts this paradigm to low-frequency VLM planning and high-frequency lightweight execution: the VLM sp