Paper recorded by Signals 4 on 2026-08-31 in cs.CV. Abstract reproduced from arXiv; link to the original below.
Published 2026-08-31 on arXiv · recorded by Signals 4 on 2026-09-01
Category: cs.CV · 计算机视觉 · first seen 2026-09-01
Recent video generators often omit audio or synthesize it in a separate stage, limiting reciprocal modeling of visual dynamics and acoustic events. We present DreamX-Creator 1.0, a compact native joint audio-video generation system centered on a 7B generator. Conditioned on a first frame and a text prompt, the generator jointly denoises modality-specialized audio and video streams. The streams are