Paper recorded by Signals 4 on 2026-08-28 in cs.CV. Abstract reproduced from arXiv; link to the original below.
Published 2026-08-28 on arXiv · recorded by Signals 4 on 2026-08-31
Category: cs.CV · 计算机视觉 · first seen 2026-08-31
Conventionally, Building Damage Assessment (BDA) is tackled either with dedicated network architectures or by fine-tuning geospatial image foundation models. In this work, we ask whether a general-purpose Vision-Language Model (VLM) can localize buildings and grade their damage through autoregressive sequence generation alone. We cast BDA as predicting a variable-length set of bounding boxes, each