NARA Discovery
Article Details
← Back to Search Results
Journal Article

Assessing the spatial accuracy of geocoding flood-related imagery using Vision Language Models

Sebastian Schmidt; Eleonor Díaz Fragachan; Dorian Arifi; David Hanny; Bernd Resch
Spatial Information Research · Vol. 33, Issue 2 · 2025

Abstract

While the capabilities of large language models and visual language models for various classification tasks have advanced significantly, their potential for location inference remains largely underexplored. Therefore, this study evaluates the performance of four prominent models — BLIP-2, LLaVA1.6, OpenFlamingo, and GPT-4o — for geocoding flood-related images from Flickr. Model inferences are compared against the original photo locations and human-labelled assessments. Our findings reveal that GPT-4o achieves the highest spatial accuracy (median deviation of 89.12 km). OpenFlamingo geocodes the highest number of images (90.7%), albeit with fluctuating quality (median 408.35 km), while still outperforming the human annotators. LLaVA1.6 geocodes only 18.9% of all images, while BLIP-2 exhibits the highest median deviation (1,781 km). We observe a spatial bias in our results, with inferences being most accurate in Central Europe. Additionally, model results improve when images feature recognisable landmarks. The proposed workflow could significantly increase the amount of geocoded web-based data available for disaster management, though further research is required to enhance accuracy across diverse geographic contexts.

Bibliographic Information

JournalSpatial Information Research
PublisherSpringer
Publication Date2025-04-01
Publication Year2025
Volume33
Issue2
Document TypeJournal Article
Print ISSN2366-3286
eISSN2366-3294
DOI10.1007/s41324-025-00609-0

Access Information

NARA Access Coverage2016-01-01~Current
Journal Homepagehttps://www.springer.com/journal/41324
Publisher PageOpen Publisher Page
Full-text access depends on NARA's subscribed coverage and institutional access.