On March 12, Google Research publicly unveiled Groundsource, a scalable data extraction framework that leverages its flagship large language model Gemini. The system automatically processes massive amounts of unstructured global news articles and converts them into structured historical disaster records. The first open-source dataset released under this framework contains 2.6 million urban flash flood events spanning over 150 countries.
Why Groundsource Matters: Filling the Data Gap
Natural disasters cause hundreds of millions of casualties and tens of billions in economic losses each year. To advance climate research, build accurate hydrological models, and issue timely warnings, scientists need robust historical baseline data. Yet such data is often scarce and scattered. Groundsource addresses this bottleneck by using Gemini's natural language processing to extract verified ground-truth data from news reports and online sources. The framework requires no manual labeling and can be applied at global scale.
First Open Dataset: 2.6 Million Urban Flash Flood Events
The initial release from Groundsource focuses on urban flash floods. The dataset covers 150+ countries with a total of 2.6 million historical flood events, all fully open-sourced. Google says the dataset provides a high-quality source for urban planning, insurance risk assessment, and emergency response. Unlike official statistics, these records were automatically extracted and deduplicated by Gemini from news articles, offering unprecedented coverage and granularity.
Scaling to Earthquakes and Wildfires
Groundsource is not limited to floods. Google's team emphasizes that the methodology is highly extensible, with potential applications to earthquakes, wildfires, and other natural hazards. As extreme weather becomes more frequent, AI-powered extraction of historical disaster footprints from news could become a critical component of global climate resilience efforts.

