News

SDSC Helps Improve AI-Ready Geospatial Data Through GeoCroissant

Published September 29, 2026

By Kimberly Bruch, SDSC Comms and Varsha Balaji, Canyon Crest Academy Student and SDSC Comms Intern

An orange starfish rests on a delicate, branching white coral against a dark deep-sea background.

Researchers like Karen Stocks at the Scripps Institution of Oceanography can now better understand—thanks to GeoCroissant providing easy access to fellow scientists’ data—what lies in the deep ocean (below 200 meters), which covers 66% of Earth’s surface. Credit: Schmidt Ocean Institute

Researchers working with Earth science data often face a basic but time-consuming challenge: Determining what datasets exist, what they contain and whether they can be used together. Doug Fils, with the University of California San Diego Halıcıoğlu School of Data Science and Computing’s San Diego Supercomputer Center (SDSC), is helping advance a metadata approach intended to make that process easier.

Fils is contributing to GeoCroissant, an extension of Croissant, the open MLCommons metadata standard for machine learning datasets. Rather than changing the underlying data, Croissant provides a machine-readable description of a dataset such as its contents, structure, provenance and conditions for use. This allows people and software tools to more easily find, understand and reuse it. Croissant builds on schema.org, a widely used web vocabulary for describing structured information.

GeoCroissant adds information that is especially important for geospatial and Earth science data, including details about location, time and spatial characteristics. Fils’s work connects GeoCroissant with geosemantic information and GeoSPARQL, an Open Geospatial Consortium standard that supports the representation and querying of geospatial information.

“Our goal is to improve the discoverability, accessibility and interoperability of Earth science datasets,” said Fils, who is an information technology specialist with SDSC’s Research Data Services Division.

The effort addresses a challenge common in data-intensive research. Earth science data may be distributed across many repositories, documented in different ways and collected at different spatial and temporal scales. Those differences can make it difficult for researchers to locate suitable data, assess whether it fits a research question and prepare it for analysis or machine learning workflows.

“Applied to geospatial data, GeoCroissant can help streamline the early stages of data discovery and preparation,” Fils said. “Instead of spending extensive time searching across databases with inconsistent metadata, researchers can more efficiently identify datasets that contain the spatial information most relevant to their models.”

The work has potential applications in areas such as environmental monitoring, geological research, oceanography and natural-hazard analysis. Better dataset descriptions can help researchers determine where data were collected, when observations were made, what resolution was used and how the data may be combined with other sources.

One application involves work with Karen Stocks, an oceanographer and data scientist at UC San Diego’s Scripps Institution of Oceanography. Stocks’s work includes the documentation, discovery, access, integration and curation of oceanographic data.

Fils and Stocks are exploring how GeoCroissant can support the description and discovery of ocean-depth data.

“Even though depth is one of the most crucial factors in determining the ecosystem and environmental conditions of a particular point in the ocean, we are often missing information,” Stocks said. “By using standardized frameworks that focus on metadata about depth, we can discover relevant information to serve our models and predictions much more efficiently.”

Croissant was developed by MLCommons as a community-built standard for describing machine learning datasets and is designed to support dataset discovery, interoperability, documentation and reproducibility. The standard includes extensions for specialized domains, including geospatial data. 

The GeoCroissant initiative is led by Rajat Shinde and Manil Maskey of NASA and the University of Alabama in Huntsville, with contributions from researchers and organizations across academia, government and industry. Fils said the project’s collaborative standards-development process is essential to making the approach useful across diverse Earth science and geospatial research communities.

Fils said that by improving how geospatial datasets are described, GeoCroissant could help researchers spend less time interpreting fragmented data documentation and more time using data to investigate the planet’s complex systems.

Archive

Media Contact

Kimberly Mann Bruch
SDSC Communications