NASA and IBM Open-Source a Lunar Foundation Model, Dataset Included

·11 min read·Evergreen Tools Team

For decades, sensors and instruments have observed the Moon, accumulating petabytes of data. Yet to study the lunar surface, scientists have had to either sift through maps and images by hand or reach for low-resolution, task-specific machine learning models that are computationally intensive and may lack the scientific accuracy a study demands. On September 10, 2026, IBM and NASA announced the open-source release of the NASA-IBM Lunar Foundation Model, which IBM describes as one of the first publicly available foundation models built specifically for scientific exploration of the Moon, now available.

1. What Was Released

Start with the result. The announcement's headline number is that the model exceeds widely used methods by up to 23 percent in identifying key geographic features on the Moon's surface, including potential ice deposits, craters, and volcanic formations. The more interesting half is the dataset. IBM and NASA scientists built the first open-source lunar dataset of its kind: a unified, machine-learning-ready dataset aggregating over 30 spatially aligned layers from nine instruments across four missions, combining tens of thousands of images and maps from NASA's Lunar Reconnaissance Orbiter and GRAIL mission with complementary data from the Japanese Aerospace Exploration Agency's SELENE/Kaguya for a multimodal view of the lunar surface and subsurface.

# The model and dataset are open on Hugging Face, so the fastest way to
# understand what it can do is to load it and look.

from transformers import AutoModel, AutoImageProcessor
from huggingface_hub import hf_hub_download

REPO = "nasa-ibm-ai4science/NASA-IBM-Lunar-Foundation-Model"

processor = AutoImageProcessor.from_pretrained(REPO)
model = AutoModel.from_pretrained(REPO)
model.eval()

# The technical report ships with the model, which is the right place to
# read before you trust any headline number.
report = hf_hub_download(REPO, "NI_LFM_Technical_Report.pdf")
print("technical report:", report)
Scientific data visualisation

The model unifies decades of multi-instrument observations

2. The Dataset Is the Underrated Half

The dataset construction is the most underrated part of this release. IBM states it plainly: despite the wealth of lunar data, no publicly available, unified dataset existed that brought multimodal, multi-resolution data into a common framework suitable for modern machine learning. That means every research team previously had to spend significant effort on data engineering before it could do any science. That work is now standardised across 30-plus aligned layers, nine instruments, and four missions. Code sample 4 shows why that matters. Spatial alignment is the prerequisite for all multi-resolution fusion, and once the layers are aligned you can concatenate features across scales instead of fighting coordinate systems and resampling.

# Running the model on a lunar tile. The architecture is multimodal and
# multi-resolution on purpose: a crater needs metre-scale detail, while
# volcanic patches are best seen at context scale around 100 metres.

import torch
from PIL import Image

def embed(tile: Image.Image, resolution: str = "context"):
    inputs = processor(images=tile, return_tensors="pt")
    with torch.no_grad():
        out = model(**inputs)
    # Keep the feature map, not just the pooled vector: downstream tasks
    # are segmentation-style, not classification-style.
    return out.last_hidden_state

metro_tile = Image.open("lro_nac_sample.tif")
context_tile = Image.open("lro_wac_sample.tif")

f_metro = embed(metro_tile, resolution="metre")
f_context = embed(context_tile, resolution="context")

# Fuse the two views instead of picking one. That fusion is where the
# reported gains over single-resolution baselines come from.
fused = torch.cat([f_metro.mean(dim=1), f_context.mean(dim=1)], dim=-1)

3. Application One: Lunar Ice, Up to 22 Percent Lower Error

Three applications come with concrete numbers. The first is potential lunar ice deposits. Permanently shadowed regions are among the Moon's most difficult environments to observe, yet they may contain lunar ice below the surface, and lunar ice means water and oxygen, resources considered essential for a future Moon base and for producing rocket fuel for missions to Mars. The model combines multimodal and multi-resolution observations to predict where ice may be present. An IBM and NASA technical paper shows the model reduced error, measured as RMSE, in identifying high-potential lunar ice areas by up to 22 percent compared with the SwinV2-B model trained on ImageNet. Code sample 5 reproduces that kind of evaluation honestly, against the baseline the paper used rather than a convenient re-split.

# Fine-tuning on your own labels. The paper notes the model reaches
# comparable crater accuracy while being more efficient and cheaper to
# fine-tune than task-specific baselines, so start small.

from transformers import Trainer, TrainingArguments

training_args = TrainingArguments(
    output_dir="./lunar-ft",
    per_device_train_batch_size=4,
    learning_rate=1e-5,
    num_train_epochs=8,
    eval_strategy="steps",
    eval_steps=200,
    save_steps=400,
    # Half the training data was enough for the context-scale crater task,
    # so resist the instinct to throw everything at it.
    max_steps=2000,
)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=my_lunar_tiles,
    eval_dataset=my_holdout_tiles,
)
trainer.train()
Model evaluation results

Up to 22 percent error reduction on lunar ice potential

4. Applications Two and Three: Volcanism and Craters

The second application is volcanic history. Scientists study lunar volcanic features called Irregular Mare Patches to better understand the Moon's volcanic history and thermal evolution, and identifying these changing regions is also strategic for future surface operations. Using imperfect labels, the model better captures the extent of the volcanic features than SwinV2-B by 3 percent while reaching comparable accuracy with greater efficiency and lower fine-tuning costs. The third is crater detection. Craters are among the Moon's most important and distinguishing features, revealing terrain age, geology, and the chemistry of the early lunar interior, and crater mapping helps NASA select safe landing sites and avoid hazards such as steep slopes and boulders. At metre-scale resolution, accuracy is comparable to state-of-the-art models like SwinV2-B while offering greater efficiency and lower fine-tuning costs; at context-scale resolution around 100 metres, it outperforms SwinV2-B by nearly 19 percent using just half the training data.

# Multi-resolution fusion in practice: aggregate over several spatial
# layers rather than trusting one zoom level. The dataset already aligns
# the layers, which is the hard part you no longer have to do yourself.

def multi_resolution_features(tile_stack):
    """tile_stack: dict of resolution -> tensor, spatially aligned."""
    features = []
    for resolution in ("1m", "10m", "100m", "1km"):
        tensor = tile_stack.get(resolution)
        if tensor is None:
            continue
        features.append(model.encode(tensor))
    return torch.cat(features, dim=-1)

# Aligned multi-modal input is what lets the model connect an ice signature
# in one instrument with terrain in another. Feeding it a single blurred
# image throws that away.
tile_stack = load_aligned("crater_region_044")
features = multi_resolution_features(tile_stack)

5. Why Science Suits Foundation Models

Why do foundation models suit science so well? NASA's chief science data officer, Kevin Murphy, put his finger on it: NASA has spent decades building an extraordinary scientific record of the Moon, but collecting data is only part of the job, and the data also has to be easier for scientists to explore and use. Juan Bernabe-Moreno, Director of IBM Research Europe, UK and Ireland, adds that the model gives scientists a foundation to explore the Moon at scale, connecting observations across instruments and revealing patterns that are difficult to see in isolation, as an open platform the global research community can build on. The paradigm shift is clear: instead of building a new algorithmic system for every scientific question, researchers start from a shared model and adapt it to new tasks.

# Evaluate honestly. The headline comparisons use RMSE for ice potential
# and accuracy or IoU for craters and volcanic patches. Reproduce the
# metric on your own holdout before you report a number.

import numpy as np

def rmse(pred, truth):
    return float(np.sqrt(np.mean((np.asarray(pred) - np.asarray(truth)) ** 2)))

def iou(pred_mask, truth_mask):
    pred_mask = pred_mask.astype(bool)
    truth_mask = truth_mask.astype(bool)
    inter = (pred_mask & truth_mask).sum()
    union = (pred_mask | truth_mask).sum()
    return float(inter / union) if union else 0.0

# Compare against the baseline the paper used, SwinV2-B trained on ImageNet,
# otherwise your "improvement" may just be a different evaluation split.
print("rmse:", rmse(pred_ice, truth_ice))
print("iou :", iou(pred_crater, truth_crater))
A research computing environment

Both model and dataset are open on Hugging Face

6. Getting Started, and One Thing to Watch

To get hands on, work in this order. First, load the model and the technical report from Hugging Face and read the report before trusting any headline number, which is where code sample 1 starts. Second, run inference and understand that the outputs are segmentation-style feature maps rather than classification vectors, as code sample 2 shows when it fuses multi-resolution features. Third, fine-tune on your own labels with code sample 3, and heed its advice: half the training data was enough for the context-scale crater task, so resist the instinct to throw everything at it. Fourth, compare against the baseline the paper used, with code sample 5 providing RMSE and IoU implementations, because an improvement manufactured by a different evaluation split is not an improvement. Fifth, note that this model belongs to the Prithvi family, which spans geospatial, weather, and heliophysics and now the Moon, so if your question crosses Earth science domains it is worth following that thread.

📌 Frequently Asked Questions

When was the NASA-IBM Lunar Foundation Model released?

IBM and NASA announced the open-source release on September 10, 2026, describing it as one of the first publicly available foundation models built specifically for scientific exploration of the Moon.

How much better is it at identifying surface features?

IBM says the model exceeds widely used methods by up to 23 percent in identifying key geographic features on the Moon's surface, including potential ice deposits, craters, and volcanic formations.

What does the accompanying dataset contain?

It is the first open-source lunar dataset of its kind, aggregating over 30 spatially aligned layers from nine instruments across four missions, combining NASA's LRO and GRAIL data with complementary JAXA SELENE/Kaguya data.

How does it perform on lunar ice detection specifically?

An IBM and NASA technical paper shows the model reduced error as measured by RMSE by up to 22 percent compared with the SwinV2-B ImageNet model when identifying areas with high potential for lunar ice.

Where can I get the model and dataset?

The model and its technical report are published on Hugging Face in the nasa-ibm-ai4science/NASA-IBM-Lunar-Foundation-Model repository, as part of the Prithvi family of open foundation models.