San Juan Riparian Watch
Ai2 OlmoEarth · an independent field evaluation · San Juan River basin

How healthy is the river's forest, and which way is it heading?

A Western river's health shows in its corridor: how much native cottonwood and willow still holds the banks, and how much ground the invasives have taken. This maps both across four decades of satellite imagery and sets up to watch them year over year, so the signs that matter (native forest shrinking, invasives spreading) surface early enough to act.

The question underneath it all is whether a river corridor is healthy, and whether it's holding steady or slipping. This works that out from decades of Landsat and Sentinel-2 imagery: mapping the native riparian forest and the invasive share, then tracking how each shifts over time. Along the way it puts a Random Forest up against Ai2's OlmoEarth foundation model in a real field test, on ground neither model has seen, to find where the foundation model genuinely earns its keep and where the two simply tie. It's just as candid about what didn't hold up: two tempting results that came apart under scrutiny, kept in rather than quietly dropped. And the agent that answers your own questions at the end runs on Ai2's OLMo. OlmoEarth sees; OLMo speaks.

23%
of the mapped corridor is invasive (Farmington, 2020)
0.557→0.889
RF→OlmoEarth AUC on the held-out arroyo
1984→
Landsat archive depth we probed, and where it breaks
2
tempting results a pre-registered test killed
OlmoEarth · Ai2 EO foundation model OLMo · Ai2 open LLM (the agent) STAC · Microsoft Planetary Computer Sentinel-2 10 m · Landsat 30 m, 1984→ Random Forest · 144-D spectro-temporal PostGIS medallion ETL · hybrid RAG
Ai2 stack

Built end-to-end on the Allen Institute for AI's open models: OlmoEarth (Earth-observation foundation model) delineates the vegetation, and OLMo (open LLM) grounds every answer in the documents. An independent test and an application of Ai2's models, not a benchmark they ran themselves.

· Why this matters

In the arid Southwest, the whole ecosystem hangs on a green ribbon

Across the desert Southwest, life crowds along the rivers. The riparian corridor (the galleries of cottonwood and willow that line the water) is a thin sliver of the land, yet a disproportionate share of the region's birds, mammals, amphibians and pollinators depend on it for water, shade, food, and safe passage between habitats. In a landscape this dry, the river's edge is the habitat, and its condition is a fair proxy for the health of everything around it.

That ribbon is under pressure. Tamarisk (saltcedar) and Russian olive, brought over from Eurasia, have spread through Western river systems and pushed out the native cottonwood-willow forest. They grow in dense monocultures, load the soil with salt, change how fire moves through the bottomlands, and leave behind poorer habitat than what they replaced. For anyone managing these lands, the first question is simply where the corridor has turned invasive, and how quickly. That's exactly what's hard to see from the ground.

There was a hopeful intervention. In the 2000s the tamarisk leaf beetle was released across the West as a biological control, meant to defoliate saltcedar and give the natives room to recover. Whether it has durably rolled the invasion back is genuinely unsettled: the beetle spread, but tamarisk often grows back, and the net effect has been hard to pin down. This project's own attempt to find that effect in the satellite record came up empty, which is a result we keep in rather than bury. That kind of uncertainty is precisely what long-term monitoring exists to resolve.

The long game The aim isn't one reach or one year. It's to map riparian extent and its invasive share across the entire San Juan watershed, reconstruct how both have shifted over the four-decade Landsat record, and then watch for change year over year: a living map that flags where the corridor is gaining ground, losing it, or quietly turning over. What follows is the working proof that the pieces hold together.
00 · What's actually new here

Extent is solved. The contribution is the honest invasive layer.

Someone has already mapped where the riparian woods are in this watershed, and mapped them well (CO-RIP, Woodward 2018, κ 0.80). So reproducing that isn't the achievement; it's the control. What doesn't exist yet is a wall-to-wall native-vs-invasive map that's honest about its own uncertainty: which species, how much, and over which stretches of the record the answer can actually be trusted. That's the target. What follows is the evidence for how close it currently gets.

The discipline that runs through all of it Every headline number here comes paired with the test that could have proven it wrong. Two of those tests came back negative, a deep-time growth curve and a beetle-driven signal flip, and both stayed in, because a result you can't break isn't really a result.
01 · Method

A bounding box becomes a phenology cube, then a decision

What separates tamarisk from native cottonwood isn't how they look on any single day, it's their timing. The two leaf out and drop their leaves on different schedules, and hold water differently in a way the shortwave-infrared bands can see. So one snapshot is the wrong input. Instead, each reach becomes a 12-month cube of median mosaics, all composited the same way, and both models get to read the whole year's rhythm rather than a single greenness number.

Ingest
STAC time-series
Sentinel-2 / Landsat via Planetary Computer
Composite
12-month cube
median mosaics, phenology-aligned
Label
NMRipMap
native vs invasive truth (κ-grade)
Model
RF vs OlmoEarth
144-D pixel · vs fine-tuned FM
Product
Corridor + invasive
GeoJSON, corridor-masked
The pipeline. Identical compositing is the experiment, a mis-composited single-scene shortcut once collapsed a cross-reach transfer to AUC 0.37.

Why not just NDVI?

Because for this particular problem, greenness is almost blind. Ask a model to tell invasive from native using the NDVI greenness index alone, and it scores AUC ≈ 0.50, a coin flip. The real signal lives elsewhere: in the shortwave-infrared and red-edge bands, followed across the season. Tamarisk simply stays green later into the fall than cottonwood does.

late-season divergence JanJulDec tamarisk cottonwood
Schematic of the discriminating signal. Plain NDVI over one date can't see it (AUC ≈ 0.50); the 12-month SWIR/red-edge trajectory can. Curves illustrative.
02 · The product

Nearly a quarter of the corridor is invasive

Green is the predicted riparian-woody corridor (7.6 km²); red is the invasive tamarisk / Russian-olive share within it (1.7 km²). Toggle either layer. This is the real prediction, on satellite imagery.

loading…
San Juan River at Farmington · 2020 · Landsat, NMRipMap-labelled. Imagery © Esri.
In-sample calibration, not independent validation The 23% masked prediction reproduces the NMRipMap label proportion (15,656 / 67,625 woody pixels invasive = 23%): well-calibrated to the 2020 labels it trained on. A spatially-held-out test is the honest next step, and it's named as such rather than dressed up as validation.
03 · The OlmoEarth field test, Random Forest vs Ai2's foundation model

On rivers they tie. One arroyo decides it.

First, what's actually being compared: two very different ways to turn satellite pixels into a map.

Random Forest (RF)
The classic workhorse. It looks at each pixel on its own, reads its colour across the seasons, and votes with hundreds of little decision trees. Simple, fast, runs on any laptop, no GPU. A genuinely strong baseline, not a straw man.
OlmoEarth (the FM)
Ai2's Earth-observation foundation model: a large network pre-trained on enormous volumes of satellite imagery, then fine-tuned for this task. Unlike the RF, it reads each pixel in context, its neighbours, the shape of the corridor. More capable, but it needs a GPU.

Why put them head to head? Because "just use the fancy model" isn't a plan. A foundation model costs real money and hardware, so it's only worth deploying where it measurably beats the cheap baseline. The goal here isn't to crown a winner; it's to find the exact conditions under which OlmoEarth earns its keep, and to say so plainly where it doesn't. That's a real deployment decision, made on evidence instead of hype.

The test that matters isn't how well a model scores on data it trained on; it's whether it transfers to a place it has never seen. So we train on some reaches and predict a held-out one. A per-pixel Random Forest, pooled across four very different reaches, transfers well to unfamiliar river corridors. It has exactly one blind spot: the lone arroyo, the only landform of its kind in the set. And that's precisely where a foundation model should have the edge, it comes down to spatial context on a shape the training data barely covers. This is the honest test of whether OlmoEarth actually delivers there.

Bars = Random Forest transfer AUC (0.5 = random, at the left edge). The green tick on the arroyo is the fine-tuned foundation model: 0.557 → 0.889.
OlmoEarth eval card. Task: riparian extent. Protocol: leave-one-reach-out transfer over 4 morphologically diverse reaches, scored once on the held-out reach (shared NMRipMap labels). Result: RF and OlmoEarth tie on the common morphology (rivers, mean 0.88); OlmoEarth wins the under-represented one, the arroyo, its single predicted opening, 0.557 → 0.889. Model: OlmoEarth (Ai2) fine-tuned on the pooled reaches. Reproducible: the LORO result · the pre-registered eval contract.

Here are all three over the labeled reach. The foundation model (green) tracks the corridor; the NMRipMap truth (outline) is ground reality; and the Random Forest (orange), dropped into an arroyo it never trained on, fires only in the wetter eastern end, the stretch whose spectra resemble the river reaches it did learn, while missing the dry western wash that both the foundation model and the truth pick up. That's genuine under-prediction, not a data gap (~0.1% of pixels, checked against a fully-covered cube). It's the 0.557-vs-0.889 gap made visible: to the eye, the foundation model is simply the more accurate map. Both layers are full-extent, leave-one-reach-out predictions.

loading…
RF vs FM over the held-out Malpais arroyo, both full-extent leave-one-reach-out predictions, viewed over the NMRipMap-labeled reach. The RF fires only in the wetter east and misses the western wash; the FM tracks the whole corridor, the 0.557-vs-0.889 gap. (An earlier version used a partial RF export and was withdrawn; this uses the regenerated full-extent RF, verified against a fully-covered cube, the catch.) Imagery © Esri.
Why AUC, not accuracy Malpais is 82% invasive; a model that guessed "invasive everywhere" would score 82% accuracy and be useless. Prevalence swings 47%→82% across reaches, so the transfer metric is AUC (ranking quality, prevalence-invariant), reported alongside F1 at the deployment threshold, never accuracy. For plain extent, RF is as good and needs no GPU; the foundation model earns its keep only on the under-represented morphology.
04 · A negative, kept in

The beetle that didn't break the classifier

We expected a tamarisk classifier trained on today's imagery to flip once you ran it back past the tamarisk beetle's arrival (~2004–07), when the plant's greenness signal changes. It didn't budge. And the negative control we'd declared up front, Russian olive, which the beetle doesn't touch, moved seven times more than the tamarisk "signal" did. That's the tell: the cross-era noise is simply too large to resolve a beetle effect at all.

taxon vs native (AUROC)202020152000 pre-beetleΔ
tamarisk (signal)0.8490.8130.862+0.013
Russian olive (control)0.8910.7670.553−0.338

Declaring the control in advance is what let it veto the result instead of us over-reading a 0.013 wiggle. The honest read: a beetle-robust discriminator, on a sample too small and too cross-sensor-confounded to say more.

05 · A negative, kept in

The "5× growth" that was an artifact

The first pass back through the Landsat archive looked like a clean ~5× rise in invasive extent since 1990. Every robustness check dissolved it. Two things sink it. The pre-2000 numbers swing with the compositing recipe, no normalization fixes it, because 1990 is pure Landsat-5 TM, spectrally too far from the training year to reconcile. And once the invasive layer is properly gated to woody vegetation, even the face-value change is small: from under 1% to about 1.4% of the area, most of it within the method's own noise.

0% 1% 2% 3% recipe-dependent → no claim 1990 2000 2010 2020 stable: ~0.8% → 1.4%
The extent-gated invasive estimate per epoch (solid). Before 2000 it's recipe-dependent, the pure Landsat-5 TM composite shifts by percentage points with the compositing recipe (dashed spread, illustrative of the documented instability), so no pre-2000 claim. From 2000 on it's stable and small: ~0.8% → 1.4% of area, barely above the ±0.6 pp method noise. Only the present-day product is asserted.

See it directly. Drag the year: the map shows the model's actual invasive prediction for each epoch, over the same corridor (green reference). Watch the flag in the corner, from ~2000 on the estimate is stable; slide to 1990 and it turns amber, because that pure Landsat-5 map moves with the compositing recipe. This isn't a growth animation. It's a demonstration of where the record stops being trustworthy.

2020present-day · calibrated
1990200020102020
Invasive-woody prediction over the Farmington corridor, per composite epoch. Green = the 2020 riparian corridor reference; red = invasive prediction for the selected year. Imagery © Esri.
06 · Why you can trust the parts that stand

The method is the deliverable

What separates this from a demo is simple: the ways it could be wrong were hunted down, not hidden. Four habits do most of that work.

Spatial CV
Leave-one-reach-out, not random pixels, the only split that measures real deployment transfer across morphology.
Pre-registered controls
The Russian-olive control was declared before the run, so it could veto the beetle result rather than be quietly dropped.
Recipe-robustness
Every deep-time number was recomputed under multiple compositing recipes; the ones that moved were retired, not reported.
Prior-art falsification
The novelty claim is actively attacked against the literature, CO-RIP already owned extent, so extent is framed as the control, not the win.
Calibrated language, on purpose You'll notice "in-sample calibration," "documented negative," "cannot resolve", not "validated," "proven," "breakthrough." The claims are sized to the evidence. That restraint is the credibility.
07 · The other method, how the work itself stays honest

AI-assisted research that catches its own errors

This was built fast, with a lot of AI assistance, which is exactly where research quietly goes wrong. Hallucinated code is the easy kind of failure; a compiler catches it in seconds. The dangerous failures are semantic: they compile, they pass the tests, and they read beautifully, while being wrong.

AI-assisted research fails by producing work that is fast, fluent, plausible, and wrong, and the wrongness is invisible precisely because the output looks finished. Exhortation doesn't fix this. Gates do.
Failure modeCaught by
Hallucinated API callthe compiler
Wrong logica unit test
A retracted result still published as factnothing
A model scored against 45%-wrong labelsnothing
A novelty claim already falsified by a 2018 papernothing

So honesty here is mechanized, not exhorted. Retired numbers, withdrawn claims and orphaned docs are caught by drift gates that run identically on a laptop and in CI. The finding, stated plainly in the method write-up: every rule that was merely documented eventually drifted; every rule that was mechanized held.

The most honest receipt The project's own merge gate was theatre for 25 of its first 29 merged PRs, a skipped AI review still posts a green check, and "no findings" is indistinguishable from "no review." It was caught only because a reviewer asked a question that couldn't be answered without checking. That failure is kept in the record, not scrubbed, which is the whole point.

How it's built

The same discipline runs through the stack, a reproducible EO pipeline, a typed spatial store, two model tracks, a services API, and the grounded agent, each behind its own gate.

Ingest
STAC ETL · Planetary Computer
Store
PostGIS medallion · bronze→gold
Model
Random Forest · OlmoEarth (Ai2)
Serve
.NET API · MVT vector tiles
Map + Ask
MapLibre · hybrid RAG on OLMo

The agent below is part of the same discipline: no source, no claim. Grounding is enforced backend-side, so every sentence it returns is pinned to a retrieved document you can open.

08 · Interrogate it

Ask the agent, it answers from the documents, and cites them

Everything above, the findings, the methods, and the watershed literature behind them, is indexed in a hybrid-retrieval system (dense embeddings + BM25). Ask it anything. The answer is written by OLMo, Ai2's open language model, grounded in the retrieved sources with clickable citations, and if you mention a reach the map flies to it. No source, no claim, so the same Ai2 stack that maps the vegetation also explains it.

Riparian document agent
Hybrid RAG over the corpus + these findings
offline
Ask about the corridor, the invasives, the beetle test, or the RF-vs-foundation-model decision. Tap a question below, or type your own, answers cite their sources, and a reach mention flies the map.
Connecting to the live agent…