San Juan Riparian Watch
Ai2 OlmoEarth · an independent field evaluation · San Juan River basin

Finding the invaders, and knowing when not to trust the map

Tamarisk and Russian olive have taken over Western riparian corridors. A green map is easy; the honest question is how much of the corridor is the wrong plant, where, and how sure can we be.

This maps riparian woody vegetation and its invasive share from Landsat / Sentinel-2 time-series and puts a Random Forest head-to-head with Ai2's OlmoEarth foundation model — a real, leave-one-reach-out field test of where the FM earns its keep and where it ties. It also shows the tests that made two tempting results fall apart. And the grounded agent that answers your questions at the end runs on Ai2's OLMo: OlmoEarth sees, OLMo speaks.

23%
of the mapped corridor is invasive (Farmington, 2020)
0.557→0.889
RF→OlmoEarth AUC on the held-out arroyo
1984→
Landsat archive depth we probed — and where it breaks
2
tempting results a pre-registered test killed
OlmoEarth · Ai2 EO foundation model OLMo · Ai2 open LLM (the agent) STAC · Microsoft Planetary Computer Sentinel-2 10 m · Landsat 30 m, 1984→ Random Forest · 144-D spectro-temporal PostGIS medallion ETL · hybrid RAG
Ai2 stack

Built end-to-end on the Allen Institute for AI's open models: OlmoEarth (Earth-observation foundation model) delineates the vegetation, and OLMo (open LLM) grounds every answer in the documents. An independent test and an application of Ai2's models — not a benchmark they ran themselves.

00 · What's actually new here

Extent is solved. The contribution is the honest invasive layer.

Basin-scale riparian extent was already mapped for this watershed (CO-RIP, Woodward 2018, κ 0.80). Reproducing it is the control, not the result. What doesn't exist yet is a wall-to-wall, native-vs-invasive product that is honest about its own uncertainty — which species, how much, and over what stretch of the record the answer is even trustworthy. That is the target, and this page is the evidence for how far it currently reaches.

The discipline that runs through all of it Every headline number here is paired with the test that could have falsified it. Two did land as negatives — a deep-time growth curve and a beetle-driven signal flip — and they are kept in, because a result you can't break isn't a result.
01 · Method

A bounding box becomes a phenology cube, then a decision

The signal that separates tamarisk from native cottonwood is phenological — they leaf out and senesce on different schedules, with a distinct short-wave-infrared water signature. So a single image is the wrong input. Each reach becomes a 12-month median-mosaic cube, identically composited, and both models see the full spectro-temporal vector — not a greenness index.

Ingest
STAC time-series
Sentinel-2 / Landsat via Planetary Computer
Composite
12-month cube
median mosaics, phenology-aligned
Label
NMRipMap
native vs invasive truth (κ-grade)
Model
RF vs OlmoEarth
144-D pixel · vs fine-tuned FM
Product
Corridor + invasive
GeoJSON, corridor-masked
The pipeline. Identical compositing is the experiment — a mis-composited single-scene shortcut once collapsed a cross-reach transfer to AUC 0.37.

Why not just NDVI?

Because for this problem, greenness is nearly blind. Trained to separate invasive from native on the NDVI index alone, AUC ≈ 0.50 — a coin flip. The separability lives in the SWIR / red-edge bands across the season: tamarisk holds green later into the fall than cottonwood.

late-season divergence JanJulDec tamarisk cottonwood
Schematic of the discriminating signal. Plain NDVI over one date can't see it (AUC ≈ 0.50); the 12-month SWIR/red-edge trajectory can. Curves illustrative.
02 · The product

23% of the mapped corridor is the wrong plant

Green is the predicted riparian-woody corridor (7.6 km²); red is the invasive tamarisk / Russian-olive share within it (1.7 km²). Toggle either layer. This is the real prediction, on satellite imagery.

loading…
San Juan River at Farmington · 2020 · Landsat, NMRipMap-labelled. Imagery © Esri.
In-sample calibration — not independent validation The 23% masked prediction reproduces the NMRipMap label proportion (15,656 / 67,625 woody pixels invasive = 23%): well-calibrated to the 2020 labels it trained on. A spatially-held-out test is the honest next step, and it's named as such rather than dressed up as validation.
03 · The OlmoEarth field test — Random Forest vs Ai2's foundation model

On rivers they tie. One arroyo decides it.

The real test isn't in-sample accuracy — it's leave-one-reach-out transfer: train on some reaches, predict a reach the model has never seen. A per-pixel Random Forest, pooled across four morphologically diverse reaches, transfers well to unseen river corridors. But it has exactly one hole — the lone arroyo, the only example of its kind. That gap is precisely OlmoEarth's predicted opening: spatial context on under-represented morphology. This is the honest field test of whether the FM earns it.

Bars = Random Forest transfer AUC (0.5 = random, at the left edge). The green tick on the arroyo is the fine-tuned foundation model: 0.557 → 0.889.
OlmoEarth eval card. Task: riparian extent. Protocol: leave-one-reach-out transfer over 4 morphologically diverse reaches, scored once on the held-out reach (shared NMRipMap labels). Result: RF and OlmoEarth tie on the common morphology (rivers, mean 0.88); OlmoEarth wins the under-represented one — the arroyo, its single predicted opening — 0.557 → 0.889. Model: OlmoEarth (Ai2) fine-tuned on the pooled reaches. Reproducible: the LORO result · the pre-registered eval contract.

All three on the ground, over the labeled reach: the foundation model (green) tracks the corridor; NMRipMap truth (outline) is reality; and the Random Forest (orange), transferred to this arroyo, fires only in the wetter eastern end — the part whose spectra resemble the river reaches it trained on — and misses the western dry wash that FM and truth both capture. That's genuine under-prediction (~0.1% of pixels; verified against a fully-covered cube, not a data gap), and it's the 0.557-vs-0.889 gap made spatial: to the eye, the FM is simply more accurate. Both model layers are full-extent leave-one-reach-out predictions.

loading…
RF vs FM over the held-out Malpais arroyo — both full-extent leave-one-reach-out predictions, viewed over the NMRipMap-labeled reach. The RF fires only in the wetter east and misses the western wash; the FM tracks the whole corridor — the 0.557-vs-0.889 gap. (An earlier version used a partial RF export and was withdrawn; this uses the regenerated full-extent RF, verified against a fully-covered cube — the catch.) Imagery © Esri.
Why AUC, not accuracy Malpais is 82% invasive; a model that guessed "invasive everywhere" would score 82% accuracy and be useless. Prevalence swings 47%→82% across reaches, so the transfer metric is AUC (ranking quality, prevalence-invariant), reported alongside F1 at the deployment threshold — never accuracy. For plain extent, RF is as good and needs no GPU; the foundation model earns its keep only on the under-represented morphology.
04 · A negative, kept in

The beetle that didn't break the classifier

We predicted a today-trained tamarisk classifier would invert before the tamarisk beetle arrived (~2004–07), when the greenness signal flips. It didn't. And a pre-registered negative control — Russian olive, which the beetle doesn't touch — moved seven times more than the tamarisk "signal," proving the cross-era noise is too large to resolve a beetle effect at all.

taxon vs native (AUROC)202020152000 pre-beetleΔ
tamarisk (signal)0.8490.8130.862+0.013
Russian olive (control)0.8910.7670.553−0.338

Declaring the control in advance is what let it veto the result instead of us over-reading a 0.013 wiggle. The honest read: a beetle-robust discriminator, on a sample too small and too cross-sensor-confounded to say more.

05 · A negative, kept in

The "5× growth" that was an artifact

The first pass back through the Landsat archive looked like a clean ~5× rise in invasive extent since 1990. Every robustness check dissolved it. The pre-2000 numbers swing with the compositing recipe — no normalization fixes it — because 1990 is pure Landsat-5 TM, spectrally too distant from the training year to reconcile.

0% 2% 4% 6% recipe-dependent → no claim 1990 2000 2010 2020 stable: 2.7% → 3.5%
Three compositing recipes (dashed = single-year & 5-yr window; solid blue = relative radiometric normalization). Before 2000 they disagree by percentage points; 2000–2020 they converge — and the real rise there is ~2.7% → 3.5%, scarcely larger than the ±0.6 pp method noise. So: no pre-2000 trajectory claim — only the present-day product.

See it directly. Drag the year: the map shows the model's actual invasive prediction for each epoch, over the same corridor (green reference). Watch the flag in the corner — from ~2000 on the estimate is stable; slide to 1990 and it turns amber, because that pure Landsat-5 map moves with the compositing recipe. This isn't a growth animation — it's a demonstration of where the record stops being trustworthy.

2020present-day · calibrated
1990200020102020
Invasive-woody prediction over the Farmington corridor, per composite epoch. Green = the 2020 riparian corridor reference; red = invasive prediction for the selected year. Imagery © Esri.
06 · Why you can trust the parts that stand

The method is the deliverable

What separates this from a demo is that the failure modes were hunted, not hidden. Four practices do the work:

Spatial CV
Leave-one-reach-out, not random pixels — the only split that measures real deployment transfer across morphology.
Pre-registered controls
The Russian-olive control was declared before the run, so it could veto the beetle result rather than be quietly dropped.
Recipe-robustness
Every deep-time number was recomputed under multiple compositing recipes; the ones that moved were retired, not reported.
Prior-art falsification
The novelty claim is actively attacked against the literature — CO-RIP already owned extent, so extent is framed as the control, not the win.
Calibrated language, on purpose You'll notice "in-sample calibration," "documented negative," "cannot resolve" — not "validated," "proven," "breakthrough." The claims are sized to the evidence. That restraint is the credibility.
07 · The other method — how the work itself stays honest

AI-assisted research that catches its own errors

This was built fast, with heavy AI assistance — which is exactly where research quietly goes wrong. Hallucinated code is the easy failure; a compiler catches it. The dangerous ones are semantic: they compile, pass the tests, and read beautifully.

AI-assisted research fails by producing work that is fast, fluent, plausible, and wrong — and the wrongness is invisible precisely because the output looks finished. Exhortation doesn't fix this. Gates do.
Failure modeCaught by
Hallucinated API callthe compiler
Wrong logica unit test
A retracted result still published as factnothing
A model scored against 45%-wrong labelsnothing
A novelty claim already falsified by a 2018 papernothing

So honesty here is mechanized, not exhorted. Retired numbers, withdrawn claims and orphaned docs are caught by drift gates that run identically on a laptop and in CI. The finding, stated plainly in the method write-up: every rule that was merely documented eventually drifted; every rule that was mechanized held.

The most honest receipt The project's own merge gate was theatre for 25 of its first 29 merged PRs — a skipped AI review still posts a green check, and "no findings" is indistinguishable from "no review." It was caught only because a reviewer asked a question that couldn't be answered without checking. That failure is kept in the record, not scrubbed — which is the whole point.

How it's built

The same discipline runs through the stack — a reproducible EO pipeline, a typed spatial store, two model tracks, a services API, and the grounded agent, each behind its own gate.

Ingest
STAC ETL · Planetary Computer
Store
PostGIS medallion · bronze→gold
Model
Random Forest · OlmoEarth (Ai2)
Serve
.NET API · MVT vector tiles
Map + Ask
MapLibre · hybrid RAG on OLMo

The agent below is part of the same discipline: no source, no claim. Grounding is enforced backend-side, so every sentence it returns is pinned to a retrieved document you can open.

08 · Interrogate it

Ask the agent — it answers from the documents, and cites them

The findings above, the methods, and the underlying watershed literature are indexed in a hybrid-retrieval RAG system (dense + BM25). Ask anything: the answer is generated by OLMo, Ai2's open language model, grounded in the retrieved sources with clickable citations, and a mention of a reach flies the map to it. No source, no claim — so the same Ai2 stack that maps the vegetation also explains it.

Riparian document agent
Hybrid RAG over the corpus + these findings
offline
Ask about the corridor, the invasives, the beetle test, or the RF-vs-foundation-model decision. Tap a question below, or type your own — answers cite their sources, and a reach mention flies the map.
Connecting to the live agent…