San Juan Riparian Watch
An independent field evaluation · San Juan River basin · not affiliated with Ai2

How healthy is the river's forest, and which way is it heading?

In the desert Southwest, most of a river's life crowds into the thin green ribbon at the water's edge. Two invaders from Eurasia, tamarisk and Russian olive, have spent decades taking that ribbon over, and from the ground you can't see how far they've gotten.

From orbit you can, if you read time instead of color. On any single day an invaded stretch looks exactly as green as a healthy one. But tamarisk holds its green for weeks after the native trees brown out, and that timing signature is written into forty years of satellite imagery. I built a system that reads it: where the river forest is, how much of it has turned invasive, and which way both are heading.

It's also a field test with real stakes. Every map here is made two ways, once by a classic Random Forest and once by Ai2's OlmoEarth foundation model, to find out where the bigger model actually earns its cost on stretches of river neither has seen. And the whole thing shows its work. Two tempting headline results failed their own tests and stayed on the page as negatives instead of quietly disappearing. Ask the agent at the bottom about any of it. It answers from the project's own documents and cites them.

23%
of the mapped corridor is invasive (Farmington, 2020)
0.89
OlmoEarth on the held-out reach where the Random Forest collapses to 0.56
1984→
how far back the satellite record reaches, and where it stops being trustworthy
2
headline results killed by their own tests, and kept on the page
OlmoEarth · Ai2 EO foundation model OLMo · Ai2 open LLM (the agent) STAC · Microsoft Planetary Computer Sentinel-2 10 m · Landsat 30 m, 1984→ Random Forest · 144-D spectro-temporal PostGIS medallion ETL · hybrid RAG
Independent

This is an independent project, not produced, sponsored, or endorsed by the Allen Institute for AI (Ai2). It is built on Ai2's open models under their open licenses: OlmoEarth maps the vegetation, and OLMo is the model the agent is built for. Their names appear here only to identify the models being evaluated.

Proof of concept

Experimental, not verified data. Every map and number on this site, from both models, is a research prototype that has not been independently validated. The work is done with real scientific rigor, and each figure says how it was measured. Still: treat every output as provisional and possibly wrong, not as fact.

Why this matters

In the arid Southwest, the whole ecosystem hangs on a green ribbon

Across the desert Southwest, life crowds along the rivers. The riparian corridor (the galleries of cottonwood and willow that line the water) is a thin sliver of the land, yet a disproportionate share of the region's birds, mammals, amphibians and pollinators depend on it for water, shade, food, and safe passage between habitats. In a landscape this dry, the river's edge is the habitat, and its condition is a fair proxy for the health of everything around it.

That ribbon is under pressure. Tamarisk (saltcedar) and Russian olive, brought over from Eurasia, have spread through Western river systems and pushed out the native cottonwood-willow forest. They grow in dense monocultures, load the soil with salt, change how fire moves through the bottomlands, and leave behind poorer habitat than what they replaced. For anyone managing these lands, the first question is simply where the corridor has turned invasive, and how quickly. That's exactly what's hard to see from the ground.

There was a hopeful intervention. In the 2000s the tamarisk leaf beetle was released across the West as a biological control, meant to defoliate saltcedar and give the natives room to recover. Whether it has durably rolled the invasion back is genuinely unsettled: the beetle spread, but tamarisk often grows back, and the net effect has been hard to pin down. My own attempt to find that effect in the satellite record came up empty. That result stays on this page, because that kind of uncertainty is exactly what long-term monitoring exists to resolve.

The long game The aim isn't one reach or one year. It's to map riparian extent and its invasive share across the entire San Juan watershed, reconstruct how both have shifted over the four-decade Landsat record, and then watch for change year over year: a living map that flags where the corridor is gaining ground, losing it, or quietly turning over. What follows is the working proof that the pieces hold together.
What's actually new here

Extent is solved. The contribution is the honest invasive layer.

Someone has already mapped where the riparian woods are in this watershed, and mapped them well (CO-RIP, Woodward 2018, κ 0.80). So reproducing that isn't the achievement; it's the control. What doesn't exist yet is a wall-to-wall native-vs-invasive map that's honest about its own uncertainty: which species, how much, and over which stretches of the record the answer can actually be trusted. That's the target. What follows is the evidence for how close it currently gets.

The discipline that runs through all of it Every headline number here comes paired with the test that could have proven it wrong. Two of those tests came back negative, a deep-time growth curve and a beetle-driven signal flip, and both stayed in, because a result you can't break isn't really a result.
Method

The trick is timing: read the whole year, not one photo

What separates tamarisk from native cottonwood isn't how they look on any single day. It's their timing. The two species leaf out and shed on different schedules, and they hold water differently in a way the satellite's infrared bands can see. Ecologists call this seasonal rhythm phenology, and it's the fingerprint everything here runs on. So instead of one photo, each stretch of river becomes a stack of twelve monthly images, one cloud-free composite per month, built the identical way every time. Both models read the year's whole rhythm instead of a single greenness number.

Ingest
STAC time-series
Sentinel-2 / Landsat via Planetary Computer
Composite
12-month cube
median mosaics, phenology-aligned
Label
NMRipMap
native vs invasive truth (κ-grade)
Model
RF vs OlmoEarth
144-D pixel · vs fine-tuned FM
Product
Corridor + invasive
GeoJSON, corridor-masked
The pipeline. Building every monthly stack the identical way is not a detail, it is the experiment: a shortcut that used single scenes once dropped a model's score on an unseen reach from 0.80 to 0.37.

Why not just measure greenness?

Because for this particular problem, greenness is almost blind. Ask a model to tell invasive from native using the standard greenness index (NDVI) alone and it scores AUC ≈ 0.50. AUC is the score used throughout this page: 1.0 is a perfect classifier, 0.5 is a coin flip. So: a coin flip. The real signal lives elsewhere, in the infrared bands followed across the season. Tamarisk simply stays green later into the fall than cottonwood does.

late-season divergence JanJulDec tamarisk cottonwood
Schematic of the discriminating signal. Plain NDVI over one date can't see it (AUC ≈ 0.50); the 12-month SWIR/red-edge trajectory can. Curves illustrative.
The product

Nearly a quarter of the corridor is invasive

Here is what that instrument sees over Farmington. Green is the predicted river-forest corridor (7.6 km²). Red is the share of it the model calls invasive (1.7 km²). Toggle either layer. This is the real prediction, drawn on satellite imagery.

loading…
San Juan River at Farmington · 2020 · Landsat, NMRipMap-labelled. Imagery © Esri.
A model agreeing with its teacher, not passing an exam The 23% reproduces the proportion in the expert 2020 labels the model trained on (15,656 of 67,625 woody pixels). In other words it's well calibrated, not independently validated. The honest next step is a test on ground the model never saw, and I name it as that instead of dressing this up as validation.
Ask the agent
The OlmoEarth field test, Random Forest vs Ai2's foundation model

On familiar ground they tie. One unfamiliar reach decides it.

First, what's actually being compared: two very different ways to turn satellite pixels into a map.

Random Forest (RF)
The classic workhorse. It looks at each pixel on its own, reads its colour across the seasons, and votes with hundreds of little decision trees. Simple, fast, runs on any laptop, no GPU. A genuinely strong baseline, not a straw man.
OlmoEarth (the FM)
Ai2's Earth-observation foundation model: a large network pre-trained on enormous volumes of satellite imagery, then fine-tuned for this task. Unlike the RF, it reads each pixel in context, its neighbours, the shape of the corridor. More capable, but it needs a GPU.

Why put them head to head? Because "just use the fancy model" isn't a plan. A foundation model costs real money and hardware, so it's only worth deploying where it measurably beats the cheap baseline. The goal here isn't to crown a winner; it's to find the exact conditions under which OlmoEarth earns its keep, and to say so plainly where it doesn't. That's a real deployment decision, made on evidence instead of hype.

The test that matters isn't how well a model scores on data it trained on. It's whether the model survives a place it has never seen. So I train both models on three stretches of river and make them predict a held-out fourth, taking turns until every reach has been the unseen one. On reaches that resemble the training set, the cheap Random Forest does fine (0.85 to 0.90) and the foundation model merely matches it. But on Malpais, the one reach whose landscape is unlike the rest (a valley of irrigated fields and dry upland), the Random Forest collapses to a coin flip at 0.56 while OlmoEarth holds at 0.89. That is the whole result: the foundation model doesn't see better, it breaks less on unfamiliar ground. And an unlabeled basin is almost entirely unfamiliar ground.

Bars = Random Forest transfer AUC (0.5 = random, at the left edge). The green tick on Malpais is the fine-tuned foundation model: 0.557 → 0.889 on the one unfamiliar reach.
OlmoEarth eval card. Task: riparian extent. Protocol: leave-one-reach-out transfer over 4 morphologically diverse reaches, scored once on the held-out reach (shared NMRipMap labels). Result: on the three familiar reaches the two are a wash (mean ~0.88; OlmoEarth even trails slightly on Kirtland, 0.812 vs 0.845). OlmoEarth wins the one unfamiliar reach, Malpais, 0.557 → 0.889. Its whole edge is robustness on ground unlike its training data; a trade you buy for the unlabeled basin, not a free win. Model: OlmoEarth (Ai2) fine-tuned on the pooled reaches. Reproducible: the LORO result · the pre-registered eval contract.
This result survived its own audit An earlier version of this page called Malpais a "desert arroyo" and credited the win to that terrain. A spatial audit showed the reach is actually a river-dominated subwatershed of the San Juan valley, so the attribution to arroyo morphology is unverified and that framing was retracted. The numbers themselves held up: the Random-Forest scores reproduced on a re-run (0.90 / 0.85 / 0.89 / 0.56), and both models were scored on the same held-out pixels. It is also not a tamarisk effect; the Random Forest is near-chance on native and invasive vegetation alike (0.59 vs 0.53). What you just read is the corrected explanation. The full trail: the post-mortem · the retraction note.

Here are all three on the same reach. The expert labels are the outline. The foundation model (green) tracks the whole corridor. The Random Forest (orange), dropped onto ground it never trained on, fires in one small eastern pocket, the stretch that most resembles the reaches it learned, and misses everything else. That is the coin-flip collapse made visible. To the eye, the foundation model is simply the more accurate map.

loading…
Both models over the held-out Malpais reach (a San-Juan-valley subwatershed, not a desert arroyo), each predicting ground it never trained on. The Random Forest fires in one small eastern pocket and misses the rest; OlmoEarth tracks the whole corridor. This is the 0.557-versus-0.889 gap, drawn on the map. Imagery © Esri. ⚠ Experimental model output, provisional, not validated ground truth.
Why AUC, not accuracy Malpais is 82% invasive, so a model that just guessed "invasive everywhere" would score 82% accuracy and be useless. The invasive share also swings from 47% to 82% across reaches. So the score here is AUC, which measures how well a model ranks its answers and gives no credit for guessing the majority, reported alongside F1 at the deployment threshold, never accuracy. For plain extent mapping, the Random Forest is as good and needs no GPU; the foundation model earns its keep only on the unfamiliar reach.
Ask the agent
A negative, kept in

The beetle that didn't break the classifier

I expected a tamarisk classifier trained on today's imagery to flip once I ran it back past the beetle's arrival around 2004, because the beetle changes how green the plant looks. It didn't budge. Better: before the run I declared a negative control, Russian olive, a plant the beetle doesn't touch. If the record were clean, its score should have stayed put. Instead it moved seven times more than the tamarisk "signal" did. That's the tell: the noise between eras is simply too large to measure a beetle effect at all.

taxon vs native (AUROC)202020152000 pre-beetleΔ
tamarisk (signal)0.8490.8130.862+0.013
Russian olive (control)0.8910.7670.553−0.338

Declaring the control in advance is what let it veto the result, instead of me over-reading a 0.013 wiggle. The honest read: the classifier looks beetle-proof, but the sample is too small, and too tangled across satellite generations, to say more.

A negative, kept in

The "5× growth" that was an artifact

The first pass back through the Landsat archive looked like a clean ~5× rise in invasive cover since 1990. Every robustness check dissolved it. Two things sink it. First, the pre-2000 numbers change depending on how the imagery is assembled, and no correction fixes that, because 1990 comes from an older satellite whose sensor is too different from the training year to reconcile. Second, once the invasive layer is limited to woody vegetation, even the face-value change is small: from under 1% to about 1.4% of the area, most of it within the method's own noise.

0% 1% 2% 3% recipe-dependent → no claim 1990 2000 2010 2020 stable: ~0.8% → 1.4%
The extent-gated invasive estimate per epoch (solid). Before 2000 it's recipe-dependent, the pure Landsat-5 TM composite shifts by percentage points with the compositing recipe (dashed spread, illustrative of the documented instability), so no pre-2000 claim. From 2000 on it's stable and small: ~0.8% → 1.4% of area, barely above the ±0.6 pp method noise. Only the present-day product is asserted.

See it directly. Drag the year: the map shows the model's actual invasive prediction for each decade, over the same corridor (green reference). Watch the flag in the corner. From about 2000 on the estimate is stable; slide to 1990 and it turns amber, because that estimate changes with how the imagery is assembled. This isn't a growth animation. It's a demonstration of where the record stops being trustworthy.

2020present-day · calibrated
1990200020102020
Invasive-woody prediction over the Farmington corridor, per composite epoch. Green = the 2020 riparian corridor reference; red = invasive prediction for the selected year. Imagery © Esri. ⚠ Experimental model output, provisional, not validated ground truth.
Why you can trust the parts that stand

The method is the deliverable

By this point you've watched two negatives and a retraction survive on the same page as the results. That isn't an accident. I hunted down the ways this work could be wrong instead of hiding them, and four habits do most of that job.

Held-out reaches
Models are tested on whole stretches of river they never saw, not on random pixels. That is the only split that predicts real deployment.
Controls declared first
The Russian-olive control was declared before the beetle run, so it could veto the result rather than be quietly dropped.
Robustness re-runs
Every deep-time number was recomputed under several ways of assembling the imagery. The ones that moved were retired, not reported.
The novelty claim gets attacked
The claim of newness is tested against the literature on purpose. Extent mapping was already solved (CO-RIP), so it is framed as the control, not the win.
Calibrated language, on purpose You'll notice the wording here: "well calibrated," "documented negative," "cannot resolve." Not "validated," "proven," "breakthrough." The claims are sized to the evidence, and that restraint is the credibility.
Ask the agent
How the work itself stays honest

AI-assisted research that catches its own errors

I built this fast, with a lot of AI assistance, and that is exactly where research quietly goes wrong. Hallucinated code is the easy kind of failure; a compiler catches it in seconds. The dangerous failures are the ones that mean the wrong thing: they compile, they pass the tests, and they read beautifully, while being wrong.

AI-assisted research fails by producing work that is fast, fluent, plausible, and wrong, and the wrongness is invisible precisely because the output looks finished. Good intentions don't fix this. Gates do.
Failure modeCaught by
Hallucinated API callthe compiler
Wrong logica unit test
A retracted result still published as factnothing
A model scored against 45%-wrong labelsnothing
A novelty claim already falsified by a 2018 papernothing

So honesty here is enforced by machinery, not by reminders. Retired numbers, withdrawn claims and orphaned docs are caught by automated checks that run identically on my laptop and in CI. The finding, stated plainly in the method write-up: every rule that was merely written down eventually drifted; every rule a machine enforced held.

The most honest receipt The project's own merge gate was theatre for 25 of its first 29 merged PRs, a skipped AI review still posts a green check, and "no findings" is indistinguishable from "no review." It was caught only because a reviewer asked a question that couldn't be answered without checking. That failure is kept in the record, not scrubbed, which is the whole point.

How it's built

The same discipline runs through the stack: a reproducible satellite pipeline, a spatial database, two model tracks, a services API, and the grounded agent. Each sits behind its own checks.

Ingest
STAC ETL · Planetary Computer
Store
PostGIS medallion · bronze→gold
Model
Random Forest · OlmoEarth (Ai2)
Serve
.NET API · MVT vector tiles
Map + Ask
MapLibre · hybrid RAG · OLMo-ready

The agent below is part of the same discipline: no source, no claim. Grounding is enforced backend-side, so every sentence it returns is pinned to a retrieved document you can open.

Ask the agent
What this demonstrates

Four disciplines, one artifact

Mapping invasive riparian vegetation honestly, and shipping it as a live product, usually takes four different people. Here it was one person, me. Every claim below is backed by something on this page or in the code.

🌿
Environmental scientist

Framed the riparian-health question, found the phenological signal that separates the species, and held the whole thing to real scientific standards: leave-one-reach-out validation, pre-registered controls, and two negative results kept in rather than buried.

🛰️
Remote-sensing / GIS engineer

Built the Earth-observation pipeline end to end: STAC access to the Landsat and Sentinel-2 archives, phenology-aligned median-mosaic cubes, a PostGIS medallion store, and the MapLibre vector-tile maps you've been scrolling through.

🤖
AI developer

Ran a rigorous field evaluation of Ai2's OlmoEarth foundation model against a Random Forest, then built the grounded agent below for Ai2's OLMo: hybrid retrieval, prompt-injection and PII guards, semantic caching, and self-correcting retrieval.

🏗️
Senior software engineer

Shipped it as a real, hardened system: a .NET / PostGIS services layer, a deployed and locked-down agent, and a novel AI-assisted engineering method that mechanizes honesty instead of hoping for it.

Interrogate it

Ask the agent, it answers from the documents, and cites them

Everything above, the findings, the methods, and the watershed literature behind them, is indexed for search by meaning and by keyword. Ask it anything. The answer is written by an open language model and grounded in the retrieved sources with clickable citations; mention a reach and the map flies to it. The agent is built for Ai2's OLMo. No provider serves a live OLMo endpoint yet, so for now it speaks through another open model, and OLMo becomes selectable the moment one does. The rule underneath is simple: no source, no claim.

Riparian document agent
Hybrid RAG over the corpus + these findings
offline

I'm the assistant for this project. Ask about the riparian science and findings, how the maps were made, the RF-vs-OlmoEarth field test, the engineering method behind it, or how this agent itself was built. Tap a question below or type your own. When I'm live, answers are grounded in the sources with citations, and a reach mention flies the map; offline, you'll get short pre-written notes.

Try asking
Offline: pre-written answers about this page. The live agent (grounded + cited) turns on when its endpoint is configured.