How healthy is the river's forest, and which way is it heading?
From orbit you can, if you read time instead of color. On any single day an invaded stretch looks exactly as green as a healthy one. But tamarisk holds its green for weeks after the native trees brown out, and that timing signature is written into forty years of satellite imagery. I built a system that reads it: where the river forest is, how much of it has turned invasive, and which way both are heading.
It's also a field test with real stakes. Every map here is made two ways, once by a classic Random Forest and once by Ai2's OlmoEarth foundation model, to find out where the bigger model actually earns its cost on stretches of river neither has seen. And the whole thing shows its work. Two tempting headline results failed their own tests and stayed on the page as negatives instead of quietly disappearing. Ask the agent at the bottom about any of it. It answers from the project's own documents and cites them.
This is an independent project, not produced, sponsored, or endorsed by the Allen Institute for AI (Ai2). It is built on Ai2's open models under their open licenses: OlmoEarth maps the vegetation, and OLMo is the model the agent is built for. Their names appear here only to identify the models being evaluated.
Experimental, not verified data. Every map and number on this site, from both models, is a research prototype that has not been independently validated. The work is done with real scientific rigor, and each figure says how it was measured. Still: treat every output as provisional and possibly wrong, not as fact.
In the arid Southwest, the whole ecosystem hangs on a green ribbon
Across the desert Southwest, life crowds along the rivers. The riparian corridor (the galleries of cottonwood and willow that line the water) is a thin sliver of the land, yet a disproportionate share of the region's birds, mammals, amphibians and pollinators depend on it for water, shade, food, and safe passage between habitats. In a landscape this dry, the river's edge is the habitat, and its condition is a fair proxy for the health of everything around it.
That ribbon is under pressure. Tamarisk (saltcedar) and Russian olive, brought over from Eurasia, have spread through Western river systems and pushed out the native cottonwood-willow forest. They grow in dense monocultures, load the soil with salt, change how fire moves through the bottomlands, and leave behind poorer habitat than what they replaced. For anyone managing these lands, the first question is simply where the corridor has turned invasive, and how quickly. That's exactly what's hard to see from the ground.
There was a hopeful intervention. In the 2000s the tamarisk leaf beetle was released across the West as a biological control, meant to defoliate saltcedar and give the natives room to recover. Whether it has durably rolled the invasion back is genuinely unsettled: the beetle spread, but tamarisk often grows back, and the net effect has been hard to pin down. My own attempt to find that effect in the satellite record came up empty. That result stays on this page, because that kind of uncertainty is exactly what long-term monitoring exists to resolve.
Extent is solved. The contribution is the honest invasive layer.
Someone has already mapped where the riparian woods are in this watershed, and mapped them well (CO-RIP, Woodward 2018, κ 0.80). So reproducing that isn't the achievement; it's the control. What doesn't exist yet is a wall-to-wall native-vs-invasive map that's honest about its own uncertainty: which species, how much, and over which stretches of the record the answer can actually be trusted. That's the target. What follows is the evidence for how close it currently gets.
The trick is timing: read the whole year, not one photo
What separates tamarisk from native cottonwood isn't how they look on any single day. It's their timing. The two species leaf out and shed on different schedules, and they hold water differently in a way the satellite's infrared bands can see. Ecologists call this seasonal rhythm phenology, and it's the fingerprint everything here runs on. So instead of one photo, each stretch of river becomes a stack of twelve monthly images, one cloud-free composite per month, built the identical way every time. Both models read the year's whole rhythm instead of a single greenness number.
Why not just measure greenness?
Because for this particular problem, greenness is almost blind. Ask a model to tell invasive from native using the standard greenness index (NDVI) alone and it scores AUC ≈ 0.50. AUC is the score used throughout this page: 1.0 is a perfect classifier, 0.5 is a coin flip. So: a coin flip. The real signal lives elsewhere, in the infrared bands followed across the season. Tamarisk simply stays green later into the fall than cottonwood does.
Nearly a quarter of the corridor is invasive
Here is what that instrument sees over Farmington. Green is the predicted river-forest corridor (7.6 km²). Red is the share of it the model calls invasive (1.7 km²). Toggle either layer. This is the real prediction, drawn on satellite imagery.
On familiar ground they tie. One unfamiliar reach decides it.
First, what's actually being compared: two very different ways to turn satellite pixels into a map.
Why put them head to head? Because "just use the fancy model" isn't a plan. A foundation model costs real money and hardware, so it's only worth deploying where it measurably beats the cheap baseline. The goal here isn't to crown a winner; it's to find the exact conditions under which OlmoEarth earns its keep, and to say so plainly where it doesn't. That's a real deployment decision, made on evidence instead of hype.
The test that matters isn't how well a model scores on data it trained on. It's whether the model survives a place it has never seen. So I train both models on three stretches of river and make them predict a held-out fourth, taking turns until every reach has been the unseen one. On reaches that resemble the training set, the cheap Random Forest does fine (0.85 to 0.90) and the foundation model merely matches it. But on Malpais, the one reach whose landscape is unlike the rest (a valley of irrigated fields and dry upland), the Random Forest collapses to a coin flip at 0.56 while OlmoEarth holds at 0.89. That is the whole result: the foundation model doesn't see better, it breaks less on unfamiliar ground. And an unlabeled basin is almost entirely unfamiliar ground.
Here are all three on the same reach. The expert labels are the outline. The foundation model (green) tracks the whole corridor. The Random Forest (orange), dropped onto ground it never trained on, fires in one small eastern pocket, the stretch that most resembles the reaches it learned, and misses everything else. That is the coin-flip collapse made visible. To the eye, the foundation model is simply the more accurate map.
The beetle that didn't break the classifier
I expected a tamarisk classifier trained on today's imagery to flip once I ran it back past the beetle's arrival around 2004, because the beetle changes how green the plant looks. It didn't budge. Better: before the run I declared a negative control, Russian olive, a plant the beetle doesn't touch. If the record were clean, its score should have stayed put. Instead it moved seven times more than the tamarisk "signal" did. That's the tell: the noise between eras is simply too large to measure a beetle effect at all.
| taxon vs native (AUROC) | 2020 | 2015 | 2000 pre-beetle | Δ |
|---|---|---|---|---|
| tamarisk (signal) | 0.849 | 0.813 | 0.862 | +0.013 |
| Russian olive (control) | 0.891 | 0.767 | 0.553 | −0.338 |
Declaring the control in advance is what let it veto the result, instead of me over-reading a 0.013 wiggle. The honest read: the classifier looks beetle-proof, but the sample is too small, and too tangled across satellite generations, to say more.
The "5× growth" that was an artifact
The first pass back through the Landsat archive looked like a clean ~5× rise in invasive cover since 1990. Every robustness check dissolved it. Two things sink it. First, the pre-2000 numbers change depending on how the imagery is assembled, and no correction fixes that, because 1990 comes from an older satellite whose sensor is too different from the training year to reconcile. Second, once the invasive layer is limited to woody vegetation, even the face-value change is small: from under 1% to about 1.4% of the area, most of it within the method's own noise.
See it directly. Drag the year: the map shows the model's actual invasive prediction for each decade, over the same corridor (green reference). Watch the flag in the corner. From about 2000 on the estimate is stable; slide to 1990 and it turns amber, because that estimate changes with how the imagery is assembled. This isn't a growth animation. It's a demonstration of where the record stops being trustworthy.
The method is the deliverable
By this point you've watched two negatives and a retraction survive on the same page as the results. That isn't an accident. I hunted down the ways this work could be wrong instead of hiding them, and four habits do most of that job.
AI-assisted research that catches its own errors
I built this fast, with a lot of AI assistance, and that is exactly where research quietly goes wrong. Hallucinated code is the easy kind of failure; a compiler catches it in seconds. The dangerous failures are the ones that mean the wrong thing: they compile, they pass the tests, and they read beautifully, while being wrong.
AI-assisted research fails by producing work that is fast, fluent, plausible, and wrong, and the wrongness is invisible precisely because the output looks finished. Good intentions don't fix this. Gates do.
| Failure mode | Caught by |
|---|---|
| Hallucinated API call | the compiler |
| Wrong logic | a unit test |
| A retracted result still published as fact | nothing |
| A model scored against 45%-wrong labels | nothing |
| A novelty claim already falsified by a 2018 paper | nothing |
So honesty here is enforced by machinery, not by reminders. Retired numbers, withdrawn claims and orphaned docs are caught by automated checks that run identically on my laptop and in CI. The finding, stated plainly in the method write-up: every rule that was merely written down eventually drifted; every rule a machine enforced held.
How it's built
The same discipline runs through the stack: a reproducible satellite pipeline, a spatial database, two model tracks, a services API, and the grounded agent. Each sits behind its own checks.
The agent below is part of the same discipline: no source, no claim. Grounding is enforced backend-side, so every sentence it returns is pinned to a retrieved document you can open.
Four disciplines, one artifact
Mapping invasive riparian vegetation honestly, and shipping it as a live product, usually takes four different people. Here it was one person, me. Every claim below is backed by something on this page or in the code.
Framed the riparian-health question, found the phenological signal that separates the species, and held the whole thing to real scientific standards: leave-one-reach-out validation, pre-registered controls, and two negative results kept in rather than buried.
Built the Earth-observation pipeline end to end: STAC access to the Landsat and Sentinel-2 archives, phenology-aligned median-mosaic cubes, a PostGIS medallion store, and the MapLibre vector-tile maps you've been scrolling through.
Ran a rigorous field evaluation of Ai2's OlmoEarth foundation model against a Random Forest, then built the grounded agent below for Ai2's OLMo: hybrid retrieval, prompt-injection and PII guards, semantic caching, and self-correcting retrieval.
Shipped it as a real, hardened system: a .NET / PostGIS services layer, a deployed and locked-down agent, and a novel AI-assisted engineering method that mechanizes honesty instead of hoping for it.
Ask the agent, it answers from the documents, and cites them
Everything above, the findings, the methods, and the watershed literature behind them, is indexed for search by meaning and by keyword. Ask it anything. The answer is written by an open language model and grounded in the retrieved sources with clickable citations; mention a reach and the map flies to it. The agent is built for Ai2's OLMo. No provider serves a live OLMo endpoint yet, so for now it speaks through another open model, and OLMo becomes selectable the moment one does. The rule underneath is simple: no source, no claim.
I'm the assistant for this project. Ask about the riparian science and findings, how the maps were made, the RF-vs-OlmoEarth field test, the engineering method behind it, or how this agent itself was built. Tap a question below or type your own. When I'm live, answers are grounded in the sources with citations, and a reach mention flies the map; offline, you'll get short pre-written notes.