How healthy is the river's forest, and which way is it heading?
The question underneath it all is whether a river corridor is healthy, and whether it's holding steady or slipping. This works that out from decades of Landsat and Sentinel-2 imagery: mapping the native riparian forest and the invasive share, then tracking how each shifts over time. Along the way it puts a Random Forest up against Ai2's OlmoEarth foundation model in a real field test, on ground neither model has seen, to find where the foundation model genuinely earns its keep and where the two simply tie. It's just as candid about what didn't hold up: two tempting results that came apart under scrutiny, kept in rather than quietly dropped. And the agent that answers your own questions at the end runs on Ai2's OLMo. OlmoEarth sees; OLMo speaks.
An independent, unaffiliated project — not produced, sponsored, or endorsed by the Allen Institute for AI (Ai2). It's built on Ai2's open models, used under their open licenses: OlmoEarth (Earth-observation foundation model) delineates the vegetation and OLMo (open LLM) grounds the answers. An outside field test and application of those models, not a benchmark Ai2 ran. "OlmoEarth" and "OLMo" are Ai2's names, used here only to identify the models evaluated.
Experimental — not verified data. Every map, number, and model output on this site (both the Random Forest and the OlmoEarth results) is an in-progress research prototype and has not been independently validated. The figures are in-sample calibration, not ground truth. This was built with genuine scientific rigor and will keep being validated and refined to represent the real world more accurately over time — but for now, treat every output as provisional and possibly wrong, not as fact.
In the arid Southwest, the whole ecosystem hangs on a green ribbon
Across the desert Southwest, life crowds along the rivers. The riparian corridor (the galleries of cottonwood and willow that line the water) is a thin sliver of the land, yet a disproportionate share of the region's birds, mammals, amphibians and pollinators depend on it for water, shade, food, and safe passage between habitats. In a landscape this dry, the river's edge is the habitat, and its condition is a fair proxy for the health of everything around it.
That ribbon is under pressure. Tamarisk (saltcedar) and Russian olive, brought over from Eurasia, have spread through Western river systems and pushed out the native cottonwood-willow forest. They grow in dense monocultures, load the soil with salt, change how fire moves through the bottomlands, and leave behind poorer habitat than what they replaced. For anyone managing these lands, the first question is simply where the corridor has turned invasive, and how quickly. That's exactly what's hard to see from the ground.
There was a hopeful intervention. In the 2000s the tamarisk leaf beetle was released across the West as a biological control, meant to defoliate saltcedar and give the natives room to recover. Whether it has durably rolled the invasion back is genuinely unsettled: the beetle spread, but tamarisk often grows back, and the net effect has been hard to pin down. This project's own attempt to find that effect in the satellite record came up empty, which is a result we keep in rather than bury. That kind of uncertainty is precisely what long-term monitoring exists to resolve.
Extent is solved. The contribution is the honest invasive layer.
Someone has already mapped where the riparian woods are in this watershed, and mapped them well (CO-RIP, Woodward 2018, κ 0.80). So reproducing that isn't the achievement; it's the control. What doesn't exist yet is a wall-to-wall native-vs-invasive map that's honest about its own uncertainty: which species, how much, and over which stretches of the record the answer can actually be trusted. That's the target. What follows is the evidence for how close it currently gets.
A bounding box becomes a phenology cube, then a decision
What separates tamarisk from native cottonwood isn't how they look on any single day, it's their timing. The two leaf out and drop their leaves on different schedules, and hold water differently in a way the shortwave-infrared bands can see. So one snapshot is the wrong input. Instead, each reach becomes a 12-month cube of median mosaics, all composited the same way, and both models get to read the whole year's rhythm rather than a single greenness number.
Why not just NDVI?
Because for this particular problem, greenness is almost blind. Ask a model to tell invasive from native using the NDVI greenness index alone, and it scores AUC ≈ 0.50, a coin flip. The real signal lives elsewhere: in the shortwave-infrared and red-edge bands, followed across the season. Tamarisk simply stays green later into the fall than cottonwood does.
Nearly a quarter of the corridor is invasive
Green is the predicted riparian-woody corridor (7.6 km²); red is the invasive tamarisk / Russian-olive share within it (1.7 km²). Toggle either layer. This is the real prediction, on satellite imagery.
On rivers they tie. One arroyo decides it.
First, what's actually being compared: two very different ways to turn satellite pixels into a map.
Why put them head to head? Because "just use the fancy model" isn't a plan. A foundation model costs real money and hardware, so it's only worth deploying where it measurably beats the cheap baseline. The goal here isn't to crown a winner; it's to find the exact conditions under which OlmoEarth earns its keep, and to say so plainly where it doesn't. That's a real deployment decision, made on evidence instead of hype.
The test that matters isn't how well a model scores on data it trained on; it's whether it transfers to a place it has never seen. So we train on some reaches and predict a held-out one. A per-pixel Random Forest, pooled across four very different reaches, transfers well to unfamiliar river corridors. It has exactly one blind spot: the lone arroyo, the only landform of its kind in the set. And that's precisely where a foundation model should have the edge, it comes down to spatial context on a shape the training data barely covers. This is the honest test of whether OlmoEarth actually delivers there.
Here are all three over the labeled reach. The foundation model (green) tracks the corridor; the NMRipMap reference labels (outline) are the comparison baseline; and the Random Forest (orange), dropped into an arroyo it never trained on, fires only in the wetter eastern end, the stretch whose spectra resemble the river reaches it did learn, while missing the dry western wash that both the foundation model and the truth pick up. That's genuine under-prediction, not a data gap (~0.1% of pixels, checked against a fully-covered cube). It's the 0.557-vs-0.889 gap made visible: to the eye, the foundation model is simply the more accurate map. Both layers are full-extent, leave-one-reach-out predictions.
The beetle that didn't break the classifier
We expected a tamarisk classifier trained on today's imagery to flip once you ran it back past the tamarisk beetle's arrival (~2004–07), when the plant's greenness signal changes. It didn't budge. And the negative control we'd declared up front, Russian olive, which the beetle doesn't touch, moved seven times more than the tamarisk "signal" did. That's the tell: the cross-era noise is simply too large to resolve a beetle effect at all.
| taxon vs native (AUROC) | 2020 | 2015 | 2000 pre-beetle | Δ |
|---|---|---|---|---|
| tamarisk (signal) | 0.849 | 0.813 | 0.862 | +0.013 |
| Russian olive (control) | 0.891 | 0.767 | 0.553 | −0.338 |
Declaring the control in advance is what let it veto the result instead of us over-reading a 0.013 wiggle. The honest read: a beetle-robust discriminator, on a sample too small and too cross-sensor-confounded to say more.
The "5× growth" that was an artifact
The first pass back through the Landsat archive looked like a clean ~5× rise in invasive extent since 1990. Every robustness check dissolved it. Two things sink it. The pre-2000 numbers swing with the compositing recipe, no normalization fixes it, because 1990 is pure Landsat-5 TM, spectrally too far from the training year to reconcile. And once the invasive layer is properly gated to woody vegetation, even the face-value change is small: from under 1% to about 1.4% of the area, most of it within the method's own noise.
See it directly. Drag the year: the map shows the model's actual invasive prediction for each epoch, over the same corridor (green reference). Watch the flag in the corner, from ~2000 on the estimate is stable; slide to 1990 and it turns amber, because that pure Landsat-5 map moves with the compositing recipe. This isn't a growth animation. It's a demonstration of where the record stops being trustworthy.
The method is the deliverable
What separates this from a demo is simple: the ways it could be wrong were hunted down, not hidden. Four habits do most of that work.
AI-assisted research that catches its own errors
This was built fast, with a lot of AI assistance, which is exactly where research quietly goes wrong. Hallucinated code is the easy kind of failure; a compiler catches it in seconds. The dangerous failures are semantic: they compile, they pass the tests, and they read beautifully, while being wrong.
AI-assisted research fails by producing work that is fast, fluent, plausible, and wrong, and the wrongness is invisible precisely because the output looks finished. Exhortation doesn't fix this. Gates do.
| Failure mode | Caught by |
|---|---|
| Hallucinated API call | the compiler |
| Wrong logic | a unit test |
| A retracted result still published as fact | nothing |
| A model scored against 45%-wrong labels | nothing |
| A novelty claim already falsified by a 2018 paper | nothing |
So honesty here is mechanized, not exhorted. Retired numbers, withdrawn claims and orphaned docs are caught by drift gates that run identically on a laptop and in CI. The finding, stated plainly in the method write-up: every rule that was merely documented eventually drifted; every rule that was mechanized held.
How it's built
The same discipline runs through the stack, a reproducible EO pipeline, a typed spatial store, two model tracks, a services API, and the grounded agent, each behind its own gate.
The agent below is part of the same discipline: no source, no claim. Grounding is enforced backend-side, so every sentence it returns is pinned to a retrieved document you can open.
Four disciplines, one artifact
Mapping invasive riparian vegetation honestly, and shipping it as a live product, usually takes four different people. This is one person doing all four, and every claim below is backed by something on this page or in the code.
Framed the riparian-health question, found the phenological signal that separates the species, and held the whole thing to real scientific standards: leave-one-reach-out validation, pre-registered controls, and two negative results kept in rather than buried.
Built the Earth-observation pipeline end to end: STAC access to the Landsat and Sentinel-2 archives, phenology-aligned median-mosaic cubes, a PostGIS medallion store, and the MapLibre vector-tile maps you've been scrolling through.
Ran a rigorous field evaluation of Ai2's OlmoEarth foundation model against a Random Forest, then built the grounded agent below on Ai2's OLMo: hybrid retrieval, prompt-injection and PII guards, semantic caching, and self-correcting retrieval.
Shipped it as a real, hardened system: a .NET / PostGIS services layer, a deployed and locked-down agent, and a novel AI-assisted engineering method that mechanizes honesty instead of hoping for it.
Ask the agent, it answers from the documents, and cites them
Everything above, the findings, the methods, and the watershed literature behind them, is indexed in a hybrid-retrieval system (dense embeddings + BM25). Ask it anything. The answer is written by OLMo, Ai2's open language model, grounded in the retrieved sources with clickable citations, and if you mention a reach the map flies to it. No source, no claim, so the same Ai2 stack that maps the vegetation also explains it.