Seven years on from the famous “data moat” essay – what’s changed?
Re-reading “The Empty Promise of Data Moats”, seven years on
In May 2019, Martin Casado and Peter Lauten published one of the most useful contrarian essays ever written about data businesses. “The Empty Promise of Data Moats” took aim at the line every AI startup put on slide three of its deck: “our data is our moat.” Mostly wishful thinking, they argued. The idea of a “data flywheel” (more users generating more data, which improves the product, which attracts more users) turns out to be rare in practice. And the idea that simply having more data keeps making your product better runs into diminishing returns: each new batch adds less than the last, until more data buys you almost nothing. On top of that, the cost of acquiring each additional piece of genuinely unique data tends to rise even as its value falls, which is the opposite of how a normal economy of scale works.
Seven years and a foundation-model revolution later, the essay is worth a proper re-read, because it has aged in a strange and instructive way. The mechanics were right, probably more right than the authors realised. But it answered the wrong question. Or rather, the market went off and asked a different one.
What held up
Let’s start with what held up, because it held up emphatically.
Casado’s core claim centred on what he called the “minimum viable corpus” eg. the smallest body of data you need to get a model working well enough to launch. His argument was that this starter dataset is cheap to assemble, whether by scraping the web, adapting an existing model or generating synthetic data, so the data an incumbent has built up rarely protects it from a determined newcomer. In 2019 that was a bold claim; in 2026 it’s just how the market works. Foundation models have made getting started almost trivial. A team of five with API access and a synthetic-data pipeline now starts roughly where a 2019 startup landed after two years of gathering data. The distance between “no data” and “enough data to ship” has never been shorter. If you’re defending a product with the pile of data you’ve collected, you’re defending it with something a competitor can increasingly reproduce by Thursday.
There’s a nice irony buried in the essay’s best-known chart, too. It came from Arun Chaganty’s study of customer-support chatbots, and it plotted what share of customer questions a bot could handle as you fed it more and more training data. The line climbed steeply at first, then flattened off at around 40% and stopped rising, no matter how much more data you added. Casado read that as proof of diminishing returns: past a certain point, more data bought you nothing.
The flattening was real, but it wasn’t a limit of the data. It was a limit of the method those 2019 bots used, called intent classification, where the bot tries to match each message to one item on a fixed list of pre-defined requests. That approach can only ever get so good, however much data you feed it. Then large language models arrived and went straight through the 40% ceiling, not by training on more support transcripts, but by working in a completely different way, trained on external data at a scale nobody in that market was even considering. The diminishing-returns curve everyone had quoted as a law of nature turned out to describe the limits of one machine-learning method, and nothing more.
The value of a dataset isn’t intrinsic. It’s set by the systems that consume it, and those systems have changed significantly.
What the essay missed
Casado framed data as defence: does your data protect your product from competitors? His answer, usually not, was right. But framing it that way meant the essay never asked the question that turned out to matter far more: might your data be worth something to someone else’s model? That’s the market that formed.
A training-data economy that barely existed in 2019 was worth roughly $3.6 billion in 2025, and is projected to reach $4.4 billion in 2026. That’s just the dataset-licensing layer, before you add the labelling and expert-data industry sitting alongside it. And the deals on record are remarkable.
OpenAI’s five-year agreement with News Corp is reportedly worth more than $250 million. Reddit disclosed $203 million of data-licensing contracts at its IPO, anchored by a roughly $60-million-a-year arrangement with Google. Amazon’s deal with the New York Times reportedly runs $20–25 million a year. Shutterstock, the quiet arms dealer of the market, expected $138 million of AI licensing revenue in 2024, selling to several buyers at $25–50 million a deal.
None of these companies built a data moat in Casado’s sense. Reddit’s data didn’t stop anyone else building a forum. News Corp’s archive didn’t protect it from digital disruption, famously the opposite. What happened instead is that a new kind of buyer turned up, one whose models are hungry for exactly the qualities Casado said wouldn’t protect you: data that spans a huge range of topics, that’s messy and unpolished rather than neatly curated, and that captures the rare, one-off ways real people write and behave. Data mostly failed as a wall, but it succeeded as a product.
That doesn’t disprove the 2019 essay so much as it effectively completes it. Casado said defensibility isn’t inherent to data. That’s true, but neither is worthlessness. Value is set by the consuming system, and between 2019 and 2023 that system changed from “your own narrow model” to “frontier models with an almost unbounded appetite”.
The section everyone skimmed
Here’s the part of the re-read I enjoyed the most. The 2019 essay has four sections on the data journey: minimum viable corpus, acquisition cost, incremental value, and the one nobody ever quotes, data freshness. Streets change, temperatures change, attitudes change. A dataset decays, and the work of keeping it fresh only grows as it gets bigger.
Everyone skimmed that section. The market priced it.
You can see it in the shape of licensing deals. The 2023–24 wave was mostly flat annual fees for training rights: one-time access to the archive. The early landmark agreements all included model training; by 2026, only around four in ten publicly announced deals do. What buyers increasingly pay for instead is live access: ongoing feeds and real-time content the model can draw on as it answers. The archive, the stock, is a one-off sale. The flow is the recurring revenue.
This is the freshness section of the essay playing out at market scale, and it turns the instinct most data-rich companies still hold on its head. That instinct is to value the pile: we’ve got fifteen years of records. The market is telling them the pile gets bought once, at a discount, because once its training value has been extracted, it’s extracted. What earns a recurring price is the thing that didn’t exist yesterday: the accumulating stream you can’t backfill. The moat, in other words, isn’t the data you own; it’s the data you’re still producing, the one thing a competitor can’t bootstrap, scrape or synthesise. The licensing market is the biggest test yet of that idea, and the flow is winning.
The buyer knows more than you do
The essay’s central idea has a commercial edge it never explored; if the value of data is set by the system that consumes it, then the buyer, the frontier lab, is also the only one who can actually measure it. Labs do it the honest way: train a model with the data, train one without it, and watch how the benchmarks move. But those results almost never reach the seller. So the first publishers to license their archives were pricing an asset with no market, against buyers who knew far more than they did about how much it was worth. The flat fees of 2023–24 weren’t really prices. They were guesses, mostly in the buyer’s favour.
That’s starting to correct, not through new transparency but through the structure of the deals. Reddit, for one, is reportedly pushing for pricing that scales with how much its data improves AI answers, using its position as the most-cited source in Google’s and Perplexity’s AI results as leverage.
So the lesson for anyone sitting on a data asset is clear: turn up to a negotiation without your own evidence of uplift and you’ll be paid a guess. A controlled study of an AI valuation agent in pharma showed it’s possible to evidence the value yourself: with access to proprietary data the system recovered 96% of a gold-standard record, against 25–38% without. That kind of with-and-without number is exactly what a buyer can’t wave away.
The 2026 verdict on the 2019 essay
So, seven years on, does the essay still hold? Yes and no.
As advice to founders defending a product, it’s still right, and more so than ever. Don’t skimp on go-to-market, verticalisation and workflow depth in the belief that the data you’ve accumulated will keep competitors out. It won’t, and foundation models have made it a lot easier for a rival to catch up
Where the essay falls short isn’t in its answer, but in its question. It judged data by a single test, whether it could keep competitors out, and concluded (rightly) that mostly it can’t. But it never looked at the other side. Instead of keeping people out, what if deliberately distributing your data could make you more competitive? Data may have failed as a moat, but 7 years on it’s succeeding as a revenue-generating product, bought by a kind of customer that, for the most part, didn’t exist in 2019.
Ironically, the same tests Casado used to argue that data wasn’t a moat, are the ones that now tell you whether your data is a worthy product:
• Is it a stock or a flow? Only one of them earns recurring revenue.
• Can you evidence its uplift? The buyer already can, and silence prices against you.
• Is its value protected by anything a consuming system can’t simply route around? The limit you’ve measured may belong to the architecture, not to your data.
Most importantly, having data worth distributing is one thing; being able to do it well (and safely) at scale is another. Can you control who uses your data and how it’s used, across many consumers (including AI)? And can you do it consistently enough that strategically distributing your data can become a competitive moat?
The 2019 essay said data was no magic moat. It was right.
The moat was never the data itself, but what you can do with it.