Databases in the Age of Agents
Almost every data analytics company I look at these days is chasing the same thing: agents. Natural language to SQL, context layers, semantic models, MCP servers, “ask your data anything.” The idea is that the way people talk to data is changing, and whoever owns that agent layer owns the customer.
The interface is changing, no doubt. But having spent a decade working on databases, I suspect a lot of these companies are optimizing for the demo. The agent layer is mostly a wrapper. It’s useful, but it’s also the easiest part to copy, and the foundation model vendors are commoditizing it anyway. The things that will actually decide who wins are down in the engine, and most of them come down to performance.
Agents change the shape of the query workload
A human analyst runs a handful of queries an hour. They think, read the results, and refine. An agent doesn’t work like that. It fans out, probing the schema, checking distributions, trying a join, backing out, and trying another one. A single question can turn into dozens of queries, most of them speculative and fired off in bursts. Multiply that across every user and every workflow running unattended, and the query volume goes up by an order of magnitude. That breaks a couple of assumptions that a lot of analytics systems were built on.
The first is query latency. An eight-second query is perfectly fine for a human who’s going to stare at the result for a minute anyway. It’s not fine for an agent that has to run twenty of them in a loop before it can say anything. The agent is blocked on every round trip, so latency that used to be acceptable suddenly becomes the dominant cost. The engines that do well here are the ones that put in the boring work (vectorized execution, caching, data pruning) instead of just throwing more hardware at the problem.
The second one is cold start, and I think it’s underrated. Agents don’t keep a warehouse warm the way a team of analysts does. They show up, do a burst of work, and disappear. If bringing up fresh compute takes thirty seconds, or a few minutes, that cost now shows up on every single interaction instead of being amortized over a workday. An engine that can go from nothing to serving queries in under a second has a real edge here, and it’s the kind of thing that’s very hard to bolt on later.
Agents need a safe place to experiment
The other big area is experimentation. If you want to actually let an agent loose on your data (to test a migration, validate a transformation, or explore some what-if), it can’t be doing any of that against production. It needs its own copy that’s isolated, instant, and cheap to throw away. And copying terabytes around for every agent obviously isn’t going to scale.
Branching is the cleanest example of what this should look like. Developers
have had cheap, instant branches in source control for two decades now, and
it completely changed how software gets built. Data has never really had an
equivalent. You either work against production and pray, or you wait around
for a full copy. What you really want is copy-on-write branching at the
storage layer: a branch created in milliseconds that costs nothing until
something actually writes to it. Think of it as git checkout -b for your
database. The nice thing is that the open table formats and the separation of
storage and compute that the industry has been building toward for years are
exactly what you need to make this work. Companies that did that homework can
offer branching almost for free; the ones that didn’t will have a hard time
bolting it on.
It also solves the safety problem, which agentic data work badly needs. Let the agent do whatever it wants on its own branch, diff the result against production, and merge only after a human signs off. That’s a lot better than handing an agent write access to production and hoping for the best.
Where I think this goes
The agent layer still matters, and the UX around it will keep getting better. But it’s converging. Everyone’s natural-language interface ends up roughly as good as everyone else’s, because they’re all sitting on the same handful of models.
The engine doesn’t converge nearly as fast. Fast queries under a workload that just got ten times heavier, near-instant cold start, cheap isolated experimentation: these are deep systems problems that take years of work in storage and execution to get right. You can’t catch up on that by shipping another wrapper. So my guess is that the database companies treating agents as a reason to finally get serious about the engine, rather than as one more feature to wrap, are the ones that will grow the fastest over the next few years.
For databases in general, this is a genuinely exciting time. The work that quietly went into query engines for years is suddenly the thing that matters most, and the agents everyone’s chasing are exactly what’s making it matter.