Three Things That Would Genuinely Surprise Me in Two Years
More useful than predictions: naming what would count as a genuine surprise. Three of them, each identifying a load-bearing belief — and what evidence would overturn it.
Most predictions about agents are either extrapolations of current trends or restatements of current frustrations. More informative is to name what would count as a genuine surprise — because that identifies which beliefs are load-bearing, and gives you something to actually watch for.
Three things, each with the reasoning that makes it surprising and what would have to be true for it to happen.
1. Agents reliably handling problems with no feedback signal
Why it would surprise me: the strength comes from iterating against a check. Compile, test, validate, reconcile — attempt, verdict, adjust. Where no verdict exists, an agent produces one confident attempt, and there's no mechanism converting that into a converged answer.
Prioritization, architecture choices whose rightness depends on the future, strategy: these are evaluated over months, confounded, sometimes never conclusively. Better reasoning produces a better guess. It doesn't produce a checked answer, because the checking apparatus isn't absent from the model — it's absent from the world.
What would have to change: either cheap reliable simulation of consequences in messy domains, or a fundamentally different learning signal than iterate-against-verdict. Both are conceivable. Neither is an extrapolation of what's currently working.
2. Organizations granting broad autonomy quickly
Why it would surprise me: autonomy is granted based on predictability and recourse, not accuracy. An agent that's right more often but fails unpredictably gets less latitude than a weaker one whose limits are legible — because you can route around known limits and can't route around surprise.
Trust also moves on a social timescale. It requires demonstrated reliability over time, a record people can point at, and a bad incident resets it substantially. Regulatory and insurance structures move slower still.
What would have to change: agents that reliably know and state their own limits — a genuine calibration advance rather than a capability one — plus accountability structures that adapt faster than they historically have. The first is plausible. The second isn't fast.
⚠️ The caveat: I'd expect autonomy to expand quickly in bounded domains where reversibility is high and blast radius is small. The surprise would be broad autonomy over consequential, irreversible actions.
3. Specification quality ceasing to be the bottleneck
Why it would surprise me: specification means converting a person's partly-formed intent into something precise. The difficulty isn't the writing — it's that the intent doesn't fully exist until questions force it into shape, and the questions require knowing what matters.
An agent can draft, surface ambiguity, enumerate edge cases. It can't know that the finance team needs the old field name, or that this customer is in a renewal, or that the requester hasn't considered what happens to existing records.
What would have to change: organizational context becoming captured as a routine byproduct of work rather than a documentation project. This is more plausible than it sounds — richer artifacts, transcripts, decision records — and it's the one of the three I'd bet on moving most, unevenly, and mostly in organizations that decide to make it move.
💡 What this exercise is for
Naming your surprises makes your model falsifiable. If any of the three happened, I'd have been wrong about something structural rather than about a timeline — specifically about the claim that feedback signals, trust dynamics, and unwritten context are the binding constraints rather than model capability.
That's a more useful position than a prediction with a date, because it says what evidence would change it.
✅ The complementary exercise, and the one that's actually actionable: name what wouldn't surprise you. Longer coherent runs, better tool selection, cheaper inference, more of the routine work automated, coding agents handling larger changes. Those are extrapolations. If you're planning around them you're planning around the consensus, which is fine — just don't mistake it for insight.
What I'd watch for
- An oracle appearing where there wasn't one — cheap, reliable experimentation with clean attribution in a messy domain. That would move item 1 and a lot else.
- A calibration advance — agents that decline correctly rather than guessing well. That would move item 2 faster than capability gains would.
- Context capture becoming a byproduct rather than a chore. That would move item 3.
None of those are model-scale stories, which is the point.
The takeaway
The load-bearing beliefs are that feedback signals gate capability, trust gates autonomy, and unwritten context gates specification. Being surprised on any of the three would mean the constraint was somewhere other than I thought. Naming your own surprises is worth more than naming your predictions — it tells you which evidence would change your plans, and most of the evidence that matters isn't a benchmark number.