You’ve seen the demo. Every AI keynote this year features the same one: an agent given a complex goal, planning, clicking, calling APIs, completing work while the presenter sips coffee and smiles. And here’s the thing: the demos are real. But the distance between demo and dependable production is ALSO real, and understanding both halves is the only way to read agent announcements sensibly. So let’s skip the keynote version. Here’s what’s actually happening where agents meet real work, based on deployment evidence rather than stagecraft.
Key takeaways
- Agents act rather than advise: the capability shift is real, and so is the reliability gap.
- Production success concentrates in coding, research and support: bounded, verifiable, reviewable.
- Compounding error across steps is the core unsolved problem; ask vendors for end-to-end success rates.
- Per-outcome pricing converts agents into labor economics; watch where it survives.
- The career move is becoming the director and verifier of agent work.
What an agent actually is
Let’s define the thing precisely, because the word gets abused. An agent is a language model in a loop with tools: it plans, acts through search, code, browsers or business software, observes results, and continues until done or stuck. The leap from chatbot to agent is the leap from advising to doing. A chatbot tells you how to reconcile the invoices. An agent reconciles them. That difference in kind, not degree, is why the category matters, and why its failures matter more.
Where agents work in production today
Strip away the hype and three domains show genuine deployed success. Software engineering leads decisively: coding agents complete bounded tasks, fix issues and write tests with human review, and adoption data shows real usage at scale. Why code first? Because code offers something rare: automatic verification through tests. Research and analysis follows: deep research agents produce cited reports in minutes that took analysts hours, imperfect but genuinely useful as first drafts. Customer support rounds out the trio: resolution agents like Intercom’s Fin handle defined ticket categories autonomously, with per-resolution pricing aligning vendor confidence with customer value. Notice the pattern across all three: bounded scope, verifiable outputs, human checkpoints on irreversible actions.
Where they still fail
Now the part the keynotes skip. The failure mode is compounding error, and the math is brutal. A step accuracy of ninety-five percent (impressive for a model!) yields barely sixty percent task success across fifteen steps. Production agents fail by drifting from goals over long horizons, misunderstanding interfaces never designed for them, and confidently persisting down wrong paths. The vendors know this. The serious ones ship evaluation dashboards, human-in-the-loop gates and rollback. The unserious ones ship demos. So when evaluating any agent claim, ask the questions that cut through: what’s the end-to-end success rate on unscripted tasks? What happens on step twelve? Who catches the mistake?
The economics arriving now
Agent pricing is settling into two models: per-seat copilot pricing and per-outcome pricing (per resolution, per completed task). The second is revolutionary when it works, converting AI from a tool cost into a labor substitute priced against the work itself. Watch this space closely: per-outcome pricing succeeding or failing in support and coding will predict how fast agents spread to less verifiable domains. When a vendor bets its own revenue on the agent completing the task, that’s confidence you can price.
What this means for jobs, honestly
The tasks agents automate first share a profile: digital, structured, verifiable, forgiving of occasional human review. Junior coding tasks, first-line support, routine research and data processing sit squarely in it. The labor market evidence so far shows task transformation outpacing job elimination, with entry-level pressure in exposed categories the clearest signal. Our analysis of the employment research examines the data in detail. The strategic response for individuals is unchanged and urgent: become the person who directs and verifies agent work, because that role is expanding as fast as the tasks beneath it contract.
How to run your first agent pilot
Convinced enough to test agents in your organization? The pilot design determines what you learn. Choose a task with four properties: digital end to end, bounded in scope, verifiable output, and reversible actions. Good first pilots: triaging inbound requests with drafted responses for approval, assembling recurring reports from defined sources, codebase maintenance with tests as the verifier. Bad first pilots: anything touching money movement, customer commitments or deletion, and anything where success is judged by vibes.
Structure it in three phases. Shadow mode first: the agent works in parallel with the human process, and you compare outputs. Assisted mode second: the agent prepares, the human approves each action. Autonomy last, only for categories that proved themselves, with logging, rollback and a named human owner. Measure end-to-end success rate and time saved, not step counts. And decide your kill criterion in advance, because that’s what separates learning from hoping.
The next two years, predicted carefully
Expect steady reliability gains, since the labs’ entire research weight sits on long-horizon performance. Expect computer-use agents to become mundane within specific enterprise software first. And expect at least one public, expensive agent failure to recalibrate the hype. The organizations extracting value share a posture: ambitious scope in experimentation, conservative scope in production, ruthless measurement of what agents actually complete. Boring consistency beats dazzling generality, as it did with spreadsheets and every automation wave before this one.
How we cover agents. Our analysis draws on vendor-published reliability data, independent evaluations, and conversations with teams running agents in production. We distinguish demonstrated capability from deployed reliability throughout. Standards on our methodology page.
The bottom line
So, back to that keynote demo. Enjoy it, then ask for the reliability numbers. If your work involves bounded digital tasks with verifiable outcomes, agent tools deserve a pilot this year, with human review designed in from the start. If it doesn’t, watch the numbers quarterly and wait. Either way, understanding the underlying architecture separates informed adoption from expensive tourism. The agents are coming for the boring parts of your job first. Honestly? Let them.