Blog
Workflows Nobody Predicted: The AI Inflection of July 2026

The 72 Hours That Rewrote the Workflow Map
On 24 July, Anthropic shipped Claude Opus 5, posting a score on ARC-AGI 3 that is three times higher than the next-best model. On the same announcement, Opus 5 also leads Zapier's AutomationBench by roughly 1.5× at equal cost, and tops OSWorld 2.0 (a computer-use benchmark) at one-third the price per completed task. On 23 July, a mathematician posted on X about closing six open Erdős problems in five days, working with OpenAI's freshly-released GPT-5.6 Sol. On the same day, Y Combinator's S26 batch launched Screenpipe, which records how you actually work and converts the footage into agents that take action on your behalf.
Three releases. One shape. The barrier between "I could automate that" and "I automated that while I slept" just collapsed. What changed is not raw IQ. What changed is the economics of delegation.
Consider the Erdős problems. Erdős was prolific, generous, and unusually good at leaving small open conjectures scattered through the literature. Most sat for decades. A single researcher with GPT-5.6 Sol — a model trained specifically for proof generation — closed six of them in five days. The previous state of the art required a research group, a university affiliation, and usually a sabbatical. The new state of the art requires a laptop and a subscription.
Or consider Screenpipe. Its pitch is almost embarrassingly mundane: it watches your screen, OCRs what you do, builds a model of your working patterns, and dispatches agents that finish the boring parts. Two years ago this would have been a research project at a frontier lab. Today it is a YC S26 launch with a working product. The unprecedented workflow is not "AI watches you work." The unprecedented workflow is "AI watches you work, learns your taste, and quietly files your expense reports."
The third workflow is the one operators should be paying closest attention to. Anthropic's announcement calls out specific Opus 5 gains on life-sciences tasks — protein-function prediction up 7.7 points over Opus 4.8, spectroscopy-to-structure inference up 10.2 points. Those are not glamorous benchmarks. They are the boring middle of every biotech pipeline. A 10-point improvement on protein-function prediction at a third of the cost per task means a two-person lab can now run the kind of structural-biology screen that previously required a compute grant and a CRO. That is not an AI story. That is a story about which labs will publish first.
The Part the Legacy Software Crowd Will Not Want to Hear
For a decade, the canonical objection to AI in the workplace was a polite one: helpful for drafts, useless for outcomes. It is a line that survived every model release. It survived the launch of copilots. It survived the first wave of agents. It survived because, until very recently, the cold truth was that models could propose but not finish.
That line is now indefensible.
Look at what Opus 5 is doing on Zapier's AutomationBench. The benchmark measures whether a model can complete a real business workflow end-to-end without a human babysitter. The previous best model completed some workflows. Opus 5 completes 1.5 times as many at the same price. That is not a marginal improvement. That is a step function. The same step function appears on OSWorld 2.0, where Opus 5 surpasses the prior best at one-third of the cost per completed task.
For an analyst who has been quietly building RPA bots for years, the math is brutal. The cost of running an automation just dropped by 3× and the success rate went up. The cost of not running an automation just dropped to zero. The legacy RPA vendor's pitch — "we have years of integration IP" — is starting to sound like "we were slow to a fight that does not need us anymore."
The predictable objection is the usual one: this is hype, benchmarks are gamed, agents hallucinate. All true. None of it matters for the trajectory. The reason is simple. The cost curve has been falling since 2023. The capability curve has been rising since 2023. At every prior intersection, the cynics were right that the demo would not generalize. At the current intersection — Opus 5 on AutomationBench, GPT-5.6 Sol on Erdős problems, Screenpipe on the personal desktop — the demos are no longer demos. They are shipping products with paying customers.
The hard part was always the last 30%. The part where the model gives you a draft and you have to fix the citations, fill in the missing cell, copy the result into the right CRM field. That last 30% is what made AI feel like a toy. As of this week, the last 30% is collapsing into the first 70%. The toy just became infrastructure.
The Quiet Part
The workflows that ship this quarter will not look like AI workflows. They will look like ordinary businesses moving faster than their competitors. A solo founder closing six Erdős problems. A two-person biotech running screens that used to take a CRO six months. An ops team whose automation backlog just cleared because the cost per task collapsed by 3×.
Nobody will call any of this "AI." They will call it Tuesday.
That is the inflection. The wins no longer announce themselves. They just show up in the results, the same way electricity stopped being a story once every factory had a meter on the wall. We are somewhere around 1930 on that curve. The thing that was a miracle last year is plumbing this year, and the year after that it will be invisible. The interesting question is no longer whether the workflows ship. It is which teams notice early enough to build on top of them before their competitors do.
The teams that move first will not be the ones with the best models. They will be the ones with the best taste about which boring, repetitive, unloved part of the business is about to become free. Watch for the next twelve weeks. The inflection is here. The only question is who is paying attention.