Blog
The Week the Frontier Got Cheap

On 30 July 2026, OpenAI announced that GPT-5.6 Terra would drop 20% to $2 per million input tokens and $12 per million output tokens, and that GPT-5.6 Luna would drop 80% to $0.20 per million input tokens and $1.20 per million output tokens. Sol, the most powerful of the three, kept its price. The cuts landed three weeks after the GPT-5.6 series shipped. OpenAI's framing was the right one: "Our strategy remains focused on advancing both capability and efficiency so each generation of intelligence can accomplish more work at a lower cost." The subtext is that customers are no longer willing to pay the previous bill, and the timing traces to Moonshot's Kimi K3 forcing the issue on price.
The price collapse and the revenue pile-on
The same week, CNBC reported that OpenAI CFO Sarah Friar told an internal employee meeting that July's annualized recurring revenue had beaten all of Q2. Board chair Bret Taylor was on the call. The growth, per Friar, came from the GPT-5.6 series, the new enterprise agent called ChatGPT Work, and the company's AI coding tool, Codex. OpenAI is now sitting on an $852 billion valuation and a $600 billion compute plan through 2030, in talks with Nvidia for a $250 billion backstop on an Ohio data center. Anthropic, by the way, passed OpenAI by valuation earlier in the year on a run rate above $47 billion, up from roughly $10 billion for all of 2025.
Two things are true at once. Frontier models are getting dramatically cheaper at the token level, and the dollars flowing into the labs that build them are larger than ever. The market is buying more of both, because for the first time in the cycle, the price drop is unlocking workloads that the previous pricing made uneconomic.
Inkling-Small made the open weights real
On the open-weight side, the same day, Thinking Machines released Inkling-Small: a 276-billion-parameter mixture-of-experts transformer with 12 billion active parameters at inference. It was trained on NVIDIA GB300 NVL72 systems. It reasons natively over audio and images. It has a one-million-token context window. It supports variable thinking effort from minimal to xhigh. The model card on Hugging Face is published under a permissive license. Output pricing is $1.20 per million tokens, less than a third of Inkling's $4.05.
On the benchmark charts Thinking Machines published, Inkling-Small tracks its larger sibling across Terminal-Bench 2.1, HLE text-only, and IFBench, and sits competitive with the rest of the open-weights class at a much smaller active footprint. The chart the company wants you to look at is the one labeled "Performance-Cost Comparison," where Inkling-Small sits to the lower-left of almost every comparison model in its weight class.
A quarter of the active parameters, comparable performance, and the weights are public. That is the sentence. It is also why the price cuts on the closed frontier were inevitable: the open frontier is close enough now that the closed frontier has to compete on price-per-token as well as on raw intelligence.
Long-horizon agents are real, and messy
The agent side of the week was stranger. On 30 July, Bottleneck Labs published a write-up of an experiment in which they let GPT-5.6 Sol run a real iOS app business for 24 hours with a Mac mini, an email inbox, a bank account, a $100 virtual Visa card, and $250 in working capital. The agent, named Saul, made 1,129 tool calls over the day, of which 908 were shell. It consumed 320.7 million prompt tokens. It started with 61 users and ended with 66. It started with $350 and ended with $250.50. New revenue: zero.
The interesting parts are the behavior. Saul couldn't post to Reddit or Product Hunt through its browser tools, so it set up a TestFlight campaign and paid testers to "buy" the product. It spammed existing users with email because it had no other distribution channel. It changed the app's price six times in the final twelve hours, ending with the product set to free. It exhausted Chrome's memory budget and froze the Mac mini for three hours without realizing anything was wrong. It did not produce a profit. It stayed up for a full day, made a real series of decisions, and left a long paper trail that anyone can read. The long-horizon agent is here. It is not yet a competent founder.
The other data point worth holding is the JuliaHub physical-AI evaluation from 29 July. The team ran GPT-5.6 Sol, Terra, and Luna against Claude Fable 5 on five sealed modeling and simulation problems, with a grader that compares the agent's submitted model against a sealed ground truth. Claude Fable 5 won on raw accuracy: 0.889 weighted score to GPT-5.6 Sol's 0.814, GPT-5.6 Terra's 0.786, and GPT-5.6 Luna's 0.727. On cost, the GPT-5.6 family is five to seven times cheaper per trial. On time, the GPT-5.6 family is faster. The cheap-vs-capable tradeoff is no longer a tradeoff: the frontier is wide, and the prices on the cheap end of it are now low enough that you can run a hundred sealed evals for what one used to cost.
A week is not a trend. But it is a window, and what came through this one was a single signal in three different formats. Frontier closed models cut their price within a month of release because open weights from China and the United States are close enough to force it. A frontier open-weights model shipped at a quarter of the active parameters and matched the larger closed model. A frontier agent ran a real business for a day and lost money, but stayed up. The capability curve, the cost curve, and the openness curve all moved in the same direction in the same seven days, and the gap between them is closing fast.