On the wire, the labs are mid-apology.
Issue #5. The frontier coding agents wobbled simultaneously this week. Tibo is somewhere asleep with a queued reset. Boris is explaining database contention. The curve bends through the meltdown anyway.
Friday, late. Tibo Sottiaux — Codex team lead at OpenAI — ships a fix and resets rate limits as the apology for forty-eight hours of degraded GPT-5.5 in the coding-agent endpoint. The community converges in real time on a tactical instruction: /fast /max — turn on fast mode, turn on the max plan, burn credits hard before the next reset cycle clips you. Saturday morning the meme is everywhere. Saturday morning, also, the other lab. Boris Cherny — Claude Code product lead at Anthropic — is forty-eight hours into explaining that the frequent Claude Code outages are growing pains: databases hitting limits, contention, the unsexy engineering reality. Not model regressions. The same week, the same shape, the same apology-reset arc playing on a different brand. Both labs are running the frontier-service equivalent of changing the engine while flying. The fact that they are pulling it off — barely, transparently, in public — is the actual signal. 124
Underneath the meltdown, the leaderboard is fragmenting. Claude Mythos Preview sits past the 16-hour autonomy ceiling that METR itself flags as unreliable — 17.41 hours at p50, three hours at p80 — alone at the top of a curve nobody has ever measured before. GPT-5.5 (xhigh) owns the daily-driver intelligence lane and the agentic-terminal lane. Claude Opus 4.7 owns multi-file code reasoning at 87.6% on SWE-bench and FrontierMath at 43.8%. Gemini 3.1 Pro owns multimodal and long-context. DeepSeek V4-Pro owns cost-performance. Five labs, five different #1s, one composite leaderboard nobody believes in. Meanwhile Cat Wu has reframed the product fight from metering to proactivity — what the agent decides to do without being asked — while Axios documents the same week as the end of free-feeling agent budgets. The customers are negotiating the price of agent initiative with the labs, in public, on X, in real time. 12131415
Structural moves under the noise. OpenAI Trusted Access for Cyber quietly went live May 13 with Deutsche Telekom, BBVA, Telefónica, Sophos, Scalable Capital, and the European Commission as inaugural partners — a same-shape mirror of Anthropic Project Glasswing from April 22. Cyber productization is a product category now, not an Anthropic frame. Same week, OpenAI launched a $4 billion Deployment Company — Tomoro acquisition, 150 Forward Deployed Engineers, 19 investors led by TPG — the enterprise-distribution mirror of Anthropic's stack. Same week, Anthropic and the Gates Foundation opened a four-year, $200 million non-revenue lane nobody else has matched yet. BBVA is named in two of these structures at once. One European bank, betting OpenAI on the cyber side and the deployment side, in the same week. 57811
Read it all at once: the agents wobble, the leaderboard fragments, the structural deals close, the curve bends anyway. Tibo and Boris are running the hardest job in AI right now: keeping a frontier service alive while shipping new features daily under a load profile nobody has ever produced before. Mythos at 17 hours p50 sits above the line that METR refuses to commit to. The frontier is past the workday at p50, past the credibility ceiling at peak, and the people building it are tweeting apologies on Saturday night because the load is real. The /fast /max meme is funny because it is also rational: when an apology cycle promises a reset, the smart move is to spend the reset before the next one. 1412
Saturday, late. Tibo's reset is queued. Boris's contention is being patched. Mythos keeps measuring.