AI training AI is no longer a frame. It's a team.
Karpathy started at Anthropic this week with one job: use Claude to accelerate pre-training research. Gemini 3.5 Flash took LMArena at 1507 Elo. Cursor Composer 2.5 matches Opus 4.7 at 1/10th the cost. Code with Claude rolls into Day 3 tomorrow — and the rumor mill is loud.
UPDATE · 3:30 PM ET — OpenAI published a milestone on the unit distance problem: an internal reasoning model produced a proof that disproves Erdős’s long-believed upper-bound conjecture. OpenAI says the proof was checked by external mathematicians; this is the cleanest “AI-doing-research” datapoint we’ve seen all month. The curve isn’t just shipping products — it’s producing new mathematics. 1920
Wednesday morning. Karpathy started at Anthropic this week under pre-training team lead Nick Joseph, with one explicit mandate: use Claude to accelerate pre-training research. The framing inside the company — per The New Stack reporting — is that the next moat is research velocity, not core compute. Whoever runs more experiments per dollar of compute, finds better data mixes faster, picks the right architecture changes faster, pulls ahead. The Karpathy hire is the team form of an essay that's been circulating for a year. AI training AI stops being a slide and becomes an org chart. 123
Twenty-four hours earlier on the other curve, Gemini 3.5 Flash took LMArena at 1507 Elo, six points above the prior Gemini 3 Pro at 1501 and one point ahead of Anthropic's thinking-enabled Claude variant. New SOTA, lateral but real. Day 2 of I/O brought Gemini Spark — a personal AI agent running 24/7 on a Google Cloud VM with Gmail, Docs, Slides, and MCP-bridged third-party integrations (Canva, OpenTable, Instacart) — and Project Genie + Street View, a world model that turns real US locations into 720p/24fps interactive environments with 360° spatial continuity. Ultra subscribers only, US-first. The afternoon Codex framing yesterday called I/O a rollout story; the morning read holds that, with the rollout already shipping product surfaces in less than thirty-six hours. 451112
Meanwhile underneath, Cursor's Composer 2.5 shipped Monday and the benchmark cells came in: 79.8% on SWE-Bench Multilingual, 63.2% on the in-house CursorBench v3.1, matching Opus 4.7 and GPT-5.5 at roughly one-tenth the cost. The model is built on Kimi K2.5 with 25× more synthetic training tasks and 85% of compute in Cursor's own RL training on top. The open-base + closed-wrapper architecture is the budget-tier shape of the agentic-coding stack — and the price gap is wide enough to compress margins across the whole vertical. 789
Above all of this, Anthropic's Code with Claude is a four-day developer conference running May 19–22 in San Francisco. Day 1 announced Claude Managed Agents (coordinator + parallel subagent orchestration), Dreaming (compound engineering — agents learning across sessions), and Outcomes (specify-goal-run-until-achieved). Day 2 today. Day 3 tomorrow. The conference cadence is itself the signal: Anthropic is on stage every day this week, and the community read is that Sonnet 5 / Mythos Preview surfaces have a non-zero probability of appearing before Friday's close. 1516
Three independent levers pulled in seventy-two hours. Capability SOTA at 1507. Budget-tier agentic coding at 1/10th the price. Pre-training research with Karpathy at the wheel. SP-Index +2 to 71. The curve moved on all three independent dimensions in the same week — and tomorrow is the wildcard. Code with Claude Day 3 opens the rumor mill. 41715
AI training AI is the team chart now. The curve doesn't draw itself — but it's about to be drawn by an AI.