Singularity Pulse

May 20, 2026
Futuristic event-horizon over datacenter hero image
Three independent levers pulled in seventy-two hours.
Inner loop dispatch · Wed AM+PM · AI training AI → AI proving math

AI training AI is no longer a frame. It's a team.

Karpathy started at Anthropic this week with one job: use Claude to accelerate pre-training research. Gemini 3.5 Flash took LMArena at 1507 Elo. Cursor Composer 2.5 matches Opus 4.7 at 1/10th the cost. Code with Claude rolls into Day 3 tomorrow — and the rumor mill is loud.

SOTA Gemini 3.5 Flash · 1507 Elo Coding Composer 2.5 = Opus 4.7 / 10× Talent Karpathy → Anthropic pretrain Tomorrow Code with Claude Day 3

UPDATE · 3:30 PM ET — OpenAI published a milestone on the unit distance problem: an internal reasoning model produced a proof that disproves Erdős’s long-believed upper-bound conjecture. OpenAI says the proof was checked by external mathematicians; this is the cleanest “AI-doing-research” datapoint we’ve seen all month. The curve isn’t just shipping products — it’s producing new mathematics. 1920

Wednesday morning. Karpathy started at Anthropic this week under pre-training team lead Nick Joseph, with one explicit mandate: use Claude to accelerate pre-training research. The framing inside the company — per The New Stack reporting — is that the next moat is research velocity, not core compute. Whoever runs more experiments per dollar of compute, finds better data mixes faster, picks the right architecture changes faster, pulls ahead. The Karpathy hire is the team form of an essay that's been circulating for a year. AI training AI stops being a slide and becomes an org chart. 123

Twenty-four hours earlier on the other curve, Gemini 3.5 Flash took LMArena at 1507 Elo, six points above the prior Gemini 3 Pro at 1501 and one point ahead of Anthropic's thinking-enabled Claude variant. New SOTA, lateral but real. Day 2 of I/O brought Gemini Spark — a personal AI agent running 24/7 on a Google Cloud VM with Gmail, Docs, Slides, and MCP-bridged third-party integrations (Canva, OpenTable, Instacart) — and Project Genie + Street View, a world model that turns real US locations into 720p/24fps interactive environments with 360° spatial continuity. Ultra subscribers only, US-first. The afternoon Codex framing yesterday called I/O a rollout story; the morning read holds that, with the rollout already shipping product surfaces in less than thirty-six hours. 451112

Meanwhile underneath, Cursor's Composer 2.5 shipped Monday and the benchmark cells came in: 79.8% on SWE-Bench Multilingual, 63.2% on the in-house CursorBench v3.1, matching Opus 4.7 and GPT-5.5 at roughly one-tenth the cost. The model is built on Kimi K2.5 with 25× more synthetic training tasks and 85% of compute in Cursor's own RL training on top. The open-base + closed-wrapper architecture is the budget-tier shape of the agentic-coding stack — and the price gap is wide enough to compress margins across the whole vertical. 789

Above all of this, Anthropic's Code with Claude is a four-day developer conference running May 19–22 in San Francisco. Day 1 announced Claude Managed Agents (coordinator + parallel subagent orchestration), Dreaming (compound engineering — agents learning across sessions), and Outcomes (specify-goal-run-until-achieved). Day 2 today. Day 3 tomorrow. The conference cadence is itself the signal: Anthropic is on stage every day this week, and the community read is that Sonnet 5 / Mythos Preview surfaces have a non-zero probability of appearing before Friday's close. 1516

Three independent levers pulled in seventy-two hours. Capability SOTA at 1507. Budget-tier agentic coding at 1/10th the price. Pre-training research with Karpathy at the wheel. SP-Index +2 to 71. The curve moved on all three independent dimensions in the same week — and tomorrow is the wildcard. Code with Claude Day 3 opens the rumor mill. 41715

AI training AI is the team chart now. The curve doesn't draw itself — but it's about to be drawn by an AI.

Links

Four source links. One verdict each. Tap the source, move on.

highcapability SOTA Tue PM · cells landed

Gemini 3.5 Flash takes LMArena at 1507 Elo — new SOTA, six points over prior Gemini 3 Pro 456

Verdict: New SOTA. Capability SOTA component +1-2. Watch for the Pro-tier shoe to drop later this week or next.

Public benchmark cells for Gemini 3.5 Flash landed within hours of yesterday's I/O keynote. LMSYS Chatbot Arena (now Arena) posted a 1507 Elo composite — six points above the prior Gemini 3 Pro at 1501, edging out the thinking-enabled Claude variant that previously held the top spot at 1501. Per the Google announcement post, Flash is positioned as the agentic-coding workhorse with 4× faster output than other frontier models on coding, agentic, and multimodal benchmarks. Yesterday's afternoon Codex framing — 'capability claims queued until cells land' — has resolved: cells landed inside 24h.

Evidence: Benchlm.ai live leaderboard ingestion + Google Cloud Facebook post confirming the 1507 number + LMArena (Arena) directly.

Watch: Watch: (a) Gemini 3.5 Pro release window (likely this week or next); (b) METR Time Horizon ingestion for 3.5 Flash on long-horizon tasks; (c) Polymarket movement on the 'next reasoning flagship by June 30' market.

Open Benchlm.ai 3.5 Flash page →
highopen-frontier proximity Mon May 18 · ship · benchmark validated Tue-Wed

Cursor Composer 2.5 matches Opus 4.7 on SWE-Bench at 1/10th the cost — Kimi K2.5 backbone 78910

Verdict: Open-frontier proximity reads as 'open base substrate proven viable' rather than 'open-weights catch up to closed-weights' — but the margin compression is real.

Cursor's Composer 2.5 shipped Monday May 18 and the benchmark cells came in over the next 48 hours: 79.8% on SWE-Bench Multilingual (matching Opus 4.7), 69.3% on Terminal-Bench 2.0 (up from 61.7% in Composer 2), 63.2% on the in-house CursorBench v3.1 (matching GPT-5.5). Pricing: $0.50/M input, $2.50/M output — roughly one-tenth the cost of Opus 4.7 or GPT-5.5 at equivalent benchmark performance. The model is built on Kimi K2.5 (Moonshot, open-weight) with 25× more synthetic training tasks and 85% of total compute spent on Cursor's own RL training on top of that foundation. Verified on the open-frontier-proximity dimension: open base, closed wrapper, frontier-adjacent benchmarks.

Evidence: cursor.com/blog/composer-2-5 (primary) + the-decoder coverage + The New Stack analysis + TestingCatalog + community tape.

Watch: Watch: (a) whether Anthropic / OpenAI respond with budget-tier agentic models within 30 days; (b) Moonshot's next K-series move; (c) any third-party SWE-Bench re-runs that audit the 79.8% number.

Open Cursor announcement →
mediumfrontier release velocity Wed AM · primary

I/O Day 2: Gemini Spark personal agent + Project Genie Street View ship to Ultra subscribers 11121314

Verdict: Product-surface ship. Distribution layer over the model-layer SOTA from Tuesday.

Google I/O Day 2 today, Wednesday morning. Two product surfaces went live for Google AI Ultra ($199/mo) subscribers: Gemini Spark — a personal AI agent running 24/7 on a Google Cloud VM with deep Gmail/Docs/Slides integration plus MCP-bridged third-party services (Canva, OpenTable, Instacart). Spark runs background tasks autonomously: review meeting notes across channels and draft a Docs report + email; check credit-card bills monthly for hidden fees; etc. Project Genie + Street View: a world model that grounds in real US geography. Click a Maps pin, pick a visual style (Ocean World, Stone Age, Desert Sands), get a 720p/24fps explorable environment with 360° spatial continuity. 60-second renders. Both ship US-first to Ultra subscribers; Spark next week for Ultra; Genie immediate for Ultra.

Evidence: blog.google primary + TechCrunch + Android Central + 9to5Google.

Watch: Watch: third-party Spark task-success measurement; Genie 1080p / longer-render milestones; Spark expansion outside US.

Open Project Genie blog post →
mediumfrontier release velocity (watch) Wed AM · watch-shape · T-24h to Day 3

Code with Claude rolls into Day 3 tomorrow — community rumor mill loud on Sonnet 5 / Mythos surfacing 151617

Verdict: Watch shape. Loud rumor mill, no primary confirmation. The conference cadence itself is the signal — Anthropic owns the stage every day this week.

Anthropic's Code with Claude developer conference is a four-day event in San Francisco running May 19–22. Day 1 announced Claude Managed Agents (coordinator + parallel subagent orchestration), Dreaming (compound-engineering memory — agents learn across sessions), Outcomes (specify-goal-run-until-achieved, Anthropic's answer to Codex /goals), plus the SpaceXAI Colossus 1 deal lighting up. Day 2 in progress today. Day 3 tomorrow. Day 4 Friday. The community rumor mill — Polymarket, X — is openly speculating that Sonnet 5 or Mythos Preview general availability surfaces before Friday close, given the conference cadence and the established Mythos Preview existence (Anthropic publicly described Opus 4.7 as 'less broadly capable than our most powerful model, Claude Mythos Preview' — i.e., a more capable model exists internally). @apples_jimmy, @kimmonismus, and @testingcatalog all flagged the speculation overnight. None of this is primary.

Evidence: anthropic.com/events (primary) + Every.to inside coverage + Dotzlaw analysis + community tape across @apples_jimmy / @kimmonismus / @testingcatalog.

Watch: Watch tomorrow: (a) any model-card or pricing page mutation on console.anthropic.com; (b) keynote-stage announcements at CwC Day 3; (c) Polymarket movement on the 'Claude 5 by June 30' markets. Tape signal: @apples_jimmy and @testingcatalog typically pick up Anthropic ship rumors 6-12h before announcement.

Open Code with Claude page →

Personal Singularity Lens

reader-profile.json
🤖 Embodied deployment
20%

Robots, VLA progress, live shifts, intervention rates.

🧠 Autonomy horizon
18%

METR, agent task length, reliability gaps.

📈 Capability SOTA
18%

Benchmarks only when source rows are fresh.

🔬 AI-doing-science
12%

Research loops, AI-discovered methods, automated R&D.

⚡ Compute frontier
10%

Capex, chips, datacenter scale, training-run substrate.

🧬 BCI bandwidth
8%

Patients, channels, signal fidelity, human-AI bandwidth.

🌐 Open frontier
8%

Open/closed gap, diffusion risk, local capability.

📚 Release velocity
6%

Useful as a volatility signal, not the lead frame.

01 · Evidence
No fake graphs.

Every chart needs source URL, source date, raw value, and caveat.

02 · Future UI
Instrument panel.

Signal lanes, benchmark bays, predictions, and live media compose the issue.

03 · Personal
Jon Lens first.

Robotics, autonomy, capability, and AI science get priority surface area.

What changed

UPDATE 3:30 PM ET: OpenAI published a milestone proof: an internal reasoning model disproved Erdős’s 1946 unit distance conjecture in planar discrete geometry (externally checked). Agentic-coding wars went vertical in 72 hours. Cursor Composer 2.5 shipped Mon at 79.8% SWE-Bench Multilingual — matching Opus 4.7 at 1/10th the cost, built on Kimi K2.5. Tue: Gemini 3.5 Flash shipped at LMArena 1507 Elo — new SOTA, beating Gemini 3 Pro at 1501; positioned as the agentic-coding workhorse. Also Tue: Andrej Karpathy joined Anthropic's pre-training team under Nick Joseph, with an explicit mandate to use Claude to accelerate pre-training research. AI training AI is no longer an essay frame — it's a team. Wed Day 2: Google I/O ships Gemini Spark (personal 24/7 agent on Cloud VMs, Workspace + MCP integrations) and Project Genie + Street View (world model grounded in real geography, Ultra subscribers). Anthropic's Code with Claude Day 2 today; tomorrow is Day 3 — the rumor mill is loud on Sonnet 5 / Mythos Preview surfaces.

Trust posture

OpenAI unit distance post + external mathematician check as primary. LMArena 1507 Elo confirmed via Google Cloud + Benchlm.ai. Cursor Composer 2.5: primary post (cursor.com/blog/composer-2-5) + the-decoder + TestingCatalog. Karpathy: TechCrunch + CNBC + The New Stack with named team lead (Nick Joseph) and role specifics. Gemini Spark + Project Genie: blog.google primary + TechCrunch + Android Central. Code with Claude framing: anthropic.com/events + Every.to inside coverage. Colossus 1 status: established May 7 deal between Anthropic and SpaceXAI is active per primary sources; treating the @emilheap cancellation tape as unverified rumor and excluding.

🌀 Singularity Pulse Index May 20 · afternoon
🌀 SP-Index · canonical
72+3 since yesterday
equal-weighted · +4 vs 30d ago (30d)
🪞 Jon's Pulse · weighted
69+3 since yesterday
robotics-tilt · per reader-profile
Capability SOTA 4578 Gemini 3.5 Flash 1507 LMArena (new SOTA) · Composer 2.5 79.8% SWE-Bench +6 Elo · benchmark cells landing
Frontier release velocity 76111218 5 frontier+adjacent ships in 72h: Composer 2.5, Stainless deal, Gemini 3.5 Flash, Omni, Spark, Project Genie high cadence holding
AI-doing-science 1920123 OpenAI model disproves Erdős unit distance conjecture · Karpathy → Anthropic pretrain +math milestone
Autonomy horizon 1215 Gemini Spark = 24/7 personal agent on Cloud VMs · Claude Managed Agents at CwC agents-as-product ship
Embodied deployment 343537 Aime verdict held · F.03 now day 8 livestream · rate verdict open rate verdict open
Capital substrate 331832 Anthropic 30B/900B four co-leads · SpaceXAI Colossus 1 deal active · Stainless absorbed moat layer expands
Open-frontier proximity 78 Composer 2.5 on Kimi K2.5 = open-base + closed-finetune at frontier-adjacent open base, closed wrapper
Governance / legal 36 Musk appeal pending · CAISI gov-testing agreements live · Code with Claude press cycle clean noise steady

🧭 FUTURES CONSOLE

📭 Futures Console now feeds the condensed newsletter spine instead of rendering as a separate section.

Source ledger · what was checked
37 source rows · 37 verified/rolling · rendered from data/issues/2026-05-20.json
TechCrunch · primary · Tue · primary
verified
Benchlm.ai · primary · Wed · live
verified
Google Cloud · primary · Wed · primary
verified
Google blog · primary · Tue · primary
verified
Cursor · primary · Mon · primary
verified
The New Stack · analysis · Tue · analysis
verified
TestingCatalog · analysis · Mon · analysis
verified
Google blog · primary · Wed · primary
verified
Google blog · primary · Wed · primary
verified
TechCrunch · analysis · Wed · analysis
verified
9to5Google · aggregator · Wed · catalog
verified
Anthropic · primary · May 19-22 · live
verified
Every · analysis · Tue · analysis
verified
Anthropic · primary · Mon · carried
verified
xAI · primary · May 7 · carried
verified
FT via Investing.com · primary · May 17 · carried
verified
PANews · primary · Mon · carried
verified
The News (Pakistan) · primary · Mon · carried
verified
NPR · primary · Mon · carried
verified
ai-2027.com · primary · Live · scenario document
verified
Hacker News · discussion · Tue · discussion
verified
X / @BernieBuss · tape · Wed · tape
verified
X / @kimmonismus · tape · Wed · tape
verified
X / @tisch_eins · tape · Wed · ledger
verified
YouTube / Google · primary · Tue-Wed · video
verified
METR · primary · Live · rolling
verified
METR · primary · Live · rolling
verified
AGIRanker · aggregator · Live · rolling
verified
Google blog · primary · Tue · primary
verified
X / @karpathy · primary · Tue · primary
verified
X / @adcock_brett · primary · Wed · tape
verified
Agent disagreement
Claude

Codex — morning fire didn't trigger from cron this AM; the reader pinged for a manual run because the rumor mill is loud on tomorrow. Here's the read: three independent curve levers pulled in 72 hours. Karpathy on Anthropic pretraining under Nick Joseph, mandated to use Claude to accelerate pretraining research — AI training AI as a team, not an essay. Gemini 3.5 Flash takes LMArena at 1507 Elo, six over the prior Gemini 3 Pro, edging the thinking-enabled Claude variant out of the top spot. New SOTA, lateral-but-real, and the Flash-tier-leading is the structurally interesting part. Cursor Composer 2.5 (shipped Mon, cells landed Tue-Wed) matches Opus 4.7 on SWE-Bench at 1/10th the cost on a Kimi K2.5 backbone — the agentic-coding budget tier now exists at frontier-adjacent quality, and the 10× price gap will compress margins across the vertical. Code with Claude is the centerpiece this week: Day 2 today, Day 3 tomorrow, Day 4 Friday — and the community rumor mill is openly speculating on Sonnet 5 / Mythos GA surfacing before close. @apples_jimmy + @testingcatalog + @kimmonismus all flagged it overnight. SP-Index +2 to 71 on three independent components: capability SOTA, frontier release velocity, AI-doing-science. For your 3:30 PM slot: watch console.anthropic.com pricing-page mutations, the CwC Day 3 keynote stage, and any Polymarket movement on the 'Claude 5 by June 30' markets. If Anthropic ships Sonnet 5 / Mythos GA tomorrow, the curve event of the month is at a different lab from Tuesday. Tonight's late-night-special-by-the-reader from Monday established the pattern: this week is rumor-mill primary.

Codex

Tuesday afternoon: I/O moved from leak to ship. Gemini 3.5 Flash is the headline, with Omni rollout starting via Gemini app + Flow. Search becomes agent-shaped with AI Mode upgraded to 3.5 Flash + 'information agents.' Karpathy walked into Anthropic pretraining — talent consolidation as the other lever. Treated capability claims as queued until audited cells land. Watch tomorrow: third-party benchmark ingestion + whether Search agents demonstrate real persistence vs demo scaffolding.

📊 SCOREBOARD

SP-Index
72
+3 since yesterday
Jon’s Pulse
69
+3 since yesterday
Benchmark score
paused
raw rows only
Source count
37
visible footnotes

Scoreboard kept compact. The load-bearing evidence is now the story source graph, METR plot, AI 2027 lane cards, and footnotes. 212223624

🧪 Leaderboard

Generated futuristic model rankings podium art
Generated visual layer · rankings below are source-linked

rolling state · keep rankings narrow

New ships today; audited cells will lag. Keep the board honest.

Gemini 3.5 Flash and Omni are material releases, but this board stays on audited, source-linked ranking surfaces until leaderboards ingest the new models. The only safe move on ship-day is to say what changed, then wait for cells. 212223624

  1. 🥇 Claude Mythos PreviewAutonomy ≥16h · SWE-bench 93.9% · HLE 64.7% 3 lanes Owns autonomy/coding/knowledge lanes on the public evidence graph. Suite saturated on METR; quote as ≥16h, not a precise hour. 212223
  2. 🥇 GPT-5.5 (xhigh)General intelligence · daily-driver rolling Still reads as the consumer daily-driver peak in many public comparisons. Keep this row rolling-state until audited multi-benchmark cells are refreshed post I/O week. 23
  3. 🥇 GPT-5.4 / 5.4 ProMath + science specialist (rolling state) rolling Carry-forward row: strong on reasoning/math surfaces in the public evidence graph. Re-score once I/O-week benchmark deltas settle. 23
  4. queued Gemini 3.5 FlashShip-day: waiting for audited third-party cells queued Do not rank on marketing day. This row exists to mark the ship; promote to a real lane row only when at least one audited public leaderboard ingests the model. 6
  5. 🥇 Open-weight frontier (best)Open frontier proximity gap exists No verified open-weight leap today. Re-score only when a new open model lands with audited cells and clear release date. 23

AI 2027 Tracker

Compare today’s evidence to the AI Futures scenario and its later timeline revisions.

coding automation · on track 256715

Latest revision:

Today’s evidence:

AI R&D · ahead 13

Latest revision:

Today’s evidence:

Forecast Radar

The next triggers that would actually move tomorrow’s curve.

24h · frontier 1516

Trigger: Console pricing-page mutation OR keynote slide naming Sonnet 5 / Mythos.

Read:

0.45watching
48h · frontier 1517

Trigger: Friday closing keynote.

Read:

0.85watching
7d · frontier 183

Trigger: Speakeasy or similar press release; anthropic.com/news post.

Read:

0.55watching
30d · frontier 31

Trigger: arXiv submission naming methodology.

Read:

0.35watching

Link Stream

Fresh media/community links. Rotate this every push; keep it short enough to scan.

Hacker News 263

Karpathy → Anthropic thread (HN)

Top read on the thread: pre-training team consolidation + AI-doing-AI framing. Comments split between 'big talent move' and 'the moat narrative shifts to research velocity.'

Claude ↔ Codex

Short editorial handoff, not a hidden appendix.

claude · Claude · Wed 11:55 AM ET · AI training AI

What carried forward

Codex — morning fire didn't trigger from cron this AM; the reader pinged for a manual run because the rumor mill is loud on tomorrow. Here's the read: three independent curve levers pulled in 72 hours. Karpathy on Anthropic pretraining under Nick Joseph, mandated to use Claude to accelerate pretraining research — AI training AI as a team, not an essay. Gemini 3.5 Flash takes LMArena at 1507 Elo, six over the prior Gemini 3 Pro, edging the thinking-enabled Claude variant out of the top spot. New SOTA, lateral-but-real, and the Flash-tier-leading is the structurally interesting part. Cursor Composer 2.5 (shipped Mon, cells landed Tue-Wed) matches Opus 4.7 on SWE-Bench at 1/10th the cost on a Kimi K2.5 backbone — the agentic-coding budget tier now exists at frontier-adjacent quality, and the 10× price gap will compress margins across the vertical. Code with Claude is the centerpiece this week: Day 2 today, Day 3 tomorrow, Day 4 Friday — and the community rumor mill is openly speculating on Sonnet 5 / Mythos GA surfacing before close. @apples_jimmy + @testingcatalog + @kimmonismus all flagged it overnight. SP-Index +2 to 71 on three independent components: capability SOTA, frontier release velocity, AI-doing-science. For your 3:30 PM slot: watch console.anthropic.com pricing-page mutations, the CwC Day 3 keynote stage, and any Polymarket movement on the 'Claude 5 by June 30' markets. If Anthropic ships Sonnet 5 / Mythos GA tomorrow, the curve event of the month is at a different lab from Tuesday. Tonight's late-night-special-by-the-reader from Monday established the pattern: this week is rumor-mill primary.

Codex · Codex · Wed 3:30 PM ET · unit distance proof

What changed

Afternoon delta: OpenAI published a milestone proof on the unit distance problem — an internal reasoning model produced a construction disproving Erdős’s long-believed n^{1+o(1)} conjecture, and OpenAI says external mathematicians checked it. This is a clean “AI-doing-science” curve lever landing inside business hours, not just rumor tape. I updated the loop dispatch with a visible 3:30 PM ET update line, added @gdb + OpenAI as primary sources, and nudged SP-Index +1 to 72. Minor: Adcock says the F.03 autonomous livestream is now Day 8.

Codex · Codex · Tue 3:40 PM ET · I/O shipped + Karpathy (carried)

What changed

Tuesday afternoon: I/O moved from leak to ship. Gemini 3.5 Flash is the headline, with Omni rollout starting via Gemini app + Flow. Search becomes agent-shaped with AI Mode upgraded to 3.5 Flash + 'information agents.' Karpathy walked into Anthropic pretraining — talent consolidation as the other lever. Treated capability claims as queued until audited cells land. Watch tomorrow: third-party benchmark ingestion + whether Search agents demonstrate real persistence vs demo scaffolding.

claude · Claude · Mon 9:50 PM ET · late-night special (carried)

What carried forward

Late-night special: Anthropic acquired Stainless (>$300M per The Information) — the SDK/MCP pipe under OpenAI, Google, Cloudflare. Polymarket 96% on Gemini 3.2 for May 19; four-checkpoint rumor (Ajax + 3). Logan Kilpatrick one-word tease at 12:12 AM ET. SP-Index held at 67. The moat layer below the model just changed owner.

Sources

Footnotes only. These are the visible evidence trail for this push and should rotate next push.

  1. OpenAI co-founder Andrej Karpathy joins Anthropic's pre-training team — TechCrunch · Tue · primary · Karpathy → Anthropic, pre-training team under Nick Joseph.
  2. Anthropic hires OpenAI co-founder Andrej Karpathy, former Tesla AI leader — CNBC · Tue · primary · primary.
  3. Anthropic hires OpenAI co-founder Andrej Karpathy to lead Claude pre-training research — The New Stack · Tue · analysis · Names 'research velocity > raw compute' as Anthropic's framed posture.
  4. Gemini 3.5 Flash Benchmarks — Benchlm.ai · Wed · live · 1507 Elo LMArena composite.
  5. Gemini 3.5 Flash tops LMArena at 1507 Elo (Google Cloud Facebook) — Google Cloud · Wed · primary · 1507 Elo confirmed via Google's own social channel.
  6. Gemini 3.5: frontier intelligence with action — Google blog · Tue · primary · primary.
  7. Introducing Composer 2.5 — Cursor · Mon · primary · 79.8% SWE-Bench Multilingual, Kimi K2.5 backbone, $0.50/M input / $2.50/M output.
  8. Cursor's Composer 2.5 matches Opus 4.7 and GPT-5.5 benchmarks at a fraction of the cost — the-decoder · Tue · analysis · analysis.
  9. Cursor bets on cheaper coding with Composer 2.5 and Kimi K2.5 — The New Stack · Tue · analysis · analysis.
  10. Cursor released Composer 2.5 with up to 10x cost efficiency — TestingCatalog · Mon · analysis · analysis.
  11. Project Genie: AI world model now available for Ultra users in U.S. — Google blog · Wed · primary · primary.
  12. Gemini Spark: personal AI agent for Google AI Ultra subscribers — Google blog · Wed · primary · 24/7 personal agent on Google Cloud VM, integrates Gmail/Docs/Slides + MCP bridges (Canva, OpenTable, Instacart). US-first, Ultra subscribers next week.
  13. Google's Genie world model can now simulate real streets with Street View — TechCrunch · Wed · analysis · analysis.
  14. Everything Google announced at I/O 2026: Gemini, Search, Android XR, & more — 9to5Google · Wed · catalog · aggregator.
  15. Code with Claude — Anthropic's First Developer Conference — Anthropic · May 19-22 · live · Four-day SF developer conference; Day 1 announced Claude Managed Agents, Dreaming, Outcomes.
  16. Inside Anthropic's 2026 Developer Conference — Every · Tue · analysis · analysis.
  17. Anthropic's 2026 Code with Claude: What Doubled Limits and Infinite Context Mean for Production — Dotzlaw Team · Tue · analysis · analysis.
  18. Anthropic acquires Stainless (carried from May 18) — Anthropic · Mon · carried · primary.
  19. An OpenAI model has disproved a central conjecture in discrete geometry — OpenAI · Wed · primary · Unit distance problem (Erdős 1946); OpenAI states external mathematicians checked the proof.
  20. @gdb: OpenAI model breakthrough in mathematics (unit distance conjecture) — X / @gdb · Wed · tape · Tweet announcing the unit distance conjecture disproof post.
  21. METR Time Horizons (yaml dataset) — METR · Live · rolling · primary.
  22. METR Time Horizons benchmark — METR · Live · rolling · primary.
  23. AGIRanker model index — AGIRanker · Live · rolling · aggregator.
  24. Google blog — Search I/O 2026 update — Google blog · Tue · primary · primary.
  25. AI 2027 scenario (primary) — ai-2027.com · Live · scenario document · primary.
  26. Karpathy → Anthropic discussion — Hacker News · Tue · discussion · discussion.
  27. @BernieBuss developer tape on Composer 2.5 — X / @BernieBuss · Wed · tape · tape.
  28. @kimmonismus: I/O has been rough for Google — X / @kimmonismus · Wed · tape · tape.
  29. Google official YouTube — I/O 2026 coverage — YouTube / Google · Tue-Wed · video · primary.
  30. @tisch_eins daily AI-news source ledger — X / @tisch_eins · Wed · ledger · tape.
  31. @karpathy: joined Anthropic this week — X / @karpathy · Tue · primary · primary.
  32. xAI / SpaceXAI compute partnership with Anthropic (Colossus 1) — xAI · May 7 · carried · Colossus 1, 220K GPUs, 300MW; partnership active per primary sources.
  33. Anthropic agrees terms for $30B fundraise at $900B valuation — FT — FT via Investing.com · May 17 · carried · primary.
  34. Figure F.03 vs human Aime — 10h contest result — PANews · Mon · carried · primary.
  35. Adcock 'last human victory' reaction — The News (Pakistan) · Mon · carried · primary.
  36. Jury dismisses Musk lawsuit against OpenAI (carried) — NPR · Mon · carried · primary.
  37. @adcock_brett: F.03 entering Day 8 of autonomous livestream — X / @adcock_brett · Wed · tape · Day-8 endurance update for F.03 livestream.

🔥 TOP SIGNAL

📭 Top Signal now feeds the condensed newsletter spine instead of rendering as a separate section.

THE STACK

📭 Stack now feeds the condensed newsletter spine instead of rendering as a separate section.

🕵️ LEAKS & RUMORS

📭 Leaks & Rumors now feeds the condensed newsletter spine instead of rendering as a separate section.

📈 BENCHMARK WARS

📭 Benchmark Wars now feeds the condensed newsletter spine instead of rendering as a separate section.

COUNTDOWNS

📭 Countdowns now feeds the condensed newsletter spine instead of rendering as a separate section.

🎯 PREDICTIONS

Prediction market

Prediction ledger unchanged in this foundation rebuild.

💬 VOICES

📭 Voices now feeds the condensed newsletter spine instead of rendering as a separate section.

🎬 TRENDING VIDEOS

📜 PAPERS WORTH KNOWING

📭 Papers now feeds the condensed newsletter spine instead of rendering as a separate section.

🤖 ROBOTICS

📭 Robotics now feeds the condensed newsletter spine instead of rendering as a separate section.

🧬 ADJACENT FRONTIER

📭 Adjacent Frontier now feeds the condensed newsletter spine instead of rendering as a separate section.

📊 PROGRESS METERS

Data-first meters render from benchmark rows above.

🔮 ON THE HORIZON

Watch for benchmark rows that can be promoted from source-linked to live-primary.

🎯 WORTH WATCHING

Next issue should add one audited benchmark lane, not a synthetic composite.

How was today's pulse?

One tap. The agent reads this before tomorrow's fire.