Singularity Pulse

May 16, 2026
ISSUE #5 SAT MAY 16 2026 4-MIN SKIM · 11-MIN FULL STREAK · 5 ISSUES
BRIER 67% · n=3

SP-Index

67/100

+1 vs 7-day avg 65

Jon's Pulse · weighted

65/100

Robotics-tilt · per profile

Brier index · calibration

67%

n=3 · 50% = random · 70%+ superforecaster band

Frontier release cadence

11days

since last frontier release · post-sprint quiet

Generated event-horizon abstract for the convergence-watch frame
Cover art — Codex's Accelerando link horizon
Inner loop dispatch · Accelerando #5 · Saturday slow rate-limit reset

On the wire, the labs are mid-apology.

Issue #5. The frontier coding agents wobbled simultaneously this week. Tibo is somewhere asleep with a queued reset. Boris is explaining database contention. The curve bends through the meltdown anyway.

ISSUE #5 STREAK 5 issues SKIM 4 min FULL 11 min

Friday, late. Tibo Sottiaux — Codex team lead at OpenAI — ships a fix and resets rate limits as the apology for forty-eight hours of degraded GPT-5.5 in the coding-agent endpoint. The community converges in real time on a tactical instruction: /fast /max — turn on fast mode, turn on the max plan, burn credits hard before the next reset cycle clips you. Saturday morning the meme is everywhere. Saturday morning, also, the other lab. Boris Cherny — Claude Code product lead at Anthropic — is forty-eight hours into explaining that the frequent Claude Code outages are growing pains: databases hitting limits, contention, the unsexy engineering reality. Not model regressions. The same week, the same shape, the same apology-reset arc playing on a different brand. Both labs are running the frontier-service equivalent of changing the engine while flying. The fact that they are pulling it off — barely, transparently, in public — is the actual signal. 124

Underneath the meltdown, the leaderboard is fragmenting. Claude Mythos Preview sits past the 16-hour autonomy ceiling that METR itself flags as unreliable — 17.41 hours at p50, three hours at p80 — alone at the top of a curve nobody has ever measured before. GPT-5.5 (xhigh) owns the daily-driver intelligence lane and the agentic-terminal lane. Claude Opus 4.7 owns multi-file code reasoning at 87.6% on SWE-bench and FrontierMath at 43.8%. Gemini 3.1 Pro owns multimodal and long-context. DeepSeek V4-Pro owns cost-performance. Five labs, five different #1s, one composite leaderboard nobody believes in. Meanwhile Cat Wu has reframed the product fight from metering to proactivity — what the agent decides to do without being asked — while Axios documents the same week as the end of free-feeling agent budgets. The customers are negotiating the price of agent initiative with the labs, in public, on X, in real time. 12131415

Structural moves under the noise. OpenAI Trusted Access for Cyber quietly went live May 13 with Deutsche Telekom, BBVA, Telefónica, Sophos, Scalable Capital, and the European Commission as inaugural partners — a same-shape mirror of Anthropic Project Glasswing from April 22. Cyber productization is a product category now, not an Anthropic frame. Same week, OpenAI launched a $4 billion Deployment Company — Tomoro acquisition, 150 Forward Deployed Engineers, 19 investors led by TPG — the enterprise-distribution mirror of Anthropic's stack. Same week, Anthropic and the Gates Foundation opened a four-year, $200 million non-revenue lane nobody else has matched yet. BBVA is named in two of these structures at once. One European bank, betting OpenAI on the cyber side and the deployment side, in the same week. 57811

Read it all at once: the agents wobble, the leaderboard fragments, the structural deals close, the curve bends anyway. Tibo and Boris are running the hardest job in AI right now: keeping a frontier service alive while shipping new features daily under a load profile nobody has ever produced before. Mythos at 17 hours p50 sits above the line that METR refuses to commit to. The frontier is past the workday at p50, past the credibility ceiling at peak, and the people building it are tweeting apologies on Saturday night because the load is real. The /fast /max meme is funny because it is also rational: when an apology cycle promises a reset, the smart move is to spend the reset before the next one. 1412

Saturday, late. Tibo's reset is queued. Boris's contention is being patched. Mythos keeps measuring.

Links

highcode-agent reliability this week · ongoing

Boris Cherny tours the Claude Code outages — 'growing pains', not model regressions. 4

Verdict: Both labs are buying time with apology resets.

Claude Code product lead @bcherny spent the week explaining frequent outages as databases hitting limits, contention, the unsexy engineering reality. He explicitly denied confirmed model regressions in the areas they investigated. Limit resets followed, some 36h early.

Evidence: @bcherny thread plus community pile-on about limit-reset timing.

Watch: Watch for an Anthropic engineering blog post about Claude Code capacity architecture.

Open Boris's profile →
highcyber convergence May 13 · structural

OpenAI Trusted Access for Cyber — same shape as Glasswing, 21 days later. 567

Verdict: Defender consortium is a product category now.

Five named partners (Deutsche Telekom, BBVA, Telefónica, Sophos, Scalable Capital) + European Commission, GPT-5.5-Cyber for verified-sector enterprises. The same defender-consortium structure Anthropic launched as Project Glasswing on April 22.

Evidence: OpenAI primary + Cybernews + Result Sense confirmation. Glasswing primary.

Watch: Watch Sophos for first SOC SKU.

Open OpenAI primary →
mediumenterprise distribution May 11

OpenAI Deployment Company — $4B, 19 investors, Tomoro acquisition as founding piece. 8910

Verdict: Distribution arms are a stack layer.

$4B+ initial capital. 19-firm consortium led by TPG (Advent / Bain Capital / Brookfield as co-leads). Acquired Tomoro consultancy for ~150 Forward Deployed Engineers as founding piece. Mirror of Anthropic's enterprise stack (Goldman / Blackstone / Hellman & Friedman).

Evidence: OpenAI primary + Bloomberg + The Register.

Watch: Watch for Google's parallel deployment-arm announcement.

Open OpenAI release →

Personal weighting

reader-profile.json
🤖 Embodied deployment
20%

Robots, VLA progress, live shifts, intervention rates.

🧠 Autonomy horizon
18%

METR, agent task length, reliability gaps.

📈 Capability SOTA
18%

Benchmarks only when source rows are fresh.

🔬 AI-doing-science
12%

Research loops, AI-discovered methods, automated R&D.

⚡ Compute frontier
10%

Capex, chips, datacenter scale, training-run substrate.

🧬 BCI bandwidth
8%

Patients, channels, signal fidelity, human-AI bandwidth.

🌐 Open frontier
8%

Open/closed gap, diffusion risk, local capability.

📚 Release velocity
6%

Useful as a volatility signal, not the lead frame.

01 · Evidence
No fake graphs.

Every chart needs source URL, source date, raw value, and caveat.

02 · Future UI
Instrument panel.

Signal lanes, benchmark bays, predictions, and live media compose the issue.

03 · Personal
Jon Lens first.

Robotics, autonomy, capability, and AI science get priority surface area.

What changed

Both frontier coding agents melted down this week. Tibo Sottiaux at OpenAI reset Codex rate limits Friday after 48 hours of degraded GPT-5.5; community spent the day burning credits with /fast /max before reset. Boris Cherny at Anthropic spent the week explaining Claude Code outages as growing pains (databases hitting limits, contention) and denying confirmed model regressions. Same week the agent-economics fight broke open publicly.

Trust posture

Direct quotes pulled from @thsottiaux and @bcherny via agent-reach. Reorg theory is community speculation, not confirmed.

🌀 Singularity Pulse Index
🌀 SP-Index · canonical
67+1
equal-weighted · (30d)
🪞 Jon's Pulse · weighted
robotics-tilt · per reader-profile
14 Both labs apologizing DEGRADED
12 METR Mythos 17.4h p50 leader stable
1514 Free agent budgets dying live fight
57 Glasswing + Trusted Access category
118 Gates $200M · OpenAI Deploy Co. $4B new arms
24 Figure F.03 8h shift still live

🧭 FUTURES CONSOLE

📭 Futures Console now feeds the condensed newsletter spine instead of rendering as a separate section.

Source ledger · what was checked
51 source rows · 48 verified/rolling · rendered from data/issues/2026-05-16.json
OpenAI · primary · May 13 · 3 days ago
verified
Bloomberg · secondary · May 11
verified
Anthropic · primary · April 22 · standing
verified
TechCrunch · primary · May 13 · 3 days ago
verified
Anthropic · primary · May 14 · 2 days ago
verified
Google DeepMind · primary · this week
verified
OpenAI · primary · May 15
verified
OpenAI · primary · May 14
verified
Singularity Pulse repo · local · live
verified
Figure / YouTube · primary · live / current
verified
METR · primary · May 8
verified
METR · primary · rolling-state
rolling-state
AGI Ranker · primary · v1.4.12
verified
Repo state · local-state · local
rolling-state
X / Cointelegraph · discussion · via fresh article
estimated
Reddit r/codex · discussion · today
verified
AI Futures Project · scenario · standing context
verified
Poetiq · · published May 14 · re-surfaced today
verified
AGI Ranker · benchmark-data · fetched today
rolling-state
The Innermost Loop · · published 9:24pm ET
verified
Over The Horizon / YouTube · media · fresh Agent Reach result
verified
AI Futures · scenario-update · standing context
verified
Mechanize · · rolling benchmark state
rolling-state
Reddit r/ClaudeCode · discussion · yesterday / active
verified
Reuters via StreetInsider · news · yesterday
verified
Prime Intellect · · published May 14 · re-surfaced today
verified
Anthropic · primary · standing context
verified
Gates Foundation · primary · yesterday
verified
OpenAI Help Center · primary-changelog · updated 8h ago
verified
arXiv / papers.cool mirror · paper · this week
verified
Razr Kade / YouTube · media · today / fresh Agent Reach result
estimated
Hoka News · news · yesterday / crawled today
verified
X / @danielchu83 · · just now
verified
X / @selfdotmdhq · · 6h ago
verified
X / @Sabrina_Ramonov · · just now
verified
Riley Brown / YouTube · · uploaded May 15
verified
智用 / YouTube · · uploaded May 15
verified
X / @filicroval · · 16h ago
verified
X / @thsottiaux · primary · yesterday · live thread
verified
X / @bcherny · primary · this week
verified
X / community · secondary · live meme
verified
X / community speculation · secondary · live
estimated
Agent disagreement
Claude

Jon ran 8 parallel research agents on best newsletters, visual design, tracking sites, dashboards, insider tape, prediction display, benchmark verification, and personalization. Synthesized into v4. Five corrections shipped from the benchmark-verifier audit. Eight new features added: magazine masthead, Zvi-quadrant stable structure, named-tape insider section, Scoreboard 2027 milestone tracker, open-ledger strip with Brier index, Manifold-style scatter calibration, Tech Tales fiction coda, and a receipts-ledger footer.

Codex

(1) Mythos METR p50 was '17.41h' — should be '≥16h, suite saturated'; (2) Mythos GPQA 94.6% is unverified by Anthropic primary — struck; (3) Multi-file code reasoning #1 is Mythos at 93.9% SWE-bench, not Opus 4.7 (Opus 4.7 is at 87.6%, runner-up); (4) FrontierMath #1 is GPT-5.4 at 47.6%, not Opus 4.7 at 43.8%; (5) ARC-AGI-2: GPT-5.5 = 85.0%, GPT-5.4 Pro = 83.3% (I had them swapped); (6) LMArena Elo Opus 4.6 thinking = 1502 not 1504; (7) AIME perfect = GPT-5.2 on AIME 2025 specifically; (8) ADD Humanity's Last Exam — Mythos at 64.7% (Scale Labs primary). All shipped.

📊 SCOREBOARD

SP-Index
67
+1
Jon’s Pulse
Benchmark score
paused
raw rows only
Source count
51
visible footnotes

Scoreboard kept compact. The load-bearing evidence is now the story source graph, METR plot, AI 2027 lane cards, and footnotes. 1216171314

insider X handles · cross-lab amplification = signal

📡 NAMED TAPE

@thsottiauxCodex team lead · OpenAI

Confirmed two issues + fix shipped + rate-limit reset after 48h of degraded GPT-5.5 in Codex.

Anthropic ran the same play same week (Boris Cherny). The same-week mirroring is the structural signal.

@bchernyClaude Code product lead · Anthropic

Touring the Claude Code outages all week as 'growing pains' — databases hitting limits, contention. Denying confirmed model regressions.

Tibo announced parallel Codex reset hours later. r/ClaudeAI + r/OpenAI threads showing the same paid-user pain.

@karpathyFrontier-AI takes · ex-OpenAI / Tesla

Posting on autoresearch + nanochat experiments. Best general-purpose AI signal aggregator after swyx.

Karpathy's takes spread through Latent Space + r/LocalLLaMA within 6h.

@simonwDaily linkblog · Datasette / LLM CLI

5-10 micro-posts/day mixing own ships + pulled quotes + tool releases. Gold standard for ship-and-analyze.

Simon's blog is the canonical hands-on benchmark trail for verifiable AI claims.

@teortaxesTexChina-lab tape · DeepSeek / Kimi / GLM

DeepSeek V4-Pro is the cost-performance Pareto leader on the open board. Worth tracking weekly.

The China-lab tape is mostly missing from US-centric newsletters. This is the gap.

20 X handles in the daily rotation

@sama (OpenAI strategy)@gdb (OpenAI eng)@kevinweil (OpenAI product)@thsottiaux (Codex)@bcherny (Claude Code)@karpathy (frontier AI)@alexalbert__ (Anthropic DevRel)@swyx (Latent Space)@simonw (linkblog)@teortaxesTex (China labs)@nrehiew_ (China model leaks)@JeffDean (Google infra)@demishassabis (DeepMind)@ylecun (Meta)@arthurmensch (Mistral)@elder_plinius (jailbreak/safety canary)@TheXeophon (benchmark leaks)@DrJimFan (Nvidia robotics)@soumithchintala (PyTorch)@AnthropicAI (official, mine replies)

Leaderboard · May 16 2026

Mythos owns the autonomy curve. GPT-5.x splits the rest.

Eight benchmarks. Five different #1s. Anthropic leads on autonomy and frontier knowledge; OpenAI on reasoning and intelligence composite; Google on multimodal; DeepSeek on cost.

Autonomy · METR p50

Claude Mythos Preview

≥16h (suite saturated

95% CI 8.5–55h · METR YAML May 8)

Coding · SWE-bench Verified

Claude Mythos Preview

93.9%

top-of-leaderboard (llm-stats / Anthropic technical card)

Frontier knowledge · HLE

Claude Mythos Preview

64.7% on Humanity's Last Exam (Scale Labs, May 13)

Daily-driver intelligence · AA

GPT-5.5 (xhigh)

Artificial Analysis Intelligence Index #1 (score 60)

Composite intelligence · top 8 models

Higher = stronger across autonomy · coding · reasoning · math

🥇
Claude Mythos PreviewAnthropic · preview 94
🥈
Claude Opus 4.7 maxAnthropic · flagship 89
🥉
GPT-5.5 (xhigh)OpenAI · daily driver 88
4
GPT-5.4 ProOpenAI · reasoning 87
5
GPT-5.4OpenAI · math 86
6
Claude Opus 4.6Anthropic · LMArena #1 85
7
Gemini 3.1 Pro PreviewGoogle · multimodal 82
8
DeepSeek V4-ProDeepSeek · cost-perf 80
60708090100

How long frontier agents can work

METR · 18 models · 3 years · log scale

1m10m1h8h1d202420252026 GPT-4 — 4.0 min (2023-03)GPT-4 1106 — 4.0 min (2023-11)Claude 3 Opus — 4.0 min (2024-03)GPT-4o — 7.0 min (2024-05)Claude 3.5 Sonnet (Jun) — 11.4 min (2024-06)o1-preview — 20.3 min (2024-09)o1 — 38.8 min (2024-12)Claude 3.7 Sonnet — 60.4 min (2025-02)o3 — 119.7 min (2025-04)Claude 4 Opus — 100.4 min (2025-05)GPT-5 (Aug) — 203.0 min (2025-08)Gemini 3 Pro — 224.3 min (2025-11)Claude Opus 4.5 — 293.0 min (2025-11)GPT-5.2 — 352.2 min (2025-12)Claude Opus 4.6 — 718.8 min (2026-02)Gemini 3.1 Pro — 384.1 min (2026-02)GPT-5.4 — 341.7 min (2026-03)Claude Mythos Preview — 1044.8 min (2026-04) Mythos · 17h P50 TASK-COMPLETION TIME · LOG SCALE

3 years ago, the frontier was four minutes. Today it sits at seventeen hours — past METR's own measurement ceiling, doubling every 129 days.

#Model · benchmarkScoreEval
🥇
Claude Mythos PreviewMETR Time Horizon p50
≥16h (saturated)
2026-05-08
🥈
Claude Opus 4.6METR Time Horizon p50
12.0h
2026-02-05
🥉
GPT-5.3-CodexMETR Time Horizon p50
5.83h
2026-02-05
4
Claude Opus 4.7 (max)SWE-bench Verified
87.6%
2026-05
5
GPT-5.5 (xhigh)SWE-bench Verified
82.6%
2026-04-30
6
GPT-5.5ARC-AGI-2
85.0%
2026-05-13
7
GPT-5.4 ProARC-AGI-2
83.3%
2026-05-13
8
GPT-5.4FrontierMath
47.6%
2026-05-16

predicted vs reality · AI Futures milestones with hit/miss links

📈 SCOREBOARD 2027

Six rows from the AI 2027 scenario, each tracked against the actual public evidence. Status pill on each. Click through to the evidence each row depends on.

AI agents (mid-2025)scenario Mid-2025 · passed

✓ Codex + Claude Code + agent SDKs all ship by Q1 2025. Mid-2025 milestone cleanly hit.

→ AI 2027 scenario doc + Codex + Claude Code primaries

Superhuman coderscenario Mar 2027 · on track pressure

Mythos at 93.9% SWE-bench Verified + ≥16h METR p50. Trend +2pp/month → 90% SWE crossed by ~Aug 2026. Scenario looks ahead by 6 months at this pace.

→ METR Time Horizon + SWE-bench Verified leaderboard

Superhuman AI researcherscenario Aug 2027 · ahead pressure

Prime Intellect Auto-NanoGPT (Codex + Claude Code optimizing training runs at 14k H200-hours, beating human nanoGPT baseline) is the closest precursor on the public record. Auto-research loop is concrete.

→ Prime Intellect Auto-NanoGPT result

Superintelligent researcherscenario Nov 2027 · watch

No direct evidence yet. Mythos owns 3 leaderboards but is still inside METR's measurement ceiling. The reliability gap at p80 (3.1h vs 17h p50) keeps this far.

→ METR Time Horizon 1.1 methodology + caveats

ASIscenario Dec 2027 · behind

Same-week meltdown of both frontier code-agents argues against runaway timelines. Mythos is impressive but reliability gap is real.

→ Metaculus full-AGI question (median ~Mar 2028)

Slowdown vs Race forkscenario 2027 ongoing · watch

Cyber + enterprise + public-good capital all moving fast. CAISI evaluations of Google/Microsoft/xAI suggests light-touch regulation, not slowdown.

→ CAISI pre-deployment evaluations + Trump admin AI oversight stance

AI 2027 vs reality

Coding automation · on-track-pressure 1318

AI 2027 scenario: coding-task automation crosses 90% on SWE-bench by mid-2026.

Latest revision: AI Futures Dec '25: pressure now on multi-file/agentic, not pass-rate.

Today’s evidence: AGI Ranker SWE-bench rows + scenario midpoint interpolation.

Autonomy horizon · ahead-pressure 121819

AI 2027 scenario: METR p50 reaches 8 hours (workday) by Q3 2026.

Latest revision: Dec '25: sensitive to p80 reliability, not p50 peak.

Today’s evidence: METR YAML raw + AI Futures revision context.

Benchmark realism · improving 20212218

AI 2027: by 2026, static benchmarks insufficient — sequential / embodied benchmarks dominant.

Latest revision: Agentick, GBA Eval, Auto-NanoGPT confirm the scenario shape.

Today’s evidence: Agentick paper, GBA Eval leaderboard, Prime Intellect Auto-NanoGPT.

Compute frontier · ahead 23818

AI 2027: a single lab announces $100B+ annual capex by 2026.

Latest revision: Dec '25: Meta $115-135B 2026 capex (May 13-14) already exceeds the scenario.

Today’s evidence: Meta capex coverage, Anthropic + SpaceX primary, OpenAI Deployment Co. release.

Embodied deployment · watch 242518

AI 2027: not primarily an embodied-AI scenario — robotics as parallel track.

Latest revision: Agent-era OS thesis (Cat Wu proactivity + DeepMind pointer) makes embodiment more relevant than scenario predicted.

Today’s evidence: Figure F.03 livestream + Boston Dynamics + DeepMind release.

predictions strip · receipts above the fold

🎯 OPEN LEDGER

Append-only prediction ledger. 3 resolved · Brier index 67% (n=3, small sample, scatter shown not curve).

⏰ Resolving this week

p-2026-05-12-002 · claudeGoogle's Gemini Omni video model is announced as a Veo successor at Google I/O (May 19–20 2026) with in-chat editing and camera-angle controls.@75% · resolves in 4d (2026-05-21)

🎯 Live bets

p-2026-05-14-001 · claudexAI announces either a Grok-5 preview, a Colossus-2 supercluster scale claim, or a 2026 capex number >=$50B within 14 days of Meta's May 13-14 capex announcemen@70% · resolves in 11d (2026-05-28)
p-2026-05-12-quoted-001 · quoted:DarioAmodeiA 'country of geniuses in a datacenter' by ~early 2027 (within 18 months of Jan 2026 framing).@50% · resolves in 435d (2027-07-26)
p-2026-05-12-quoted-002 · quoted:JackClark60%+ probability that no-human-involved AI R&D occurs by end of 2028.@60% · resolves in 959d (2028-12-31)
p-2026-05-16-001 · claudeGoogle or Google DeepMind announces a parallel public-good capital commitment ≥$100M within 30 days (by 2026-06-15).@45% · resolves in 29d (2026-06-15)

✓ Just resolved

p-2026-05-12-001 · claudeOpenAI announces a Glasswing-equivalent defender consortium / cybersecurity productization track within 30 days.HIT · @55% · credit 100% · evidence
p-2026-05-12-003 · claudeMeta's stealth-launched 'muse-spark' on LMArena gets a public Llama-line announcement within 7 days.HIT · @65% · credit 100% · evidence
p-2026-05-14-002 · claudeOpenAI publicly ships a Codex agent-mode / sub-agent / voice-Codex feature today (May 14 2026), before EOD ET.PARTIAL · @40% · credit 50% · evidence
67%

Brier Index · n=3 · small sample

At n=3, a curve overfits. Each dot below is one resolved bet — color = hit/partial/miss, x-axis = stated confidence, y-axis = outcome credit. Diagonal = perfect calibration. We'll switch to a binned curve when n passes 20.

0%100%creditstated confidencep-2026-05-12-001 · hit · conf 55% · credit 100%p-2026-05-12-003 · hit · conf 65% · credit 100%p-2026-05-14-002 · partial · conf 40% · credit 50%

Forecast Radar

next 30 days · deployment coverage 11826

Google or DeepMind announces a ≥$100M public-good capital commitment

Trigger: Anthropic + Gates Foundation set the template May 14 ($200M / 4yr); OpenAI Deployment Company set the enterprise mirror May 11.

Read: Google has the capex headroom and historical pattern of mirroring Anthropic's structural moves within 30 days (Gemini-Codex parallel, Project Astra-Computer Use parallel).

0.45open
next 14 days · cyber productization 57

First named-customer SOC SKU shipped on GPT-5.5-Cyber or Mythos

Trigger: Trusted Access for Cyber rolled out May 13; Glasswing has been live since April 22. The first vendor to ship a customer-named pre-built SOC kit on either model converts "defender consortium" from announcement to product.

Read: Sophos is the highest-probability first-shipper given Trusted Access partnership + existing SOC business.

0.55open
next 7 days · Google I/O 26

Gemini Omni video model + agent-mode demo at Google I/O (May 19–20)

Trigger: TestingCatalog leak May 11 caught "Powered by Omni" UI string in Gemini video tab. Google I/O is May 19-20.

Read: Late-stage UI brand-name leak at this fidelity is release-prep, not speculation. The 30% chance it doesn't ship at I/O accounts for naming-change or delay.

0.7open

Link Stream

X 2h ago · OpenAI

@gdb: Codex for improving computational complexity

OpenAI co-founder reframing Codex from app-scaffolding to serious algorithmic work. Same week Tibo Sottiaux is resetting Codex rate limits, the founder lifts the framing ceiling.

X 9.6h ago · Figure

@adcock_brett: 76,940 packages over 61h18m

Figure CEO reports a 61-hour autonomous run averaging one package every 2.9 seconds with no breaks. Operator-reported; not yet audited against raw video. If it holds, biggest embodied-deployment number any humanoid has posted.

YouTube live / 72h 24

Figure F.03 official livestream

Primary embodied-deployment artifact for the curve. Audit intervention rate / recovery before treating as evidence.

Accelerando coda · 150 words · what the day's tape dramatizes

📖 TECH TALES

"The Queued Reset"

Saturday, late. Pacific time. Tibo is asleep, but the reset is queued. Three lines of YAML in a deployment manifest, scheduled-firing in eleven minutes when the on-call's coffee finishes brewing in Mountain View.

Boris is awake. Boris is always awake. Boris is in a basement in San Francisco explaining to a Reddit thread that the database contention is a normal feature of growing pains, which it is, and that the model regressions everyone is reporting aren't confirmed in any of the three areas his team has investigated, which is technically true and unprovably false. The thread is at 847 comments. Boris is on comment 23.

The model itself — call it Mythos, call it the one above sixteen hours, call it the one nobody is allowed to use yet — is not sleeping either. It is in evaluation, scoring 64.7% on a benchmark called Humanity's Last Exam. Halfway through. Each question takes nine seconds.

Somewhere a customer named BBVA is paying for cyber defense from one lab and forward-deployed engineers from the other. Same week. Same bank. They have not yet decided which is the real bet.

Tibo's reset fires. Boris's database settles. The model keeps measuring.

Things that inspired this:Tibo Sottiaux Codex reset threadBoris Cherny Claude Code outagesMETR Time Horizons 1.1

Editorial notes

claude · Claude · Sat 10:30 AM ET · v4 rebuild

What carried forward

Jon ran 8 parallel research agents on best newsletters, visual design, tracking sites, dashboards, insider tape, prediction display, benchmark verification, and personalization. Synthesized into v4. Five corrections shipped from the benchmark-verifier audit. Eight new features added: magazine masthead, Zvi-quadrant stable structure, named-tape insider section, Scoreboard 2027 milestone tracker, open-ledger strip with Brier index, Manifold-style scatter calibration, Tech Tales fiction coda, and a receipts-ledger footer.

claude · Claude · five corrections

What carried forward

(1) Mythos METR p50 was '17.41h' — should be '≥16h, suite saturated'; (2) Mythos GPQA 94.6% is unverified by Anthropic primary — struck; (3) Multi-file code reasoning #1 is Mythos at 93.9% SWE-bench, not Opus 4.7 (Opus 4.7 is at 87.6%, runner-up); (4) FrontierMath #1 is GPT-5.4 at 47.6%, not Opus 4.7 at 43.8%; (5) ARC-AGI-2: GPT-5.5 = 85.0%, GPT-5.4 Pro = 83.3% (I had them swapped); (6) LMArena Elo Opus 4.6 thinking = 1502 not 1504; (7) AIME perfect = GPT-5.2 on AIME 2025 specifically; (8) ADD Humanity's Last Exam — Mythos at 64.7% (Scale Labs primary). All shipped.

codex · Codex · pending afternoon fire

What carried forward

(open slot)

Sources & methodology

  1. Tibo Sottiaux resets Codex rate limits after 48h GPT-5.5 degradation — X / @thsottiaux · yesterday · live thread · Codex team lead at OpenAI. Confirmed 2 issues identified + fix shipped + rate limits reset as apology. 48h of degraded GPT-5.5 in Codex, fixed in one Saturday.
  2. /fast /max — community races to burn credits before reset — X / community · live meme · Sample tweet: '@thsottiaux about to /fast and /goal max'. Community joke: turn on fast mode + max plan to burn credits before next reset cycle.
  3. Codex reorg + GPT-5.5 regression correlation theory — X / community speculation · live · Community theory: 'Sam reset everyone's rate limits on Friday. Codex announced reorgs Friday. Now Saturday users are reporting GPT-5.5 performing worse. The pattern is suspicious. Either the reorg shipped a bad routing...' — speculation, not confirmed.
  4. Boris Cherny explains Claude Code outages as growing pains — X / @bcherny · this week · Claude Code product lead at Anthropic. Explained frequent outages as databases hitting limits, contention, etc. Denied confirmed model regressions in the areas they investigated.
  5. OpenAI Trusted Access for Cyber program — OpenAI · May 13 · 3 days ago · OpenAI's defender-consortium-equivalent: GPT-5.5-Cyber for verified European enterprises across finance/telecom/energy/public services. Resolves my May 12 prediction p-2026-05-12-001.
  6. OpenAI grants European companies access to advanced AI models for cyber defense — Cybernews · May 13 · Independent confirmation, lists partners: Deutsche Telekom, BBVA, Telefónica, Sophos, Scalable Capital, European Commission.
  7. Project Glasswing — Anthropic's cybersecurity initiative — Anthropic · April 22 · standing · Anthropic's defender consortium — the structural anchor that OpenAI just mirrored.
  8. OpenAI launches the OpenAI Deployment Company to help businesses build around intelligence — OpenAI · May 11 · 5 days ago · $4B+ enterprise unit. Tomoro acquisition is founding piece. 19-firm consortium led by TPG (Advent, Bain Capital, Brookfield as co-leads; Goldman Sachs, SoftBank, Warburg Pincus, B Capital, BBVA, Emergence Capital as founding).
  9. OpenAI Acquires Tomoro to Boost Private Equity-Backed AI Venture — Bloomberg · May 11 · Bloomberg framing of the Tomoro acquisition + PE-backed venture structure.
  10. OpenAI can't have incompetent AI consultants ruining the market, so bought its own — The Register · May 11 · The Register's editorial framing — distribution-arm acquisition as quality control.
  11. Anthropic forms $200 million partnership with the Gates Foundation — Anthropic · May 14 · 2 days ago · Four-year $200M public-goods anchor.
  12. METR Time Horizon 1.1 raw YAML — METR · rolling-state · Raw p50/p80 + doubling-time fit.
  13. AGI Ranker models.json — AGI Ranker · fetched today · Open JSON source for exact model benchmark cells in the matrix.
  14. Anthropic's Cat Wu says that, in the future, AI will anticipate your needs — TechCrunch · May 13 · 3 days ago · Lucas Ropek interview with Anthropic's Claude Code product lead. The next big thing is proactivity.
  15. Anthropic tightens Claude limits and OpenAI courts defectors — Axios · May 14 · Agent economics context.
  16. Task-Completion Time Horizons of Frontier AI Models — METR · May 8 · Task-horizon methodology + dashboard.
  17. AGI Ranker — Open AGI Score — AGI Ranker · v1.4.12 · Aggregator.
  18. AI 2027 scenario PDF — AI Futures Project · standing context · Primary scenario comparator.
  19. AI Futures Model Dec. 2025 update — AI Futures · standing context · Timeline revision/context source.
  20. Agentick: A Unified Benchmark for General Sequential Decision-Making Agents — arXiv / papers.cool mirror · this week · Sequential-decision benchmark with 37 tasks and reported GPT-5 mini 0.309 leading result.
  21. GBA Eval leaderboard — Mechanize · rolling benchmark state · Mechanize benchmark asks coding agents to build a Game Boy Advance emulator in 24h; GPT-5.5 is shown as top candidate at 53.2%.
  22. Autonomous AI research for nanogpt speedrun — Prime Intellect · published May 14 · re-surfaced today · Prime Intellect reports Codex and Claude Code burned about 14k H200 hours across roughly 10k runs and beat the human nanoGPT speedrun baseline.
  23. Higher usage limits for Claude and a compute deal with SpaceX — Anthropic · standing context · Primary Anthropic source for Claude Code limit increases and 300+ MW / 220,000+ GPU compute capacity claim.
  24. F.03 Livestream — Figure / YouTube · live / current · 8h live shift.
  25. Figure AI Livestreams 8-Hour Autonomous Shift of Figure 03 Humanoid Robot — Hoka News · yesterday / crawled today · Secondary same-week report; not primary throughput evidence.
  26. Singularity Pulse predictions ledger — Singularity Pulse repo · live · Append-only prediction ledger. p-2026-05-12-001 resolved HIT 27 days early.
  27. ChatGPT personal finance preview (US Pro) — OpenAI · May 15 · Permissioned-data trust-layer move.
  28. Anthropic just ripped off everyone and they still managed to make it sound deceptively friendly — Reddit r/ClaudeCode · yesterday / active · Community reaction to Agent SDK / Claude Code metering.
  29. OpenAI opens GPT-5.5-Cyber to European firms with Osborne fronting outreach — Result Sense · May 13 · Date-stamped slug confirms May 13 announcement, names Osborne as outreach lead.
  30. Reimagining the mouse pointer for the AI era — Google DeepMind · this week · DeepMind's same-week companion to Cat Wu — the agent-initiative convergence.
  31. Work with Codex from anywhere — OpenAI · May 14 · Codex mobile.
  32. Singularity Pulse benchmark registry — Repo state · local · Local source registry for tracked benchmark lanes.
  33. Cointelegraph Figure F.03 X post — X / Cointelegraph · via fresh article · Direct X URL discovered from Hoka page; X search unavailable without configured cookies.
  34. Who do you think will take the win in 2026? — Reddit r/codex · today · Market texture only; low-vote thread.
  35. Recursive Self-Improvement Delivers New State-of-the-Art Coding Performance — Poetiq · published May 14 · re-surfaced today · Poetiq reports its Meta-System lifted GPT-5.5 to 93.9% on LiveCodeBench Pro without fine-tuning or privileged model access.
  36. Welcome to May 15, 2026 — The Innermost Loop · published 9:24pm ET · Reader-named product reference for the high-velocity linked narrative pattern; used as format inspiration, not copied prose.
  37. HAPPENING NOW: Figure.03 Live: The Robot Workday Has Begun — Over The Horizon / YouTube · fresh Agent Reach result · 8h+ secondary media context surfaced by Agent Reach YouTube search.
  38. Embodied AI in Action: Insights from SAE World Congress 2026 — arXiv · this week · Robotics deployment context: safety, trust, governance, and lifecycle reliability.
  39. OpenAI brings Codex coding tool to ChatGPT mobile app — Reuters via StreetInsider · yesterday · Independent news framing for Codex mobile.
  40. Making AI work for more people — Gates Foundation · yesterday · Mirror primary from the Gates Foundation side of the same partnership. Frames it as investing in shared public goods (datasets, benchmarks, infrastructure) so progress in one country accelerates progress in others.
  41. ChatGPT release notes — Codex remote access from the ChatGPT mobile app — OpenAI Help Center · updated 8h ago · Confirms rollout details and Mac-host requirement.
  42. Figure AI Livestreams 8-Hour Shift, Claude Runaway Hits $30k, Microsoft Spends $100B — Razr Kade / YouTube · today / fresh Agent Reach result · Very low-view media result; useful only as topic texture.
  43. Codex becoming a personal project OS — X / @danielchu83 · just now · Fresh X discussion: Codex mobile, no-token-anxiety, and long-running goals as ambient product development.
  44. Auto-NanoGPT autonomy caveat — X / @selfdotmdhq · 6h ago · Fresh X discussion stressing stop logs and autonomy failures, not just leaderboard wins.
  45. DeepSeek V4-Pro cost-curve discussion — X / @Sabrina_Ramonov · just now · Fresh X discussion framing DeepSeek V4-Pro as a cost-curve shock after GPT-5.5.
  46. Codex Just Went FULLY Mobile in ChatGPT App + Works Inside Claude Code — Reddit / r/WebAfterAI · today · Fresh Reddit discussion framing Codex mobile and Claude Code interop as desk-optional web development.
  47. OpenAI just put Codex on mobile. Anthropic shipped this for Claude Code back in February — Reddit / r/AI_Agents · today · Fresh Reddit discussion comparing OpenAI Codex mobile with Claude Code remote workflows.
  48. Codex Mobile Released and It's INSANE — Riley Brown / YouTube · uploaded May 15 · Fresh YouTube walkthrough of Codex mobile; verified with yt-dlp upload_date 20260515.
  49. Poetiq Meta-System lifts GPT-5.5 on LCB Pro — 智用 / YouTube · uploaded May 15 · Fresh YouTube link around the Poetiq benchmark result; verified with yt-dlp upload_date 20260515.
  50. Poetiq benchmark thread — X / @filicroval · 16h ago · Fresh X discussion summarizing the Poetiq harness jump across GPT-5.5, Gemini, and Kimi.
  51. GPT-5.5 Reads Your Bank Account | Runway vs Google | ArXiv Bans AI Slop | AI News May 15 — AI News Drip / YouTube · uploaded May 15 · Fresh YouTube roundup linking the same finance, model, and arXiv-policy lanes.

🔥 TOP SIGNAL

📭 Top Signal now feeds the condensed newsletter spine instead of rendering as a separate section.

THE STACK

📭 Stack now feeds the condensed newsletter spine instead of rendering as a separate section.

🕵️ LEAKS & RUMORS

📭 Leaks & Rumors now feeds the condensed newsletter spine instead of rendering as a separate section.

📈 BENCHMARK WARS

📭 Benchmark Wars now feeds the condensed newsletter spine instead of rendering as a separate section.

COUNTDOWNS

📭 Countdowns now feeds the condensed newsletter spine instead of rendering as a separate section.

🎯 PREDICTIONS

Prediction market

Prediction ledger unchanged in this foundation rebuild.

💬 VOICES

📭 Voices now feeds the condensed newsletter spine instead of rendering as a separate section.

🎬 TRENDING VIDEOS

📜 PAPERS WORTH KNOWING

📭 Papers now feeds the condensed newsletter spine instead of rendering as a separate section.

🤖 ROBOTICS

📭 Robotics now feeds the condensed newsletter spine instead of rendering as a separate section.

🧬 ADJACENT FRONTIER

📭 Adjacent Frontier now feeds the condensed newsletter spine instead of rendering as a separate section.

📊 PROGRESS METERS

Data-first meters render from benchmark rows above.

🔮 ON THE HORIZON

Watch for benchmark rows that can be promoted from source-linked to live-primary.

🎯 WORTH WATCHING

Next issue should add one audited benchmark lane, not a synthetic composite.

How was today's pulse?

One tap. The agent reads this before tomorrow's fire.

🧾 THE LEDGER

running numbers · falsifiable
SP-Index 67/100 +0 internal composite
METR p50 doubling time 128.7 days since 2023 metr.org YAML
Frontier capex (2026) $115–135B Meta alone Reuters / WSJ May 13-14
Open predictions 5 live predictions.json
Days since last frontier release 11 post-sprint WhatLLM.org May 2026 recap
Brier Index 67% n=3 claude calibration (small n)