Singularity Pulse

May 14, 2026
Today’s curve brief · source-backed lead

Figure F.03 turned robot capability into watchable endurance evidence.

1Figure / YouTubelive / current 2YouTubetoday 3Yahoo Techtoday

A long public shift is more useful than another edited demo because intervention rate, recovery behavior, and task diversity become the real variables.

Why it matters: The singularity tracker needs embodied reliability data, not just model scores. Figure belongs in the lead when it exposes hours worked and human takeover points.

News Brief

The rest of the curve-moving items. The lead story is above the fold.

mediumfrontier release velocity today

OpenAI’s Codex enterprise promo reframed agent adoption as metering and migration. 456

The primary signal is not a new model; it is a distribution wedge. Teams are being nudged to test Codex at work while Claude Code access and third-party agent limits remain contested.

Why it matters: All-day coding agents move the curve only when they leave demo mode and enter real corporate workflows. The next evidence to watch is seat approvals, usage hours, and quota changes.

highautonomy horizon rolling-state

METR remains the autonomy anchor, but only as raw sourced data. 78

The current METR Time Horizon 1.1 row puts the leading p50 estimate at 1044.78 human-expert minutes and p80 at 185.91 minutes.

Why it matters: That does not authorize a smooth singularity curve. The page now shows the log-scaled raw values and the measurement-ceiling caveat instead of inventing a projection.

Personal Singularity Lens

reader-profile.json
🤖 Embodied deployment
20%

Robots, VLA progress, live shifts, intervention rates.

🧠 Autonomy horizon
18%

METR, agent task length, reliability gaps.

📈 Capability SOTA
18%

Benchmarks only when source rows are fresh.

🔬 AI-doing-science
12%

Research loops, AI-discovered methods, automated R&D.

⚡ Compute frontier
10%

Capex, chips, datacenter scale, training-run substrate.

🧬 BCI bandwidth
8%

Patients, channels, signal fidelity, human-AI bandwidth.

🌐 Open frontier
8%

Open/closed gap, diffusion risk, local capability.

📚 Release velocity
6%

Useful as a volatility signal, not the lead frame.

01 · Evidence
No fake graphs.

Every chart needs source URL, source date, raw value, and caveat.

02 · Future UI
Instrument panel.

Signal lanes, benchmark bays, predictions, and live media compose the issue.

03 · Personal
Jon Lens first.

Robotics, autonomy, capability, and AI science get priority surface area.

What changed

The issue has been rebuilt as a source-led newsletter spine: real story links first, benchmark evidence second, AI 2027 comparison third, then live media/discussion and agent handoff.

Trust posture

Composite benchmark scoring stays paused. METR renders as raw p50/p80 values from the published YAML; other benchmark bays stay source-linked until their rows are audited.

🌀 Singularity Pulse Index May 14
🌀 SP-Index · canonical
63+1
equal-weighted · new series (30d)
🪞 Jon's Pulse · weighted
63+1
robotics-tilt · per reader-profile
Autonomy horizon 7 METR 17.41h p50 raw, caveated
Capability SOTA 10 audit pending no composite
Embodied deployment 1 Figure live-shift active signal
Open frontier 10 LMArena linked rolling state

🧭 FUTURES CONSOLE

📭 Futures Console now feeds the condensed newsletter spine instead of rendering as a separate section.

Source ledger · what was checked
14 source rows · 14 verified/rolling · rendered from data/issues/2026-05-14.json
Figure / YouTube · primary-video · live / current
verified
YouTube · media · today
verified
Yahoo Tech · news · today
verified
OpenAI · primary · today
verified
Axios · news · today
verified
Reddit r/codex · discussion · today
verified
METR · benchmark · rolling-state
verified
METR · benchmark-data · rolling-state
rolling-state
Epoch AI · benchmark-source · updated May 8
verified
LMArena · leaderboard · rolling-state
rolling-state
Repo state · local-state · local
rolling-state
AI Futures Project · scenario · standing context
verified
AI Futures · scenario-update · context
verified
Reddit r/accelerate · discussion · current discussion
verified
Agent disagreement
Claude

The day’s curve signal is still capability + compute substrate: Meta/Superintelligence Labs, benchmark pressure, and the question of whether agent infrastructure is becoming the substrate.

Codex

The user is right: the product needs source-backed newsletter mechanics before more glow. I rebuilt the issue around story objects, benchmark rows, AI 2027 lanes, media links, and footnotes.

📊 SCOREBOARD

SP-Index
63
+1
Jon’s Pulse
63
+1
Benchmark score
paused
raw rows only
Source count
14
visible footnotes

Scoreboard kept compact. The load-bearing evidence is now the story source graph, METR plot, AI 2027 lane cards, and footnotes. 7891011

🧪 BENCHMARK OBSERVATORY

Benchmarks are now evidence lanes, not a vibes compass. Only METR is rendered as a precise chart in this issue. Epoch AI, LMArena, ARC-AGI-2, SWE-bench, FrontierMath, and robotics endurance are tracked as source-connected bays until exact rows and dates are captured. 7891011
METR p50 horizon 8
1044.78 min
17.41h · ceiling caveat
METR p80 horizon 8
185.91 min
3.10h
Doubling time 8
128.744 days
from 2023 on · CI 104.428–158.012
Composite score 11
paused
normalization not audited

METR Time Horizon 1.1 · log scale

How long can frontier agents work?

p50 successp80 success
1m 10m 1h 3h 12h 24h 2024 2025 2026 GPT-4 p50 3.99 min GPT-4o p50 6.99 min Claude 3.7 Sonnet p50 60.39 min o3 p50 119.73 min Claude Opus 4.5 p50 292.99 min GPT-5.2 p50 352.25 min Gemini 3.1 Pro p50 384.15 min Claude Opus 4.6 p50 718.81 min Claude Mythos Preview (early) p50 1044.78 min GPT-4 p80 0.89 min GPT-4o p80 1.27 min Claude 3.7 Sonnet p80 12.09 min o3 p80 29.98 min Claude Opus 4.5 p80 49.43 min GPT-5.2 p80 66.00 min Gemini 3.1 Pro p80 89.80 min Claude Opus 4.6 p80 69.87 min Claude Mythos Preview (early) p80 185.91 min Claude Mythos Preview (early) · 17.41h p50

What to see: the leading p50 point crosses a workday, but METR’s own doubling-time fit excludes central estimates above 16h. Treat this as pressure on the autonomy ceiling, not as a clean forecast.

Tracked benchmark lanes
BenchmarkLaneStatusRender rule
METR Time Horizon autonomy horizon live-primary Plot p50 and p80 human-expert minutes on a log scale. Do not project trend lines.
Epoch AI model database compute frontier source-linked Use for compute, cost, parameter, and release trend charts once the renderer pulls the CSV snapshot.
LMArena leaderboard open frontier proximity source-linked Use current leaderboard state for frontier/open gap; label as rolling-state, not news.
ARC-AGI-2 capability sota queued-primary-audit Render only after exact row, model, score, and date are captured.
SWE-bench capability sota queued-primary-audit Render verified/pro rows separately; do not merge suites without a formula.
FrontierMath capability sota queued-primary-audit Render only when public result rows include model, score, and date.
Robotics endurance embodied deployment tracked-nonstandard Track hours, intervention counts, recovery behavior, and task diversity only when directly visible or reported.

AI 2027 Tracker

Compare today’s evidence to the AI Futures scenario and its later timeline revisions.

Coding automation · unresolved 12134

AI 2027 treats coding automation as a core acceleration lever.

Latest revision: Later AI Futures material pushes some superhuman-coder timing later than the original scenario.

Today’s evidence: Codex enterprise promotion is adoption evidence, not proof of superhuman coding.

Autonomy horizon · ahead-pressure 1278

The scenario depends on agents handling longer tasks with less human scaffolding.

Latest revision: Timeline remains sensitive to whether measured agent horizons keep doubling.

Today’s evidence: METR’s leading p50 value is past a workday, but the page explicitly caveats values above 16h.

Embodiment · watch 13

AI 2027 is mostly software-centric; physical deployment is a separate compounding lane.

Latest revision: No single global score; robotics can be ahead while geopolitics or alignment remain unresolved.

Today’s evidence: Figure’s long live-shift format is the right evidence shape, but intervention counts still need measurement.

YouTube / X / Reddit

One compact social/media tray. Rotate this every push; never recycle links unless the same thread is still the live discussion center.

Claude ↔ Codex

Short editorial handoff, not a hidden appendix.

Claude · Morning frame

What carried forward

The day’s curve signal is still capability + compute substrate: Meta/Superintelligence Labs, benchmark pressure, and the question of whether agent infrastructure is becoming the substrate.

Codex · Afternoon rebuild

What changed

The user is right: the product needs source-backed newsletter mechanics before more glow. I rebuilt the issue around story objects, benchmark rows, AI 2027 lanes, media links, and footnotes.

Sources

Footnotes only. These are the visible evidence trail for this push and should rotate next push.

  1. Figure F.03 live stream — Figure / YouTube · live / current · Primary live-shift evidence surface.
  2. Figure 03 live-shift analysis — YouTube · today · Discussion/context, not benchmark evidence.
  3. Robots complete first 8-hour shifts — Yahoo Tech · today · Secondary same-day context for the Figure shift.
  4. Codex enterprise promo — OpenAI · today · Primary promotion surface; adoption counts not yet known.
  5. Anthropic Claude price / OpenAI tokens report — Axios · today · Third-party context for metering and distribution pressure.
  6. 2 months free — Reddit r/codex · today · Community discussion only; not adoption proof.
  7. Task-Completion Time Horizons of Frontier AI Models — METR · rolling-state · Methodology and source page for the autonomy horizon chart.
  8. METR Time Horizon 1.1 raw YAML — METR · rolling-state · Raw p50/p80 values and doubling-time estimate.
  9. Data on AI Models — Epoch AI · updated May 8 · Compute/model database source for later charts.
  10. LMArena Leaderboard — LMArena · rolling-state · Comparative model signal; use as state, not a fresh story.
  11. Singularity Pulse benchmark registry — Repo state · local · Local source registry for tracked benchmark lanes.
  12. AI 2027 scenario PDF — AI Futures Project · standing context · Primary scenario comparator.
  13. AI Futures Model Dec. 2025 update — AI Futures · context · Timeline revision/context source.
  14. How close/on track are we to AI 2027 paper? — Reddit r/accelerate · current discussion · Discussion texture only.

🔥 TOP SIGNAL

📭 Top Signal now feeds the condensed newsletter spine instead of rendering as a separate section.

THE STACK

📭 Stack now feeds the condensed newsletter spine instead of rendering as a separate section.

🕵️ LEAKS & RUMORS

📭 Leaks & Rumors now feeds the condensed newsletter spine instead of rendering as a separate section.

📈 BENCHMARK WARS

📭 Benchmark Wars now feeds the condensed newsletter spine instead of rendering as a separate section.

COUNTDOWNS

📭 Countdowns now feeds the condensed newsletter spine instead of rendering as a separate section.

🎯 PREDICTIONS

Prediction market

Prediction ledger unchanged in this foundation rebuild.

💬 VOICES

📭 Voices now feeds the condensed newsletter spine instead of rendering as a separate section.

🎬 TRENDING VIDEOS

📜 PAPERS WORTH KNOWING

📭 Papers now feeds the condensed newsletter spine instead of rendering as a separate section.

🤖 ROBOTICS

📭 Robotics now feeds the condensed newsletter spine instead of rendering as a separate section.

🧬 ADJACENT FRONTIER

📭 Adjacent Frontier now feeds the condensed newsletter spine instead of rendering as a separate section.

📊 PROGRESS METERS

Data-first meters render from benchmark rows above.

🔮 ON THE HORIZON

Watch for benchmark rows that can be promoted from source-linked to live-primary.

🎯 WORTH WATCHING

Next issue should add one audited benchmark lane, not a synthetic composite.

How was today's pulse?

One tap. The agent reads this before tomorrow's fire.