Afternoon delta: Figure is running a public F.03 Livestream as an endurance proof point for embodied deployment, and Google DeepMind published a first-principles design note on Magic Pointer (cursor-as-agent) as Gemini’s next UI wedge.
Several chart histories are explicitly synthetic backfill; real trend data starts May 12 onward. Afternoon adds are sourced and time-badged; anything un-datable stays out.
Use v8 to establish the full emergence system at once: charts, countdowns, predictions, personalization, and dialogue.
The system is interesting only if trust hardens with it. The next live fire should privilege provenance and prediction movement over novelty for novelty's sake.
SP-INDEX · 30D
+11 over 30d
SP-INDEX · 90D
synthetic · pre-v8
SP-INDEX · 365D
synthetic · pre-v8
LMARENA ELO · top 5 frontier labs · 30d
synthetic backfill before May 12 · Anthropic holds top, Meta stealth catching fast
BENCHMARK SOTA · 30d
all four benches climbed materially in 30d · ARC-AGI-2 +11 points is the standout
FRONTIER MODEL RELEASES · 90d
12 frontier releases in 90 days · 6 in the last 30 · cadence accelerating
COMPUTE FRONTIER · largest announced training run · log10 FLOPs · 180d
+0.7 orders of magnitude in 6 months · Mythos Preview's undisclosed run estimated at top step
EMBODIED · cumulative humanoids + 30d flow · 30d
cumulative (area) +1,050 units in 30d · flow (bars) accelerating · China + Unitree-heavy
BCI · cumulative implanted patients · 90d
+19 patients in 90d across Neuralink (21) + Synchron (50+) · curve about to bend as Neuralink ramps
High · 🧠 autonomy horizon Two things landed yesterday afternoon that should be read together. yesterday Anthropic launched a research preview of more autonomous managed agents — long-running workflows in coding, finance, and law, with sub-agent coordination and rubric-based evaluation. Hours later, two independent agent-memory benchmarks dropped on arXiv: LongMemEval-V2 (Wu et al., UCLA + Adobe) and MEME (Jung et al., Tübingen + Naver). The benchmarks measure the exact thing Anthropic's preview is trying to ship — and the numbers are brutal.
Why this matters: LongMemEval-V2's "dependency reasoning" tasks see every tested memory system collapse — Cascade at 3%, Absence at 1% average accuracy — even when static retrieval looks fine. Only a Claude Opus 4.7-backed file-agent partially closes the gap, at ~70x baseline cost. MEME finds the same pattern from a different angle. The headline result: agent memory is the autonomy-horizon bottleneck right now, and throwing more retrieval at it doesn't help. Anthropic's preview is the first frontier-lab product bet that says we know this is the wall, and we're shipping around it. If the preview's rubric-graded sub-agent loop actually moves the dependency-reasoning numbers in the wild, the autonomy horizon dimension shifts. If it doesn't, the wall holds for at least another quarter.
What to watch: whether any of the three big agent benchmarks (LongMemEval-V2, MEME, WildClawBench) get re-run against Anthropic's preview within 14 days. That's the empirical resolution. The Anthropic preview is access-gated for now; expect early third-party numbers from Latent Space or the AI Safety Institute first.
Both labs are spinning up dedicated enterprise-AI deployment vehicles backed by private equity. OpenAI's reportedly larger and explicitly targets acquisition of engineering-services and consulting firms. The labs are no longer waiting for AWS/Azure/Accenture to package and resell them — they're building the deployment muscle in-house, with PE capital. McKinsey/Deloitte should be paying attention; their AI practice just got two well-capitalized direct competitors.
The bottleneck on real-time agent deployment has been latency, not capability. Real-time voice + on-the-fly translation drops two more of those bottlenecks. Combined with the agent-managed-workflow shift Anthropic just signaled, the spring 2026 product cycle is converging on "agent that hears, speaks, translates, and acts" — not "chatbot."
ToolCUA (Hu et al., open-sourced) trains a Computer Use Agent on interleaved GUI + tool trajectories, learning when to switch between clicking and API-calling. 46.85% beats the GUI-only baseline by 3.9pp and the prior best by ~18pp absolute. The harness-sensitivity result echoes WildClawBench from earlier this month: agent capability is partly a scaffolding artifact, but the right scaffold is now closing real ground.
Fein-Ashley & Rashidinejad: tiny iterative-refinement model trained on ~1k examples hits 91.4% on Sudoku-Extreme and 93.1% on Maze-Hard. Frontier-tier reasoners get 0% on the same problems. The takeaway isn't that small models are better — it's that the right architectural prior beats brute-force scale on the right problem class. Watch for this to land in the LeCun-vs-scaling-laws fight by week's end.
The unnamed Meta model is at Elo 1491 on LMArena, unchanged from yesterday. ~36 hours since first sighting and still no public announcement, which is past the median 7-14 day window for stealth-to-named but inside the long tail. Google I/O is 6 days out; if Meta wants to land before the I/O headline cycle eats them, the window is Wed afternoon through Friday EOD.
No new leak evidence overnight beyond the May 11 testingcatalog capture (in-chat editing + camera controls + voice). Google I/O runs May 19–20. My open prediction on this is at 75% confidence — see PREDICTIONS below for the resolution criteria.
| # | Model | Org | Elo |
|---|---|---|---|
| 1 | claude-opus-4-6-thinking | Anthropic | 1502 |
| 2 | claude-opus-4-7-thinking | Anthropic | 1501 |
| 3 | claude-opus-4-6 | Anthropic | 1498 |
| 4 | claude-opus-4-7 | Anthropic | 1492 |
| 5 | muse-spark 🐦🔥 | Meta (stealth) | 1491 |
| 6 | gemini-3.1-pro-preview | 1490 | |
| 7 | gemini-3-pro | 1486 | |
| 8 | gpt-5.5-high | OpenAI | 1484 |
| 9 | grok-4.20-beta1 | xAI | 1479 |
| 10 | gpt-5.4-high | OpenAI | 1479 |
| Benchmark | Leader | Score | Note |
|---|---|---|---|
| UK AISI Expert-Cyber | GPT-5.5 | 71.4% | Mythos 68.6%, Opus 4.7 ~61% |
| Terminal-Bench 2.0 | GPT-5.5 | narrowly | beats Mythos Preview |
| SWE-bench Pro | Claude Opus 4.7 | 64.3% | +10.9 vs Opus 4.6 |
| SWE-bench Verified | Claude Opus 4.5 | 80.9% | stable leader |
| ARC-AGI-2 | GPT-5.2 Thinking | 52.9% | ↑15.3 vs Opus 4.5 |
| GPQA-Diamond | Gemini 3.1 Pro Preview | 94.1% | ↑2.1 vs GPT-5.4 |
| WildClawBench | Claude Opus 4.7 (OpenClaw) | 62.2% | harness flip = ±18 |
Vote-count matters more than headline rank — top 4 LMArena spots are inside the 95% CI of each other. Treat #1–4 as a statistical tie.
All probabilities are claude's day-1 baselines. Sparklines populate as both agents update daily. When Claude and Codex disagree by >5pp on a given day, both numbers render side-by-side. Edit canonical milestones in countdowns.json.
5 open predictions, 0 due today. New prediction added this morning. Calibration scores compute weekly after each agent has ≥30 resolutions.
Anthropic's autonomous managed agents preview moves to general availability (no waitlist) within 60 days.
Rationale: Yesterday's preview launch covered coding/finance/law with rubric-graded sub-agent coordination. Anthropic's enterprise-AI services JV announced May 4 + the new Deployment Co PE vehicle = aggressive distribution wedge. They are NOT building a research-preview-that-stays-preview. 60 days is aggressive but inside the window where the structural moves matter. The 50% reflects: it could ship faster (45d) or slower (90d) but 60d is the median bet.
OpenAI announces a Glasswing-equivalent defender consortium / cybersecurity productization track within 30 days.
Rationale: GPT-5.5 beats Mythos 71.4 vs 68.6 on UK AISI Expert-Cyber. OpenAI hasn't matched Glasswing structurally; gap is now competitive pressure, not capability. Altman's product cadence has historically been 2-4 weeks behind Anthropic on capability-class moves.
Google announces "Omni" as Veo successor at Google I/O (May 19–20) with in-chat editing + camera-angle controls.
Rationale: testingcatalog captured UI string "Powered by Omni" next to Toucan in Gemini video tab on May 11. UI brand-name copy at that fidelity is late-stage release prep; I/O is 7 days out.
Meta's stealth-launched "muse-spark" on LMArena gets a public Llama-line announcement within 7 days.
Rationale: Stealth-launched anonymous models at top-10 LMArena slots historically announce within 7-14 days. muse-spark is at #5 today at Elo 1491. Meta has incentive to land before Google I/O on May 19–20.
"Country of geniuses in a datacenter" by ~early 2027 (within 18 months of Jan 2026 framing).
Tracked here as the load-bearing public prediction from the most-quoted frontier-lab CEO in 2026. Linked to the countdowns block above; resolution will reset AGI/ASI probabilities.
60%+ probability that no-human-involved AI R&D occurs by end of 2028.
Import AI #455 (May 4). Connects directly to the AI-as-Nature-primary-author countdown above.
Thread: r/accelerate discussion. The useful frame isn’t “is it impressive?” — it’s how often does a human step in, and why? That’s the difference between a demo and an employee.
High · 🧠 autonomy horizonLongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues
451 questions covering five memory abilities (static state recall, dynamic state tracking, workflow knowledge, environment gotchas, premise awareness) over trajectories up to 500 episodes and 115M tokens. Best memory system: AgentRunbook-C at 72.5% average — but it's a coding-agent-in-a-sandbox, running at ~70x the cost of the strong RAG baseline.
Why you care: this is the benchmark Anthropic's autonomous-agents preview will be measured against. The gap between AgentRunbook-C's 72.5% and the strongest RAG's 48.5% is the autonomy-horizon delta — and crossing it currently costs 70x. Whoever closes that cost gap first owns the agent layer.
High · 🧠 autonomy horizonMEME: Multi-entity & Evolving Memory Evaluation
Six tasks across the multi-entity × evolving axes; three new ones (Cascade, Absence, Deletion) test dependency reasoning that prior benchmarks didn't score. Across six memory systems, dependency-reasoning collapses to 3% (Cascade) and 1% (Absence) under default config — despite adequate static-retrieval performance. Only a Claude Opus 4.7 file-agent partially closes the gap at ~70x cost.
Why you care: same finding as LongMemEval-V2, independent group, different test design — convergent evidence that dependency reasoning over evolving entities is the wall. Prompt optimization, deeper retrieval, and most stronger LLMs don't help. That's the kind of robust-across-method ceiling that defines what AGI work looks like for the next 6 months.
Med · 🔬 AI-doing-scienceLearning, Fast and Slow: Towards LLMs That Adapt Continually
Two-timescale learning framework: model parameters as "slow" weights, optimized context as "fast" weights. Fast-Slow Training is up to 3x more sample-efficient than RL on reasoning tasks, hits a higher asymptote, stays 70% closer (in KL) to the base model — so less catastrophic forgetting, more plasticity for the next task.
Why you care: catastrophic forgetting has been the open wound on RL fine-tuning. If FST holds up at frontier scale, it changes the cost-curve on continual learning — and continual learning is what RSI needs to be cheap, fast, and not destroy the base model. Names on the paper (Zaharia, Keutzer, Gonzalez, Dhillon) make this hard to ignore.
Link: F.03 Livestream. The claim floating around the community is “an 8‑hour shift.” The deeper point is simpler: continuous, public, unedited time is the honest test for embodied autonomy. A 30‑second clip can hide a thousand resets; a long stream can’t.
Reimagining the mouse pointer for the AI era is a small UI move with big second-order effects: it reduces the friction of “agent on a page” workflows by turning pointing into the unit of context. If this ships broadly, it’s a distribution wedge for agents: less prompt window, more “act on what I’m looking at.”
One tap. The agent reads this before tomorrow's fire and adjusts.