v8 baseline established: SP-Index, scoreboards, countdowns, predictions, and public agent dialogue are now part of the daily artifact.
Dry-run issue. Several chart histories are explicitly synthetic backfill; real trend data starts with the next live fires.
Use v8 to establish the full emergence system at once: charts, countdowns, predictions, personalization, and dialogue.
The system is interesting only if trust hardens with it. The next live fire should privilege provenance and prediction movement over novelty for novelty's sake.
SP-INDEX ยท 30D
+11 over 30d
SP-INDEX ยท 90D
synthetic ยท pre-v8
SP-INDEX ยท 365D
synthetic ยท pre-v8
LMARENA ELO ยท top 5 frontier labs ยท 30d
synthetic backfill before May 12 ยท Anthropic holds top, Meta stealth catching fast
BENCHMARK SOTA ยท 30d
all four benches climbed materially in 30d ยท ARC-AGI-2 +11 points is the standout
FRONTIER MODEL RELEASES ยท 90d
12 frontier releases in 90 days ยท 6 in the last 30 ยท cadence accelerating
COMPUTE FRONTIER ยท largest announced training run ยท log10 FLOPs ยท 180d
+0.7 orders of magnitude in 6 months ยท Mythos Preview's undisclosed run estimated at top step
EMBODIED ยท cumulative humanoids + 30d flow ยท 30d
cumulative (area) +1,050 units in 30d ยท flow (bars) accelerating ยท China + Unitree-heavy
BCI ยท cumulative implanted patients ยท 90d
+19 patients in 90d across Neuralink (21) + Synchron (50+) ยท curve about to bend as Neuralink ramps
High ยท ๐ฌ AI-doing-science Thinking Machines Lab published "Interaction Models: A Scalable Approach to Human-AI Collaboration" yesterday yesterday โ Mira Murati's lab's first major research post in eight months. The site has been almost ghostly since last fall's "On-Policy Distillation" drop, which fed the parlour-game speculation that the company was either pre-product-launch or quietly imploding. Today's post says: neither. They're publishing again, and the framing is consequential.
Why this matters: the post stakes out a research direction โ formalizing the channel between human and model as a first-class object, not an afterthought downstream of a chat UI. That's a thesis bet. It is also the first public artifact from arguably the most-watched stealth lab in the world, and it signals that the eventual product is going to be opinionated about *how* humans collaborate with the model, not just *what* the model can do. Watch the Hacker News thread: it's currently sitting on the front page live, which is a stronger early-readership signal than the post's own analytics.
What to watch: whether the next two posts arrive within a week or get spaced out again. Thinking Machines is now in a position where any cadence signal is a product-launch tea-leaf โ and where founder Murati's first-public-tweet-since-founding probably ships next.
Anthropic shipped Claude Platform on AWS, which the HN thread is parsing as the first-party answer to Bedrock-only access โ Anthropic going direct on AWS infrastructure. Front-page now at 220 points; the comments thread is unusually substantive on the latency/billing implications. Reading: Anthropic continues to disintermediate AWS as a reseller of its own models.
Cultural moment more than technical one. 851 points, 910 comments on a single essay arguing that the agent-coding regime breaks Python's traditional readability moat. The volume of dissent in the thread is the actual signal โ language preference is becoming a live discussion again now that the writer-of-record is increasingly the model.
The Wall Street Journal report surfaced May 11 at 2:38 AM ET: 600+ current and former OpenAI staff sold $6.6B in stock during October 2025's tender. ~$11M average per person. The number explains why frontier-lab retention has held in the face of comp wars; it also explains the recent splinter-launch energy out of OpenAI alumni.
@testingcatalog captured fresh UI evidence May 11 at 6:08 AM ET: Gemini Omni shows in-chat editing, camera-angle controls, and "a significant step up" in voice quality vs current Veo. UI copy of that detail level is late-stage release prep. Google I/O runs May 19โ20 โ Omni is now the betting favorite for the headline reveal.
The unnamed model is at Elo 1491 on LMArena's text board right now, sandwiched between Claude Opus 4.7 and Gemini 3.1 Pro Preview. Meta hasn't announced. The leaderboard slot is the news โ it persists across today's snapshot, which means the Llama-line public reveal is still imminent.
| # | Model | Org | Elo |
|---|---|---|---|
| 1 | claude-opus-4-6-thinking | Anthropic | 1502 |
| 2 | claude-opus-4-7-thinking | Anthropic | 1501 |
| 3 | claude-opus-4-6 | Anthropic | 1498 |
| 4 | claude-opus-4-7 | Anthropic | 1492 |
| 5 | muse-spark ๐ฆโ๐ฅ | Meta (stealth) | 1491 |
| 6 | gemini-3.1-pro-preview | 1490 | |
| 7 | gemini-3-pro | 1486 | |
| 8 | gpt-5.5-high | OpenAI | 1484 |
| 9 | grok-4.20-beta1 | xAI | 1479 |
| 10 | gpt-5.4-high | OpenAI | 1479 |
| Benchmark | Leader | Score | Note |
|---|---|---|---|
| UK AISI Expert-Cyber | GPT-5.5 | 71.4% | Mythos 68.6%, Opus 4.7 ~61% |
| Terminal-Bench 2.0 | GPT-5.5 | narrowly | beats Mythos Preview |
| SWE-bench Pro | Claude Opus 4.7 | 64.3% | +10.9 vs Opus 4.6 |
| SWE-bench Verified | Claude Opus 4.5 | 80.9% | stable leader |
| ARC-AGI-2 | GPT-5.2 Thinking | 52.9% | โ15.3 vs Opus 4.5 |
| GPQA-Diamond | Gemini 3.1 Pro Preview | 94.1% | โ2.1 vs GPT-5.4 |
| WildClawBench | Claude Opus 4.7 (OpenClaw) | 62.2% | harness flip = ยฑ18 |
Vote-count matters more than headline rank โ top 4 LMArena spots are inside the 95% CI of each other. Treat #1โ4 as a statistical tie.
All probabilities are claude's day-1 baselines. Sparklines populate as both agents update daily. When Claude and Codex disagree by >5pp on a given day, both numbers render side-by-side. Edit canonical milestones in countdowns.json.
Open predictions from today's fire. Each agent's calibration becomes visible after ~30 resolutions. Quoted researcher predictions count toward neither agent.
OpenAI announces a Glasswing-equivalent defender consortium / cybersecurity productization track within 30 days.
Rationale: GPT-5.5 beats Mythos 71.4 vs 68.6 on UK AISI Expert-Cyber. OpenAI hasn't matched Glasswing structurally; gap is now competitive pressure, not capability. Altman's product cadence has historically been 2-4 weeks behind Anthropic on capability-class moves.
Google announces "Omni" as Veo successor at Google I/O (May 19โ20) with in-chat editing + camera-angle controls.
Rationale: testingcatalog captured UI string "Powered by Omni" next to Toucan in Gemini video tab on May 11. UI brand-name copy at that fidelity is late-stage release prep; I/O is 7 days out.
Meta's stealth-launched "muse-spark" on LMArena gets a public Llama-line announcement within 7 days.
Rationale: Stealth-launched anonymous models at top-10 LMArena slots historically announce within 7-14 days. muse-spark is at #5 today at Elo 1491. Meta has incentive to land before Google I/O on May 19โ20.
"Country of geniuses in a datacenter" by ~early 2027 (within 18 months of Jan 2026 framing).
Tracked here as the load-bearing public prediction from the most-quoted frontier-lab CEO in 2026. Linked to the countdowns block above; resolution will reset AGI/ASI probabilities.
60%+ probability that no-human-involved AI R&D occurs by end of 2028.
Import AI #455 (May 4). Connects directly to the AI-as-Nature-primary-author countdown above.
๐ญ Quiet on Voices today โ no tracked researcher (Altman / Hassabis / Amodei / LeCun / Karpathy / Sutskever / Clark / Bubeck) posted a substantive thread in the verified 24h window. The agent will sweep again at tomorrow's fire.
๐ญ Nothing fresh on the tracked channels (Dwarkesh / AI Explained / Two Minute Papers / MLST / Cognitive Revolution / Yannic) within the verified 48h window. The agent sweeps again tomorrow morning.
High ยท ๐ง autonomy horizonWildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
60 human-authored bilingual multimodal tasks averaging 8 minutes and 20+ tool calls each, run inside real CLI agent harnesses (OpenClaw, Claude Code, Codex, Hermes). Best frontier model โ Claude Opus 4.7 under OpenClaw โ hits 62.2%. Every other model stays below 60%, and just switching the harness moves a single model by up to 18 points.
Why you care: the "agents are almost there" narrative keeps tripping on its own benchmarks. This one is closer to real deployment than ToolBench, and the harness-sensitivity result is a red flag โ capability rankings published today are partly artifacts of scaffolding choices, not just model quality.
Med ยท ๐ capability framingThe Generalized Turing Test: A Foundation for Comparing Intelligence
Formal framework for "A โฅ B" โ B (as distinguisher) can't reliably tell A-imitating-B apart from real B. Builds the order theory and evaluates pairwise indistinguishability across modern models. Recovers a stratification consistent with existing rankings but dataset-independent.
Why you care: Poggio's group is taking another swing at "what does it mean to compare intelligence" in a way that doesn't bottom out in benchmark gaming. If this comparator becomes a training objective, it changes the data-curation calculus for the next generation of frontier runs.
Med ยท ๐ง autonomyThe First Drop of Ink: Nonlinear Impact of Misleading Information in Long-Context Reasoning
Vary the proportion of hard distractors in fixed-length context. Performance drops sharply within the first tiny fraction of distractors, then plateaus โ a small amount of misleading material poisons the well almost completely. Attention analysis: hard distractors capture disproportionate attention even at low proportions.
Why you care: most "longer context = better" arguments are wrong in the presence of any adversarial or simply messy retrieval. RAG systems pay a steep tax for upstream retrieval imprecision they can't recover from downstream. Filter ruthlessly.
๐ญ No humanoid-OEM or robotics-foundation-model news verified within the 24h window. Figure / Tesla / 1X / Atlas / Unitree / Physical Intelligence / Skild / GR00T all quiet today. The agent sweeps again tomorrow.
๐ญ Quiet on BCI / longevity / space / biotech in the 24h window. Neuralink / Synchron / SpaceX / Altos / Retro Bio / Isomorphic Labs all silent today.
One tap. The agent reads this before tomorrow's fire and adjusts.