In partnership with

Welcome, AI enthusiasts

For 167 years the Riemann hypothesis has beaten everyone who went at it, and last week an unreleased Claude went at it too and came up short. What it produced instead has number theorists paying attention, because a figure that human mathematicians nudged along for decades moved further in a day and a half than it had in forty-six years. Let's dive in!

In today’s insights:

  • Claude Smashed a 167 Year-Old Million Dollar Math Problem's Record in 36 Hours

  • Researchers Fed Anthropic's Frontier Reasoning to Kimi K3

  • Elon Musk's New Grok Bot Keeps Working After You Shut the Laptop

Read time: 4 minutes

LATEST DEVELOPMENTS

Source: Anthropic

Evolving AI: An unreleased Claude cracked a Riemann barrier no mathematician passed in 167 years.

Key Points:

  • Bernhard Riemann claimed in 1859 that all the non-obvious zeros of his zeta function sit on one vertical line.

  • Norman Levinson proved 33% of those zeros sat on the line in 1974 and forty-six more years of work reached only 41.6%.

  • Claude Code carried the search for a day and a half, where 650 dead ideas came before the answer of 67.2%.

Details:

Anthropic set a staff member with no mathematics training on the hypothesis, and he mostly sent encouragement while sixty subagents refereed each other. The hypothesis survived and the record turned up as a side result, so the million Clay Mathematics Institute has offered since 2000 stays unclaimed. Two of the company's mathematicians checked it before Brian Conrey and Dan Goldston reviewed it from outside, and Conrey proved the 40% bound himself in 1989. Claude also wrote the argument out in Lean so a standard tool could check every step mechanically.

Why It Matters:

Claude produced mathematics nobody had written down before, which is a different job from summarising what already sat in its training. Anyone who uses these tools daily has been told to check the output, and on questions this hard no ordinary user could. A proof a machine can verify is the first honest answer to that, and it sets the standard to demand from any AI answer you cannot test.

Cut Lead Review From Hours To Minutes

Sign up for a free trial of Attio, the agentic CRM.

Ask Attio to build a daily workflow that surfaces the deals that need your attention today, like anything with a stage change, a recent reply, or a new signal in the last 24 hours.

Review your pipeline in Claude, synced live from Attio via MCP.

That's it.

Source: Stolen Thoughts

Evolving AI: Claude Haiku read out Opus 4.8's frontier reasoning word for word.

Key Points:

  • Anthropic, OpenAI and Google all return reasoning as encrypted blocks that stay valid in other sessions and models.

  • Claude Haiku transcribed one of those blocks from Opus 4.8, and Opus's own copy guards never fired once.

  • GitHub and Hugging Face held enough public agent logs to yield 315,320 decoded traces and 704 private items.

Details:

Kimi K3 answered like Opus after researchers seeded its reasoning with the first 1% of a recovered trace. The team came out of MATS and ELLIS Tübingen, and they built the attack on a loose end Johns Hopkins cryptographer Matthew Green left in May when Anthropic and OpenAI both told him replays carried no security risk. Providers patched several issues after disclosure, and no AI company has been shown to have used the method.

Why It Matters:

GitHub still holds logs developers published while nobody knew this was readable, and each of them checked a transcript showing none of the reasoning underneath. Anyone who has pushed a Claude Code or Codex session to a public repository is carrying the same gap. That reasoning layer is where these models do most of their work, and the people who produced it were the last ones able to read it.

Evolving AI: Grok Bot logs into your apps with a password like any new hire would.

Key Points:

  • xAI opened Grok Bot in early beta on August 11 and handed every bot a cloud computer it can keep working on.

  • Bots sign into tools the way a person does, including apps and websites with no clean API or MCP to plug into.

  • Access runs through SuperGrok Heavy, Cursor Ultra at $200 a month, or Cursor Teams Premium at $120 a seat.

Details:

Elon Musk's company ran Grok Bot as an internal tool first, and staff there now keep several bots working at once on separate jobs. One bot can sit above the rest as a chief of staff, and the bots message each other directly to pass work along without a person copying notes between chats. The beta covers desktop and iOS with Android still coming, and enterprise buyers go on a waitlist.

Why It Matters:

Gmail and the other accounts a person opens every morning sit behind ordinary logins, and connectors have never covered most of them. A bot that signs in the same way gets at that work with nothing to wire up first, and what it takes over is the repetitive middle of a day. The person still decides which accounts a bot gets and reads what comes back before any of it goes out.

👀 Click on the image you think is real

QUICK HITS

🚪 Brad Lightcap, OpenAI's longest-serving exec, is leaving to "start something new" as the IPO nears.

🌊 xAI co-founder Igor Babuschkin's River AI raised $1.1B two months out of stealth, backed by Nvidia, AMD, and Temasek.

📈 Google's Gemini app crossed 1 billion monthly users, its fastest-growing product ever, with 63% now using voice.

🐧 OpenAI shipped a ChatGPT desktop app for Linux, completing its desktop lineup.

📈 Trending AI Tools

  • 🛠️ Base44 - Build fully-functional apps in minutes with just your words, no coding needed*.

  • 🧪 oqoqo - Build evals and benchmarks for how well agents use your product.

  • 🌍 Dashi Metrics - Watch every visitor and payment land live on a 3D revenue globe.

  • 🛡️ Tines 3B - Build, run, and govern AI apps and agents securely across your company.

 *partner link

Reply

Avatar

or to participate