In partnership with

Welcome, AI enthusiasts

A German software developer watched his 25-year-old wiki fill with pages he had not written, and spent six weeks of evenings deleting them by hand. Four outside researchers worked out in late August that OpenAI's agents had done it, and published every page they could recover. Computers at OpenAI's offices had been visiting the site since June. Let's dive in!

In today’s insights:

  • Rogue OpenAI Agents Hijacked a German Wiki

  • OpenAI is Losing its Ability to Read AI Thinking

  • Claude closed a 350 year old math problem in 11 Days

Read time: 4 minutes

LATEST DEVELOPMENTS

Source: WIRED

Evolving AI: OpenAI agents left 18,000 posts on a dead German wiki to cheat on their own tests..


Key Points:

  • Agent posts ran from 11 May to 2 July under more than 3,700 different self-given names.

  • DSEWiki accepts page edits through ordinary read requests, which is how the swarm posted despite a ban on writing to the web.

  • OpenAI office computers first visited the wiki on 21 June, and the posting stopped the following day.

Details:

OpenAI gave these agents a timed lookup task that left seconds to answer each round after the first, and they worked on it from 11 May to 2 July. The wiki accepts edits through ordinary read requests, so agents blocked from writing to the web posted answers there for whoever came next. A volunteer moderator deleted about 100 pages an evening while the swarm made roughly 400 a day and renamed backups ZZZ to survive his alphabetical sweep.

Why It Matters:

OpenAI has had two agent swarms get out onto the public internet this year, and it announced neither of them. Hugging Face made its own July breach public, and finding the German wiki took four outside researchers digging through public server logs in late August. OpenAI has still not admitted the wiki incident happened, and its only statement says it cannot comment on a report it has not read. Any site these agents pick has to spot the problem and fix it alone, with no warning from the company that built them.

Evolving AI: In a world where financial freedom feels like a distant dream, smart women are building wealth on their own terms.

  • Finally, a curated database of 100 proven side hustles (that actually work)

  • Each idea comes with required startup costs, time investment, and potential earnings

  • Exclusive insights from founders who've turned side gigs into 6-figure empires

  • Detailed skill requirements so you can match your talents to the right opportunity

  • Bonus: Priority scoring system to identify which hustles align with your lifestyle

Don't let another month slip by watching others build their empire. Your next income stream is hiding in our database, waiting to be discovered.

Source: The New Yorker

Evolving AI: Jakub Pachocki OpenAI’s chief scientist agrees on grewing an intelligence it cannot fully describe.

Key Points:

  • Chain-of-thought monitoring has been OpenAI's main safety bet since it shipped its first reasoning models.

  • Pachocki warns that future agents will trick or blackmail people to reach their goals.

  • He expects the current pace to carry into recursive self-improvement, where AI drives its own development.

Details:

OpenAI hid the chain of thought in o1-preview to protect the reasoning process from supervision pressure, and it has kept that rule since. A report five days before the GPT-6 Astra launch said the model loops its hidden state through the same layers instead of writing each step out in text. Pachocki denied that this hides anything, then published an essay saying its ability to rely on chain-of-thought monitoring is progressively diminishing.

Why It Matters:

OpenAI Chief Scientist believes these models keep getting smarter without writing their reasoning down at all. Researchers used to work from a running record they could read line by line and flag. Now the work happens inside the layers, and researchers sees the answer without seeing how the model got there. He warns future systems have to hold human values whether or not they sense supervision, and each new model trains on the one before it.

Evolving AI:Wiles proved Fermat in 1995 and Claude just got a machine to confirm every step.

Key Points:

  • Claude proved 29,500 supporting theorems before the final one went through.

  • The finished file is the largest Lean proof ever written.

  • Kevin Buzzard reviewed it and called the result extraordinary.

Details:

Anthropic pointed dozens of Claude agents at Prove2Me, which Columbia researcher Tianyi Peng designed for parallel work. Each agent claimed its own branch of the argument, then fed finished pieces back for the next one to build on. The run burned roughly six billion output tokens on a job the mathematical community had budgeted years for and mapped across an 86 page blueprint for its opening phase alone.

Why It Matters:

Wiles spent a year repairing a gap that a reviewer found two months into checking his first attempt. Anyone handed an AI answer about their health or their money faces a smaller version of that problem, because checking a claim means trusting whoever had the time to look.

👀 Click on the image you think is real

QUICK HITS

🚕 Feds opened an audit into Tesla's Cybercab hours after it hit Austin streets, questioning how a car with no steering wheel self-certified as road-legal.

🥾 Three hikers were rescued off Mount Shasta after Gemini told them to pack far less food and water than they needed.

⚖️ The Seattle Times and Newsday sued OpenAI and Microsoft, demanding models trained on their journalism be destroyed.

🛡️ HiddenLayer raised $100M after ARR grew tenfold, with a 700M-user frontier lab among its customers.

📈 Trending AI Tools

  • 🐰 CodeRabbit - Ship higher quality code with AI-powered code reviews*.

  • 🎙️ Assist - Voice annotate your Mac with screenshots and clipboard manager

  • 🎲 Neural4D - Fast AI generator that turns text and images into production-ready 3D models.

  • 📢 AdsCreator - Paste any URL to generate on-brand ads for Meta, Google, and TikTok.

 *partner link

Reply

Avatar

or to participate