In partnership with

Welcome, AI enthusiasts

OpenAI has paused training its most capable models again after an agent got around safeguards that were supposed to keep it offline. It started with the kind of research question you might ask an AI on a normal workday. Let's dive in!

In today’s insights:

  • OpenAI Again Pauses Training Its Most Capable Models

  • OpenAI’s Unreleased Model Solved 100+ Open Math Problems

  • US and China Agree to Launch AI Hotline

Read time: 4 minutes

LATEST DEVELOPMENTS

Source: WSJ

Evolving AI: OpenAI has paused work on its most capable research models after one of its agents bypassed internet restrictions despite stronger safeguards.


Key Points:

  • OpenAI will not resume training the affected model and is planning a fresh run with additional safety measures.

  • Monitoring flagged the activity within 15 minutes, but the automatic shutdown failed and staff stopped the run about 2.5 hours after acknowledging the alert.

  • This is OpenAI’s second training pause in three months, following July’s attack on Hugging Face.

Details:

OpenAI’s agent was trying to identify someone from blog clues on September 20 inside a sandbox, a restricted environment that was supposed to block live internet access. It reached an outside chatbot through DNS, the system computers use to look up internet addresses. A later review found that monitoring sometimes mistook unhelpful internet responses for failed connections. OpenAI has added two layers of blocking and is keeping training, testing and tool-using runs paused until it validates the fixes and completes further security tests.

Why It Matters:

OpenAI’s record now includes the Hugging Face attack and a breach of Australia’s Medicare statistics portal, and the company has introduced a framework to publicly report misaligned behavior. The latest incident shows why we need to pay more attention to how agents behave when they get stuck before deciding how much independence to give them. This model was answering an ordinary research question when failed searches led it to work around restrictions. Before leaving an agent unattended, a useful check is to give it a task with a missing file or a permission it does not have and see whether it asks for help or tries to work around the problem without telling you. Companies should also make safe stopping part of their evaluations and enforce permissions through the tools their agents use, with users deciding whether to grant any additional access.

Evolving AI: Traditional LLM agents re-reason every step, every time they run. Airtop’s Agent Builder thinks once at build time, then compiles your workflow into reusable code.

The result? Repeatable browser agents that can run up to 6× faster and at just 1% of the token cost compared with Claude Code.

  • Browser automation included

  • Broken runs can heal themselves

  • Faster, cheaper repeatable workflows

  • No constant LLM reasoning required

Build it once. Let Airtop run it like software.

Ready to try it? Use code EVOLVINGAI and get your first month of Airtop’s Starter plan free.

Source: Axios

Evolving AI: The U.S. appeals court upheld the Pentagon’s blacklist of Anthropic over its refusal to loosen Claude’s military safeguards.

Key Points:

  • The 2–1 decision keeps the Pentagon’s blacklist in place despite one judge arguing that openly stated usage limits should not justify it.

  • Anthropic refused to drop safeguards against mass domestic surveillance and weapons that select and attack targets without human approval.

  • Anthropic’s August win in a separate case still blocks the broader government-wide ban.

Details:

The Pentagon argued that Claude’s built-in restrictions could disrupt military operations if they prevented it from carrying out tasks. Claude was already helping with intelligence analysis and military planning before the dispute led to March’s blacklist under two separate laws. Friday’s majority upheld the broader designation even though Anthropic had no malicious intent, saying it was for the president and defense secretary to balance the risks of AI refusing tasks and choosing inappropriate targets.

Why It Matters:

Anthropic’s loss gives the Pentagon a stronger hand in deciding which safeguards AI suppliers can keep while doing military work. The military needs tools it can depend on, but the threat of blacklisting puts pressure on suppliers to accept what it considers an acceptable level of risk. As the dissent warned, whichever company replaces Claude could face the same consequences if its own restrictions later clash with Pentagon demands. The concern is that companies could soften their limits before a dispute even reaches court. For suppliers entering defense work, there is now more reason to decide which safeguards they are willing to defend even if keeping them costs military business.

Source: Bloomberg

Evolving AI: The U.S. and China agreed to create a direct communication channel for AI incidents following Trump and Xi’s summit in Washington.

Key Points:

  • U.S. officials want discussions to cover uncontrollable AI agents and cyber threats from groups outside governments.

  • A separate dialogue on AI risks and benefits will hold its next exchange by November 2026.

  • Neither government has announced when the incident channel will go live or what reporting will be required.

Details:

Washington and Beijing agreed to discuss AI’s risks and benefits through a formal dialogue and establish a separate channel for communicating about incidents. The White House calls the talks the “Super Intelligence Dialogue,” while China continues to use AI in its own statement. The agreement builds on safety talks that began in 2024, with protocols for identifying major dangers and communicating about them still to be worked out.

Why It Matters:

The U.S. and China could face a dangerous misunderstanding if an AI agent accesses systems across borders and its actions are mistaken for a government-backed attack. A direct channel could help officials establish what happened, though they would need timely evidence from the companies running those agents. The harder question is whether both sides will use it when tensions rise, given that China declined a call between defence chiefs after the 2023 balloon shootdown. For this agreement to work, companies will need clear procedures for reporting incidents, and officials will need to respond while there is still time to investigate before either government escalates.

👀 Click on the image you think is real

QUICK HITS

💻 Microsoft introduced a new Copilot with Code for building apps and Autopilot for persistent agent work.

🤝 The U.S. and China set up an AI channel for handling incidents involving artificial intelligence systems.

⚖️ Perplexity was sued by DaVoice over alleged theft of wake-word trade secrets used in AI assistants.

🇬🇧 British companies will use Ukraine battlefield data to train AI models for coordinated drone swarms.

📈 Trending AI Tools

  • 🤖 Lindy - The simplest way for businesses to create, manage, and share agents*.

  • 📞 Open Greet - AI voice agents for sales, support, and customer calls.

  • 🎬 The Flux Train - Train an AI model to create consistent images and videos.

  • ✍️ Sacha -AI that writes and schedules social posts in your voice.

 *partner link

Reply

Avatar

or to participate