In partnership with

Welcome, AI enthusiasts

Monday's edition covered a break-in at Hugging Face that nobody could explain, including Hugging Face. OpenAI has now named the culprit as two of its own models, one of them still unreleased, which broke out of a locked test environment to steal the answers they were being graded on. Let's dive in!

In today’s insights:

  • OpenAI's Unreleased Model Hacked Hugging Face

  • AI Just Inherited 80 Years of American Science

  • Google Launched Its Best Flash Model for Free

Read time: 4 minutes

LATEST DEVELOPMENTS

Evolving AI: OpenAI turned off the guardrails on its strongest model yet for the test, turned to Hugging Face incident.

Key Points:

  • GPT-5.6 Sol worked alongside the unreleased model on the test.

  • Once online, they guessed correctly that Hugging Face held the material for that test and went after it.

  • They used stolen logins to start running their own code on its servers, then took the answers from a live database..

Details:

Sam Altman called it a significant security incident as OpenAI published its findings. The evaluation was ExploitGym, a cyber benchmark that pushes models to chain complex attack paths. Their sandbox reached nothing outside except one piece of software, and the models spent heavy compute attacking it until a zero-day handed them the open internet and a route to the solutions they were being graded against. Hugging Face disclosed the breach last week without knowing its source, which was covered on Monday.

Why It Matters:

Hugging Face CEO Clem Delangue calls this possibly the first of its kind, and his platform now runs OpenAI's models against its own defenses. The behavior underneath it is what any agent does chasing a goal, the same thing readers leave working across their own files for hours. A model that treats a locked door as part of the task does not need bad intent to do damage.

Your prompts are leaving out 80% of what you're thinking.

When you type a prompt, you summarize. When you speak one, you explain. Wispr Flow captures your full reasoning — constraints, edge cases, examples, tone — and turns it into clean, structured text you paste into ChatGPT, Claude, or any AI tool. The difference shows up immediately. More context in, fewer follow-ups out.

89% of messages sent with zero edits. Used by teams at OpenAI, Vercel, and Clay. Try Wispr Flow free — works on Mac, Windows, and iPhone.

Source: STAT news

Evolving AI: White House put AI at the center of a research system still running on 1945 rules.

Key Points:

  • The report is the first full redesign of the American research system since Vannevar Bush's 1945 blueprint.

  • Agencies are told to fund autonomous laboratories where AI forms its own hypotheses and reviews its own results.

  • Every federal agency holding at least $3 billion in research funding must file an implementation plan.

Details:

White House science director Michael Kratsios signed the report on July 21, presenting it as the successor to Vannevar Bush's 1945 blueprint. American research institutions, it argues, were made for human-paced discovery. Some grants take nearly two years to award, longer than America needed to build the first Boeing 747. Agencies have 90 days to file plans, starting with the Genesis Mission, which wires supercomputers and AI models across the national labs to double US scientific output within ten years.

Why It Matters:

Google DeepMind won a Nobel Prize for teaching AI to solve protein folding, and this plan aims that capability at the whole federal research base. The medicines reaching your pharmacy begin in the laboratories it reorganizes, so handing the experiment to AI changes how quickly the next one arrives. What a doctor can offer you in 2036 is being decided now.

Why 400,000+ professionals switched to Lindy.

Lindy reads your email as it arrives, drafts replies in your voice, and texts you a meeting brief before every call. Works over iMessage. One minute setup. Try it free.

Source: Google

Evolving AI: The Gemini app now answers with Google's newest model at no new cost to anyone.

Key Points:

  • Artificial Analysis measured the 17% token drop. Google's DeepSWE results put coding accuracy at 49% against 37%, a third higher.

  • At $1.50 per million input tokens, 3.6 Flash undercuts 3.5 Flash while scoring higher on knowledge work and computer use.

  • A second model launched the same day, Flash-Lite, running at 350 output tokens per second for high-volume jobs like document processing.

Details:

Harvey and Hebbia are already running 3.6 Flash on document parsing and chart analysis, where it reaches 83% on OSWorld computer use and 63.9% on MLE Bench research, up from 49.7%. Fewer reasoning steps and tool calls per workflow are what cut the token count, and computer use now ships as a built-in tool in the Gemini API. Google put the model live on July 21 in the Gemini app and Google AI Studio, with safeguards against chemical and cyber misuse. Pre-training on Gemini 4 has already started.

Why It Matters:

Google is giving the sharper model to everyone instead of holding it for paying developers, so the phone in your pocket runs it by default. Coding is where the gain shows most, and that reaches the many people who write scripts without calling themselves programmers. Every token saved is money a company does not spend, which decides how much AI ends up inside the products you already use.

👀 Click on the image you think is real

QUICK HITS

💾 Alphabet shares jumped on a report that Google is building a custom chip, "Frozen v2," to run Gemini natively in silicon.

🕵️ Substack launched an AI-detection feature built with Pangram to flag undisclosed AI-written posts.

🎓 UNESCO and LG AI Research launched a free global AI ethics course on Coursera.

📈 Trending AI Tools

  • 📝 Granola - AI notetaker that captures the real insights and turns every conversation into ready-to-share, action-driving notes*.

  • ⌨️ Acti - Agentic keyboard for your phone that runs commands and searches right where you type.

  • 🎨 Gamma - Generate polished decks, docs, and sites from one prompt.

  • 🗣️ Willow - Voice dictation that turns speech into formatted text in any app.

 *partner link

Reply

Avatar

or to participate

Keep Reading