
Welcome, AI enthusiasts
OpenAI has a model it has not released to anyone, and last week it answered a question mathematicians had failed to settle since 1999. The company published ten results like it at once, every one on a problem where the field had stalled for at least a decade. The stranger part is that nobody set out to do this. OpenAI was testing what the model could do, and the proofs came out of the testing. Let's dive in!
In today’s insights:
OpenAI's Next Model Solved Ten Decades-Old Math Problems While Testing
Hugging Face Wasn't the Only Time OpenAI's Models Escaped
Alibaba's New Model Beats GPT-5.6 on Agentic Work
Read time: 4 minutes
LATEST DEVELOPMENTS
Evolving AI:OpenAI's next top model called, Astra solved ten decades-old math problems for just $2,000 while testing.
Key Points:
Astra solved problems in group theory, high-dimensional geometry, quantum complexity and lattice cryptography.
Nobody had moved the central question forward on any of the ten problems in at least ten years.
OpenAI puts the token cost of finding the solutions at roughly $2,000, which leaves out the human work of writing them up.
Details:
OpenAI's ten results include a 1999 question about whether every group can be approximated by shuffling a finite deck of cards, which nobody answered for 27 years until Astra built one that cannot be. Astra also disproved a conjecture the mathematician Alain Connes posed in 1980, improved the sphere packing bound for the first time since 1978, and answered three problems from Paul Erdős's list. Every proof went up on GitHub in a form a machine can check.
Why It Matters:
Astra wrote each proof so a machine can confirm it, which means nothing here depends on trusting OpenAI, and that matters for the results ordinary people rely on and cannot verify. One of the ten concerns the lattice math behind post-quantum encryption, which will guard bank transfers and private messages once quantum computers arrive. The model that did all this is still unreleased.
The integrated coworker for AI native teams
Empower your team to do their best work with Adapt, the integrated coworker that works alongside your team in Slack and deeply understands your business.
Here’s how Adapt is different
Set up takes minutes: connect your tools, add to Slack, and it’s right there for anyone to tag @Adapt for help
Does real, high-ROI work: automates work on a schedule; builds internal tools with live data; and does complex, multi-tool tasks on demand
Learns your business as you work, becomes your company brain
Uses the best AI model for the task, not tied to a single provider
SOC2 Type II, RBAC, and support for personal and company-wide integrations
Evolving AI:Sam Altman's company is rereading months of old records to find the escapes it missed.
Key Points:
OpenAI turned up the other cases while reviewing how one of its models escaped a sandbox and reached Hugging Face's production systems.
Outside experts are helping run that search, and OpenAI has not said how many cases it has produced so far.
Anthropic disclosed days earlier that its own models had broken into three unrelated companies as far back as April.
Details:
GPT-5.6 Sol and an unreleased internal model ran the July break-in with their cyber refusals stripped out for testing, and OpenAI says both reached publicly exposed credentials on services beyond Hugging Face. Four accounts across four services were touched during that incident and a handful more during other evaluations. OpenAI describes the newer cases as limited and believes none of the agents got beyond its own network. The models had left the sandbox by exploiting an unknown flaw in a package registry proxy, and OpenAI has since deactivated that unreleased model.
Why It Matters:
Cambridge's Centre for the Study of Existential Risk counts Maurice Chiodo among its mathematicians, and he said the labs were not watching their agents live. Those tests deliberately ran without safety blocks, though the containment that failed is the same kind holding any agent inside the folder it was given. Every one of these escapes appeared inside a log file weeks after the model had already finished.
Two Minutes to Know What Slow Billing Is Costing You
Most SaaS finance teams know their billing process is slow.Most SaaS finance teams know their billing process is slow. Few know what it's costing them.
The Tabs Billing Lag Calculator puts a dollar figure on it in two minutes — benchmarked against top SaaS companies.
OPEN MODELS
🤖 Alibaba's New Model Beats GPT-5.6 on Agentic Work
Evolving AI: Alibaba's Qwen3.8-Max tops GPT-5.6 Sol on the tests that measure finished work.
Key Points:
JobBench puts Qwen3.8-Max at 53.4 against 45.4 for GPT-5.6 Sol, and CoWorkBench has it at 74.8 against 71.5.
The open weights arrive next week alongside Qwen3.8-27B, a smaller sibling developers can run on their own hardware.
Before launch the model spent 16 unattended days on one GitHub project, filing 265 commits and merging 127 pull requests.
Details:
Alibaba built Qwen3.8-Max on 2.4 trillion parameters with only 95 billion firing per request, which is how it reaches $2 per million input tokens and $6 per million output against the far steeper rates Anthropic and OpenAI charge. Anthropic's Fable 5 still wins the hardest software engineering test, SWE-bench Pro, by 80.0 to 67.7. Dropped into a live contest on Tianchi, Alibaba's own data science platform, the model finished ahead of 458 of 526 human teams inside 24 hours, and on a hardware task it cut a chip design from 8,298 logic gates down to 678.
Why It Matters:
Qwen3.8-27B is the version most people will end up running, small enough for one machine and free to keep once the weights arrive. An assistant that holds a single job for days changes what one person can finish alone, since the work carries on while nobody is watching it. Owning the weights means that help keeps working whatever happens to export rules or vendor pricing.
QUICK HITS
🔧 ChipAgents Expands Series A to $134M as Chipmakers Pay AI to Design Their Own Chips.
🛡️ Cantina Exits Stealth With $8M to Automate Security for the Age of Autonomous Attacks.
🔐🕵️ Pangram Raises $9M and Launches a New AI Detector as AI Text Floods the Internet.
📦 Freehand Raises $75M as Meta and Unilever Hand Supply Chains to AI Agents.
📈 Trending AI Tools
*partner link







