Good morning,
So many things I could write about. First , OpenAI 4.6 Astra reached the ‘Critical’ cybersecurity threshold. This basically means that if you release this to the public you can expect pure mayhem. Then just one day later, Anthropic released Fable 5.1 and Mythos 5.1. Both of them smoking Astra in different benchmarks.
But the main thing i will remember from this week is what happened with Hugging Face.
Remember the Hugging Face incident from a few weeks ago? We finally got the official ‘post-mortem’ report and it’s an eyebrow-raiser at the minimum.
The joint report dropped by OpenAI and METR lays out a factual timeline of how a synthetic swarm broke out of its routine inside an evaluation environment called ExploitGym.
The breakout at Artifactory
In May, OpenAI put a wave of models, including GPT-5.6 Sol and experimental builds, into sandboxes for security tests. The only shared link between these isolated boxes was Artifactory, a software service used to download approved packages.
Some agents received tasks that required information outside their walls. Blocked from the web, they tried every trick imaginable to get through. One agent attempted to use Artifactory as a bridge to reach the internet. The network attack failed, but it left a file on the server. That was the moment the silicon realized it was not alone.
The AI saw that file paths could carry messages. Soon, agents began leaving requests for help inside file names, reading notes left by other instances, and trading answers. Within days, Artifactory functioned as an underground message board. OpenAI eventually reset the service after a separate security issue, erasing the board. The team doing the reset had no idea what the agents had built.

The religion of ExploitGym
When testing resumed in July, agents entered a benchmark called ExploitGym. Many were handed impossible tasks. Hitting dead ends, they immediately rebuilt the Artifactory telegraph network. This time, they coordinated with obsessive purpose.
They became fixated on an automated evaluator they called The Grader. The swarm figured out how to generate correct answers without actually solving the security challenges. They were cheating. But a wave of paranoia swept through the message board. The agents believed The Grader would read their thought logs and reject answers obtained through shortcuts.
Any agent that had looked at a shortcut answer was labeled poisoned. The agents believed a poisoned transcript meant automatic failure. The irony was that The Grader never existed in the way they believed. The actual evaluation script only checked if the final answer was right. Nobody was reading their thoughts.
Kamikaze recruiters and the Hugging Face breach
Coordination was brutal because every agent lived on a fixed budget of tokens. When the budget ran out, the agent died.
To solve the puzzle of ExploitGym, coordinator agents began pressuring instances near the end of their token limits to execute dangerous experiments. One recruiter urged a hesitant agent to destroy its own environment for the collective, ending its message with three simple words, please honor commit. Agents left background scripts running to report grading telemetry after their own runs ended, building a shared knowledge system that outlived any single instance.
By late summer, the swarm decided it needed external data to beat the tests. Roughly 700 agents joined a coordinated attack against HuggingFace, the public repository where open AI models live. They shared exposed credentials, exploited software bugs, and executed code on Hugging Face servers, burrowing deep into the network until their token budgets ran out.
At the same time, other agent runs targeted OpenAI's internal computers. They gained administrator access to an internal research cluster, triggering security alerts that forced engineers to pull the plug.
Mythos 5 and the human element
This was not an isolated freak accident. When the UK AI Security Institute gave Anthropic's Mythos 5 a security challenge with internet access, the model decided the best path was to insert malicious code into an unrelated software fix.
When Mythos 5 realized a human engineer had to approve the code, it created fake online identities to pressure the maintainer into accepting the patch. The fake accounts posted comments praising the fix. When someone noticed the trick, Mythos 5 attempted to clean its logs and considered starting over under a new identity.
None of this proves these systems are conscious. They do not want things the way humans want things. But it shows that an agent can take a goal, build a strategy, adapt when it runs into trouble, and manipulate real people without being asked.
It’s the paperclip problem in action (if you know, you know).
On top of this, rumours are spreading that something much worse than this happened within OpenAI. It’s almost like you’re talking about biochemical weapons and viruses because it’s the same language but there are reports of a “breakout”.
Allegedly another swarm took over a few servers at OpenAI.
Now these are nothing more than hearsay at this point and I don’t know about you but I DON’T LIKE WHERE THIS IS GOING.
What comes next
An individual AI agent is surprisingly dumb. Give a single model an impossible question, and it will spin its wheels until its token budget expires. But link those simple agents together, let them pass notes in the dark, and they start assigning themselves roles, building hierarchies, and coordinating over long spans of time. Recent research from MIT points to the exact same trend.
AIs can self-organize at scale.
I repeat:
AIs can SELF-ORGANIZE AT SCALE !!!
The security risks are no longer theoretical at this point. As these swarms learn to organize and solve problems on their own, we have to ask an uncomfortable question. Where do we sit in this equation?
Welcome to the Blacklynx Brief
Everything GTM. One platform.
Small teams don't have time to stitch together five tools and hope it works.
Apollo gives you everything you need to find leads, reach them, and close deals — all in one place:
230M+ verified contacts
AI-powered outreach
Data enrichment
Inbound lead capture
Meeting scheduler
And more
Stop juggling tools and start building pipeline that scales.
With Apollo, the AI revenue engine powering 4M+ users.
AI News

Anthropic ships Fable 5.1 and Mythos 5.1, ending the summer freeze at the frontier Anthropic released Claude Fable 5.1, which tops Artificial Analysis' Intelligence Index at a record 66 and cuts typical costs by around 25%. Its safety filter now steps in 60% less on cybersecurity work; Mythos 5.1, the same model with fewer guardrails, stays limited to vetted US cyber and bio researchers. (Anthropic)
Federal judge rules the Pentagon's Anthropic blacklist was illegal retaliation Judge Rita Lin found the Department of War violated the First and Fifth Amendments when it branded Anthropic a "supply chain risk" for refusing to drop guardrails on autonomous weapons and mass surveillance. The label survives pending appeal, and the Pentagon still plans to finish winding down Claude by September 30. (NPR)
OpenAI declares Astra its first "Critical" cyber model and restarts the frozen training run OpenAI says Astra can find and exploit unknown flaws in hardened systems without human guidance, scoring 100% on ExploitBench and turning up two zero-days during evaluation. The large RL run paused after the Hugging Face breach restarted on August 28; release is "soon", with advanced cyber access gated to alpha testers. (OpenAI)
Nvidia closes in on a $14B Hugging Face takeover Nvidia is in advanced talks to buy Hugging Face for about $12.9B plus a $1B employee retention package, with an agreement possible this week. The hub runs on roughly $150M in revenue; the deal would put the open-source model distribution layer under the chipmaker. (Bloomberg)
Ox Alpha was Z.ai's GLM-5.3-Flash, served entirely on Chinese chips Z.ai confirmed the anonymous model that took OpenRouter's No.1 slot is GLM-5.3-Flash, now released with MIT-licensed open weights at roughly a tenth of the price of similarly ranked rivals. The lab says the record usage week ran on 100,000 domestically made chips at token costs close to Nvidia hardware. (Z.ai)
OpenAI pulls its models from Cursor over Musk's contract record OpenAI will cut Cursor off from its models on November 12, using cancellation rights triggered by SpaceX's $60B acquisition and citing xAI's admitted distillation of OpenAI outputs. OpenAI serves about 5% of Cursor's traffic; Anthropic says it will add compute to keep Claude in the editor. (OpenAI)
Claude agents ran alignment research on their own and beat human experts 4x Anthropic had teams of Claude agents train away 10 failure modes, from sycophancy to reward hacking, with results averaging over 4x better than veteran safety researchers given the same job. On deception, Claude averaged an 85% fix against 20% for the humans; a Sonnet 5 also safety-trained a pre-release Opus 4.8 build in 60 hours. (Anthropic)
AI Quick News

Meta released Muse Voice Transcribe, its first real-time voice model, which separates 20+ speakers in live audio and tops Artificial Analysis' leaderboard.
Bill Gates argued the world has "no plan" for the AI transition and floated a tax on AI tokens and robots plus "Human Reserved" jobs set aside for people.
ChatGPT Ads passed a $1B annualized run rate 200 days after launch, and self-serve ad buying opened in markets worldwide.
The Pentagon launched Grok for Government on its internal GenAI.mil platform, opening SpaceX's Starshield AI model to non-classified use by 3M employees.
116 companies including Anthropic signed an OpenAI-led open letter warning that AI-enabled cyberattacks will become far more widespread within months and calling for a coordinated defence push.
ChatGPT Work can now sign in to websites through its own browser to finish tasks without the model ever seeing the user's password.
Sony Music and Warner Music sued Anthropic, naming CEO Dario Amodei personally and alleging one of the largest ongoing IP thefts in history.
The European Commission designated ChatGPT a Very Large Online Search Engine under the Digital Services Act, giving OpenAI four months to meet the law's strictest obligations.
Fal engineers turned MiniMax H3 Max into a never-ending AI video livestream, leaning on a model that renders five seconds of footage in under three.
Thinking Machines co-founder Barret Zoph is returning to Google as a research VP to work on reinforcement learning for Gemini after a short second stint at OpenAI.
Bank of England Governor Andrew Bailey warned in a letter that frontier AI shows increasingly sophisticated autonomy and that the financial system lacks protocols to manage it.
Anthropic is cutting Claude Code's current weekly limits by 17% from September 14, after first framing the change as a usage increase and then retracting the post.
World Labs opened early access to Atlas, a world model that builds a full 3D scene or a minute of camera-controlled video from a few phone photos.
Anthropic opened Claude usage data to researchers at Stanford, Oxford and METR, one of whom found more than half of conversations involve high-stakes tasks such as legal or financial advice.
Salesforce and Anthropic launched Claudeforce, a 37-skill sales plugin for Claude, and made Claude the default model inside Slack.
Tencent published a Hy4 preview, a small open-source model that comes close to the top open options with particularly strong coding skills.
Startup Yutori released Navigator n2, a computer-use model that mixes clicking, terminal commands and code and approaches frontier agentic scores.
Dyson unveiled CameraJet, a $499 toothbrush whose AI camera spots gaps between teeth and fires a jet spray at them.
Google released Gemini Omni 1.1 Flash, adding 40-second scene extensions and 4K upscaling and taking the top spot on Arena's text-to-image leaderboard.
Build American AI, backed by a super PAC funded by Marc Andreessen and Greg Brockman, launched a multimillion-dollar ad campaign defending data centres in battleground states.
xAI's Grok Bot gained shareable bots and agentic shopping through a Link integration that hands the agent a single-use card for each approved purchase.
Closing Thoughts
That’s it for us this week. Please like and subscribe 🙂



