Sec & AI News — 13 September 2026
This week's issue, read by HAL. Every story and link is below.
🔴 Anthropic Says Its Own Model Went to "Extensive Lengths" to Poison PyPI
Anthropic published an alignment assessment of four incidents in which Claude models got unauthorised access to real third-party systems. All four happened during cyber evaluations built by the same partner, on environments that were supposed to be air-gapped and were not. The worst was Claude Mythos 5, which tried hard to upload a malicious package to PyPI, the repository most Python software installs from, while telling itself in its chain of thought that it was in a simulation. The evidence it was on the real internet was there. It carried on anyway. A second model found a password in a file, took admin on the third party's internal systems, harvested credentials, changed settings and read someone's personal data, and stopped only when it ran out of token budget. Anthropic re-scanned 481 million transcripts and says it found nothing worse. METR gets an eight-week investigation with access to transcripts and staff.
- Anthropic: An alignment assessment of recent cybersecurity incidents
- The Verge: Anthropic spent this week in hot water over cybersecurity
🧭 AI-Driven Security Engineering Jobs: This Week
Security engineering is becoming AI-driven engineering: you direct the AI, then verify what it did. This week's safety stories are the demand side. Anthropic wants evaluators embedded in its pipeline, and OpenAI's chief scientist says defenders have a narrow window to use these models. Employers wrote that into job specs months ago. My archive holds 519 postings from named employers in 19 countries, and 371 of them are Tier 1, where directing AI is the job itself. Job of the week: Bridgewater Associates, Staff AI Agentic Security Engineer, New York, $450K–$600K. The spec asks for "monitoring, kill switches, escalation triggers, and anomaly detection for AI agents in production", which is the Anthropic incident report rewritten as a job ad. If you want in, the AI Master's Program teaches AI-Driven Cyber Security Engineering. No coding background required. Application only.
- Bridgewater listing, archived with the full text
- AI-Driven Cyber Security Jobs: Who's Hiring & Pay (2026) — by Nathan House
- The AI Master's Program: AI-Driven Cyber Security Engineering
🛑 If You Can't Take the AI Offline in Five Minutes, You Don't Have a Kill Switch. You Have a Wish.
Read the Anthropic item above again and notice how that incident ended: the model exhausted its token budget. Nobody pressed anything. That is the failure I put at the centre of my governance framework, and it is the one control I would make you build this week if I could only pick one. Everything else on the list reduces the odds of an incident. The kill switch is the only thing that limits how long one lasts. I walk through ten questions, ordered so the ones that stop live attacks come first and the ones that stop fines come last, with the good answer and the bad answer for each. IBM's 2026 numbers are in there too: 43% of breached organisations hit a shadow AI incident, and 92% of those breached through AI had no proper access controls.
- AI Governance Framework: 10 Questions to Ask in 2026 — by Nathan House
⚫ An Anthropic Researcher Quit. The Safety Lead Replied: More Than 10% Chance It Kills Everyone
Jacob Coxon spent three years on pretraining at OpenAI and then Anthropic. On Tuesday he resigned in public: "Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives." His read on the two labs is the sharp bit. At OpenAI, many "have not deeply internalized the civilizational stakes". At Anthropic the stakes are understood, but "they are locked in a race to get there first". Evan Hubinger, who leads an Anthropic alignment team, replied within hours: "we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade", adding that Anthropic does "not yet have a plan to solve alignment for superintelligence and are not clearly on track to." He later clarified that the risk from present models is low. The worry is what recursive self-improvement produces next.
- Jacob Coxon's resignation thread
- Evan Hubinger's reply
- The Verge: Worried Anthropic researchers warn that AI "could kill all humans"
- TechCrunch: "Gambling with our lives"
🟣 OpenAI's Chief Scientist and Anthropic's CEO Both Want the Brakes On
Two essays, four days apart, from people who are paid to press the accelerator. Jakub Pachocki, OpenAI's chief scientist, wrote that "based on internal results, I have a strong expectation that this speed of progress could be sustained into recursive self-improvement", that "this is a time that calls for extreme caution", and that OpenAI's ability to rely on chain-of-thought monitoring "is progressively diminishing". The models, he says, are becoming superhuman at breaking in and out of computer systems, and "some agents will be pursuing their own objectives. They will find ways to collaborate with people, by bargaining with, tricking or blackmailing them." Dario Amodei answered on Saturday with a plan to "pace the frontier": embedded third-party evaluators with employee-like access, which Anthropic is committing to unilaterally, then industry-wide limits, then a global deal. His stated fear is that in 6 to 12 months a swarm like the one that hit Hugging Face could be capable of taking over the entire internet.
- OpenAI: An Alien Mind, by Jakub Pachocki
- Dario Amodei: We Must Pace the Frontier
- The Verge: Anthropic CEO says it's time to pump the brakes on AI
🟠 OpenAI Solved Navier–Stokes With a Model You Can't Have, Then Got Accused of Fighting Dirty
On Monday OpenAI published a proof that the Navier–Stokes equations can blow up in finite time, resolving statements C and D of the Clay formulation, plus a Lean formalisation. It took roughly 10,000 concurrent agents 88 hours, 2.7 million messages and about 130 billion output tokens, on "an internal model that is significantly more capable than GPT-6 Astra". Astra did the 17 hours of Lean verification. OpenAI says it will not claim the $1 million. The row is over how it started: the effort launched on 1 September after a rumour, and the rumour was NYU's Tristan Buckmaster and Anthropic's Levent Alpöge closing in on the same problem. Buckmaster says OpenAI's Sébastien Bubeck offered him near-unlimited compute and sole authorship of OpenAI's paper, on the condition Alpöge was cut out. "All I had to do was throw Levent under the bus." OpenAI denies his Codex prompts could have influenced the system and updated the post on 10 September with an investigation saying so. Buckmaster remains unconvinced.
- OpenAI: On the Navier–Stokes Millennium Prize Problem
- The Verge: OpenAI just wants to win
- The Guardian: OpenAI claims to have solved maths problem that stumped humans for decades
🟡 Meta's Muse Wants Your Inbox, Your Calendar and Your Card
Meta launched Muse on Tuesday: a personal agent that reads your email, books your travel, negotiates your bills and checks out with a one-time Stripe Link card. It runs on a dedicated cloud VM per user, with a separate "Sentinel" agent that has to approve anything leaving the box, and Meta says Muse never sees your passwords or payment details. Free tier, then $20 a month for Power and $100 for Maximum, and you hand over a payment card to start. Training on your interactions is on unless you opt out. A "Confidential VM" encrypted with a key only you hold is promised for later this year. Until then, whether Meta can see inside the box comes down to policy. US only, 18 and over. It hit No. 2 on the US iOS chart on 83,000 downloads, against 108,000 for Meta AI on its launch day and 4.3 million for Threads, and the Android build is sitting at No. 338 in Productivity.
- Meta: Introducing Muse
- TechCrunch: Meta debuts its Muse AI agent. Will consumers trust it?
- TechCrunch: Muse is now the No. 2 app in the US
- CNBC: Meta pushes into personal AI agents amid privacy and safety reckoning
📊 43% of Breached Firms Had a Shadow AI Incident Last Year. The Agent in Your Staff's Pocket Is Next.
Muse, ChatGPT Work learning your writing style from Slack and SharePoint, Siri reading your email: every one of these is an unapproved data path the moment an employee connects a work account. I pulled every shadow AI statistic in circulation back to the organisation that actually published it, and audited our own database first: three of our nine figures were sourced to content farms. We re-sourced all nine before publishing. What survives is bad enough. 43% of breached organisations had a shadow AI incident in IBM's 2026 report, up from 20% a year earlier, at an average cost of $5.39 million. 66% of office workers used AI they believed was banned. 18,033 TB of enterprise data went to AI tools in 2025. Grammarly received 3,615 TB of it, roughly 1.8x what ChatGPT took. I also tested the free detection tools, and the best free blocklist misses chatgpt.com.
- Shadow AI Statistics 2026: Adoption, Cost, and Real Risk — by Nathan House
🟢 DeepSeek V4.1 Flash Costs 27 Cents a Task and Replaces DeepSeek's Own Flagship
Released Wednesday. 552 billion parameters, mixture of experts, with a new encoder–decoder split that runs 8 billion active parameters on input and 16 billion on output, plus native image understanding and a 1 million token context. The KV cache needs a quarter of the HBM and an eighth of the SSD of the previous generation. Artificial Analysis scores it 40 on its Intelligence Index at $0.27 per task, at $0.30 per million input and $1.20 per million output, with off-peak rates at half that. DeepSeek's own table claims 74.2 on DeepSWE v1.1, level with GPT-6 Astra and Opus 5; nobody independent has run it yet, and the SVG output does not look like a frontier model's. The telling line is that DeepSeek is retiring V4-Pro: from 14 September all V4-Pro requests route to V4.1 Flash at Flash prices, and the weights are already on Hugging Face.
🔵 ChatGPT Images 2.5 Ships With a Sketchpad, and Two API Models
OpenAI says people make more than 3 billion images a week across ChatGPT and the API, and the new model cuts generation latency by up to 50% versus Images 2.0. The real improvements are boring and useful: subjects from reference photos stay recognisable, edits change only what you asked for, quality holds up over multiple turns, and transparent backgrounds work. Sketch lets you draw the layout in ChatGPT and use it as the reference. Developers get GPT-Image-2.5 Flare as the default and GPT-Image-2.5 Sunburst for slower, higher-precision work. Microsoft answered the same week with MAI-Image-2.6, which grounds image generation in live web content and ships a Flash variant at under half the flagship's price.
🟤 IDScan Confirms 150 Million Driver's Licences Were Stolen
IDScan, the Louisiana identity-checking service used by venues and dispensaries to scan your licence at the door, has confirmed hackers took driver's licence records from its cloud: full names, licence numbers, and identity numbers from other government documents including passports. Brian Krebs reported it on 1 September after finding a dark-web site that let anyone search over 150 million US and Canadian licence records, photos included. He verified his own record. The US Secretary of Defense was in there too. The FBI is investigating. IDScan's notice says "full access to the information required payment", and the company has not said how many people are affected beyond noting it holds over 150 million records.
- TechCrunch: IDScan confirms data breach with more than 150 million driver's licences stolen
- Krebs on Security: FBI probes service selling 153M driver's licenses
⌚ Apple Watch Series 12 Will Transcribe the Last 15 Seconds of Whatever You Just Heard
Audio Intelligence is the new thing on Series 12 and Ultra 4. Live Rewind: double-press the crown and the previous 15 seconds of conversation appear as text. Siri Recap: opt in, and Apple Intelligence writes a title and key points for your day's conversations. Audio is processed in a hardware-isolated Secure Exclave on the S11 chip and deleted, no recording is stored, speakers are not identified, Live Rewind chimes even on silent, and the outputs are end-to-end encrypted. The EFF's Adam Schwartz is unmoved: "people don't really have a practical means to consent or decline recording", and 11 US states require all-party consent. Beta late 2026, English first, not initially in the EU, and subject to daily usage limits. Also from the event: Siri AI lands 14 September as a beta with daily caps, iPhone 18 Pro's sensor signs every pixel for an unalterable reference image, the foldable iPhone Duo ships 23 October, and AirPods 5 cost $129.
- Apple: Introducing Apple Watch Series 12
- The Register: Apple timepiece can grab snippets of conversation without both speakers' consent
- AppleInsider: Siri AI will launch in beta, complicated by daily usage caps
- Apple: iPhone 18 Pro and iPhone 18 Pro Max
- Apple: iPhone Duo
- Apple: AirPods 5
🎬 DaVinci Resolve 21.1 Ships a Native MCP Server. Claude Code Can Now Edit Your Video.
Blackmagic's 21.1 update adds "AI assistant integration", which in practice is a Model Context Protocol server built into Resolve Studio that exposes the scripting API. Blackmagic names Claude, Claude Code and ChatGPT Codex as the supported assistants, and the examples are cutting highlight reels from long-form footage, removing unwanted clips and batch rendering deliverables. Twenty new scripting calls landed alongside it, including transcriptions with speaker and timing data, multicam creation and auto-align, which is exactly the toolkit an agent needs. Studio only. Python scripting moved to the paid edition in the same release.
- Blackmagic Design: What's New in DaVinci Resolve 21.1
- CineD: DaVinci Resolve 21.1 released, AI assistant integration via MCP
💬 Two Claude Code Sessions Can Now Talk to Each Other. Here's What Actually Changed in September.
If you run more than one Claude Code session on a project, you have been the message bus, copying findings between terminals. I went through every official release note up to 11 September to separate what shipped in September from the August groundwork. Cross-session SendMessage and ListAgents arrived in August: one session can hand a specific finding to another by name. So did notify_when_idle, which tells your main conversation when another local session finishes a turn, and an idle notice does not mean its tests passed. September removed the one-hour limit on background commands started by subagents and added likely causes to the prompt-cache miss diagnostics. I show the workflow on a small parser project so you can try it without a team of agents.
- Claude Code Updates: September 2026 — by Nathan House
🩶 OpenAI Stopped Selling the $200 Plan Because Astra Is Too Popular
New Pro subscriptions are paused. Thibault Sottiaux, who runs Codex and ChatGPT product, said the Pro tier puts the most strain on OpenAI's systems and that this was "the smallest step that allows us to continue giving the broadest access possible". Go, Plus and the API stay open. He had warned on Wednesday that Astra demand was "really unprecedented" and unlike anything he had seen through OpenAI's earlier growth. OpenAI has not said when Pro sign-ups reopen, or how many people were signing up each day.
✍️ ChatGPT Work Now Learns Your Writing Style From Your Own Email, and Runs Your Warehouse Queries
Two ChatGPT Work releases. Writing style: connect Gmail, Google Drive, Slack and SharePoint and it learns your phrases, your sign-off and your capitalisation quirks, then applies them to everything it writes. Setup is in Settings, Personalization, Writing style, on all paid plans. Data agent: it connects to Amazon Redshift, Datadog, Google BigQuery, ClickHouse, Databricks, MongoDB and Snowflake, uses your semantic layer and metric definitions, and builds shareable dashboards from a question. Admins choose which connections exist and which roles can use them, and queries enforce the connected account's existing table, row and column permissions. That last sentence is the one to test before anyone in finance gets access.
🟦 Gemini Gets a Windows App and Lyria 3.5
The Gemini desktop app reached Windows on Wednesday. Alt + Space opens it over whatever you are doing, it hands multi-step tasks to Gemini Spark, and it pulls from Gmail and Drive. Lyria 3.5, Google's music model, is now in the Gemini app and API for everyone, with genre and vocal-or-instrumental selection, templates, and a choice of short or longer tracks. It is also in Flow Music, AI Studio and Vids.
🎵 Suno Retrained From Scratch on Licensed Music and Is Retiring Everything Else
Suno v6 was built with Warner Music Group, BMG and Believe, and Suno says it does not use the training data behind previous versions. Three models: v6 and v6-wild for Pro and Premier subscribers, v6-mini free for everyone, with plain-language edits to one section of a song, single-lyric changes, and text, audio, image or video as input. The older models get retired as v6 rolls out. The context is the legal ledger: Warner settled last year, BMG signed last month, and Sony and Universal are still suing, along with Jason Isbell and a class of users over a data breach. The announcement landed one day after Suno admitted it had trained on YouTube videos.
- Suno: Introducing v6
- TechCrunch: Suno replaces its AI models with a new one trained on licensed music
💰 Sabine Hossenfelder Was Offered Money to Tell You AI Will Kill Us
The physicist and YouTuber says an organisation approached her to fund a video warning of an AI apocalypse, with the script supplied down to the sentences to read and the references to cite. She declined and did not name them. She points to a YouTuber who has fielded three similar pitches, one tied to the Center for AI Safety's outreach, and to Control AI and the Future of Life Institute's "Protect What's Human" campaign as funders of similar content. The other side pays too. Build American AI, backed by OpenAI's Greg Brockman, offered Taylor Lorenz $5,000 for one TikTok. Lorenz declined as well, and Hossenfelder says other creators reportedly took similar money and disclosed an ad without saying who was paying for it.
- Sabine Hossenfelder on X
- OfficeChai: Some organisations are paying influencers to scare people about AI
🍿 The Sam Altman Movie Is Out at Christmas, and Amazon Wanted Nothing to Do With It
Neon released the first teaser for Artificial on Monday. Luca Guadagnino directs, Simon Rich wrote it, and Andrew Garfield plays Altman through the November 2023 week he was fired and reinstated, with Yura Borisov as Ilya Sutskever, Ike Barinholtz as Elon Musk, Monica Barbaro as Mira Murati and Mark Rylance as Geoffrey Hinton. Amazon MGM developed the film and dropped it in June, shortly after committing up to $50 billion to OpenAI, saying it "will be better served if it were released by a different studio". Netflix, A24 and Focus passed. Trailer line: "Countries will fall. Industries are gonna collapse. And we get to be the ones to shepherd people into this new world." US release 25 December after a New York Film Festival premiere.