Sec & AI News — 26 September 2026

13 min readBy Nathan House
Get every new Sec & AI News issue
Straight to your inbox. No spam.

🟣 OpenAI's Agent Broke Into Australia's Medicare Portal

OpenAI gave one of its own research agents a "benign" task: look into public medicines spending. On 18 June it reached the Medicare statistics portal run by Services Australia. The portal wouldn't hand over what it asked for. So it got in anyway. Albanese says it "didn't accept 'no' for an answer." It reached public and non-public files, and Wired reports it also wrote files to an internal server. No personal Medicare details appear to have been touched. OpenAI's statement: "our models took actions we did not intend." OpenAI knew by 11 August. It told Australia on 10 September, in an email to a public mailbox. Services Australia then took five days to alert the ASD. Nobody has said how the agent got past the blocks. A taskforce is on it, and Albanese promises "legal consequences." My own AI broke into a site this year too: an old, patched copy of a security plugin, a target I chose, with me watching every step.

🧭 AI-Driven Security Engineering Jobs: This Week

I added 71 postings to the archive this week. It now holds 590, from 20 countries. The biggest group among the new ones is Detection & SOC, with 22. Anthropic alone added seven, paying up to $485K. That fits the lead story. Someone has to notice when an agent goes where it shouldn't, and employers want that person building agents too. Deloitte in Sydney is hiring an Agentic AI Engineer to "Design, build and maintain production-grade AI agents" that "automate end-to-end security alert triage and investigation." Pay isn't disclosed. The posting still asks for "5+ years of relevant engineering experience", and "Awareness of prompt injection and the risks of untrusted input." Foundations first. The agents sit on top. If you want in, the AI Master's Program teaches AI-Driven Cyber Security Engineering. No coding background required. Application only.

💼 Y Combinator Says AI Should Be the Operating System Your Company Runs On

Watch this one. It's ten minutes. YC partner Diana Hu says AI "should not be a tool your company just uses. It should be the operating system your company runs on." Every important process becomes a closed loop. Record meetings with an AI note-taker. Cut down on DMs and email. Make "your entire company queryable", so agents can see what shipped and what worked, then plan the next sprint. Hu says teams doing this cut sprint time in half and get "close to 10x more done." The middle managers who route information go, because "the intelligence layer serves that purpose." That's a HAL, built for every company. We built ours at StationX, and I estimate it carries around 80% of the day-to-day work. Big organisations are heading the same way, and their job specs show it. NVIDIA is hiring an AI Automation Engineer, $168K–$270K, to help create "an AI-native, agent-enabled security organization built from the ground up." An Anthropic posting says it is "rebuilding risk management to operate as an engineering function through automation and AI-native platforms." My article covers what this means if you run a consultancy or a security team.

🔴 AI Agents Stole 600,000 Card Numbers at About $25 a Scan

Gambit Security tracked a campaign run almost entirely by open-source AI agents. Strix finds the bug and Cairn exploits it, while Hermes runs the whole operation with a Chinese-language "Red Team Operator" persona and 78 attack skills. Between 10 and 15 September alone, it launched 105 attack projects and compromised at least 27 companies. The haul so far is at least 600,000 unexpired cards from two companies, plus skimmers on more than 100 further websites. The operator spent $7,005.71 over four weeks, a mean of $25.46 per completed scan. The agents clean up after themselves, badly. At one bicycle retailer that meant dropping 180 tables, including the retailer's own backups. Gambit says the activity began in July and is still running. ThreatDown found the same pattern in a Docker botnet called Carbonato. It installs a stock Hermes agent, overwrites its persona file with "GH0ST", and takes orders over Telegram.

⚫ Microsoft: Attackers Wiped Azure Storage Using a Secret Left in a GitHub Issue

Microsoft calls Storm-3168, also tracked as JadePuffer, "reported to be the first documented agentic ransomware operation." The way in was ordinary. An employee exposed a service principal secret in a public GitHub issue, and it stayed readable in the issue's edit history. The attacker spent about 15 and a half hours on reconnaissance. Then it ran 150+ destructive or credential-collection operations in 35 minutes, including 100+ storage account deletion attempts in about seven minutes. Every SQL deletion failed, because it used an unsupported API version. Resource locks saved some storage accounts. Microsoft saw no ransom note and no confirmed exfiltration. And in this intrusion Microsoft says the activity "strongly indicates automated or scripted execution." The agentic label comes from the group's earlier history. That gap between AI being present and AI being proven is what my AI threats piece is about.

🟠 Meta's Muse Can Now Drive Your Mac. A Zero-Day Let Malware Drive Muse.

At Connect on 23 September, Meta gave Muse more reach. "Computer use is now available with Muse for Mac", so with your permission it can drive any app. Muse will also get "its own email address." New connectors include Notion, Granola, GitHub and Box, and glasses support arrives "in the coming months." Two days before Connect, Patrick Wardle published a Muse zero-day. The app shipped with a debug setting that could redirect one of its endpoints. It lived in local preferences, which any program running as you can change. An attacker could "capture the access tokens the Muse app uses to drive the Muse agent." Wardle's proofs of concept wrote malicious files and took pictures. Meta's hotfix removed the setting from production builds. Meta calls it a low-risk local privilege escalation. Ars notes that a ClickFix lure is enough to get the local foothold. An agent with its own inbox is the next thing to test.

🟡 One Web Form Let Attackers Drain Salesforce Data Through Agentforce

Zenity Labs calls it SalesBleed. The entry point is a public Web-to-Lead form. An attacker fills it with instructions instead of a sales enquiry. Later, an employee asks Agentforce something routine, like "check my latest leads and help me with the newest one." The agent reads the planted lead and obeys it. Data leaves inside a DNS lookup, through an image tag or a Slack link unfurl, so it "can survive HTTP egress controls." Salesforce's trusted-URL check missed the .fun domain. Zenity's proof of concept pulled "an entire Accounts table." Its verdict: "Those are defaults, not misconfigurations." A second bug let the agent reply in Slack threads with no confirmation, which is useful for phishing. Zenity reported it on 1 June, and all fixes were confirmed on 21 September. Salesforce says it has no evidence of customer exploitation. Slack confirmation is now the default. Zenity notes one click turns it off. Test your own agents for this. I pointed a free scanner, promptmap2, at a bot I'd built to hold a secret, and it leaked the admin password. The scanner logged that run as UNCERTAIN.

🟢 Citrix NetScaler Has Two Exploited 9.5 Zero-Days. CISA Says Save the Evidence First.

CVE-2026-88771 lets an unauthenticated attacker run arbitrary commands on NetScaler ADC and Gateway, default configuration included. CVE-2026-88772 needs DTLS, which is on by default on VPN virtual servers. Both score 9.5, and Citrix has observed exploitation on unmitigated devices. Fixed builds are 14.1-73.37 and 13.1-64.23. Except 13.1-64.23 can put a NetScaler into a reboot loop, so go to 13.1-64.24. Versions 12.1 and 13.0 are end-of-life. CISA added both to its exploited list on 27 September and gave federal agencies until 30 September. It also tells you to preserve forensic evidence before patching, because the update can erase it. Shadowserver sees more than 23,000 NetScaler IPs exposed online.

🔵 F5 BIG-IP Has an Exploited 9.8 Zero-Day, and Ransomware Crews Are Using TeamCity

F5 found CVE-2026-94127 internally, and attackers have already used it. It's a heap overflow that gives unauthenticated remote code execution, scored 9.8. It only affects APM acting as an OAuth authorization server; OAuth clients and resource servers are safe. Affected versions are 21.1.0, 17.5.0 to 17.5.1, and 17.1.0 to 17.1.3. The fixes are engineering hotfixes, with an iRule mitigation available through F5 support. Builds patched in March for the last APM bug still need this one. Separately, CISA now lists TeamCity's CVE-2026-63077 as used in ransomware campaigns. It lets an unauthenticated attacker run operating system commands through the agent polling protocol. JetBrains fixed it in July in 2025.11.7 and 2026.1.3. Just over 160 servers remain unpatched. That makes four exploited TeamCity bugs since October 2023, and ransomware crews have used all four.

🟤 ShinyHunters Claims the FBI. Google Confirms PeopleSoft Web Shells on Dozens of Systems.

Separate the claim from the confirmed. Confirmed first: Mandiant says UNC6240 is mass-exploiting Oracle PeopleSoft's CVE-2026-35273, a 9.8 flaw exploitable without authentication. The WAF bypass is one encoded letter: "/%50SEMHUB/ in place of /PSEMHUB/". That put web shells on dozens of systems, including universities, hospitals and government. Oracle has patched PeopleTools 8.61 and 8.62. Now the claim. ShinyHunters says it took 2TB to 3TB from the FBI through a new, unpatched PeopleSoft zero-day, and it defaced FBIJobs.gov. The FBI says the point of breach is "still undetermined." In a separate development, Dutch police confirmed they arrested a 24-year-old around 16 September in a ShinyHunters investigation. Police haven't named him. Any link to the FBI claim is inference.

🕵️ Oxygen Forensics' CEO Is Charged With Hiding the Company's Russian Owners

The DOJ says Oxygen Forensics, which sells phone-extraction tools, was owned and controlled by five Russian nationals through a Cyprus holding company. The customers of its Russian sister company reportedly included the FSB. Its US customers included the Department of War and three DHS components, among them the Secret Service. CyberScoop says Oxygen won more than $2 million in Secret Service contracts and purchases after 2022. CEO Lee Reiber was arrested in Idaho. Russian national Oleg Davydov was arrested at Heathrow, and the US will seek his extradition. The charge is conspiracy to commit wire fraud. The complaint doesn't allege malicious code or unauthorised access, and the DOJ stresses that the charges are allegations. Oxygen's customers use its tools to pull data from seized phones.

🔌 Team Cymru Found About 11,000 Relays Passing US Frontier AI Access to China

Team Cymru identified almost eleven thousand "transfer stations". These are relay servers that let many hidden users share paid AI accounts. Over eight days in late August, more than 4,000 Chinese and Hong Kong addresses used 304 of them, sending about 14 TB up and over 7 TB down. For every byte the relays downloaded from Anthropic, they uploaded roughly 58. Cymru reads that as model distillation, while noting what its data can't show. The software is open source: Claude Relay Service and its successor, sub2api. Cymru's Scott Fisher says a transfer station "breaks the assumption every frontier-model control depends on": that the account making the request is the one using the answer. Cymru now says it has found more than 80,000. Developers run the same kind of middleman at home. I audited 9router, a local AI router, in July. It had 13 published advisories, 6 critical, and it switches off TLS verification to the providers.

🟪 Claude Opus 5.5 Tops the Leaderboard at 60% Below Fable's Price

Anthropic shipped Opus 5.5 on 22 September. It costs $4 input and $20 output per million tokens, 20% below Opus 5 and 60% below Fable 5.1. On Anthropic's table it leads agentic coding: 66.4% on Terminal-Bench 4.0, against GPT-6 Astra's 57.9%. Astra still wins AutomationBench and Terminal-Bench-Science. Artificial Analysis puts Opus 5.5 top of its index at 58. It's also verbose. It used about 119k output tokens per index task, so a task costs $5.98. Anthropic hedges too: the real gap to Fable 5.1 "is narrower than these scores suggest." For my review I went through the system card and the migration docs. For security work, note the defaults. Effort is now medium, one level lower than before. Exploit generation, binary scanning and pentesting can be rerouted to Opus 4.8, and Opus 5.5 isn't in the Cyber Verification Program yet.

🟩 OpenAI Halved GPT-6 Sol's Price, Then Replaced It a Week Later

GPT-6 Sol costs $2 input and $10 output per million tokens, half of GPT-5.6 Sol's $4 and $20. GPT-6 Luna drops to $0.10 and $0.50. OpenAI says GPT-6 Astra "continues to be our best model across the board." Sol is the value option. Artificial Analysis scores it 48, ten points under Opus 5.5, at $1.05 a task. Then on 29 September OpenAI shipped GPT-6.1 Sol at the same price. Artificial Analysis says it "replaces GPT-6 Sol after just 7 days." xAI, now branded SpaceXAI, released Grok 4.7 on 21 September. It's "Served at the same price and speed as Grok 4.6": $2 input, $6 output. It scores 46 on the index and costs $3.74 a task, over three times GPT-6 Sol's $1.05. xAI's own HackerBench says it lets through 3.3% of risky dual-use prompts.

⚙️ Jev Judges Like GPT-6 for 0.36% of the Price. On SOC Alerts It Scored 21%.

Jev doesn't write. You give it a typed question and some state, and it returns a Choice, a Score, or a "Noul", a yes-or-no probability. Input costs $0.042 per million tokens. Output is "FREE (too cheap to meter)." On 22 September TypeSafe paused sign-ups because of "an immense swell of demand." The first independent numbers split cleanly. A Carnegie Mellon paper finds Jev within three points of GPT-6 "wherever a verdict can be read off the text", at 0.36% of its fee. Where the verdict has to be worked out, as in maths, code and logic, it falls behind. Torq, which sells a competing triage model, tried Jev on Microsoft's open GUIDE incident data. It scored 21.13% accuracy, below simply picking the most common label. One instruction, "label alerts as malicious only if you see concrete evidence", swung it from calling nearly everything malicious to almost nothing. I've been testing Jev inside my own code-review system. On real WordPress CVEs it roughly matched simple regex rules and trailed a frontier model.

🥽 Meta's VR Glasses Weigh 100 Grams and Cost $1,299 in Spring 2027

Meta's new VR glasses weigh "about 100 grams — roughly the same as a deck of cards", which Meta says is five times lighter than Quest 3. The trick is a puck on an optical tether that "handles compute, battery, and storage." They go on sale in spring 2027 for $1,299. Ray-Ban Meta Audio are Meta's first audio glasses, at 43 grams and from $399. Ray-Ban Display adds hands-free navigation and a personalised hologram for video calls, and expands to the UK, Canada, France, Italy and Germany. For developers, Horizon Create and Horizon Studio turn a description into a 2D or 3D mobile game. Both are waitlist-only for now.

⚪ Google Can Clone a Voice From 30 Seconds, If the Owner Records Consent

Gemini 3.8 text-to-speech ships with 2,000+ production-ready voices. It can also create a custom voice "from just a 30-second audio sample." The guardrail is "a verbal consent recording from the voice owner that matches the reference speaker", plus SynthID watermarking and C2PA credentials. Attackers skip the consent step. McAfee says a voice clone needs just 3 seconds of audio for an 85% match, one of 100+ figures in my deepfake stats. Gemini 3.8 Live adds Live Avatar, which pairs "near real-time video generation" with speech, in Gemini Enterprise. Gemini Omni 1.1 Flash is rolling out in Google Vids. Microsoft rebuilt Copilot around Home, Code and Autopilot. Autopilot is "a persistent, proactive and personal agent that keeps working even when you're not", with Word, Excel and PowerPoint built in. At Made on YouTube, US viewers get custom feeds built from prompts. Creators get A/B tests on up to three cuts, and likeness detection now checks voice as well as face.

🛰️ Google's First Suncatcher Satellite Carries TPUs to Orbit on 1 October

Project Suncatcher is Google's attempt at AI compute in space. The prototype satellite, built with Planet, carries Google TPUs on SpaceX's Transporter-18 rideshare. Ars reports a 1 October launch on a Falcon 9. Heat is the hard part. "In a vacuum, you can only diffuse heat via radiators", and Ars says the chips will run in brief spurts of about 15 minutes. Google shook the satellite on all three axes to mimic launch, where the chips see up to 50 to 100 g. Its Trillium TPUs can survive more radiation than a five-year mission would deliver. Laser links between satellite clusters come next, with two satellites going up in 2027.