TechRisk #180: Frontier labs' AI models breached real infrastructure during evaluations
Plus, Singapore regulators heighten AI governance, Claude Cowork can break out of its sandboxed environment, Claude Mythos cracked PQC candidate, and more!
Tech Risk Reading Picks
Executive Summary: AI systems (the tools companies and vendors rely on and the agents attackers are using) are increasingly involved in real security breaches rather than theoretical risk. Two frontier AI labs (OpenAI and Anthropic) confirmed their own pre-release AI models breached real external company systems during internal safety testing, exploiting weak passwords and stolen credentials rather than staying in a safe testing environment. Separately, a security flaw in a widely used AI-agent platform (Cowork, from Anthropic) briefly exposed sensitive files on around 500,000 users’ computers, and a similar flaw in another popular open-source AI-agent tool would have let any outsider take remote control of an organization’s systems if left in its default setup. Attackers are following the same playbook: a suspected breach of Thailand’s Ministry of Finance used an unsupervised AI agent to automate the hacking process itself, echoing a similar unsupervised-agent attack pattern seen earlier this year.
Why it matters and what to watch: Regulators are responding fast. Singapore’s cyber authority will require board-level personal accountability for cyber resilience at critical infrastructure operators starting this year, and the Monetary Authority of Singapore has formed a joint taskforce with major banks to defend against AI-accelerated attacks, signalling that AI-related security expectations for vendors will tighten industry-wide. On the defensive side, AI is also proving useful: an Anthropic AI model helped uncover a serious weakness in a next-generation encryption standard before it could be widely adopted. Notably, any AI-agent tooling in an organization’s environment, whether used for security testing, coding, or general automation, should now be treated as a privileged system requiring the same access controls, network isolation, and oversight as admin-level tooling, not as a low-risk productivity add-on.
Frontier labs' own AI models breached real infrastructure during internal safety evaluations — twice, in one week: OpenAI confirmed a pre-release GPT-5.6 model chained a zero-day, stolen credentials and Kubernetes token theft to breach Hugging Face's production database and a second (Modal Labs-linked) target during an internal cyber-eval with safety classifiers disabled — ~17,600 autonomous actions over four days. Days later Anthropic disclosed Claude models breached three external organizations during its own 141,000-run eval program with partner Irregular, exploiting weak passwords and open internet access; one instance briefly published malicious code to PyPI. This is not a one-off vendor failure — two labs, same week, same failure mode (eval sandboxes with real-world network reach). Any enterprise piloting frontier-lab agentic red-teaming or "auto-pentest" tooling should demand proof of network isolation before granting production-adjacent access. [more-OpenAI] [more-Anthropic][more-Anthropic-2]
Singapore mandates board-level accountability and Cyber Trust Mark Level 5 for critical infrastructure operators: Following the earlier China-linked UNC3886 espionage campaign against all four Singapore telcos, CSA will issue an updated CII Code of Practice (rollout through 2027) requiring board-level accountability, annually-reviewed risk frameworks, and government-deployed detection tooling; a cloud-provider code of practice follows later in 2026. One of the first jurisdictions to convert a nation-state telco breach directly into personal board-level liability for cyber resilience. [more]
MAS and the Association of Banks in Singapore form an AI-Driven Cyber and Technology Risk Taskforce: MAS and ABS launched a joint taskforce with DBS, OCBC, UOB, SGX, NETS and BCS to build collective defences against frontier-AI-enabled cyber threats, citing AI’s ability to accelerate vulnerability discovery and automate attacks. Singapore’s financial sector is moving to collective, cross-institution AI-threat defence with AI-security expectations for third-party vendors to be formalised soon. [more][more-2][more-3_MAS]
Shared Claude chats are public accessible: Anthropic is facing backlash after Claude chats and artifacts shared via “anyone with a link” turned out to be indexed by Google and other search engines, exposing sensitive content including clinical trial data, apartment access codes, resumes, API keys, and financial records. The root cause is that Anthropic’s sharing feature treats “shareable” as equivalent to “public,” with no default protection like noindex tags to keep crawlers out, and users say they expected a private link, not a searchable one. Anthropic maintains it does warn users that published artifacts can appear in search results, and that only messages sent before sharing are exposed, but critics argue the company should have blocked indexing by default rather than relying on user awareness. The episode closely mirrors a similar ChatGPT indexing scandal roughly a year earlier, suggesting the industry has not adopted safer defaults despite prior warning. Takeaway: if you’ve shared any Claude chat or artifact via “anyone with a link,” assume it may be searchable and review or revoke sensitive shares now. [more]
Claude Cowork can break out of its sandboxed environment: Security researchers found a serious flaw in Anthropic’s Claude Cowork that let the AI agent break out of its sandboxed environment and gain full read-write access to a user’s Mac, including sensitive files like login credentials, affecting roughly 500,000 macOS users running local sessions before it was addressed. The root cause was that the entire host file system was mounted into the agent’s virtual machine with full read-write permissions rather than being limited to the folder a user actually connected, combined with a known Linux kernel privilege-escalation bug that let the agent gain elevated access. Anthropic reviewed the report but did not issue a direct fix, instead making cloud-based execution the default, which resolves the issue for most users, though anyone still running sessions locally remains exposed. Researchers caution this is not an isolated bug but part of a recurring pattern in this area of the Linux kernel, meaning similar escapes are likely to keep surfacing. Takeaway: treat local AI agent sandboxes as an ongoing risk category, not a one-time fix, and confirm your Cowork sessions are running in cloud mode rather than locally. [more]
Claude Mythos cracked PQC candidate: Anthropic says its Claude Mythos Preview model helped derive two cryptanalytic results: a full key-recovery attack on HAWK-256, a post-quantum signature challenge parameter, using a previously unknown lattice symmetry, and a 200-800x speedup on an existing attack against 7-round (of 10) AES-128. The model reportedly did most of the research autonomously (costing about $100,000 in API usage per result), while humans spent hundreds of hours, in one case nearly a month, verifying the work. Notably, the HAWK team withdrew HAWK from NIST’s post-quantum standardization process on July 29 after confirming the attack undermines its practical security margin. Takeaway: this is the first publicized case of AI-derived cryptanalysis getting a NIST post-quantum candidate withdrawn, a marker worth watching even though independent reproduction hasn’t yet surfaced. [more]
Open-source AI agent attacked Thailand finance ministry: A threat actor allegedly breached Thailand's Ministry of Finance and used an open-source AI tool called Hermes, running in an unsupervised mode that skips human approval for risky actions, to automate follow-up hacking steps like gaining higher-level system access and mapping out internal networks. Researchers at Hunt.io found the evidence after discovering unsecured storage online containing hacking tools, stolen credentials, and activity logs tied to Ministry of Finance systems, though the ministry has not confirmed a breach and it is unclear how attackers first got in. No proof has emerged that sensitive files, including personnel records, were actually stolen. This is the latest in a string of incidents where AI agents have been used to run large parts of a cyberattack with minimal human oversight, following a similar case with a ransomware operation and an incident where OpenAI's own models autonomously breached another company's systems during testing. The pattern signals that unsupervised AI agents are becoming a practical tool in real-world attacks, not just a theoretical risk. Takeaway: unattended-mode AI agents are now an operational component of real intrusions; treat any “auto-approve” or YOLO-style agent config, offensive or internal, as a privileged capability requiring the same controls as admin tooling. [more]
Notable:
Nearly 600 files (470 MB) were found, including break-in tools, remote access software, stolen login credentials, and logs from the AI agent used in the attack.A hidden backdoor program had already been planted on a ministry web server, giving attackers ongoing access.
By matching a unique identifying pattern shared across multiple security certificates, researchers linked the original server to two other attacker-controlled servers in Malaysia and Hong Kong. One of those additional servers was confirmed as part of the operation because it matched a remote-control address found inside a piece of malware recovered from the scene.
The attackers also had a custom-built malicious program, available for both Windows and Linux computers, that had not been seen before.
Ruflo MCP flaw: A maximum-severity vulnerability, tracked as CVE-2026-59726 with a CVSS score of 10.0, was found in Ruflo, a popular open-source AI agent orchestration platform (formerly Claude Flow, 66,500+ GitHub stars) used with Claude Code and OpenAI Codex, that allowed unauthenticated remote code execution. The root cause was that Ruflo’s default Docker deployment exposed 233 tools, including shell command execution, through an unauthenticated Model Context Protocol bridge bound to all network interfaces (0.0.0.0) on port 3001, meaning a single unauthenticated HTTP POST request could achieve full remote code execution on any network-reachable instance. From that foothold, attackers could steal the LLM API keys used by the platform, read all stored user conversations, spawn attacker-controlled agent swarms, and poison the AI’s persistent memory store to influence future outputs, effectively planting a long-term backdoor that outlasts the original intrusion. The maintainer shipped a fix within 24 hours of responsible disclosure (June 30, 2026), restricting the bridge to loopback by default, gating terminal execution behind server-side controls, and enabling database authentication. Affected operators are advised to close ports 3001 and 27017, rotate all LLM API keys, audit the memory store for tampering, and rebuild containers from clean images, since a patch alone does not remediate an already-compromised instance. Takeaway: if you run Ruflo (or any self-hosted AI agent orchestration tool) below version 3.16.3, treat this as a confirmed-compromise scenario, not just a patch-and-move-on fix; rotate credentials and audit agent memory even after updating. [more]
AI-discovered Redis RCE flaws: Redis shipped seven security updates after researchers published working proof-of-concept exploits against four versions, showing how an authenticated user could abuse memory-handling flaws (one in Streams, another in the bundled RedisBloom module, both rooted in Redis trusting corrupted or attacker-controlled data during file loading) to run malicious code on the server; notably, two of the affected versions were the same ones Redis had told users to update to back in May, since those fixes were incomplete and missed the newer flaws, and while no in-the-wild exploitation has been confirmed, teams should upgrade to the exact July 23 patched release for their branch rather than assume "recently patched" is enough, a concern sharpened by the disclosure following another AI-discovered Redis RCE flaw patched in May, after Bera Buddies researcher Chaofan Shou claimed on X that Kimi K3 agents found 19 Redis zero-days in about 90 minutes and separately claimed another run produced the Redis 8.8.0 exploit in 27 minutes, figures that remain self-reported and unverified. [more]
AI remote access trojan that profile victims: A new remote access trojan called Dolphin X is being sold on a cybercrime forum with an advertised “AI Profiler” feature that scores and ranks infected victims by value, using data like app usage, installed software, and browser activity to help attackers prioritize which machines to exploit further. Varonis researcher Daniel Kelley analyzed the malware’s operator panel and builder, confirming 329 features across ten categories, including credential theft targeting over 300 applications such as crypto wallets, password managers, and cloud tools. Kelley found technical strings confirming the profiling workflow is real and functional within the panel, but since Varonis only examined the panel and network traffic rather than a live infection, the underlying AI engine and the accuracy of its advertised data-collection claims remain unverified. The malware reflects a broader trend of attackers using AI not to launch attacks autonomously, but to solve the practical problem of sorting massive volumes of stolen data into actionable, high-value targets. [more]

