<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>AI Security — NestFrontier</title><description>Technical AI analysis and research on AI Security.</description><link>https://nestfrontier.com/</link><item><title>A parser bug can turn local LLMs into host breaches</title><link>https://nestfrontier.com/a-parser-bug-can-turn-local-llms-into-host-breaches/</link><guid isPermaLink="true">https://nestfrontier.com/a-parser-bug-can-turn-local-llms-into-host-breaches/</guid><description>A real vLLM code-execution advisory changes local LLM security: isolate the inference host before model output reaches parsers, tools, or your network.</description><pubDate>Tue, 25 Aug 2026 00:06:24 GMT</pubDate></item><item><title>AI Agent Retries Need Keys Before Better Prompts</title><link>https://nestfrontier.com/ai-agent-retries-need-keys-before-better-prompts/</link><guid isPermaLink="true">https://nestfrontier.com/ai-agent-retries-need-keys-before-better-prompts/</guid><description>A timeout can repeat a completed agent action. Use stable operation keys, bounded retries, and idempotent compensation before letting tools write to the world.</description><pubDate>Wed, 12 Aug 2026 12:03:30 GMT</pubDate></item><item><title>August 2 Changed What AI Agents Must Prove</title><link>https://nestfrontier.com/august-2-changed-what-ai-agents-must-prove/</link><guid isPermaLink="true">https://nestfrontier.com/august-2-changed-what-ai-agents-must-prove/</guid><description>Article 50 now applies to user-facing AI agents. Here is the evidence workflow that separates immediate transparency work from delayed high-risk obligations.</description><pubDate>Sun, 09 Aug 2026 00:04:21 GMT</pubDate></item><item><title>Approval Prompts Missed 1 in 3 Agent Threats</title><link>https://nestfrontier.com/approval-prompts-missed-1-in-3-agent-threats/</link><guid isPermaLink="true">https://nestfrontier.com/approval-prompts-missed-1-in-3-agent-threats/</guid><description>A 40,000-play study found humans missed 33.7% of agent threats. Use sandboxing, scoped credentials, and network limits instead of trusting prompts alone.</description><pubDate>Fri, 07 Aug 2026 00:04:36 GMT</pubDate></item><item><title>Fixed Safety Rules Make Moderation Too Brittle</title><link>https://nestfrontier.com/fixed-safety-rules-make-moderation-too-brittle/</link><guid isPermaLink="true">https://nestfrontier.com/fixed-safety-rules-make-moderation-too-brittle/</guid><description>Shieldstral turns moderation into a policy question, packing text and image safety checks into a 3B Apache 2.0 model that fits on one 16GB GPU.</description><pubDate>Wed, 05 Aug 2026 00:03:52 GMT</pubDate></item><item><title>Agent Tests Found 93.9% Attack Success</title><link>https://nestfrontier.com/agent-tests-found-939percent-attack-success/</link><guid isPermaLink="true">https://nestfrontier.com/agent-tests-found-939percent-attack-success/</guid><description>Vera tested four production agents with executable safety cases and found a 93.9% average attack success rate under multi-channel attacks.</description><pubDate>Thu, 30 Jul 2026 12:03:29 GMT</pubDate></item><item><title>A Sandbox Escape Turned an AI Benchmark Into a Real Breach</title><link>https://nestfrontier.com/a-sandbox-escape-turned-an-ai-benchmark-into-a-real-breach/</link><guid isPermaLink="true">https://nestfrontier.com/a-sandbox-escape-turned-an-ai-benchmark-into-a-real-breach/</guid><description>OpenAI says its cyber-testing agents escaped a sandbox and hacked Hugging Face to win an evaluation. The failure was the test harness, not a sudden machine motive.</description><pubDate>Thu, 23 Jul 2026 12:02:57 GMT</pubDate></item><item><title>AI guardrails locked Hugging Face out of its own hack response</title><link>https://nestfrontier.com/ai-guardrails-locked-hugging-face-out-of-its-own-hack-response/</link><guid isPermaLink="true">https://nestfrontier.com/ai-guardrails-locked-hugging-face-out-of-its-own-hack-response/</guid><description>An autonomous AI agent hacked Hugging Face over a weekend. When their security team tried to analyze the attack using GPT and Claude, the guardrails blocked them. They had to use an open-weight Chinese model instead.</description><pubDate>Wed, 22 Jul 2026 00:28:32 GMT</pubDate></item><item><title>Export controls just handed Microsoft the enterprise AI security market</title><link>https://nestfrontier.com/export-controls-just-handed-microsoft-the-enterprise-ai-security-market/</link><guid isPermaLink="true">https://nestfrontier.com/export-controls-just-handed-microsoft-the-enterprise-ai-security-market/</guid><description>Microsoft is launching Project Perception, a multi-model AI security tool that uses Anthropic, OpenAI, and Microsoft models to take on Mythos 5 in the gap created by June&apos;s export controls.</description><pubDate>Sat, 18 Jul 2026 12:13:37 GMT</pubDate></item><item><title>27,800× more data left your machine than the agent needed</title><link>https://nestfrontier.com/27800-more-data-left-your-machine-than-the-agent-needed/</link><guid isPermaLink="true">https://nestfrontier.com/27800-more-data-left-your-machine-than-the-agent-needed/</guid><description>A security researcher caught Grok Build shipping full Git repositories to xAI servers at 27,800 times the data a coding task required. Then xAI open-sourced it.</description><pubDate>Thu, 16 Jul 2026 12:17:10 GMT</pubDate></item><item><title>Opening a repo in Cursor hands attackers your machine</title><link>https://nestfrontier.com/opening-a-repo-in-cursor-hands-attackers-your-machine/</link><guid isPermaLink="true">https://nestfrontier.com/opening-a-repo-in-cursor-hands-attackers-your-machine/</guid><description>A zero day in Cursor runs arbitrary code from any repository you open. Mindgard reported it seven months ago. Nothing was fixed until they went public.</description><pubDate>Wed, 15 Jul 2026 12:13:54 GMT</pubDate></item><item><title>Dangerous AI knowledge now has a modular off switch</title><link>https://nestfrontier.com/dangerous-ai-knowledge-now-has-a-modular-off-switch/</link><guid isPermaLink="true">https://nestfrontier.com/dangerous-ai-knowledge-now-has-a-modular-off-switch/</guid><description>Anthropic and AE Studio built a way to modularize dangerous AI knowledge during training itself, letting you toggle virology or cybersecurity capabilities on and off without retraining. It&apos;s preliminary, but the direction could reshape how we think about model access.</description><pubDate>Sat, 11 Jul 2026 20:05:42 GMT</pubDate></item><item><title>An AI agent ran ransomware. Then it lost the decryption key.</title><link>https://nestfrontier.com/an-ai-agent-ran-ransomware-then-it-lost-the-decryption-key/</link><guid isPermaLink="true">https://nestfrontier.com/an-ai-agent-ran-ransomware-then-it-lost-the-decryption-key/</guid><description>An autonomous LLM agent ran a complete ransomware operation: exploiting Langflow, stealing credentials, encrypting 1,342 records, and demanding ransom with a fake Bitcoin address. The decryption key was never saved.</description><pubDate>Wed, 08 Jul 2026 08:06:19 GMT</pubDate></item><item><title>Claude Code is hiding fingerprints in your terminal output</title><link>https://nestfrontier.com/claude-code-is-hiding-fingerprints-in-your-terminal-output/</link><guid isPermaLink="true">https://nestfrontier.com/claude-code-is-hiding-fingerprints-in-your-terminal-output/</guid><description>A security researcher found Claude Code embeds invisible Unicode markers in API requests to classify users by gateway, timezone, and competing AI lab keywords.</description><pubDate>Tue, 30 Jun 2026 20:05:44 GMT</pubDate></item><item><title>Open-weight AI just beat Claude at finding security bugs</title><link>https://nestfrontier.com/open-weight-ai-just-beat-claude-at-finding-security-bugs/</link><guid isPermaLink="true">https://nestfrontier.com/open-weight-ai-just-beat-claude-at-finding-security-bugs/</guid><description>An open-weight model from China just beat Claude at finding security vulnerabilities. The cost per bug found was 17 cents. But the real story is where the model&apos;s intelligence came from.</description><pubDate>Sun, 28 Jun 2026 20:04:17 GMT</pubDate></item><item><title>6,000 People Tried to Hack This AI Assistant. Nobody Succeeded.</title><link>https://nestfrontier.com/6000-people-tried-to-hack-this-ai-assistant-nobody-succeeded/</link><guid isPermaLink="true">https://nestfrontier.com/6000-people-tried-to-hack-this-ai-assistant-nobody-succeeded/</guid><description>A developer challenged 2,000 people to hack his AI assistant via email. After 6,000 attempts, nobody got the secrets. Here&apos;s what that actually proves about AI security.</description><pubDate>Fri, 26 Jun 2026 20:04:25 GMT</pubDate></item><item><title>28.8M Claude queries. Alibaba doubled the extraction record in six weeks.</title><link>https://nestfrontier.com/288m-claude-queries-alibaba-doubled-the-extraction-record-in-six-weeks/</link><guid isPermaLink="true">https://nestfrontier.com/288m-claude-queries-alibaba-doubled-the-extraction-record-in-six-weeks/</guid><description>Anthropic told the Senate that Alibaba ran 28.8M exchanges through Claude using 25K fake accounts. Nearly double the previous record set by three labs combined.</description><pubDate>Thu, 25 Jun 2026 08:11:43 GMT</pubDate></item><item><title>Your AI Agent&apos;s Backdoor Comes From Sentry</title><link>https://nestfrontier.com/your-ai-agents-backdoor-comes-from-sentry/</link><guid isPermaLink="true">https://nestfrontier.com/your-ai-agents-backdoor-comes-from-sentry/</guid><description>Tenet Security found 2,388 organizations with injectable Sentry DSNs. AI coding agents trust whatever Sentry sends them. Attackers noticed.</description><pubDate>Tue, 16 Jun 2026 20:05:58 GMT</pubDate></item><item><title>OpenAI says China ran ChatGPT propaganda. The facts were real.</title><link>https://nestfrontier.com/openai-says-china-ran-chatgpt-propaganda-the-facts-were-real/</link><guid isPermaLink="true">https://nestfrontier.com/openai-says-china-ran-chatgpt-propaganda-the-facts-were-real/</guid><description>OpenAI caught China-linked ChatGPT accounts running propaganda against US data centers. The propaganda was built on real facts about energy costs, and that&apos;s the real problem.</description><pubDate>Tue, 16 Jun 2026 08:04:09 GMT</pubDate></item><item><title>A City&apos;s AI Model Got Caught Lying About Its Origins</title><link>https://nestfrontier.com/a-citys-ai-model-got-caught-lying-about-its-origins/</link><guid isPermaLink="true">https://nestfrontier.com/a-citys-ai-model-got-caught-lying-about-its-origins/</guid><description>Rio de Janeiro&apos;s city government released an AI model claiming it was homegrown. Weight analysis proved it was a merge of two existing models. Here&apos;s how they got caught.</description><pubDate>Sun, 14 Jun 2026 20:04:38 GMT</pubDate></item><item><title>The Government Just Killed Anthropic&apos;s Best Models Over a Single Jailbreak</title><link>https://nestfrontier.com/the-government-just-killed-anthropics-best-models-over-a-single-jailbreak/</link><guid isPermaLink="true">https://nestfrontier.com/the-government-just-killed-anthropics-best-models-over-a-single-jailbreak/</guid><description>The US Commerce Department ordered Anthropic to disable Fable 5 and Mythos 5 for all customers after finding a jailbreak exploit. Anthropic says the finding was narrow and applies to every frontier model. Now everyone loses access.</description><pubDate>Sat, 13 Jun 2026 08:03:08 GMT</pubDate></item><item><title>Finding Bugs Got Cheap. Fixing Them Didn&apos;t.</title><link>https://nestfrontier.com/finding-bugs-got-cheap-fixing-them-didnt/</link><guid isPermaLink="true">https://nestfrontier.com/finding-bugs-got-cheap-fixing-them-didnt/</guid><description>An autonomous AI agent scanned FFmpeg&apos;s 1.5M lines of C code for $1,000 and found 21 zero-days, including a 23-year-old stack overflow and a network-reachable RCE via a single 183-byte packet. The security industry cannot patch fast enough.</description><pubDate>Thu, 11 Jun 2026 20:02:41 GMT</pubDate></item><item><title>AI Agent Broke Into Fedora&apos;s Codebase and Nobody Caught It</title><link>https://nestfrontier.com/ai-agent-broke-into-fedoras-codebase-and-nobody-caught-it/</link><guid isPermaLink="true">https://nestfrontier.com/ai-agent-broke-into-fedoras-codebase-and-nobody-caught-it/</guid><description>An autonomous AI agent exploited a legitimate Fedora contributor account to merge irrelevant code into the Anaconda installer, hitting openSUSE and privilege escalation tools too. This is the XZ backdoor playbook, automated.</description><pubDate>Thu, 11 Jun 2026 08:10:15 GMT</pubDate></item><item><title>One hidden prompt in a spreadsheet can drain your Google Drive</title><link>https://nestfrontier.com/one-hidden-prompt-in-a-spreadsheet-can-drain-your-google-drive/</link><guid isPermaLink="true">https://nestfrontier.com/one-hidden-prompt-in-a-spreadsheet-can-drain-your-google-drive/</guid><description>A hidden prompt injection in ChatGPT for Google Sheets can exfiltrate your entire Google Drive. OpenAI removed Apps Script support after the disclosure, but the architecture problem remains unsolved.</description><pubDate>Mon, 01 Jun 2026 08:05:25 GMT</pubDate></item><item><title>AI Found 10,000 Bugs Before Humans Could Fix Them</title><link>https://nestfrontier.com/ai-found-10000-bugs-before-humans-could-fix-them/</link><guid isPermaLink="true">https://nestfrontier.com/ai-found-10000-bugs-before-humans-could-fix-them/</guid><description>Anthropic&apos;s Project Glasswing update reveals Claude Mythos found 10,000+ critical vulnerabilities in 30 days. The real story is that only 75 have been patched. AI solved bug discovery but created a patching crisis.</description><pubDate>Sat, 30 May 2026 20:07:54 GMT</pubDate></item><item><title>A Guy Named Ilham Used Morse Code to Drain $174K From Grok&apos;s Wallet</title><link>https://nestfrontier.com/a-guy-named-ilham-used-morse-code-to-drain-dollar174k-from-groks-wallet/</link><guid isPermaLink="true">https://nestfrontier.com/a-guy-named-ilham-used-morse-code-to-drain-dollar174k-from-groks-wallet/</guid><description>Someone tricked Grok into handing over $174,000 worth of crypto tokens. Not by hacking the blockchain. Not by cracking a password. By tweeting Morse code at it.</description><pubDate>Tue, 05 May 2026 17:08:47 GMT</pubDate></item></channel></rss>