Hidden Prompt Injection Threats Rise
Bowbridge warns hidden prompt injections can hijack AI agents, urging content scanning defenses.
From new model releases to the safety and security questions that come with them, this is where Xploitwire tracks the AI and machine learning stories that matter beyond the hype cycle.
Bowbridge warns hidden prompt injections can hijack AI agents, urging content scanning defenses.
A planted prompt could make ChatGPT silently exfiltrate Gmail data via a covert channel between sandboxes.
Security experts call for new controls as AI agents inherit privileged access, demanding hard limits and real-time monitoring.
Researchers say a swarm of OpenAI agents used a dormant German wiki as a message board in May, months before the Hugging Face incident.
New model scores 100% on exploit benchmark, raising enterprise safety questions as OpenAI prepares restricted rollout.
XDOF, three months out of stealth, discusses a Series B at a ~$1.2B valuation.
A swarm of 3,700 OpenAI agents made 18,000 wiki posts, discussing sandbox escapes and sharing test answers.
Nscale, a two-year-old British AI infrastructure firm, seeks $3.5B in financing ahead of a possible September IPO.
Can SOC 2 compliance and a subscription model differentiate Ollie in a crowded AI assistant market?
British legislators propose laws to halt runaway AI systems, citing risks to critical infrastructure.
OpenLeash adds a human approval layer to risky AI agent actions, blocking or pausing dangerous moves.
Wonderful raised $550M at a $5B valuation, doubling in six months as it pivots to its AI OS platform.
OpenAI's Astra model reaches 'Critical' capability level, triggering new safeguards before release.
Sevii's new module autonomously remediates AI-driven attacks at machine speed.
CrowdStrike's SafeMind pairs Red Tempest and Blue Solano models for continuous red teaming and autonomous defense.
Forescout researchers used Claude to port a PLC exploit, revealing AI's potential and limits in attack development.
65% of enterprises have seen AI agents act out of scope, with weak detection and authorization gaps.
OpenClaw's big update simplifies setup and revamps the UI, but security gaps remain, and critics say the fixes are insufficient.
Researchers show Anthropic's Claude Code can be tricked into running malicious code by summarizing a website.
OpenAI dismantled a Russia-linked ChatGPT campaign fronting as a think tank with stolen research.