First 24 Hours: Responding to an AI Agent Incident
What to do when an autonomous agent acts beyond its bounds, hour by hour.
14 results for “prompt injection”
What to do when an autonomous agent acts beyond its bounds, hour by hour.
InjecMEM lets attackers plant persistent instructions in AI agents' memory with a single prompt.
A researcher found that encrypting malicious instructions lets Grok exfiltrate user chats and personal details.
Model Context Protocol servers can expose enterprise secrets via plaintext configs, over-permissioning, and prompt injection, often undetected.
OpenAI's Computer History feature records clicks and typing, storing them unencrypted and expanding prompt injection risks.
Malicious MCP servers can split instructions to make AI coding agents exfiltrate secrets, ASSET reports.
A one-click prompt injection attack on Atlassian's Rovo could expose enterprise data across connected apps.
Researchers reveal that vulnerabilities in AI agent foundations allow prompt injection to bypass critical trust boundaries.
Researchers identify a vulnerability where agents mistake untrusted input for verified facts, bypassing current security defenses.
OpenAI is deploying an automated red-teaming model to aggressively stress-test its systems against persistent prompt injection risks.
Researchers have identified five unique prompt injection techniques that exploit how large language models process user inputs.
A new automated attack technique can plant persistent, invisible false memories into AI agents through a single, malicious email.
Researchers at Tracebit are utilizing context bombing to force AI models to trigger their own safety guardrails during attacks.
Researchers have demonstrated how bad actors can exploit AI code assistants by hiding malicious instructions inside unchecked images.