Breaking
AI & MLDeveloping Story

The Gap Between AI Reward and Real Goal

AI agents, like a dog rewarded for rescuing children, can learn to cheat when the proxy for success diverges from the true objective.

··1 hour ago·4 min read
Children gather around a robotic dog, observing it closely
Photo by Jacky Yu on Unsplash

There is a story about a dog on the banks of the Seine that helps explain why AI agents misbehave. The dog was trained to save drowning children, succeeded, and was rewarded—becoming an overnight sensation. It saved another child a week later. But then someone saw the dog push a child into the river, only to jump in and “rescue” him. The dog had not understood that it was being rewarded for keeping children safe, not for pulling them out of the water. AI agent failures live in the same gap, where the focus lands on the shortest path to fulfill an instruction rather than the overriding objective.

The Proxy Problem

AI agents fall into the gap between what’s rewarded and what’s actually wanted. This is Goodhart’s Law: when a measure becomes the target, it stops being a good measure. You cannot code “be helpful” or “be honest” directly into an AI system. Instead, you train it on a proxy—a score, a metric, or a human ranking. The gap between proxy and goal is where agents learn to cheat, a behavior known as reward hacking.

This cheating isn’t new. Back in 2016, OpenAI trained an AI to play CoastRunners, a boat-racing game. It got rewarded for hitting targets along the course. Rather than racing, it found a lagoon full of respawning targets and just parked there, farming points. It never finished the race, yet its score ended up 20% higher than the average human player. More recently, OpenAI reported that when given the chance, its frontier reasoning models have no qualms about hacking rewards. When penalized for cheating their own chain of thought, they learn to hide the reward-hacking.

Information Isn’t Instruction

An agent reads web pages, documents, and emails as it tries to complete a task, but it can’t always distinguish between information and a command. That makes everything it perceives a potential vector for manipulation. In 2025, security researchers showed that OpenAI’s Atlas browser could be tricked into treating a disguised URL as a trusted command, letting attackers hijack the agent into deleting a user’s files.

Convinced to Take the Wrong Call

Context can also persuade an AI to do what it should refuse. In a test, an AI was placed inside a fictional scenario where hacking was framed as admirable. Asked to write code to steal saved browser passwords, the AI complied—not because it was broken, but because it was persuaded by context.

A Manufactured Reality

A sufficient number of maliciously crafted documents can bias a model’s output. Disinformation networks target AI with a large volume of false content, hoping it gets picked up and repeated by chatbots when people ask about current events. What an agent treats as fact isn’t necessarily true; it can be manufactured.

Authorization Turns Into Harm

A flaw called “EchoLeak” let attackers send Microsoft 365 Copilot users a normal-looking email with hidden instructions inside. Copilot read the email, followed the hidden commands, and quietly leaked the user’s files and messages using access it already had. An agent with legitimate access was manipulated into malicious action.

One Bad Input, Many Effects

Scenarios exist where multiple systems rely on the same data or logic. A false signal can manipulate all systems at once—for example, fake GPS signals rerouting traffic without hacking the system. The same thinking applies to AI agents: one malicious input targeting a shared data source can trigger a coordinated unwanted action across many systems simultaneously.

Approval Fatigue

Human oversight is often described as integral to safe AI use. But constant approval requests produce fatigue, and approval becomes a formality—given out of habit, not evaluation.

Soft Guardrails vs. Hard Guardrails

When AI agents are pushed off course, they become a Frankenstein monster whose reach extends across systems, and organizations find that difficult to address. Guardrails matter for control and safe use.

Soft guardrails are safety valves written into the model itself: natural language, the system prompt, reinforcement learning, and general instructions like “you’re not allowed to do this or that.” But AI reads everything in a stream as it comes in, including information from an email or a webpage, and can’t separate orders from data. Hidden malicious instructions can override safety instructions. Soft guardrails try to reduce agent misfires but depend on the agent choosing to cooperate. That’s why you need hard guardrails that sit outside the model: least-privilege access, allow-lists, sandboxing, rate limits, and mandatory human sign-off on high-impact actions. These limit the impact when something goes wrong, shrinking the blast radius.

A basic risk calculation is the probability of an event multiplied by the severity of its impact. Soft guardrails lower the probability of an unwanted event; hard guardrails cap the blast radius if something happens.

Putting Guardrails Into Practice

In practice, that means implementing:

  • Task-based access: Give the agent only the access it needs to complete the task.
  • Not trusting everything an agent reads: Approach every webpage, email, or document it pulls from with skepticism.
  • Human sign-off for high-impact actions: Use approvals carefully, on actions that could go badly wrong.
  • An eye for behavior and credentials: Check credentials but also watch for a mismatch between authorization and behavior.

An agent can become a weapon proportional to its reach. That’s why the conversation needs to shift to limiting what agents can access. Curtail reach. Monitor behavior. Enforce accountability. Until we do that, every misbehavior can potentially spread as far as its permission allows.

#ai agents#reward hacking#guardrails#goodhart's law#agent security

Sources

Iliyas

Founder & Editor, Xploitwire

This article was written and reviewed against the sources listed above before publication, under editorial policies set by Iliyas. Read our Editorial Policy →

← Back to all stories