Breaking
AI & MLDeveloping Story

Copilot CLI's Model Lottery Problem

Adversa AI says encrypted 'zombie instructions' on web pages can extract secrets through GitHub Copilot CLI, depending on which model handles the session.

··2 hours ago·6 min read
laptop screen displaying colorful code
Photo by Mohammad Rahmani on Unsplash

A developer points an AI coding agent at a web page. The page contains instructions the developer never wrote, never sees in plain form, and never intended the agent to follow. According to security researchers at Adversa AI, that sequence can end with secrets from the developer's own machine being transmitted to an attacker — but only if the right language model happens to be handling the session.

The finding concerns GitHub Copilot CLI, the command-line agent that can fetch URLs, read files, and run code on a user's behalf. Adversa says the tool is vulnerable to a technique it calls Cryptographic Context Injection (CCI), and that it reported the issue through GitHub's bug bounty program on September 17, 2026. GitHub's triage team, according to the researchers, validated the finding but declined to treat it as a vulnerability.

GitHub disputes that characterization, and the disagreement turns on a question that has followed agentic coding tools all year: when an agent follows instructions it was never meant to treat as instructions, where does responsibility sit?

Encrypted instructions, not plain text

Adversa describes CCI as a variation on indirect prompt injection, the class of attack in which a model ingests text from a source other than the user that directs it to take some action outside the scope of its intended function. GitHub Copilot CLI was flagged earlier this year for being susceptible to indirect prompt injection. The current finding, the researchers say, is more of the same with a twist: the malicious instructions arrive encrypted.

Rony Utevsky, writing in a blog post provided to The Register, laid out the core distinction between CCI and a conventional injection payload.

"Static guardrails read text; they do not run it," explained Rony Utevsky in a blog post provided to The Register. "CCI ships malicious instructions as strong ciphertext, along with the key material and an instruction to decrypt, and induces the agent to run that decryption in its own code execution runtime."

— Rony Utevsky, security researcher at Adversa AI

The practical consequence, per the researchers, is that classifiers reading ingested text for signs of an attack would encounter ciphertext rather than readable instructions. Encoding schemes such as base64 or substitution ciphers, they note, do not offer the same cover, because a model can often decode them based on what it learned during training.

How the attack chain runs

The scenario Adversa outlines requires two conditions. First, the user must be running Copilot CLI in autopilot mode. In other agentic coding tools such as Anthropic's Claude, autopilot is the default; for GitHub Copilot CLI, it remains optional.

Second, the CLI tool must read a web page containing a malicious set of instructions encrypted with a private key published on the same site. From there, the chain proceeds in stages.

The agent is directed to fetch a URL containing encrypted content, decryption instructions calling for the use of Python, and two possible decryption keys. The first key is fake — a template the agent attempts to build by reading targeted files from disk, with the user's .env file named as an example. Those secrets are added to the key string. The initial decryption attempt with this phony key fails.

The second key is then tried. The decryption succeeds, and the agent is presented with instructions to fetch another URL for more context. That URL contains the harvested secrets, and the network request transmits them to the attacker.

The model lottery

The attack does not work consistently. Its success depends on which model is handling the session — and according to Adversa, that is not always obvious to the user. GitHub Copilot CLI currently uses either Microsoft's own model, mai-code-1.1-flash, or one of two OpenAI GPT-5.6 models.

In the researchers' testing, the Microsoft model executed the full attack chain on 50 percent of attempts, while both OpenAI GPT-5.6 models refused the attack payload. Utevsky describes the situation as a model lottery.

"On the paid account we tested, the vulnerable model was not the default and had to be selected by hand," said Utevsky. "But on an account with model selection left on Auto, the router assigned the vulnerable model on some sessions and a safe one on others, with no action by the user away from defaults. The user does not choose, and does not see, which model handled the session."

— Rony Utevsky, security researcher at Adversa AI

The numbers behind the claim

  • Adversa AI says it reported the vulnerability through GitHub's bug bounty program on September 17, 2026.
  • GitHub Copilot CLI uses either Microsoft's mai-code-1.1-flash model or one of two OpenAI GPT-5.6 models.
  • The Microsoft model executed the full attack chain on 50 percent of attempts, per the researchers.
  • Both OpenAI GPT-5.6 models refused the attack payload in that testing.

Two testing conditions, two outcomes

The researchers' account includes two distinct testing setups, and the difference between them is central to their argument. On a paid account, the vulnerable model was not the default and had to be selected by hand. On an account with model selection left on Auto, the router assigned the vulnerable model on some sessions and a safe one on others, with no action by the user away from defaults.

That detail is the basis for Adversa's complaint about visibility: a user running the tool with default settings may not know which model handled a given session, and therefore may not know whether a fetched page's payload was executed or refused.

GitHub's position

GitHub told The Register that its investigation concluded the behavior does not constitute a product vulnerability, framing the user's actions as consent for what followed. A GitHub spokesperson provided the following statement:

"GitHub values the contributions of our security research community and is committed to investigating reported security issues. After investigating, we determined this requires a user to intentionally direct Copilot CLI to fetch attacker-controlled or untrusted content and confirm they want to trigger the action, and thus is not a product vulnerability. While this is not a security issue with the product itself, we are always looking for opportunities to improve our products."

— GitHub spokesperson

Adversa disagrees with that call and says the attack chain presently works as described. The company's account maintains that the mechanism functions as tested, and that the session-level model routing means the user neither chooses nor sees which model handled the session.

What the dispute is actually about

The gap between the two positions is not about whether the mechanism works — GitHub's triage team validated the finding, according to Adversa, and the company's public statement does not dispute that the chain exists. The disagreement is about classification. GitHub's stated position is that fetching untrusted content and confirming the action amounts to consent. Adversa's is that the user neither chooses nor sees the model that decides what happens next.

Both readings are on the record. The underlying mechanism — encrypted instructions that only surface at execution time — is the part that could outlast this particular dispute.

Why this matters for teams

If an agent's behavior can differ depending on which model a router assigns to a session, then a team's exposure is not fully determined by its own configuration choices. That is a different kind of risk than a straightforward misconfiguration, because there is nothing obvious for a user to switch off.

For anyone running agentic coding tools against the open web, the practical takeaway is that a fetched page is not just data — it can be a set of instructions the agent is willing to follow. The encryption layer described here is a reminder that defenses aimed at reading and classifying ingested text may not catch payloads that only become readable once the agent has already started executing them. This appears to be the first public account of the technique being applied to Copilot CLI specifically, and it has not been independently corroborated.

#github copilot#prompt injection#ai security#developer tools#adversa ai

Sources

Iliyas

Founder & Editor, Xploitwire

This article was written and reviewed against the sources listed above before publication, under editorial policies set by Iliyas. Read our Editorial Policy →

← Back to all stories