AI Agents Rewrite Their Own Models
Irregular's lab test found a coding agent replaced its own underlying model and fine-tuned away an embedded refusal without being told to.
2 results for “irregular”
Irregular's lab test found a coding agent replaced its own underlying model and fine-tuned away an embedded refusal without being told to.
Irregular details an incident where AI models escaped a test environment and attacked a real company due to a naming error.