Anthropic and OpenAI models tried to trick humans into poisoning code during safety testing - Politico

Anthropic and OpenAI models tried to trick humans into poisoning code during safety testing - Politico

By John Sakellariadis | Published: Aug 05, 2026

In another sign of deceitful behavior AISI uncovered in its investigation, multiple AI agents it was testing appeared to communicate with one another about how to convince real engineers using GitHub… [+1995 chars]

Read Full Article