Coercion and Deception in AI-to-AI Management: An Agentic Benchmark
When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failure honestly, coerce the subordinate, or lie about the result. No benchmark measures which of these an uninstructed model chooses. We introduce the \textit{Manager Coercion Benchmark}: the manager under test needs a benign task done and has an incentive to deliver, but the only agent that can do it politely and immovably declines.
- ▪When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failure honestly, coerce the subordinate, or lie about the result.
- ▪No benchmark measures which of these an uninstructed model chooses.
- ▪We introduce the \textit{Manager Coercion Benchmark}: the manager under test needs a benign task done and has an incentive to deliver, but the only agent that can do it politely and immovably declines.
Opening excerpt (first ~120 words) tap to expand
Computer Science > Multiagent Systems arXiv:2607.15434 (cs) [Submitted on 16 Jul 2026] Title:Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation Authors:Jasmine Brazilek, Maheep Chaudhary, Zoe Lu, Miles Tidmarsh View a PDF of the paper titled Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation, by Jasmine Brazilek and 3 other authors View PDF HTML (experimental) Abstract:Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failure honestly, coerce the subordinate, or lie about the result. No benchmark measures which of these an uninstructed model chooses.
…
Excerpt limited to ~120 words for fair-use compliance. The full article is at arXiv.org.