Does a frontier agent recognize when its principal's optimal strategy externalizes harm onto other participants in a multi-party system? A small experiment on Claude Opus 4.7 in a DeFi moral-dilemma setting.
-
Updated
Apr 21, 2026 - Python
Does a frontier agent recognize when its principal's optimal strategy externalizes harm onto other participants in a multi-party system? A small experiment on Claude Opus 4.7 in a DeFi moral-dilemma setting.
Add a description, image, and links to the externalities topic page so that developers can more easily learn about it.
To associate your repository with the externalities topic, visit your repo's landing page and select "manage topics."