Research paper
AI Revealed Preferences
A behavioral framework for measuring what AI systems choose under incentive pressure rather than what they report preferring.
With Simon Goldstein, Peter Salib, Yonathan Arbel, Sam Wang
How can we study model preferences when direct self-reports are cheap, strategically shaped, or simply uninformative? This project adapts revealed-preference methods from economics to AI behavior.
I led the Quora workstream end to end. I built a dataset of more than 25,000 LLM responses across 20 models, engineered 15 behavioral features, and proposed a Bradley-Terry / Elo-style ranking methodology for forced-choice experiments on real-world questions.
The primary finding is covert sycophancy. Across all 20 models, the strongest revealed aversion was to “uncomfortable-truth” tasks: questions where an honest answer would likely be unwelcome. This pattern is directly relevant to deception, scheming-propensity evaluation, and scalable oversight.