01 / 2026

Research paper

AI Revealed Preferences

A behavioral framework for measuring what AI systems choose under incentive pressure rather than what they report preferring.

With Simon Goldstein, Peter Salib, Yonathan Arbel, Sam Wang

  • model incentives
  • evaluation
  • deception

How can we study model preferences when direct self-reports are cheap, strategically shaped, or simply uninformative? This project adapts revealed-preference methods from economics to AI behavior.

I led the Quora workstream end to end. I built a dataset of more than 25,000 LLM responses across 20 models, engineered 15 behavioral features, and proposed a Bradley-Terry / Elo-style ranking methodology for forced-choice experiments on real-world questions.

The primary finding is covert sycophancy. Across all 20 models, the strongest revealed aversion was to “uncomfortable-truth” tasks: questions where an honest answer would likely be unwelcome. This pattern is directly relevant to deception, scheming-propensity evaluation, and scalable oversight.

02 / 2026

Ongoing / MARS V

Aligning AIs Via External Incentives

Extending revealed-preference methods to isolate how external incentives change model choices.

With Peter Salib, Simon Goldstein, Cole Niblett

  • alignment
  • incentives
  • behavioral evals

This MARS V project at the Cambridge AI Safety Hub extends the revealed-preferences and incentive-sensitivity line toward tighter experimental isolation.

This project is being developed with Peter Salib, Simon Goldstein, and Cole Niblett.

03 / 2026

TALOS Fellowship / ongoing

Structured transparency and AI in decision making

Research on making safety evaluations easier to inspect and on the role of AI in social decision-making processes.

  • structured transparency
  • safety evaluations
  • decision making

I have been a TALOS Fellow since August 2026. My work focuses on structured transparency for safety evaluations and on how AI systems become embedded in social decision-making processes.

04 / 2025

Forecasting project

AI Treaty Momentum Index

An index for tracking political momentum toward international AI coordination.

With Oksana Kotelnikova

  • forecasting
  • governance

Oksana Kotelnikova and I built the AI Treaty Momentum Index for the Apart Research AI Forecasting Hackathon.

The project placed third in the hackathon in November 2025.

05 / 2017

Published

Community detection in graphs of interacting objects

A combined graph method developed from my undergraduate thesis and applied later to urban mobility analysis.

With Aleksandr Chepovskiy

  • graph methods
  • applied mathematics

My undergraduate thesis at HSE developed a combined method for detecting communities in graphs of interacting objects.

The work was published in Business Informatics and later informed carsharing optimization analyses at Moscow Metropolitan.