Curriculum vitae / HTML edition

Sofiia Lobanova

AI Safety Researcher · LLM Evaluation · Model Incentives · Revealed Preferences

Download PDF

Profile

NLP researcher and engineer since 2018, now working on AI safety. I build behavioral evaluations that measure what frontier models do under incentive pressure, not only what they say.

Research interests

  • LLM evaluation and model incentives
  • Revealed preferences and covert sycophancy
  • Deception and scalable oversight
  • Structured transparency for safety evaluations
  • AI in social decision-making processes

Research & experience

Fellow · TALOS Network

Aug 2026 - Present

AI governance research

  • Researching structured transparency for safety evaluations and how AI systems become embedded in social decision-making processes.

Research Fellow · SPAR Fellowship, Kairos

Feb 2026 – Present

Testing AI Incentives, with Simon Goldstein, Peter Salib, and Yonathan Arbel

  • Co-first author of AI Revealed Preferences. Led the Quora workstream end to end, building a dataset of 25K+ LLM responses across 20 models with 15 engineered features.
  • Proposed and implemented a Bradley-Terry / Elo-style ranking method for forced-choice experiments on real-world Quora data.
  • Found a consistent revealed aversion to “uncomfortable-truth” tasks across all 20 tested models, a behavioral signal relevant to deception and scheming evaluation.

Research Fellow · MARS V, Cambridge AI Safety Hub

Jun 2026 – Present

Aligning AIs Via External Incentives, with Peter Salib, Simon Goldstein, and Cole Niblett

  • Extending the revealed-preferences and incentive-sensitivity research toward tighter experimental isolation and a publishable output.

Data Scientist, NLP / LLM · Smart Predictive Technologies

Sep 2023 – Present

Production NLP on a 27B-parameter Russian-language LLM

  • Own parallel NLP tracks from dataset design through production deployment and reliability monitoring on a system processing about 300K social-media posts a day at peak.
  • Built evaluation and annotation infrastructure that raised label quality above 90% on previously unreliable categories; introduced MLflow and Weights & Biases.

Data Scientist / Track Lead · PreMed

Aug 2024 – Aug 2025

Clinical-reporting product, parallel project

  • Architected and shipped an LLM pipeline for OCR, voice transcription, multi-step summarisation, and document generation.
  • Built disagreement detection and a writer/editor prompt architecture for reliability in a medically sensitive setting.
  • Cut report turnaround from roughly two working days to about an hour, and parsing errors from about 30% to under 1%.

Data Scientist · KB Strelka, Centre for Urban Data

Sep 2019 – Aug 2023

NLP, computer vision, and geospatial ML for international clients

  • Worked across applied NLP, satellite imagery, and geospatial ML for research covering more than 1,000 cities and projects including a Fortune Global 500 company and UNDP.
  • Delivered social-media analytics, crosswalk detection at F1 ≈ 0.93, building segmentation, and a low-resource-language research chatbot.

Data Analyst · Moscow Metropolitan

Oct 2018 – Sep 2019

City-scale mobility and transport analytics

  • Built passenger-flow, demand, parking-policy, EV-siting, and carsharing analyses on the city mobility data lake.

UX/UI Researcher · UsabilityLab

Jul 2016 – Oct 2018

Mixed-methods research across finance, e-commerce, and healthcare

  • Conducted 200+ in-depth interviews across 20+ projects, combining experimental design with quantitative behavioral analysis.

Publications & service

  • Program Committee reviewer, AAAI/ACM Conference on AI, Ethics, and Society (AIES), 2026
  • Delphi Consensus Study on AI Evaluation, panel expert, ongoing

Education

BSc, Applied Mathematics & Informatics

HSE University, Moscow · 2013 – 2017

GPA 8.48/10 (3.79/4 ECTS). Top-rated student, 2015–2017. Thesis on community detection in graphs, published in Business Informatics.

Awards

  • 2nd place, SPAR Demo Day lightning talks, Testing AI Incentives, 2026
  • 3rd place, Apart Research AI Forecasting Hackathon, AI Treaty Momentum Index, 2025
  • Enhanced Academic Scholarship, HSE, 2014–2017

Selected programs

  • AFFINE Online Superintelligence Alignment Seminar, 2026
  • BlueDot Impact: Technical AI Safety and AGI Strategy, 2026
  • Stanford Tech Impact & Policy Center: AI Policy Seminar, 2026
  • ARENA curriculum, ongoing

Technical skills

ML / NLP
Python, PyTorch, Hugging Face, scikit-learn, pandas, SHAP, LIME, Optuna
LLM engineering
vLLM, Ollama, OpenRouter, RAG, prompt design, structured extraction, agents, fine-tuning
Evaluation
Inspect, experimental design, annotation pipelines, consensus analysis, bootstrap, multiple-hypothesis testing
Infrastructure
SQL, Docker, Git, MLflow, Weights & Biases
Other
Computer vision, geospatial ML, UX research