About
I study how to make large models more reliable, more honest, and more useful. Currently at Anthropic on the model evaluations team.
Work Experience
Apr 2023 — Present
London, UK
Designing evaluations for model honesty and harmlessness. First-author on three papers including the widely-cited 'Sycophancy in Language Models' work.
Jan 2020 — Mar 2023
London, UK
Worked on the AlphaFold protein structure team and contributed to two Nature publications.
Education
Oct 2016 — Jan 2020
DPhil, Computer Science at University of Oxford
Oxford, UK
Sep 2011 — Jun 2015
BS, Mathematics at MIT
Cambridge, MA
Contact
Google Scholar
GitHub