I’m currently working on AI training, evaluation, and scaled human feedback as an AI research engineer at Prolific. I previously worked at the UK AI Security Institute in the Science of Evaluations group, and at Faculty as a data scientist.
I did a year on a PhD in statistics at Imperial College London, which is also where I studied for my MSc in statistics. My undergraduate degree was in design engineering at the University of Bristol, with electives in maths and computer science.
Updates
- Paper Is This a Bot? AI Models Lie About Being Human, Even When Not Asked To accepted to the TAIGR workshop at ICML 2026.
- Paper Quantifying Frontier LLM Capabilities for Container Sandbox Escape, with the UK AI Security Institute, accepted to ICML 2026.
- Attended ICLR in Rio de Janeiro.
- Submitted the MetaLoop benchmark to Measuring Progress Toward AGI.
- Presented Who Does Your AI Serve? Manipulation By and Of AI Assistants at the AIMII workshop, IASEAI 2026, Paris.
- Gave the talk People, Please! Why & Where to Loop People into Agent Evals at AI DevWorld.
- Released a new episode of Prolific’s The Frontier podcast.
- Runner-up at Black Forest Labs’ hackathon with Art of the Deal.
- Joined Prolific as a research engineer, focused on AI training, evaluation, and scaled human feedback.
- Coz presented our work on automated agent inventories at ERA’s Technical AI Governance Forum in London: Towards a Technical Roadmap to Govern Frontier AI.
- Mario presented our work on automated agent activity inventories to the Human-Centred AI Network at the University of Aberdeen.
- Coz presented our transcript-analysis work at the TAIG workshop at ICML 2025: Transcript Analysis and How it Relates to Technical Governance.
- We published an elicitation checklist to guide elicitation during LLM safety evaluations.
- Joined AISI’s Science of Evaluations team to work on tools and methods for analysing LLM transcripts used in pre-deployment testing.
- Joined the Frontier AI Taskforce as Research Strategy and Delivery Manager for its Safeguards Analysis team.
Papers and articles
- Is This a Bot? AI Models Lie About Being Human, Even When Not Asked To. TAIGR workshop, ICML 2026.
- Seven simple steps for log analysis in AI systems. Dubois, Zorer, Hamin, Skinner, Souly, Wynne, Coppock et al.
- Quantifying Frontier LLM Capabilities for Container Sandbox Escape. Marchand, O Cathain, Wynne, Giavridis, Deverett, Wilkinson, Gwartz, Coppock. ICML 2026.
- Who Does Your AI Serve? Manipulation By and Of AI Assistants. Petrova, Wynne. AIMII workshop, IASEAI 2026.
- Assuring Agent Safety Evaluations By Analysing Transcripts. UK AI Security Institute, Science of Evaluations.
- AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents. Andriushchenko, Souly, Dziemian, Duenas, Lin, Wang, …, Wynne et al. ICLR 2025.
Talks and posters
- Who Does Your AI Serve? Manipulation By and Of AI Assistants. AIMII workshop, IASEAI 2026, Paris.
- People, Please! Why & Where to Loop People into Agent Evals. AI DevWorld, San Jose.
- Kick-off panel, Black Forest Labs FLUX.2 hackathon, London.
- Panel with Nora Petrova: What 25,000 Humans Really Think About AI. Prolific Meetup #3, San Francisco.
- Gazing at the Ceiling. Highgate Literary and Scientific Institute, London.
Dissertations
- Automating Model Criticism for Linear Regression. MSc in Statistics, Imperial College London.
- Autonomous Learning for Adaptive Agents — Autonomous Systems for Independent Living. BEng in Engineering Design, University of Bristol.
Competitions
- Submitted the MetaLoop benchmark to Measuring Progress Toward AGI.
- Runner-up at Black Forest Labs’ hackathon with Art of the Deal, London.