Edward Beeching

I am a Machine Learning Research Scientist at Hugging Face. My recent work focuses on reinforcement learning and post-training for language models, open reasoning systems, and efficient training infrastructure.

Previously, I completed a PhD at INSA Lyon with the INRIA CHROMA team, where I studied deep reinforcement learning for planning, navigation, and embodied agents.

Edward Beeching

Research & Writing

My research spans language-model post-training, reasoning, embodied learning, and reinforcement learning. Publications and technical articles are listed together below.

2026

2025

Open R1 #3Hugging Face Blog

Open R1: Update #3

Guilherme Penedo, Lewis Tunstall, Anton Lozhkov, Hynek Kydlíček, Edward Beeching, Loubna Ben Allal, Quentin Gallouédec, Leandro von Werra, Agustín Piqueres Lajarín, Nathan Habib

Hugging Face Blog, March 11, 2025

Open R1 #1Hugging Face Blog

Open R1: Update #1

Leandro von Werra, Lewis Tunstall, Quentin Gallouédec, Guilherme Penedo, Edward Beeching, Anton Lozhkov, Brigitte Tousignant, Daniel van Strien

Hugging Face Blog, February 2, 2025

2024

ZephyrLanguage-model alignment

Zephyr: Direct Distillation of LM Alignment

Lewis Tunstall*, Edward Beeching*, Nathan Lambert, Nazneen Rajani, Kashif Rasul, et al.

Conference on Language Modeling (COLM), 2024

Distilling supervised and preference signals to align smaller open language models.

Constitutional AIHugging Face Blog

Constitutional AI with Open LLMs

Shengyi Costa Huang, Lewis Tunstall, Edward Beeching, Leandro von Werra, Omar Sanseviero, Kashif Rasul, Thomas Wolf

Hugging Face Blog, February 1, 2024

2023

StarChatHugging Face Blog

Creating a Coding Assistant with StarCoder

Lewis Tunstall, Nathan Lambert, Nazneen Rajani, Edward Beeching, Teven Le Scao, Sheon Han, Philipp Schmid, Leandro von Werra, Alexander Rush

Hugging Face Blog, May 9, 2023

2022

2021

Godot Reinforcement Learning Agents icon

Godot Reinforcement Learning Agents

Edward Beeching, Jilles Dibangoye, Olivier Simonin, Christian Wolf

AAAI Workshop on Reinforcement Learning in Games, 2021

An open interface connecting the Godot game engine to common reinforcement-learning frameworks.

2020

2014

Open Source

TRL

Post-training models with supervised fine-tuning, preference optimization, and reinforcement learning.

Open R1

An open reproduction of reasoning-model training recipes and datasets.

NuminaMath

Open models and 860,000 competition-mathematics problem-solution pairs.

JAT

A generalist Transformer agent for text, vision, robotics, and game environments.

Godot RL Agents

Tools for training reinforcement-learning agents inside the Godot game engine.

PhD Thesis

Large-Scale Automatic Learning of Autonomous Agent Behavior with Structured Deep Reinforcement Learning

University of Lyon / INSA Lyon, 2022. Supervised by Olivier Simonin, Christian Wolf, and Jilles Dibangoye.

Embodied Agent Demos

Behaviors learned with Godot RL Agents.

Image-goal navigation in a 3D scan without a provided map.

An agent learning to collect an ordered sequence of objects.