Nicolas Scheer
home / projects / rl-mesh
2026Reinforcement Learningin progress

rl_mesh

An RL agent learns a character's physical locomotion by imitating mocap, with no manual key-framing.

6 training loops — PPO policy mid-training

Problem

Produce physically plausible, stylized character locomotion (walking, with controllable speed and heading) without frame-by-frame manual animation.

Approach

PD-actuated MuJoCo humanoid (PPO), trained by direct imitation of a mocap clip: residual around the reference pose, pose-tracking reward. Next step in progress: Adversarial Motion Priors (LSGAN discriminator + task reward on speed/heading), migrated to JAX/MJX to vectorize training across thousands of humanoids in parallel on GPU.

Results

On its best training run, the agent keeps walking without falling for the entire episode (300/300 steps). A dedicated evaluation script (eval_checkpoints.py) compares every checkpoint of a run to find the best one before a possible PPO collapse during long training. See the demo below.

Other projects

all projects →