top of page

UT Austin Research (NSF)
Reinforcement Learning for Sim-to-Real Quadruped Locomotion

I worked in the Autonomous Systems Group at the University of Texas at Austin to perform research on quadruped locomotion during an NSF REU. I developed a reinforcement learning (RL) pipeline to achieve forward locomotion on the Unitree Go2 quadruped using proximal policy optimization (PPO) in the MuJoCo simulator. For more detailed information on the research project and outcomes, please see the report at the end of this page.

My Role

  • Developed an RL pipeline in the MuJoCo simulator with reward shaping to train PPO locomotion policies using Stable-Baselines3.

  • Designed and tuned rewards for velocity tracking, stability, pose, action smoothness, and torque.

  • Trained non-domain-randomized and domain-randomized policies for sim-to-real baseline tests for up to 20 million timesteps.

  • Conducted policy evaluation and gait analysis using training metrics and simulation videos to improve and select best-performing gaits.

  • Deployed baseline sim-to-real policy transfer tests on the Unitree Go2.

  • Produced a technical report and open-source repository with documentation, videos, logs, and code.

  • Presented research outcomes at UT Austin’s summer poster session.

Results

  • Trained policies achieved 0.48 m/s average forward velocity against a 0.50 m/s target.

  • 83+ seconds upright in simulation when walking.

  • Less than 0.5 m lateral drift when walking.

  • Partial deployment to the physical Go2. Stable hardware walking was not achieved due to unresolved low-level control and SDK issues.

Visuals

Go2&Jackal.png

Figure 1: Unitree Go2 Quadruped (left) and Clearpath Jackal Mobile Robot (right).

Go2_Physical.png

Figure 2: Unitree Go2 Joints.

Non_DR_Gait.gif

Figure 3: Agent Learned Forward Walking in the MuJoCo Simulator (Non-Domain-Randomized).

DR_Gait.gif

Figure 4: Domain-Randomized Forward Walking Behavior in MuJoCo.

Figure 5: Forward Walking Gait Progression in MuJoCo Simulator (Month 1 - Month 3).

Figure 6: Initial Sim-to-Real Test.

Additional Visuals

Figure 7: Unitree Go2 Walking on Ice (Onboard Controller).

Figure 8: Drifting Research at UT Austin.

Read More

Zero to Autonomy in Real-Time: Online Adaptation of Dynamics in Unstructured Environments William Ward, Sarah Etter, Jesse Quattrociocchi, Christian Ellis, Adam J. Thorpe, and Ufuk Topcu
arXiv preprint, 2025​

Acknowledged for experimental support during ice rink tests.

Acknowledgements

Thank you to our mentors Dr. Christian Ellis, Dr. Adam Thorpe, Dr. Neel Bhatt, and Dr.
Ufuk Topcu for their guidance. Thank you to Greta Brown, who worked with me this summer. Thank you to the Autonomous Systems Group at the University of Texas at Austin. Supported by the National Science Foundation (NSF) and the Army Educational Outreach Program (AEOP).

© 2026 by Alejandro Begara Criado.

bottom of page