Policy Training
I trained the locomotion actor with Proximal Policy Optimization using the RSL-RL implementation across 8,192 parallel Isaac Sim environments. Each rollout collected 196,608 samples before five learning epochs over four mini-batches.
A hands-on robot-learning program at Mila that connected reinforcement learning and sim-to-real transfer with deployment on physical Unitree Go1 quadrupeds.
The week culminated in a three-stage competition progressing from locomotion stability to obstacle traversal and fully autonomous navigation. I developed the locomotion training strategy and hierarchical navigation system. I placed in every stage and was the only participant to complete the final autonomous task.

Unitree Go1 · Final Arena
Physical deployment and evaluation under competition conditions.
Technical Focus
I trained the locomotion actor with Proximal Policy Optimization using the RSL-RL implementation across 8,192 parallel Isaac Sim environments. Each rollout collected 196,608 samples before five learning epochs over four mini-batches.
Friction, payload, center-of-mass, push, observation-noise, and terrain randomization improved robustness before deploying 12 desired joint-position offsets to the Go1 at 50 Hz.
RGB-D perception, tagged-landmark localization, occupancy maps, route or frontier selection, and collision-checked local motion generated safe velocity commands for the learned actor.
System Architecture
RealSense RGB-D, AprilTag observations, and metric depth rays
ID-only landmark localization and rolling/persistent occupancy maps
Frontier exploration or Tag-0 route with collision-checked local commands
Five-frame neural actor producing 12 Go1 joint targets at 50 Hz
The architecture deliberately separated responsibilities: learning handled contact-rich locomotion, while deterministic geometry handled visibility, localization, mapping, route commitment, and collision safety. The final competition checkpoint was Policy 3000, with a randomized S2-L checkpoint retained as a physically tested fallback.
Training Decisions
Uniform sampling underrepresented exact backward and lateral motion. Policy 3000 explicitly sampled stand, forward, backward, lateral, yaw, arc, and single-axis reversal modes to make command transitions reliable.
A low second-order action penalty complemented first-order regularization, encouraging fluid transitions without forcing every action to change slowly. Deployment also slew-limited high-level commands and forced yaw reversals through zero.
Unknown space remained distinct from observed free space, and translation stopped when depth became stale. Swept robot footprints constrained local commands to collision-free motion.
A conservative speed and risk schedule, bounded recovery, and stricter yaw limits favored stability and task completion over aggressive speed under limited physical testing time.
Locomotion Policy
One actor frame contained 45 proprioceptive and command values: base angular velocity (3), projected gravity (3), commanded planar velocity and yaw rate (3), relative joint positions (12), joint velocities (12), and the previous action (12). I flattened five consecutive frames into a 225-value observation, giving the feed-forward actor short-term motion history without recurrent-state reset problems during deployment.
ELU Actor
5 frames × 45 values = 225 inputs 225 → 512 → 256 → 128 → 12 outputs q_des(t) = q_default + 0.25 × a_t control frequency = 50 Hz
The twelve outputs were desired Go1 joint-position offsets. The navigation packet carried the current 45-value frame plus two telemetry fields; the robot-side wrapper maintained the five-frame history expected by the actor.
PPO Configuration
The reward combined planar velocity tracking and yaw-rate tracking with penalties for vertical motion, roll/pitch angular velocity, torque, joint acceleration, foot slide, zero-command drift, and action variation. A low second-order penalty targeted abrupt changes in action acceleration rather than simply making every command slow.
Δa_t = a_t - a_(t-1) Δ²a_t = a_t - 2a_(t-1) + a_(t-2) smoothness penalty = -0.1 ||Δa_t||² - 0.0025 ||Δ²a_t||²
I ablated low, medium, and high second-order coefficients. The low value, 0.0025, produced the best balance and became the S2-L checkpoint. Policy 3000 then added explicit backward, lateral, yaw, and single-axis reversal coverage instead of assuming uniform random commands would expose those transitions often enough.
Friction
Static 0.60–1.25 · dynamic 0.45–1.00
Mass & CoM
Trunk mass −1 to +3 kg · CoM perturbation
External pushes
±0.35 m/s every 12–20 seconds
Sensor noise
Angular velocity, gravity, joint position and velocity
Initial state
±0.5 m position · yaw sampled over [−π, π]
Terrain
Curriculum with stairs, boxes, rough fields and challenge-style courses
Commands
vx/vy ±1 m/s · yaw ±1 rad/s during training
Deployment
Conservative scaling, slew limits, yaw gating and tilt stop
Navigation System
The learned policy did not map the arena or detect the goal. It was a command-conditioned locomotion layer beneath a deterministic autonomy stack. Navigation ran on an external x86 laptop, while Policy 3000 ran on the Go1. Telemetry and normalized velocity commands were exchanged over Ethernet using TCP port 9292.

AprilTags supplied IDs but no surveyed coordinates. Tag 0 was unique and mounted on the target bowl, while nonzero IDs could repeat on four faces of one obstacle. A RealSense D435 supplied aligned 640×480 RGB-D at 30 FPS because the Go1 fisheye lacked a dependable depth contract.
The arena was small, static, and rich in fiducials. A sparse tag pose graph provided semantic loop closures with lower deployment risk than introducing ROS, RTAB-Map, ORB-SLAM3, and Nav2 immediately before competition. Fresh depth remained the collision authority instead of allowing pose drift to become confidently fused geometry.
Physical Deployment
I deployed the trained locomotion policy on a physical Unitree Go1 and integrated it with the external navigation process. The policy consumed a five-frame observation history and generated twelve desired joint-position offsets at 50 Hz. High-level forward, lateral, and yaw commands were normalized before being sent to the robot over Ethernet, while Go1 telemetry returned to the navigation laptop through the same TCP interface.
A calibrated Intel RealSense D435 provided aligned RGB and depth streams at 640×480 and 30 FPS. RGB frames supported AprilTag detection and target identification, while aligned metric depth converted image observations into obstacle and landmark geometry in the robot frame.
Fresh depth was also the immediate collision authority: valid rays identified free and occupied space, invalid measurements were never treated as free space, and stale frames prevented translational motion rather than allowing the robot to proceed without current geometric evidence.

For the compact competition arena, I used odometry, repeated AprilTag observations, and RGB-D obstacle geometry to maintain the spatial state required for exploration and goal navigation. Tag 0 uniquely identified the target bowl. A lightweight local occupancy representation retained enough free-space and obstacle context for planning while fresh RealSense depth remained the authority for collision avoidance.
Planning & Safety
Before Tag 0 was initialized, the robot selected reachable occupancy frontiers using information gain, distance, revisit count, and failed-frontier memory. Once Tag 0 was geometrically initialized, eight-connected cost-aware A* planned a direct route with footprint inflation, graded clearance, unknown-space cost, and diagonal corner-cut prevention. Route-improvement hysteresis preserved a safe committed path through temporary map corrections.

The local planner rolled STOP, ROTATE, STRAIGHT, LEFT, RIGHT, STRAFE, and REVERSE command families through a fitted S2-L response model containing delay, first-order lag, acceleration limits, and a growing uncertainty envelope. An oriented Go1 footprint was swept over every predicted pose. Candidates were rejected when they intersected fresh depth or local occupancy, and only the first command of the winning tube was executed before replanning.
Oracle Geometry
15 / 15
0 contacts · 0.208 m mean final distance
Probabilistic Map
15 / 15
0 contacts · 0.211 m mean final distance
Full Stack
14 / 15
0 contacts · 0.211 m mean final distance
The one full-stack miss stopped safely about one centimetre outside the evaluator threshold because of localization residual; it did not collide or become trapped.
Challenge 01
Trained and deployed a locomotion policy that completed the straight-line course reliably. The policy prioritized stability over raw speed, securing third place.
Deployed walking policy on the Unitree Go1
Challenge 02
Extended the walking behavior to handle obstacles, direction changes, and rotation while maintaining robust physical execution, finishing second in the challenge.
Physical obstacle run · Trial 01
Physical obstacle run · Trial 02
Challenge 03
Deployed the only complete autonomous navigation system in the final challenge. The robot navigated the arena and located the target bowl without manual control, making me the sole finisher and first-place winner.
Technical Evidence
Selected experiments document terrain robustness, directional control, autonomous goal acquisition, obstacle avoidance, blocked-shortcut rejection, and route commitment. Each video is labeled with its exact evaluation scope and result.
Randomized S2-L checkpoint across multiple terrain changes.
Clockwise and counterclockwise yaw reversal gate.
Hard mixed gate completed without a goal coordinate supplied by the simulator; 0.274 m final distance.
Full maze routing and obstacle avoidance without a supplied goal coordinate; 0.297 m final distance.
Predictive planner rejected an attractive blocked shortcut and completed the route at 0.248 m.
Figure-eight route completed with stable path commitment; 0.250 m final distance.
Cohort & Faculty
Assistant Professor
Université de Montréal
Assistant Professor
Brown University
Assistant Professor
McGill University
Assistant Professor
Concordia University
Assistant Professor
Polytechnique Montréal
Professor
McGill University
Professor
Polytechnique Montréal
Incoming Assistant Professor
ÉTS

MRSS 2026 Cohort & Speakers
Mila – Quebec Artificial Intelligence Institute · Montreal, Canada
Final Outcome
The final result demonstrated more than a successful locomotion policy. It required the physical robot, autonomous decision pipeline, and arena interaction to operate together reliably enough to navigate and find the target bowl under live evaluation.