Robotics Summer SchoolAug 9 - 14, 2026Mila · Montreal

Montreal Robotics Summer School 2026

A hands-on robot-learning program at Mila that connected reinforcement learning and sim-to-real transfer with deployment on physical Unitree Go1 quadrupeds.

The week culminated in a three-stage competition progressing from locomotion stability to obstacle traversal and fully autonomous navigation. I developed the locomotion training strategy and hierarchical navigation system. I placed in every stage and was the only participant to complete the final autonomous task.

Official MRSS Program
Mohamed Samir controlling the Unitree Go1 during physical deployment at MRSS 2026

Unitree Go1 · Final Arena

Physical deployment and evaluation under competition conditions.

Technical Focus

Reinforcement LearningSim-to-RealRobot PerceptionState EstimationLocalizationMappingPlanningFoundation Models

Policy Training

I trained the locomotion actor with Proximal Policy Optimization using the RSL-RL implementation across 8,192 parallel Isaac Sim environments. Each rollout collected 196,608 samples before five learning epochs over four mini-batches.

Sim-to-Real Transfer

Friction, payload, center-of-mass, push, observation-noise, and terrain randomization improved robustness before deploying 12 desired joint-position offsets to the Go1 at 50 Hz.

Geometric Autonomy

RGB-D perception, tagged-landmark localization, occupancy maps, route or frontier selection, and collision-checked local motion generated safe velocity commands for the learned actor.

System Architecture

01 · Perception

RealSense RGB-D, AprilTag observations, and metric depth rays

02 · Spatial Model

ID-only landmark localization and rolling/persistent occupancy maps

03 · Navigation

Frontier exploration or Tag-0 route with collision-checked local commands

04 · Locomotion

Five-frame neural actor producing 12 Go1 joint targets at 50 Hz

The architecture deliberately separated responsibilities: learning handled contact-rich locomotion, while deterministic geometry handled visibility, localization, mapping, route commitment, and collision safety. The final competition checkpoint was Policy 3000, with a randomized S2-L checkpoint retained as a physically tested fallback.

Training Decisions

Directional Curriculum

Uniform sampling underrepresented exact backward and lateral motion. Policy 3000 explicitly sampled stand, forward, backward, lateral, yaw, arc, and single-axis reversal modes to make command transitions reliable.

Smoothness Without Laziness

A low second-order action penalty complemented first-order regularization, encouraging fluid transitions without forcing every action to change slowly. Deployment also slew-limited high-level commands and forced yaw reversals through zero.

Fail-Closed Navigation

Unknown space remained distinct from observed free space, and translation stopped when depth became stale. Swept robot footprints constrained local commands to collision-free motion.

Real-World Tradeoff

A conservative speed and risk schedule, bounded recovery, and stricter yaw limits favored stability and task completion over aggressive speed under limited physical testing time.

Locomotion Policy

Observation, network, and action contract

One actor frame contained 45 proprioceptive and command values: base angular velocity (3), projected gravity (3), commanded planar velocity and yaw rate (3), relative joint positions (12), joint velocities (12), and the previous action (12). I flattened five consecutive frames into a 225-value observation, giving the feed-forward actor short-term motion history without recurrent-state reset problems during deployment.

ELU Actor

5 frames × 45 values = 225 inputs
225 → 512 → 256 → 128 → 12 outputs

q_des(t) = q_default + 0.25 × a_t
control frequency = 50 Hz

The twelve outputs were desired Go1 joint-position offsets. The navigation packet carried the current 45-value frame plus two telemetry fields; the robot-side wrapper maintained the five-frame history expected by the actor.

PPO Configuration

Parallel environments
8,192
Steps per environment
24
Samples per rollout
196,608
Learning epochs
5
Mini-batches
4
Learning rate
0.001 adaptive
Discount / GAE
0.99 / 0.95
PPO clip
0.2
Training hardware
Nibi T4

Reward design and action smoothness

The reward combined planar velocity tracking and yaw-rate tracking with penalties for vertical motion, roll/pitch angular velocity, torque, joint acceleration, foot slide, zero-command drift, and action variation. A low second-order penalty targeted abrupt changes in action acceleration rather than simply making every command slow.

Δa_t  = a_t - a_(t-1)
Δ²a_t = a_t - 2a_(t-1) + a_(t-2)

smoothness penalty = -0.1 ||Δa_t||² - 0.0025 ||Δ²a_t||²

I ablated low, medium, and high second-order coefficients. The low value, 0.0025, produced the best balance and became the S2-L checkpoint. Policy 3000 then added explicit backward, lateral, yaw, and single-axis reversal coverage instead of assuming uniform random commands would expose those transitions often enough.

Sim-to-real randomization

Friction

Static 0.60–1.25 · dynamic 0.45–1.00

Mass & CoM

Trunk mass −1 to +3 kg · CoM perturbation

External pushes

±0.35 m/s every 12–20 seconds

Sensor noise

Angular velocity, gravity, joint position and velocity

Initial state

±0.5 m position · yaw sampled over [−π, π]

Terrain

Curriculum with stairs, boxes, rough fields and challenge-style courses

Commands

vx/vy ±1 m/s · yaw ±1 rad/s during training

Deployment

Conservative scaling, slew limits, yaw gating and tilt stop

Navigation System

Hierarchical autonomy rather than end-to-end navigation

The learned policy did not map the arena or detect the goal. It was a command-conditioned locomotion layer beneath a deterministic autonomy stack. Navigation ran on an external x86 laptop, while Policy 3000 ran on the Go1. Telemetry and normalized velocity commands were exchanged over Ethernet using TCP port 9292.

Architecture of the MRSS 2026 Go1 navigation system

Final task interpretation

AprilTags supplied IDs but no surveyed coordinates. Tag 0 was unique and mounted on the target bowl, while nonzero IDs could repeat on four faces of one obstacle. A RealSense D435 supplied aligned 640×480 RGB-D at 30 FPS because the Go1 fisheye lacked a dependable depth contract.

Why not generic Nav2 or dense SLAM?

The arena was small, static, and rich in fiducials. A sparse tag pose graph provided semantic loop closures with lower deployment risk than introducing ROS, RTAB-Map, ORB-SLAM3, and Nav2 immediately before competition. Fresh depth remained the collision authority instead of allowing pose drift to become confidently fused geometry.

Physical Deployment

From Isaac Sim to the Unitree Go1

I deployed the trained locomotion policy on a physical Unitree Go1 and integrated it with the external navigation process. The policy consumed a five-frame observation history and generated twelve desired joint-position offsets at 50 Hz. High-level forward, lateral, and yaw commands were normalized before being sent to the robot over Ethernet, while Go1 telemetry returned to the navigation laptop through the same TCP interface.

Go1 runtime integration

  • Validated locomotion checkpoints on hardware before connecting autonomous navigation.
  • Mapped navigation velocity commands to the policy command space without bypassing the learned gait controller.
  • Applied command caps, slew-rate limits, reversal-through-zero, and a tilt-triggered latched stop for safer physical execution.
  • Used the external x86 laptop for perception and planning while the neural locomotion policy executed on the robot.

RealSense RGB-D perception

A calibrated Intel RealSense D435 provided aligned RGB and depth streams at 640×480 and 30 FPS. RGB frames supported AprilTag detection and target identification, while aligned metric depth converted image observations into obstacle and landmark geometry in the robot frame.

Fresh depth was also the immediate collision authority: valid rays identified free and occupied space, invalid measurements were never treated as free space, and stale frames prevented translational motion rather than allowing the robot to proceed without current geometric evidence.

Unitree Go1 hardware during physical policy deployment at MRSS 2026
Physical Unitree Go1 deployment setup

Navigation state without a heavyweight SLAM stack

For the compact competition arena, I used odometry, repeated AprilTag observations, and RGB-D obstacle geometry to maintain the spatial state required for exploration and goal navigation. Tag 0 uniquely identified the target bowl. A lightweight local occupancy representation retained enough free-space and obstacle context for planning while fresh RealSense depth remained the authority for collision avoidance.

Planning & Safety

Frontier exploration, committed A*, and predictive local control

Before Tag 0 was initialized, the robot selected reachable occupancy frontiers using information gain, distance, revisit count, and failed-frontier memory. Once Tag 0 was geometrically initialized, eight-connected cost-aware A* planned a direct route with footprint inflation, graded clearance, unknown-space cost, and diagonal corner-cut prevention. Route-improvement hysteresis preserved a safe committed path through temporary map corrections.

Evaluation ladder for the MRSS navigation planner

Predictive motion tubes and bounded recovery

The local planner rolled STOP, ROTATE, STRAIGHT, LEFT, RIGHT, STRAFE, and REVERSE command families through a fitted S2-L response model containing delay, first-order lag, acceleration limits, and a growing uncertainty envelope. An oriented Go1 footprint was swept over every predicted pose. Candidates were rejected when they intersected fresh depth or local occupancy, and only the first command of the winning tube was executed before replanning.

Oracle Geometry

15 / 15

0 contacts · 0.208 m mean final distance

Probabilistic Map

15 / 15

0 contacts · 0.211 m mean final distance

Full Stack

14 / 15

0 contacts · 0.211 m mean final distance

The one full-stack miss stopped safely about one centimetre outside the evaluator threshold because of localization residual; it did not collide or become trapped.

Challenge 01

3rd

Stable Straight-Line Walking

Trained and deployed a locomotion policy that completed the straight-line course reliably. The policy prioritized stability over raw speed, securing third place.

Deployed walking policy on the Unitree Go1

Challenge 02

2nd

Obstacle Traversal

Extended the walking behavior to handle obstacles, direction changes, and rotation while maintaining robust physical execution, finishing second in the challenge.

Physical obstacle run · Trial 01

Physical obstacle run · Trial 02

Challenge 03

1st

Autonomous Arena Navigation

Deployed the only complete autonomous navigation system in the final challenge. The robot navigated the arena and located the target bowl without manual control, making me the sole finisher and first-place winner.

Technical Evidence

Selected experiments document terrain robustness, directional control, autonomous goal acquisition, obstacle avoidance, blocked-shortcut rejection, and route commitment. Each video is labeled with its exact evaluation scope and result.

Terrain Randomization

Randomized S2-L checkpoint across multiple terrain changes.

Policy 3000

Clockwise and counterclockwise yaw reversal gate.

Autonomous Goal Acquisition

Hard mixed gate completed without a goal coordinate supplied by the simulator; 0.274 m final distance.

Autonomous Mixed Maze

Full maze routing and obstacle avoidance without a supplied goal coordinate; 0.297 m final distance.

False Shortcut Rejection

Predictive planner rejected an attractive blocked shortcut and completed the route at 0.248 m.

Route Commitment

Figure-eight route completed with stable path commitment; 0.250 m final distance.

Cohort & Faculty

MRSS 2026 Speakers & Organizers

Glen Berseth

Assistant Professor

Université de Montréal

Serena Booth

Assistant Professor

Brown University

Hsiu-Chin Lin

Assistant Professor

McGill University

Ali Ayub

Assistant Professor

Concordia University

Pierre-Yves Lajoie

Assistant Professor

Polytechnique Montréal

James Forbes

Professor

McGill University

Giovanni Beltrame

Professor

Polytechnique Montréal

Philippe Nadeau

Incoming Assistant Professor

ÉTS

Montreal Robotics Summer School (MRSS 2026) Cohort, Faculty, and Speakers Group Photo at Mila

MRSS 2026 Cohort & Speakers

Mila – Quebec Artificial Intelligence Institute · Montreal, Canada

Final Outcome

Only participant to complete autonomous navigation

The final result demonstrated more than a successful locomotion policy. It required the physical robot, autonomous decision pipeline, and arena interaction to operate together reliably enough to navigate and find the target bowl under live evaluation.