Muso Action
/
Back to blog

Evaluating Imitation-Learning Policies with Isaac Lab Arena

Hi, I'm Watanabe, an AI robotics engineer at Muso Action.

At Muso Action we build robots for logistics sites, and we teach them by imitation learning: a policy watches the feed from a camera on the arm and predicts the next motion. Training those policies and running them on real hardware is our day job.

This post is about a narrower question — how we evaluate them.

The short version: putting NVIDIA's Isaac Lab Arena at the center of our evaluation loop turned a process that used to require the real robot into something we can run overnight.

Evaluating one policy across many environments in parallel

Evaluating on real hardware every time was the bottleneck

For a manipulation policy, the training loss tells you much less than you would like. A loss that refuses to go down does tell you the model is bad — but a loss that goes down does not tell you the model is good. Closed-loop control is sequential, and small errors compound over the episode in ways the offline objective never sees. To decide which of two checkpoints is better, you really do have to run them and count how often the task succeeds.

Which means evaluating on the real robot. And that gets expensive fast.

  • It ties up the robot and the workspace. While an evaluation runs, nobody else can use either.
  • You cannot reproduce the initial conditions. Resetting an object to the same position and orientation by hand, to the millimetre, every time, is not realistic. And if the conditions differ, a difference in success rate could be the policy or could be how you placed the object.
  • The lighting changes with the day and the hour. Sunlight through the window, shadows cast by nearby equipment — none of it is under your control.
  • You cannot afford many trials. At tens of seconds per episode, a few dozen episodes for a single condition burns half a day.

The last two are what really hurt. Improving an imitation-learning policy means swapping the vision backbone, changing the image preprocessing, changing how actions are represented. You might have ten candidates, and on real hardware you cannot try them all. So you pick the one that "looks about right," and the reasoning behind your choice gets thinner and thinner. Evaluation had become the rate-limiting step, and it was eating into the experimentation and hyperparameter tuning we actually wanted to do.


Making Isaac Lab Arena our evaluation backbone

So we moved evaluation into a physics simulator. The one we use is Isaac Lab Arena.

We chose the Isaac stack for three reasons that matter specifically for imitation learning: photorealistic image rendering, the ability to step many robots in parallel on a GPU, and a deep library of assets to build environments from.

Within that stack, Isaac Lab Arena is the piece aimed squarely at evaluation. It is an open-source framework from NVIDIA, and the documentation describes it like this:

an open-source framework for scalable benchmark authoring and robot policy evaluation in simulation

Authoring benchmarks for robot policy evaluation is stated as the goal itself — that is why we picked it. It is not a benchmark, it is the scaffolding you write benchmarks on, and it is clear about that.

The documentation's front page has a diagram of the software stack. From the bottom: the physics solvers (PhysX / Newton), the simulation framework Isaac Lab that builds on them, the policy evaluation framework Isaac Lab-Arena that extends it, and at the top, Your benchmarks — whatever you write yourself.

Here is that top box filled in with ours.

The Isaac Lab Arena software stack and where our code sits

Drawn following the structure and terminology of the stack diagram in the Isaac Lab Arena documentation (as of September 2026), with the top row replaced by our own contents. The runtime underneath it all is Isaac Sim.

What we appreciate is that the environment is decomposed into parts. The scene (shelving and fixtures), the embodiment (robot and gripper), the task (what goes where), and the policy are all independently swappable. We registered a shelf scene modeled on a real site plus our own robot, and it came up as an evaluation environment.

A cell in the simulator: arm, gripper, shelf, target object

We extend Arena by registering our scenes, robots, and tasks into its registry — the contact surface is about a dozen lines of imports.

The loop we run looks like this.

  1. Collect data inside the simulator. A scripted demonstration emits end-effector pose targets, and Isaac Lab's differential IK converts them into joint commands. Episodes are recorded in the LeRobot dataset format.
  2. Train. The LeRobot datasets go through the same training infrastructure we use for real-robot data.
  3. Evaluate in the simulator. The trained checkpoint goes back into the Arena environment and runs closed-loop, so the compounding drift that separates success from failure over an episode shows up here too.
Replaying an episode collected in the simulator

The three stages are a single pipeline: name the task and the run, and it goes end to end.


Win #1: Changing the lighting shows how robust the vision stack really is

The best thing about moving into simulation is that we can change the lighting freely.

On real hardware this was never practical. You cannot switch off the lights at a working site, and you cannot choose the time of day. We would wonder "is this policy weak in low light?" or "will it fall apart when the sun moves and the shadows change direction?" — with no way to find out.

In the simulator, lighting is one setting away. The scripted motion, the physics, and the success criterion are all untouched; only the image the policy sees changes. These are the four conditions we set up.

How the wrist camera sees the scene under each lighting condition

Nominal matches the training data. The other three were never seen during training. The most useful of them is side lighting: the mean brightness of the image is essentially unchanged from nominal, and only the direction of the light moves. That lets us separate "it failed because it was too dark to see" from "it fell apart when the shading cues changed" — which is exactly the distinction we care about.

Running this surfaced things the real robot never showed us. A couple of the patterns we found:

  • Backbones that look equivalent under nominal lighting swap places the moment the light direction changes. Compare only under the training distribution and everything looks about equally good, which gives you nothing to choose on. The gaps open up once you step outside it.
  • With the same backbone, changing the action representation changes robustness to lighting. Predicting joint angles directly turned out to be more sensitive to a change in light direction than predicting end-effector poses, and the effect pointed the same way across multiple tasks. Same backbone, same training data — the only difference is the action representation. It reads as the visual features being used differently depending on how actions are parameterized.

Below is the wrist camera feed the policy actually receives, with the same policy and the same initial conditions and only the lighting changed. Nominal on the left, side lighting on the right.

Wrist camera for the same policy and seed, with only the lighting changed

Note that the right side is not darker. All that changed is the direction of the light — you can see the cube's cast shadow disappear. That alone is enough to make this policy misjudge where to close the gripper.

From the overhead view, the same episode splits into a success and a failure.

The same episode ends in success or failure depending on the lighting

The important part is that we learned all of this without stopping the real robot for a single second. Weak candidates get eliminated before they ever reach hardware, and real-robot evaluation is now reserved for confirming the one or two that survive.


Win #2: Running many robots in parallel to sweep evaluations

The other win is parallelism.

We have one physical robot, so evaluation happens one episode at a time. In simulation, the same policy runs across many environments at once — environments tiled on a single GPU, all stepped together. Arena's documentation calls this out explicitly as a goal: evaluate concurrently across many environments instead of sequential rollouts.

That is what the clip at the top of this post shows. Each cell is initialized independently and the same policy drives all of them. It is not limited to the shelf scene — a task handling parcels moving along a conveyor parallelizes the same way.

Evaluating the conveyor cell in parallel

What made parallelism valuable was not the raw speedup so much as the fact that comparisons became meaningful.

Because the random seed is fixed per episode, changing a training condition still evaluates against an identical set of initial states. Arena's documentation on sensitivity analysis puts the failure mode well:

If bright-light episodes also happened to use an easy object, a per-factor light chart cannot tell which one drove success.

If the brightly lit episodes happened to use easy object placements, lining up success rates per lighting condition tells you nothing about which factor actually mattered. That is precisely what was happening when we ran thirty episodes on hardware. Fixing seeds in simulation means never creating that confound in the first place.

Operationally, results land in Slack when an evaluation finishes: the success rate, plus overhead and wrist-camera video for both a success and a failure. Queue a batch of conditions in the evening and everything is waiting in the morning. Having the failure videos right there helps too — you can see whether the policy never reached the object or grasped it and dropped it, before you look at a single number.

The upshot is that choosing a backbone and narrowing down hyperparameters no longer requires taking the robot offline. The original problem — ten candidates, no way to try them all — simply went away.


What we learned about evaluating in simulation

Three things worth passing on.

1. Use the same number of parallel environments for collection and evaluation. Cameras across parallel environments are rendered into a single tiled image, so changing the environment count changes the tiling — and the camera image itself shifts subtly with it. We had a policy trained on data collected with a small number of environments score completely differently when evaluated with many. We now record the environment count at collection time and have evaluation read it back.

2. When you change a lighting value, render it once and look at it. Scaling a light source to 0.35× does not make the image 0.35× as bright. Tone mapping and indirect light mean the setting and the appearance are not proportional. Push it too far and you are no longer measuring robustness to lighting, you are measuring "the frame is black so the policy fails." Before running an evaluation we line up the wrist-camera images the policy will actually receive and check them. The four-panel figure above is exactly that check.

3. A success rate in simulation is not a success rate on hardware. Obvious, but worth stating: we use it for ranking and for cutting weak candidates, and the final call is always made on the real robot. That said, because simulation has already narrowed the field to policies that work reasonably well, the amount of real-robot evaluation we need has dropped substantially.


Wrapping up

Imitation learning gets discussed mostly in terms of training, but in practice we found evaluation is what limits the pace.

Putting Isaac Lab Arena at the center of evaluation let us vary the lighting to probe how robust the vision stack is, and run many environments in parallel to compare conditions in one go. Real-robot evaluation is now focused on final confirmation. Most of all, "we can't decide because we can't test it" became "we can decide because we can test it."

If you are working on evaluating imitation-learning policies in simulation, I hope some of this is useful.

References

Isaac Lab Arena asks to be cited as follows (documentation). Please include it if you publish a benchmark.

@misc{isaaclab-arena2025,
   title   = {Isaac Lab-Arena: Composable Environment Creation and Policy Evaluation for Robotics},
   author  = {{NVIDIA Isaac Lab-Arena Contributors}},
   year    = {2025},
   url     = {https://github.com/isaac-sim/IsaacLab-Arena}
}

We Are Hiring!

Muso Action is hiring robotics and machine-learning engineers. If building robots that work on real sites — touching both the hardware and the simulator — sounds like your kind of problem, get in touch through our careers page.

A casual chat is more than welcome.