Status as of August 21, 2026: training and the OOD validation alpha sweep are complete. The full main evaluation is paused, so the evaluation figures below are an auditable interim checkpoint rather than a final comparison with the paper.
This fork contains an independent LIBERO reproduction of RETAIN. It covers 117-task pretraining, Task-FT and co-finetuning (coFT) on three LIBERO-10 target tasks, parameter merging, alpha selection, and ID/OOD/generalist evaluation. The merge follows
The DROID hardware experiments are outside the scope of this server-only reproduction.
| Stage | Status | Verified evidence |
|---|---|---|
| Data and environment validation | Complete | 353/353 fixed-revision dataset payloads passed SHA-256 verification; the base checkpoint passed 24/24 object size and MD5 checks. |
| Training | Complete | 7/7 runs completed: one 10,000-step pretraining run, three Task-FT runs, and three coFT runs. All recorded loss and norm values are finite. |
| OOD validation alpha sweep | Complete | 54 merged policies, 270 jobs, and 2,700 episodes completed with zero episode errors. |
| Main evaluation | Paused | 246 independent main-evaluation jobs and 2,550 episodes completed, with 817 successes and zero episode errors, before the GPU was released for another user. |
Alpha was selected on the OOD_MEDIUM small-translation validation scene using 50 episodes per alpha and target. Ties were resolved in favor of the larger alpha. Here, alpha is the weight of the task-finetuned or co-finetuned checkpoint in the merge.
| Merge source | Target task | Selected alpha | Validation successes | Validation success rate | Alpha reported in the paper |
|---|---|---|---|---|---|
| Task-FT | Stove | 0.8 | 30/50 | 60% | 0.9 |
| Task-FT | Mugs | 0.7 | 21/50 | 42% | 0.8 |
| Task-FT | Basket | 0.9 | 12/50 | 24% | 0.9 |
| coFT | Stove | 0.8 | 33/50 | 66% | 0.7 |
| coFT | Mugs | 0.8 | 26/50 | 52% | 0.9 |
| coFT | Basket | 0.8 | 20/50 | 40% | 0.9 |
The complete alpha curves, per-seed counts, and artifact hashes are available in the alpha-sweep snapshot.
Training ran on a single NVIDIA RTX 6000D (sm_120). The paper configuration uses a batch size of 64, but the reproduction used a physical batch size of 16 without gradient accumulation after batch-64 attempts exceeded the single-GPU memory budget. The completed training runs and validation sweep are reproducible evidence for this hardware-adapted setting, but they should not be presented as a strictly hyperparameter-identical reproduction.
The main evaluation is not yet complete. Its aggregate progress spans different policies, tasks, and evaluation types and therefore should not be interpreted as one overall success rate. See the pause snapshot for the exact resumable state.
- Reproduction workspace and scope
- Locked experimental protocol
- Verified findings and documented deviations
- Machine-readable research state
This is the official code release for the paper:
Robust Finetuning of Vision-Language-Action Robot Policies via Parameter Merging by Yajat Yadav, Zhiyuan Zhou, Andrew Wagenmaker, Karl Pertsch, Sergey Levine
See the project page also.
This codebase is built on top of openpi. The openpi README and setup instructions are preserved below.
We provide code for:
- Model merging via weight interpolation
- Finetuning VLA policies (pi0, pi0-FAST) on new tasks (LIBERO, DROID)
- OOD Evaluation: generating out-of-distribution (OOD) task variations in LIBERO and evaluating on them.
Model merging is implemented in src/openpi/policies/model_merging.py. The key function is linear_interpolation, which takes two or more checkpoints and produces a single merged model:
# Core merging logic (from src/openpi/policies/model_merging.py)
params_list = [restore_params(ckpt / "params") for ckpt in checkpoint_dirs]
trees = [jax.tree.map(lambda x: x * w, p) for w, p in zip(mixing_coefficients, params_list)]
merged_params = jax.tree.map(lambda *x: jnp.sum(jnp.stack(x), axis=0), *trees)
model = config.model.load(merged_params)The model_mixing_coefficients control the weight given to each checkpoint (must sum to 1).
Follow the openpi installation instructions below to install dependencies via uv. For running LIBERO evaluations, you will additionally need the LIBERO environment -- see the LIBERO README for Docker-based setup (recommended) or manual installation steps.
Finetune a base model on LIBERO data (see training configs in src/openpi/training/config.py):
# Compute normalization statistics
uv run scripts/compute_norm_stats.py --config-name pi0_libero
# Launch finetuning
XLA_PYTHON_CLIENT_MEM_FRACTION=0.9 uv run scripts/train.py pi0_libero \
--exp-name=my_experiment --overwriteStart a policy server (either a single checkpoint or a merged model), then run evaluation:
# Serve a single finetuned checkpoint
uv run scripts/serve_policy.py policy:checkpoint \
--policy.config=pi0_libero \
--policy.dir=checkpoints/pi0_libero/my_experiment/20000
# OR serve a merged model
uv run scripts/merging_experiments.py \
--port 8000 --config pi0_libero \
--merging_fn linear_interpolation \
--merging_fn_kwargs '{"model_mixing_coefficients": [0.5, 0.5]}' \
--checkpoint_dirs /path/to/finetuned/checkpoint gs://openpi-assets/checkpoints/pi0_baseThen run the standard LIBERO evaluation (from inside the LIBERO Docker environment):
python examples/libero/test_on_libero_task.py \
--host <server_host> --port 8000 \
--task_suite_name libero_90 \
--task_name "close the microwave" \
--num_trials_per_task 20 \
--exp_name my_evalOOD evaluation tests policy robustness by perturbing the LIBERO environment. The implementation lives in examples/libero/:
OOD_eval_helper.py-- Core logic for generating OOD perturbations:- Translation: shift object spawn regions
- Expansion: enlarge spawn regions so objects appear in novel positions
- Distractors: add randomly chosen unseen objects to the scene
test_on_OOD_libero_task.py-- Evaluation script that applies perturbations and rolls out the policygenerate_all_arg_combinations.py-- Generates shell commands for all OOD evaluation variants
Run a single OOD evaluation:
python examples/libero/test_on_OOD_libero_task.py \
--host <server_host> --port 8000 \
--task_suite_name libero_90 \
--task_name "close the microwave" \
--num_trials_per_task 20 \
--exp_name my_ood_eval \
--do_expand --expansion_option range_expand_doubleOr generate all evaluation commands at once:
python examples/libero/generate_all_arg_combinations.py \
--host <server_host> --port 8000 \
--checkpoint_name my_merged_model \
--task_name "close the microwave" \
--do_id --do_ood_easy --do_ood_hardFor DROID experiments, finetune on DROID data using the RLDS data loader, then merge and serve:
# Finetune
XLA_PYTHON_CLIENT_MEM_FRACTION=0.9 uv run scripts/train.py full_FT_whiteboard_RLDS \
--exp-name=my_droid_experiment --overwrite
# Serve merged model for evaluation
XLA_PYTHON_CLIENT_MEM_FRACTION=0.9 uv run scripts/merging_experiments.py \
--port 8100 \
--config pi0_fast_droid \
--merging_fn linear_interpolation \
--merging_fn_kwargs '{"model_mixing_coefficients": [0.5, 0.5]}' \
--checkpoint_dirs /path/to/finetuned/checkpoint gs://openpi-assets/checkpoints/pi0_fast_droidopenpi holds open-source models and packages for robotics, published by the Physical Intelligence team.
Currently, this repo contains two types of models:
- the π₀ model, a flow-based diffusion vision-language-action model (VLA).
- the π₀-FAST model, an autoregressive VLA, based on the FAST action tokenizer.
- the π₀.₅ model, an upgraded version of π₀ with better open-world generalization.
For all models, we provide base model checkpoints, pre-trained on 10k+ hours of robot data, and examples for using them out of the box or fine-tuning them to your own datasets.
This is an experiment:
- [Jun 2025]: We have added instructions for using
openpito train VLAs on the full DROID dataset. This is an approximate open-source implementation of the training pipeline used to train pi0-FAST-DROID.
To run the models in this repository, you will need an NVIDIA GPU with at least the following specifications. These estimations assume a single GPU, but you can also use multiple GPUs with model parallelism to reduce per-GPU memory requirements by configuring fsdp_devices in the training config. Please also note that the current training script does not yet support multi-node training.
| Mode | Memory Required | Example GPU |
|---|---|---|
| Inference | > 8 GB | RTX 4090 |
| Fine-Tuning (LoRA) | > 22.5 GB | RTX 4090 |
| Fine-Tuning (Full) | > 70 GB | A100 (80GB) / H100 |
The repo has been tested with Ubuntu 22.04, we do not currently support other operating systems.
When cloning this repo, make sure to update submodules:
git clone --recurse-submodules git@github.com:Physical-Intelligence/openpi.git
# Or if you already cloned the repo:
git submodule update --init --recursiveWe use uv to manage Python dependencies. See the uv installation instructions to set it up. Once uv is installed, run the following to set up the environment:
GIT_LFS_SKIP_SMUDGE=1 uv sync
GIT_LFS_SKIP_SMUDGE=1 uv pip install -e .NOTE: GIT_LFS_SKIP_SMUDGE=1 is needed to pull LeRobot as a dependency.
Docker: As an alternative to uv installation, we provide instructions for installing openpi using Docker. If you encounter issues with your system setup, consider using Docker to simplify installation. See Docker Setup for more details.
We provide multiple base VLA model checkpoints. These checkpoints have been pre-trained on 10k+ hours of robot data, and can be used for fine-tuning.
| Model | Use Case | Description | Checkpoint Path |
|---|---|---|---|
| Fine-Tuning | Base diffusion π₀ model for fine-tuning | gs://openpi-assets/checkpoints/pi0_base |
|
|
|
Fine-Tuning | Base autoregressive π₀-FAST model for fine-tuning | gs://openpi-assets/checkpoints/pi0_fast_base |
We also provide "expert" checkpoints for various robot platforms and tasks. These models are fine-tuned from the base models above and intended to run directly on the target robot. These may or may not work on your particular robot. Since these checkpoints were fine-tuned on relatively small datasets collected with more widely available robots, such as ALOHA and the DROID Franka setup, they might not generalize to your particular setup, though we found some of these, especially the DROID checkpoint, to generalize quite broadly in practice.
| Model | Use Case | Description | Checkpoint Path |
|---|---|---|---|
|
|
Inference |
|
gs://openpi-assets/checkpoints/pi0_fast_droid |
|
|
Fine-Tuning |
|
gs://openpi-assets/checkpoints/pi0_droid |
|
|
Inference |
|
gs://openpi-assets/checkpoints/pi0_aloha_towel |
|
|
Inference |
|
gs://openpi-assets/checkpoints/pi0_aloha_tupperware |
|
|
Inference |
|
gs://openpi-assets/checkpoints/pi0_aloha_pen_uncap |
By default, checkpoints are automatically downloaded from gs://openpi-assets and are cached in ~/.cache/openpi when needed. You can overwrite the download path by setting the OPENPI_DATA_HOME environment variable.
Our pre-trained model checkpoints can be run with a few lines of code (here our
from openpi.training import config
from openpi.policies import policy_config
from openpi.shared import download
config = config.get_config("pi0_fast_droid")
checkpoint_dir = download.maybe_download("gs://openpi-assets/checkpoints/pi0_fast_droid")
# Create a trained policy.
policy = policy_config.create_trained_policy(config, checkpoint_dir)
# Run inference on a dummy example.
example = {
"observation/exterior_image_1_left": ...,
"observation/wrist_image_left": ...,
...
"prompt": "pick up the fork"
}
action_chunk = policy.infer(example)["actions"]You can also test this out in the example notebook.
We provide detailed step-by-step examples for running inference of our pre-trained checkpoints on DROID and ALOHA robots.
Remote Inference: We provide examples and code for running inference of our models remotely: the model can run on a different server and stream actions to the robot via a websocket connection. This makes it easy to use more powerful GPUs off-robot and keep robot and policy environments separate.
Test inference without a robot: We provide a script for testing inference without a robot. This script will generate a random observation and run inference with the model. See here for more details.
We will fine-tune the
- Convert your data to a LeRobot dataset (which we use for training)
- Defining training configs and running training
- Spinning up a policy server and running inference
We provide a minimal example script for converting LIBERO data to a LeRobot dataset in examples/libero/convert_libero_data_to_lerobot.py. You can easily modify it to convert your own data! You can download the raw LIBERO dataset from here, and run the script with:
uv run examples/libero/convert_libero_data_to_lerobot.py --data_dir /path/to/your/libero/dataNote: If you just want to fine-tune on LIBERO, you can skip this step, because our LIBERO fine-tuning configs point to a pre-converted LIBERO dataset. This step is merely an example that you can adapt to your own data.
To fine-tune a base model on your own data, you need to define configs for data processing and training. We provide example configs with detailed comments for LIBERO below, which you can modify for your own dataset:
LiberoInputsandLiberoOutputs: Defines the data mapping from the LIBERO environment to the model and vice versa. Will be used for both, training and inference.LeRobotLiberoDataConfig: Defines how to process raw LIBERO data from LeRobot dataset for training.TrainConfig: Defines fine-tuning hyperparameters, data config, and weight loader.
We provide example fine-tuning configs for π₀, π₀-FAST, and π₀.₅ on LIBERO data.
Before we can run training, we need to compute the normalization statistics for the training data. Run the script below with the name of your training config:
uv run scripts/compute_norm_stats.py --config-name pi05_liberoNow we can kick off training with the following command (the --overwrite flag is used to overwrite existing checkpoints if you rerun fine-tuning with the same config):
XLA_PYTHON_CLIENT_MEM_FRACTION=0.9 uv run scripts/train.py pi05_libero --exp-name=my_experiment --overwriteThe command will log training progress to the console and save checkpoints to the checkpoints directory. You can also monitor training progress on the Weights & Biases dashboard. For maximally using the GPU memory, set XLA_PYTHON_CLIENT_MEM_FRACTION=0.9 before running training -- this enables JAX to use up to 90% of the GPU memory (vs. the default of 75%).
Note: We provide functionality for reloading normalization statistics for state / action normalization from pre-training. This can be beneficial if you are fine-tuning to a new task on a robot that was part of our pre-training mixture. For more details on how to reload normalization statistics, see the norm_stats.md file.
Once training is complete, we can run inference by spinning up a policy server and then querying it from a LIBERO evaluation script. Launching a model server is easy (we use the checkpoint for iteration 20,000 for this example, modify as needed):
uv run scripts/serve_policy.py policy:checkpoint --policy.config=pi05_libero --policy.dir=checkpoints/pi05_libero/my_experiment/20000This will spin up a server that listens on port 8000 and waits for observations to be sent to it. We can then run the LIBERO evaluation script to query the server. For instructions how to install LIBERO and run the evaluation script, see the LIBERO README.
If you want to embed a policy server call in your own robot runtime, we have a minimal example of how to do so in the remote inference docs.
We provide more examples for how to fine-tune and run inference with our models on the ALOHA platform in the following READMEs:
We will collect common issues and their solutions here. If you encounter an issue, please check here first. If you can't find a solution, please file an issue on the repo (see here for guidelines).
| Issue | Resolution |
|---|---|
uv sync fails with dependency conflicts |
Try removing the virtual environment directory (rm -rf .venv) and running uv sync again. If issues persist, check that you have the latest version of uv installed (uv self update). |
| Training runs out of GPU memory | Make sure you set XLA_PYTHON_CLIENT_MEM_FRACTION=0.9 before running training to allow JAX to use more GPU memory. You can also use --fsdp-devices <n> where <n> is your number of GPUs, to enable fully-sharded data parallelism, which reduces memory usage in exchange for slower training (the amount of slowdown depends on your particular setup). |
| Policy server connection errors | Check that the server is running and listening on the expected port. Verify network connectivity and firewall settings between client and server. |
| Missing norm stats error when training | Run scripts/compute_norm_stats.py with your config name before starting training. |
| Dataset download fails | Check your internet connection. For HuggingFace datasets, ensure you're logged in (huggingface-cli login). |
| CUDA/GPU errors | Verify NVIDIA drivers are installed correctly. For Docker, ensure nvidia-container-toolkit is installed. Check GPU compatibility. You do NOT need CUDA libraries installed at a system level --- they will be installed via uv. You may even want to try uninstalling system CUDA libraries if you run into CUDA issues, since system libraries can sometimes cause conflicts. |
| Import errors when running examples | Make sure you've installed all dependencies with uv sync. Some examples may have additional requirements listed in their READMEs. |
| Action dimensions mismatch | Verify your data processing transforms match the expected input/output dimensions of your robot. Check the action space definitions in your policy classes. |
| Diverging training loss | Check the q01, q99, and std values in norm_stats.json for your dataset. Certain dimensions that are rarely used can end up with very small q01, q99, or std values, leading to huge states and actions after normalization. You can manually adjust the norm stats as a workaround. |
