Skip to content

Repository files navigation

LIBERO-Para

A Diagnostic Benchmark and Metrics for Paraphrase Robustness in VLA Models

Chanyoung Kim1*   Minwoo Kim1*   Minseok Kang1   Hyunwoo Kim2   Dahuin Jung2†

1Soongsil University    2Chung-Ang University
* Equal contribution    † Corresponding author

Paper Project Page HuggingFace Dataset Daily Papers Featured on X

🔥 Our paper is ranked #3 on Hugging Face Daily Papers! (Apr 7, 2026)


Paraphrase Taxonomy

Robotic manipulation instructions are fundamentally structured around what to act on and how to act. LIBERO-Para decomposes paraphrases along these two essential axes: object-referring expressions and action-referring expressions. Each axis is independently varied, and composing both yields compositional paraphrases. See the paper for details.


Highlights

  • Unified benchmark: LIBERO-Para extends LIBERO with 4,000+ paraphrased instructions across 10 evaluation scenarios, all in a single repository.
  • Multi-model evaluation: 6 VLA models integrated with per-model conda environments, standalone eval scripts, and step-by-step guides — clone, install, and run.
  • Original LIBERO compatible: All original LIBERO suites (Spatial, Object, Goal, LIBERO-10, LIBERO-90) are preserved and can be evaluated from the same codebase.
  • PRIDE metric: A new difficulty-aware metric that goes beyond binary success rate, giving more credit for succeeding on harder paraphrases.

Evaluation Guides

Each model is evaluated using a custom standalone script under eval_scripts/examples/, which directly interfaces with the model's inference server or loads the model directly. Follow each model's guide for environment setup and evaluation.

Model Params Architecture Release Guide Script Status
OpenVLA-OFT (Goal) 7.5B Parallel Decoding 2025.03 Guide eval_openvla_oft.py
OpenVLA-OFT (Mixed) 7.5B Parallel Decoding 2025.03 Guide eval_openvla_oft.py
Pi 0.5 (Base) 3.3B VLM + Action Expert 2025.09 Guide eval_pi05.py
Pi 0.5 (Expert-Only) 3.3B VLM + Action Expert 2025.09 Guide eval_pi05.py
X-VLA 0.9B Soft-prompted 2026.01 Guide eval_x_vla.py
VLA-Adapter 0.6B Bridge-based 2025.09 Guide eval_vla_adapter.py
Xiaomi-Robotics-0 4.7B VLM + Action Expert 2026.02 Guide eval_xiaomi_robotics_0.py
More coming soon...

Adding a new model? Each eval script follows the same pattern: pre-create 10 LIBERO envs, swap in paraphrased instructions, and query the model. Clone the model repo into eval_scripts/, write a lightweight eval script in eval_scripts/examples/, and add a guide in eval_guides/. See any existing script as a template. We plan to continuously add more VLA models.


PRIDE Metric

PRIDE (Paraphrase Robustness Index in Robotic Instructional DEviation) evaluates how robustly a VLA model handles paraphrased instructions.

It computes a Paraphrase Distance (PD) from keyword similarity (SK) and structural similarity (ST), then measures the ratio of PD-weighted successes to total possible PD, normalized to 0–100. Unlike plain success rate, PRIDE gives more credit for succeeding on harder, more deviated paraphrases.

Details: metrics/README.md  |  Interactive: PRIDE_metric_playground.ipynb


Metrics & Analysis

Setup

conda create -n libero-para python=3.10 -y
conda activate libero-para
pip install -r metrics/requirements.txt
python -m spacy download en_core_web_sm

Quick Start

Example results (Xiaomi-Robotics-0, seed7) are included in logs_para/example_xiaomi-robotics-0/.

Note: Action trajectories are stripped from the example logs to reduce file size.

python metrics/analyze_results.py \
    --model_path logs_para/example_xiaomi-robotics-0

See metrics/README.md for multi-model comparison, PRIDE sweep, and more.


Project Structure

LIBERO-Para/
├── libero/                        # Benchmark core (LIBERO-based)
├── metrics/                       # PRIDE metric & analysis tools
│   ├── analyze_results.py
│   ├── PRIDE_metric_playground.ipynb
│   └── libero_para_metadata.csv
├── eval_guides/                   # Per-model setup guides
├── eval_scripts/
│   ├── examples/                  # Eval scripts (one per model)
│   ├── openvla-oft/               # Clone: github.com/moojink/openvla-oft
│   ├── x-vla/                     # Clone: github.com/huggingface/lerobot
│   ├── vla-adapter/               # Clone: github.com/OpenHelix-Team/VLA-Adapter
│   └── xiaomi-robotics-0/         # Clone: github.com/XiaomiRobotics/Xiaomi-Robotics-0
├── logs_para/                     # Evaluation results
│   └── example_xiaomi-robotics-0/ # Example data (seed7)
├── images/
├── benchmark_scripts/
└── scripts/

TODO

  • Release the paraphrase generation pipeline (LLM-based generator + verifier used to produce the 4 092 LIBERO-Para paraphrases)
  • Release CALVIN-Para — paraphrase-robustness benchmark for CALVIN, for cross-benchmark transferability tests of PRIDE

Acknowledgement

This project is built upon LIBERO by Bo Liu, Yifeng Zhu, Chongkai Gao, Yihao Feng, Qiang Liu, Yuke Zhu, and Peter Stone.

About

Official code for "LIBERO-Para: A Diagnostic Benchmark and Metrics for Paraphrase Robustness in VLA Models" (arXiv 2603.28301).

Topics

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages