One CLI, one SDK, one workflow layer for physical-AI workloads on Nebius — data curation, simulation, synthetic data, policy training, evaluation, observability, and cluster orchestration.
Quickstart · Guides · Workbench docs · CLI reference · Cookbooks · Contributing
npa is the CLI and SDK for Nebius Physical AI. Workbench is its primary
solution: one command surface that composes data curation, simulation, synthetic
data, policy training, evaluation, export, observability, and declarative
workflows on Nebius object storage, managed Kubernetes, vLLM-compatible serving,
and GPU clusters (H100 · H200 · L40S · B300 · RTX PRO 6000).
You bring a robot, a dataset, or a pipeline idea. npa brings the containers,
the orchestration, and the preflight checks that catch a missing token before it
stalls your run.
Workbench is meant to be operated by your own coding agent. Connect the
agent to this checkout with terminal access, attach your cloud and model
credentials through its private environment or secret manager, and ask it to
configure, validate, provision, submit, and inspect workflows with you. The
prompts below are a starting point; you do not need to learn every npa command
before running something real.
| What you can do | Curate datasets · train and evaluate policies · render synthetic data · run sim-to-real loops · serve models |
| Who it's for | Robotics teams, physical-AI researchers, and partners shipping on Nebius |
| Where it runs | Nebius S3, managed Kubernetes, and GPU clusters |
| How you extend it | Declarative npa.workflow/v0.0.1 YAML specs and reusable Workbench tool refs |
Partners integrate independently. Teams assemble from open blueprints. Nebius owns the infrastructure layer and compute substrate.
Three steps from a clone to a real result on Nebius.
Python 3.10+. npa is not on PyPI — install it editable from the clone:
git clone https://github.com/nebius/nebius-physical-ai.git
cd nebius-physical-ai
python3 -m venv .venv
source .venv/bin/activate
pip install -e npa
npa --versionWindows: use WSL2 Ubuntu. Per-platform steps are in docs/install.md.
Managed deployments (such as
npa agent fresh-setup) also need Terraform 1.x onPATH—pip install -e npadoes not install it. Check withterraform version; agent bootstrap installs the tested 1.13.3 baseline only when Terraform is missing entirely.
Sign up, create a
tenant and project, install the
Nebius CLI, then let npa configure write ~/.npa/credentials.yaml and
~/.npa/config.yaml for you:
curl -fsSL https://storage.eu-north1.nebius.cloud/cli/install.sh \
| NEBIUS_CLI_VERSION=0.12.254 bash
export PATH="${HOME}/.nebius/bin:${PATH}" # add to ~/.zshrc or ~/.bashrc
npa configurenpa is tested with Nebius CLI 0.12.254 (recommended) and 0.12.227
(compatible, with a warning). Anything else is blocked before provider calls,
and the error prints the exact install command to run.
npa configure also prompts for optional model and inference tokens, linking
each setup guide inline: Hugging Face ·
NVIDIA NGC ·
Nebius Token Factory.
Its Hugging Face and NGC access summary is informative: missing, rejected,
gated, or temporarily unreachable providers do not prevent otherwise-valid
configuration from being saved. Use npa workbench health access when access
must be an enforcing gate.
If an agent is operating the checkout, attach these values to its private environment instead of pasting token values into chat:
NEBIUS_TENANT_ID=<your-tenant-id>
NEBIUS_PROJECT_ID=<your-project-id>
NEBIUS_REGION=<your-project-region>
HF_TOKEN=<your-hugging-face-read-token>
NGC_API_KEY=<your-ngc-api-key>
NEBIUS_TOKEN_FACTORY_KEY=<your-token-factory-key>The Hugging Face account behind HF_TOKEN must already have access to the
models used by the workflow. A fine-grained token must include those repos.
Gated terms can only be accepted by you on Hugging Face; the access check below
prints every exact page still requiring action. NGC is part of the general
Workbench setup even though the PAIDF + Cosmos 3 path below currently obtains
its model weights from Hugging Face. Token Factory is required by that path for
captioning and evaluation.
The active Nebius CLI identity must also be allowed to administer tenant IAM
during first-time storage setup. npa configure creates a project service
account and access key plus a tenant IAM group and bucket-scoped permit; project
admin alone is not sufficient. Remove any temporary broad role after setup and
teardown are complete.
Then give your agent this prompt:
Set up Nebius Physical AI Workbench in this checkout. Read AGENTS.md and
skills/index.yaml first and follow the relevant first-run, credential-preflight,
and GPU guidance.
Use NEBIUS_TENANT_ID, NEBIUS_PROJECT_ID, NEBIUS_REGION, HF_TOKEN, NGC_API_KEY,
and NEBIUS_TOKEN_FACTORY_KEY from the private process environment. Never print
secret values, put them in command arguments, or write them into the repository.
Never run `env`, `printenv`, `set`, `export -p`, or another command that dumps
the process environment; inspect only allowlisted names and report present or
missing. Do not read credential files except through npa's credential APIs.
Use NPA_PROJECT_ALIAS if it is set; otherwise use "workbench" as the local alias.
Install or verify npa and its reported host prerequisites, including Terraform
and, for SkyPilot Kubernetes on Debian/Ubuntu, socat. Verify the active Nebius
CLI identity and configure the known tenant, project, and region
non-interactively. Persist supported environment credentials with npa configure
--save-env-credentials and use explicit --provision to create or reuse writable
project object storage. Without --provision, known-project setup must remain
provider-free and leave storage unselected.
Confirm the active identity can create the tenant IAM objects that secure that
storage. Then run npa configure --show,
npa workbench health preflight --json, and
npa workbench health access --capability paidf,cosmos3 --json.
Do not bypass a failed gate or provision GPU resources yet. If Hugging Face
access is missing, give me the exact model-page links, wait for me to accept the
terms, and rerun the check. Finish only when project storage, credentials, and
the PAIDF/Cosmos 3 model access checks pass, then guide me straight into my first
workflow.
Creating projects from the CLI, SSO/federation profiles, non-interactive automation, and the full credential model live in docs/quickstart.md.
Check your credentials, then put them to work:
npa workbench health preflightOne PASS/WARN/FAIL/SKIP sweep over the credentials nearly every job needs —
Hugging Face, NVIDIA NGC, Nebius object storage, and Token Factory. Add --json
for machine-readable output.
Once it comes back green, launch something real on Nebius GPUs. The flagship is NVIDIA Cosmos, and any robot guide below will take you from a public dataset to a trained and evaluated policy.
Do not stop at setup. Keep the same agent in the loop and have it guide the
first workflow from input selection through the final artifact. For the
Physical AI Data Factory with real source-video-conditioned Cosmos 3, attach a
local H.264 MP4 or set PAIDF_INPUT_URI to one private s3:// MP4, then paste:
Run my first Workbench workflow with me: the PAIDF Cosmos 3 video-conditioning
workflow at npa/workflows/workbench/npa-workflows/paidf-cosmos3.yaml. Follow
docs/workbench/guides/paidf-cosmos3.md and the repository skills. Use the
configured Nebius project, region, writable bucket, and credentials. Use the
attached local H.264 MP4, or PAIDF_INPUT_URI if it is set; if neither is
available, ask me only for the input video before continuing. Keep all input and
artifact locations private.
Never run `env`, `printenv`, `set`, `export -p`, or another command that dumps
the process environment. Inspect only allowlisted variable names and report
present or missing; do not print secret values or read credential files directly.
Re-run the credential and model-access gates for paidf,cosmos3. Validate and
plan the spec with the real bucket and input, starting with one variant and one
supported GPU. Honor configured TF_VAR_* topology and reserved-capacity settings.
Bootstrap and verify SkyPilot, discover the accelerator name the target cluster
advertises, and provision the required CPU/GPU resources if they are absent.
Show me the validated plan and explain the resources it will create, then stage
the input and preflight the selected images. If a selected image fails the
SkyPilot bootstrap contract, build the repository's current compliant image,
push it to an authorized private project registry at an immutable digest, and
repeat preflight. Then submit with --runtime.
Forward only secret names through --secret-env: HF_TOKEN,
NEBIUS_TOKEN_FACTORY_KEY, AWS_ACCESS_KEY_ID, and AWS_SECRET_ACCESS_KEY; never put
secret values in YAML or command arguments.
Stay with the run until it reaches a terminal state. If it fails, diagnose the
recorded stage and resume safely rather than starting an unrelated run. If it
succeeds, show me the generated and curated artifacts and load the final Rerun
recording when an agent viewer is available. A terminal quality rejection after
the workflow's bounded refinement loop is a valid fail-closed result: do not
lower the threshold or force promotion. Show me the generated video, evaluator
report, quality disposition, and Rerun evidence, and explain that labeling and
curation were intentionally skipped.
The workflow uses the independent
paidf-cosmos3.yaml composition; the
original Physical AI Data Factory workflow continues to use Cosmos Transfer
2.5.
Short, copy-paste walkthroughs. Pick whichever sounds fun — they are independent.
| I want to… | Go here | Needs |
|---|---|---|
| Pick and place with a Franka arm | Franka + Genesis | GPU cluster |
| Teach a robot to push a T | PushT sim-to-real | GPU cluster |
| Train a Reachy 2 humanoid policy | Reachy 2 + LeRobot | GPU cluster |
| Make a Unitree G1 walk | G1 + SONIC | GPU cluster |
| Train a quadruped to run | Quadruped + Isaac Lab | RT-core GPU |
| Run the flagship GPU workload | NVIDIA Cosmos | GPU cluster |
| Augment robot video with PAIDF + Cosmos 3 | PAIDF with Cosmos 3 | GPU cluster + S3 |
| Run a real data pipeline, not a robot | Physical AI Data Factory | GPU cluster + S3 |
| Rebuild a real scene in 3D | Neural reconstruction | RT-core GPU |
| Get a browser workbench with a Rerun viewer | Deploy the npa agent |
Terraform + S3 (~20 min) |
Full index: docs/workbench/guides/README.md. Longer end-to-end recipes (BDD100K + LanceDB, Isaac-Lab BYOF, LeRobot GPU benchmarks): cookbooks.
Every tool lives under npa workbench (there is no solutions namespace).
A few highlights:
token-factory— hosted inference, captioning, and reasoning against your own frames.vlm-eval— scores rollouts with API or self-hosted vLLM backends; seevlm-eval-single.yaml.health preflight— validates HF / NGC / S3 / Token Factory before a deploy or a GPU job.foxglove— packs run frames, metrics, and logs into MCAP for the embedded viewer (CLI · export contract).golden-eval— runs per-container hello-world reruns as a CI gate.trigger— watches S3-compatible prefixes and retriggers workflows.sonic export— converts locomotion checkpoints to ONNX.
Browse the full command inventory by category
| Category | Workbench commands |
|---|---|
| Data curation | npa workbench fiftyone curate, eval, load-dataset, datasets list; npa workbench lancedb deploy, create-table, import-lerobot, import-bdd100k, backfill, create-mv, refresh-mv, query-table, query; npa workbench detection-training train, eval, status, list |
| Synthetic data | npa workbench cosmos infer, train, serve, status; npa workbench cosmos2 transfer; npa workbench cosmos3 reason; npa workbench genesis generate-demos; specs such as bdd100k-pipeline.yaml |
| Simulation | npa workbench isaac-lab train, eval, export-lerobot, export-onnx; npa workbench leisaac launch, status, destroy (browser teleoperation); npa workbench genesis train-teacher, generate-demos, eval-teacher, eval-student, diagnose, tune; npa workbench sonic retargeting run, workflow |
| Eval | npa workbench vlm-eval run, benchmark, workflow, status, list; npa workbench mjlab eval, workflow; npa workbench sonic eval; npa workbench fiftyone eval; npa workbench isaac-lab eval; npa workbench genesis eval-student; npa workbench golden-eval run, run-all, validate |
| Robot policy | npa workbench lerobot train, eval, serve, infer, list-checkpoints, benchmark, profile-train, train-student; npa workbench groot download, finetune, eval, serve, infer, convert; npa workbench sonic train, serve, export, eval, status, list |
| World models | npa workbench cosmos deploy, serve, infer, train, finetune, optimize, autoscale, status, system-info |
| Hosted LLM | npa workbench token-factory caption, generate, reason, verify, models, workflow, status |
| Workflows | npa workbench workflow validate-spec, plan-spec, run-spec, submit; workbench workflows under npa-workflows/ |
| Observability | Tool-level status, list, and system-info commands; npa workbench workflow status, logs; npa workbench health preflight; npa workbench foxglove convert-run, inspect, install-sdk, config; npa rerun host, share, list-shares, revoke; npa cluster status, list |
| Platform utils | npa configure / init, npa provision-if-absent; npa agent, npa skypilot bootstrap/status/verify, npa soperator, npa burst, npa cluster, npa network, npa adapter convert, npa convert lerobot-to-rrd/-mp4, npa viz, npa demo |
Full CLI reference: docs/cli/README.md.
Author pipelines as declarative npa.workflow/v0.0.1 specs — a state graph of
Workbench toolRef steps with S3 handoffs, gates, and loops. The same YAML is
what you validate, plan, and submit.
# Validate and plan (no submit)
npa workbench workflow validate-spec npa/workflows/workbench/npa-workflows/vlm-eval-single.yaml
npa workbench workflow plan-spec npa/workflows/workbench/npa-workflows/vlm-eval-single.yaml --run-id demo
# Launch on Nebius (after npa configure)
npa workbench workflow submit npa/workflows/workbench/npa-workflows/vlm-eval-single.yaml \
--run-id demo --registry registry.example/customer
# Inspect the plan without launching
npa workbench workflow submit npa/workflows/workbench/npa-workflows/token-factory-caption.yaml \
--plan-only --run-id demo| Format | apiVersion: npa.workflow/v0.0.1 |
| CLI | validate-spec · plan-spec · run-spec · submit |
| Workbench workflows | npa/workflows/workbench/npa-workflows/ |
| Tool catalog | npa-workflow-tool-catalog.md |
| Authoring guide | npa-workflow-guide.md |
| What submit does | Run lifecycle — gates, run identity, restart safety, status |
Prefer these specs for new pipelines. Parallel fan-out and a few specialized
paths remain outside v0.0.1 scope — see the catalog README for exceptions. The
Sim2Real 14-stage engine is a separate path
(skill) using sim2real/runbook.yaml
plus Python stage glue.
Teardown is an ordered sequence — cancel jobs, destroy the agent, remove the shared controller, destroy the cluster, delete the bucket, remove storage IAM, then clear local state. Skipping a step leaves something billing.
npa cleanup # report + the exact runbook for your machine
npa destroy --project <alias> --all # read-only until you add --yesEvery command, every guard, and how to finish a teardown you can no longer address by alias: docs/teardown.md.
Every Workbench tool ships as a container image. The publicly redistributable subset is mirrored to GHCR, so the easiest path is to pull instead of build:
export NPA_REGISTRY=ghcr.io/nebius/nebius-physical-ai
docker pull "${NPA_REGISTRY}/npa-retargeting:0.1.1"The GHCR mirror is the runtime default; no registry setup is required in
npa configure. Use your own Nebius registry only when you need private or
locally modified images, and select it with NPA_REGISTRY or an explicit image:
docker login registry.example
export NPA_REGISTRY=registry.example/customer
npa/docker/workbench/lerobot/build.sh --registry "$NPA_REGISTRY" --push| Reference | What it tells you |
|---|---|
| Public image catalog | Exact GHCR names, tags, pull commands, and intentional exclusions |
| Image ↔ GPU compatibility matrix | Every image against every Nebius GPU platform, and which cells are hardware-verified |
| Packaging contract | Tiers, non-root users, ports, and redistribution classes |
| Golden evals | The real capability test each image must pass — not an import probe |
| Blackwell compatibility | B200 / B300 build, tag, and validation runbook |
| SONIC image catalog | Manifest-driven SONIC variant routing per GPU |
| Image reproducibility | The two-tag strategy (cuda12, cuda13-b300) and how tags are pinned |
Each image declares a redistribution class that decides whether it may leave
the owning org. Public images may be mirrored to GHCR; restricted images stay
build-your-own (cosmos3-serving is restricted because its pinned base embeds a
runtime under NVIDIA's Deep Learning Container License). Set the class when you
add an image — the packaging-contract test fails a build that bakes a
non-redistributable runtime while claiming public.
Eight Workbench tools are validated end to end on Nebius today: LanceDB, FiftyOne, LeRobot, Genesis, Isaac Lab, Cosmos, GR00T, and SONIC.
| Reference | What it tells you |
|---|---|
| B300 validation matrix | Which tools pass on B300 vs which are vendor-paced or upstream-blocked |
| LeRobot GPU benchmarks | Steps/s across H200 · B300 · L40S · RTX PRO 6000 by policy type |
| NVIDIA architecture coverage | CUDA 12.8 x86_64 vs CUDA 13 aarch64 tool coverage |
| Partner roadmap | NVIDIA Omniverse / CAD-to-SimReady capabilities on the way — not yet shipped |
npa/ # Python package (CLI + SDK); install with `pip install -e npa`
src/npa/cli/ # Typer entry point and every top-level command
src/npa/workbench/ # Per-tool implementations (cosmos, lerobot, sonic, ...)
workflows/workbench/
npa-workflows/ # Workbench npa.workflow/v0.0.1 specs (author + submit these)
sim2real/ # Staged 14-stage sim2real runbook
docs/ # Quickstart, architecture, workbench guides, cookbooks
skills/ # SKILL.md files for agents and contributors (source of truth)
deploy/ # Terraform + cluster provisioning (uses Nebius solutions library)
research/ # LeRobot deploy research (older reference)
workbench/mlflow/ # MLflow tracking-server compose stack
User secrets live in a versioned, exact-project map in ~/.npa/credentials.yaml;
machine-managed config lives in ~/.npa/config.yaml. The repo supports multiple
top-level solution namespaces, and Workbench is the current primary one — future
solutions are additive and never rename or nest it. See
solutions model ·
CLI namespaces ·
contributor context.
| Topic | Where to look |
|---|---|
| Install & auth | quickstart.md · install.md |
| Workbench setup | getting-started.md |
| Beginner robot guides | guides/README.md |
| Physical AI Data Factory | deploy runbook · concepts |
| Cookbooks | cookbooks/README.md — incl. BDD100K + LanceDB and Isaac-Lab BYOF |
| Workflow authoring | npa-workflow-guide.md · tool catalog |
What submit does |
run-lifecycle.md |
| Self-hosted agent | agent.md · operator skill · fresh-operate |
| Teardown & cost | teardown.md |
| Container images | catalog · packaging contract |
| Preemptible GPU VMs | preemptible-vms.md |
| Troubleshooting | known-footguns.md · FIXME.md · FTUE audit |
| CLI reference | cli/README.md |
| Everything else | docs/ |
We welcome PRs, issues, and workflow contributions.
pip install -e "npa[dev]"
make testRead CONTRIBUTING.md for the review checklist,
skill-maintenance requirements, and repo hygiene rules. New behavior should have
a matching root skills/ entry — see skills/index.yaml.
Security disclosures: SECURITY.md. Support and community happen
through GitHub Issues and
Pull Requests.
Licensed under the Apache License 2.0. Built by Nebius and the physical-AI community.