A Multi-Modal Search and Curation Platform for Physical AI
How to Install ✦ Detailed Documentation ✦ Tutorials & Recipes ✦ Technical Report
SIL-Wheel is a framework for building searchable, curated video datasets for Physical AI. With SIL-Wheel, we can discover clips that match specific criterion, validate retrieval results, construct training and evaluation slices and analyze model behavior on those slices.
The core idea is to keep search, curation, and evaluation in the same loop. Users can retrieve clips using captions, embeddings, visual content, ego-motion patterns, object filters, metadata, and classifier predictions. These signals can be composed in a single query, so users can express targeted scenarios such as hard braking at intersections, dense pedestrian scenes, or rare trajectory patterns. Retrieved clips can then be reviewed, annotated, refined with classifiers or clustering, exported as curated datasets, or used directly for slice-based model evaluation.
-
Composable multi-modal search: Search large video corpora using caption full-text search, caption embeddings, video embeddings, visual embeddings, ego-trajectory shape matching, motion-pattern filters, perception-derived object filters, classifier scores, and metadata constraints. Multiple signals can be combined in the same query to retrieve clips that satisfy all active conditions.
-
Interactive validation and annotation: Review retrieved clips in the web interface, inspect captions and trajectories, add manual labels, edit temporal spans, and convert validated results into reusable annotations or curated slices.
-
Dataset and benchmark slice construction: Build targeted slices of data for training, evaluation, or failure analysis purposes. Candidate clips can be expanded with learned classifiers, filtered based on various criterion, and exported for downstream applications.
-
Targeted model evaluation: Evaluate models on curated slices rather than only on broad aggregate datasets. SIL-Wheel supports slice-level metrics, leaderboard-style comparisons, per-clip inspection, and human-powered pairwise preference evaluation workflows.
-
Flexible raw video ingestion: Process raw video datasets from the local filesystem, S3 object storage, or the Hugging Face Hub using the same preparation pipeline. Local and S3 inputs can be individual
.mp4files or.tararchives, while Hugging Face datasets can be provided as.taror.zipshards. -
End-to-end preprocessing pipeline: Scripts under
scripts/process raw videos and metadata into the artifacts needed by SIL-Wheel, including captions, embeddings, metadata tables, search indices, and launch configurations.Seedocs/data-preparation.mdfor the full data preparation pipeline. -
Web UI and Python clients: Use the web interface for browsing, searching, annotating, curating, and evaluating models, or query SIL-Wheel programmatically through
WheelClientandWheelHTTPClient. Both interfaces share the same search composition, ranking, and caching logic.
SIL-Wheel runs on Python 3.12. After cloning the repository, the simplest way to make sure that all dependencies are properly installed is to create the provided conda environment:
git clone https://github.com/nv-tlabs/sil-wheel.git
cd sil-wheel
conda env create -f environment.yml
conda activate wheel
python setup.py build_ext --inplace
pip install -e .Note
The preprocessing stages that extract embeddings depend on the core module from perception_models.
After you have created the wheel conda environment please install it as well as follows:
pip install --no-deps git+https://github.com/facebookresearch/perception_models.gitNote that we use --no-deps so that the package does not replace the
versions of dependencies already installed dependencies.
If you plan to read videos from S3 paths, install and configure the AWS CLI:
pip install awscliModels are pulled from Hugging Face Hub, so make sure that you have it properly authenticated
pip install --upgrade huggingface_hub
hf_transfer_login # or: huggingface-cli loginflash-attn is optional. SIL-Wheel works completely fine without it by using the standard PyTorch attention fallback.
Warning
It may not build cleanly with some CUDA and PyTorch combinations, including CUDA 13.0 and PyTorch 2.10.
Prefer containers? You can build and run SIL-Wheel from Docker instead of the
conda environment. See docker/ for the data preparation and
server images.
The quickest way to try SIL-Wheel is to run one of the example walkthroughs.
Each walkthrough downloads a dataset, runs the preparation pipeline,
builds the search indices, and writes a ready-to-run config.yaml with the
available search modalities populated. These examples use the same scripts as
the general data preparation pipelines (see scripts/), so they can also serve as references for
preparing your own datasets.
nuScenes is the simplest starting point. The public nuScenes mini split
contains 10 scenes, requires no account, and can be processed in a few minutes
on a single RTX 4090. Prepare the data, then launch the server from the
config.yaml the setup writes:
python examples/getting-started-nuscenes/setup_nuscenes.py
python scripts/launch_server.py wheel-data/config.yamlPhysical AI Autonomous Vehicles runs the same pipeline on a slice of
NVIDIA's Physical AI Autonomous Vehicles
dataset, streamed from the Hugging Face Hub as .zip shards. Prepare it, then
launch the server from the config.yaml the setup writes:
python examples/getting-started-physical-ai-autonomous-vehicles/setup_physical_ai.py \
--workdir ./wheel-data-physical-ai \
--camera camera_front_wide_120fov \
--max-clips 500
python scripts/launch_server.py wheel-data-physical-ai/config.yamlSee the nuScenes and Physical AI READMEs for prerequisites and the full set of options.
For your own data, prepare it with the same scripts/, point a copy of
config/wheel_launch_dev_server_config.yaml at your artifacts, and launch:
python scripts/launch_server.py config/wheel_launch_dev_server_config.yamlOpen the bind address printed at startup. This is the server.bindto value in
the YAML configuration.
Note
The examples bind to 127.0.0.1:8012, which only accepts connections from the
machine running the server. To reach it from anywhere else, point bindto at
the address of the host you are launching from, in any of these ways:
# at data preparation time, baked into config.yaml
python examples/.../setup_physical_ai.py --host 10.0.0.5
# at launch time, without touching config.yaml
python scripts/launch_server.py wheel-data-physical-ai/config.yaml \
--override server.bindto=10.0.0.5:8012Or edit server.bindto in config.yaml directly. The address in use is
printed at startup as Listening at ....
SIL-Wheel can also be driven directly from Python through two clients:
WheelClientfor local access, when your process can read the datasets, indices, metadata, and artifacts referenced in the launch YAML.WheelHTTPClientfor remote access, when you connect to a running SIL-Wheel server over HTTP.
Use WheelClient when running in the same environment as the indexed
artifacts.
from sil_wheel.client import WheelClient
client = WheelClient.from_config("config/wheel_launch_dev_server_config.yaml")
# Single-modality search via a convenience helper. Caption search is keyword
# based, so a bare multi-word query matches that exact phrase; use single terms
# or explicit AND/OR/NOT to combine them.
result = client.search_caption("intersection")
print(len(result), "clips matched")
# first 10 clip_ids
print(result.head(10))
# Compose multiple modalities and filters in one query, fused with RRF.
result = client.search(
# caption FTS: keyword/phrase matching against the caption text
search="intersection AND pedestrian",
# text->video embedding: free-form description, no keyword overlap needed
semantic_search_text="hard braking at intersection",
data_source=["nuscenes"],
rank_mode="rrf",
)
# per-clip per-modality scores
df = result.as_dataframe()
print(df.head(10))Use WheelHTTPClient when connecting to an already running SIL-Wheel server. Its
surface is identical to WheelClient, except that result.scores is empty for
remote results (the server returns clip IDs only), so use result.clip_ids
directly.
from sil_wheel.http_client import WheelHTTPClient
client = WheelHTTPClient(
server_url="http://wheel-host:8012",
username="alice",
password="...",
)
result = client.search_caption("intersection")
print(result.clip_ids[:10])
# Or copy a URL straight out of the UI's address bar:
result = client.search_from_url(
"http://wheel-host:8012/?search=intersection&data_source=nuscenes"
)Full client surface, search composition, and ranking modes are in the online documentation.
Full user and developer documentation, including guides, tutorials, and the complete API reference, can be found at SIL-Wheel documentation site.
For quick in-repo references:
| Topic | Reference |
|---|---|
| Data preparation: video processing, embeddings, metadata, and S3 upload | docs/data-preparation.md |
| Arena: blind model evaluation, manifest format, and Glicko-2 ratings | docs/arena.md |
| Bug reporting: the Google Sheets backed form | docs/bug-reporting.md |
If you use SIL-Wheel in your research, please cite:
@misc{sil-wheel,
title = {SIL-Wheel: A Multi-Modal Search and Curation Platform for Physical AI},
author = {NVIDIA},
year = {2026},
url = {https://github.com/nv-tlabs/sil-wheel}
}Maintained by the NVIDIA SIL team. This project is currently not accepting external contributions.
