How to Train Your Robot: The LanceDB Edition

Community Article
Published September 24, 2026

Robot learning data used to come in as many shapes as there were labs. Every group had its own recording format, its own episode layout, its own way of pairing camera video with joint states and actions. Reusing someone else's dataset meant writing a converter first, and reusing someone else's training code meant writing another one.

LeRobot helped standardize that ecosystem around a common dataset format and training stack. Now it is opening its dataset layer to support new storage formats natively, with Lance as the first.

That means LeRobotDataset can read Lance datasets directly while preserving the same training APIs and tooling. Teams can train from object storage, globally shuffle across the dataset, and search, curate, and add features without maintaining separate copies of the data.

LeRobot <> LanceDB integration

LeRobot reads datasets stored as LanceDB tables natively. It picks the right dataloader based on the source dataset. Nothing about your recording or training code changes.

What changes is what the dataset can do:

  • Train from wherever the data lives. LanceDB reads random frames straight from the Hugging Face Hub or any popular object store like Hugging Face Storage Buckets, fetching only the bytes each batch needs. No download step, and every batch can be a true global shuffle across the whole dataset.
  • Converting costs about as much as a file copy. Videos are stored as blobs without re-encoding, so items are bit-identical to the source.
  • The training table is also the index. Embeddings, derived scores, vector and full-text indexes live as columns and indexes on the same table the DataLoader reads. Search, curation and mining are queries on that table, at a version the trainer can pin.
  • Add features without copying the data. A new column is a declaration plus a backfill. Old data is never copied.
  • Look at any episode where it sits. LeRobot's dataset viewer serves episodes to Foxglove straight from the table, so inspecting one episode of a TB-scale dataset moves a few megabytes.

We first test the setup on a small dataset from a Hugging Face Storage Bucket, then scale the same approach to DROID.

Train directly from Hugging Face Storage Buckets at NVMe speed, with a global shuffle

Hugging Face Storage Buckets bring simple, Hub-native object storage to large robotics datasets. Keep any file layout in a bucket, address your data directly through hf://buckets/<namespace>/<bucket>/<path>, and manage it with the familiar hf CLI. It's an easy way to keep datasets that are too large to replicate across every machine in one place. And Buckets plug directly into the Hugging Face ecosystem, including LeRobot.

With the LanceDB integration, just point the trainer at it with --dataset.repo_type=bucket. No download, no local disk, no dataset-specific code. The reader fetches the bytes each batch needs, and the shuffle is a full global shuffle across every frame in the dataset, not a window.

# train straight from the bucket
lerobot-train \
  --dataset.repo_id=lance-format/lerobot-bench/koch_pick_place_5_lego-lance \
  --dataset.repo_type=bucket \
  --policy.path=lerobot/smolvla_base --policy.empty_cameras=1 \
  --batch_size=32 --num_workers=8

You can also point directly at the bucket URI with --dataset.root:

lerobot-train \
  --dataset.repo_id=lance-format/lerobot-bench/koch_pick_place_5_lego-lance \
  --dataset.root=hf://buckets/lance-format/lerobot-bench/koch_pick_place_5_lego-lance \
  --policy.path=lerobot/smolvla_base --policy.empty_cameras=1 \
  --batch_size=32 --num_workers=8

lerobot/koch_pick_place_5_lego, 37,972 frames from two cameras, SmolVLA, two H100s. We ran the training job comparing remote LanceDB dataset and default local NVMe after a download. Same model, same seed, same 1,500 steps.

koch_pick_place_5_lego, SmolVLA, 2×H100 LanceDB, HF Storage Bucket, global shuffle Upstream, local NVMe, global shuffle
Steady rate 212 samples/s 212 samples/s
Step time 0.30 s 0.30 s
Wall clock, 1,500 steps 503 s 484 s
Loss, step 100 to 1,500 0.324 to 0.107 0.324 to 0.107

Two H100s on this model consume about 210 samples per second, and the bucket delivered it. The GPUs were as busy reading from a Hugging Face Storage Bucket as they were reading a local copy, at the same step time and with an identical loss curve.

Loader throughput across datasets

The koch run is one dataset and one model. We also tested the reader on its own across six public datasets on the same machine, using 8 workers and a batch size of 64, with results averaged over two runs.

The shuffle window was 0.05% of the dataset, with a minimum of 5,000 frames, giving 13,815 on DROID. We tested the LeRobot streaming reader in PR #4342 with multi-threaded video decoding. This implementation is under active development, so the comparison reflects the version tested.

On the same batches, LanceDB was 1.7× faster than streaming on DROID while using a fraction of the RAM, and up to 11.6× faster on the other tested datasets. With nothing downloaded, its global shuffle is 1.1× to 1.7× faster than upstream reading a local copy. The difference depends on the size and complexity of the dataset too, such as the number of camera angles and the metadata.

The next sections show what that means end to end on DROID. For our DROID-scale experiments we put the dataset in a remote object store, and that is the setup for the rest of this post.

The experiment setup: DROID dataset from object storage

Next we trained SmolVLA on all of lerobot/droid_1.0.1: 27,630,375 frames across 95,658 episodes, 369 GB of video, read directly from a remote object store with a global shuffle over every frame. We ran the same 10,000-step job twice on 8 H100s. Seed, sample order, batch size, workers and a cold page cache were identical; only --dataset.root changed. One run used the standard reader on a local NVMe copy after a 384 GB download. The other used LanceDB over the object store with nothing downloaded.

Held-out episode 226 shown in Foxglove, streamed from object storage. A Franka arm clears paper into a recycling bin. Each panel overlays three lines for one action dimension: the human teleoperator (black), the checkpoint after 80 steps (red), and the checkpoint after 10,000 steps (teal). The gripper panel is the clearest example. Red swings across the range while teal follows the human. This episode is in the 98th percentile of arm motion across the held-out set, so it is harder than average.

The loss curves matched to four decimal places at every step, so the 33-minute gap is time the GPUs spent waiting for data. Across 48 held-out episodes, all 48 improved. Next-action mean absolute error (MAE) fell from 0.3259 to 0.1004, and the improvement held across different levels of arm motion.

10,000 steps, SmolVLA, 8×H100 LanceDB, remote object store LeRobot, local NVMe
Wall clock 1 h 27 m 29 s 2 h 00 m 22 s
Difference — +33 min (1.38×)
Steady rate 495 samples/s 361 samples/s
Time waiting on data 1.7% 37.4%
Final loss 0.2380 0.2380

The training table is also the search index

Faster training solves the read path. Robot teams spend as much time deciding what should go into the next run, and that work usually means exporting the dataset into other systems. Parquet stores columns well but has no secondary indexes, such as no vector index, no full-text index, nothing to find frames with except a scan. Because the training data here is already a LanceDB table, the same table can be searched, scored and curated in place.

Area Parquet + MP4 One LanceDB table
Embeddings and vector search Separate vector database Vector index on the embedding column
Text and full-text search Separate search service Full-text index on the instruction column
Derived scores Separate metrics table Ordinary columns on the same rows
Joining back to frames Manifest file The row itself
What the trainer reads The files, not the derived systems The same table, at a pinned version
Systems to keep in sync 4 1

You can try every query in this section on the live DROID table in our interactive demo.

Add a feature without copying the dataset

Most curation workflows start by computing a score from the data and writing it back as a column that can be used for filtering or search. With LanceDB Enterprise, a score is declared as a column and filled in later. Nothing is rewritten and the video is never touched.

@udf(data_type=pa.float32(), input_columns=["action_joint_velocity"])
def jerk_score(v):                      # per-frame motion roughness
    ...

tbl.add_columns({"jerk_score": jerk_score})   # nothing computed yet
tbl.backfill("jerk_score", concurrency=8)

The table grew by 7% instead of being copied, and readers pinned to the old version kept working. Pixel features work the same way, where a GPU UDF reads frames straight from the video blob column, decoding only the byte ranges it needs. We embedded 610,403 frames with SigLIP2 at 712 frames/s on 2 GPUs, with no local copy and no special dataset class.

Search, filter and mine on the same rows

The table now brings together a vision embedding, a derived motion score, and the success flag that shipped with the dataset. One query can combine them.

tbl.search(vec("a gripper closing on an object"), vector_column_name="emb_siglip2") \
   .where("jerk_score > 1.2875")          # the roughest 1% of frames
   .limit(4)

The examples below combine a text query against the image embeddings with filters on other columns in the same table, so results can match both what a frame looks like and what happened in the episode.

Search for "a gripper closing on an object" with WHERE jerk_score > 1.2875 to find grasp-like frames where the motion is in the roughest 1%.

jerk

Search for "a cluttered kitchen counter" with WHERE is_episode_successful = false to find visually similar frames from failed episodes.

success

Curation filters the dataset down to the examples you want to train on. Mining helps find more of those examples, even when they were never labeled.

The same query mines behavior that has no label. Pick one frame as the seed, search for its nearest neighbors in the embedding index, and narrow with SQL on the same rows.

tbl.search(seed_frame_embedding, vector_column_name="emb_siglip2") \
   .where("episode_index != 2 AND is_episode_successful AND jerk_score > 1.0") \
   .limit(600)                       # then keep the nearest frame per episode

query

The frame with the blue border is the query. Its instruction reads "move the bottom right tip of the duvet to the left". The other five are the nearest frames from five other episodes, found on pixels alone across 610,403 embedded frames.

The results show the same behavior, moving soft fabric such as a duvet, pillow, towel, or bedding, in different homes. But the task text is different even when the action is similar.

  • 0/8 are exact text matches.
  • 5/8 show the same real-world behavior.

Text search for duvet would miss most of these examples. Visual similarity finds them even when people use different words.

What one full scan revealed

We scanned all 27,630,375 rows once. Three findings stood out:

  • OpenPI's "idle filter" removes 11.66% of DROID. We reimplemented it from the released source, which hard-codes the post-filter dataset length instead of computing it.
  • The default train/eval split leaks information. A simple 80/20 split puts all 72 buildings and all 78 collectors in both the training and validation sets.
  • The usual "smoothness" quality filter points the wrong way. Across 95,603 episodes, successful episodes were jerkier than failed ones (AUROC 0.402). Filtering out jerky episodes would remove good data.

Does curation make a better robot?

Curation only matters if it trains a better policy. DROID has no simulator for closed-loop success, so we tested this on LIBERO with 1,693 successful demonstrations across 40 tasks. We corrupted 30% of the episodes on purpose, split equally across three realistic logging problems, and left the camera frames untouched.

What we broke How we detected it Score
Swapped the instruction for another task in the same suite The final frame looked unusual next to other episodes with that label goal_dist
Added action noise, a spike on 4% of frames and gripper flips on 2% Commands changed too sharply between frames jerk_score
Shifted actions 0.6 seconds ahead of observations Commands aligned best after a shift of at least two frames, with a fit gain of at least 0.05 over zero lag act_lag

The three scores are columns on an episodes table, declared and backfilled the same way as jerk_score above. An episode that trips any check gets a non-ok quality_flag, and training reads the rest through one filter:

episodes.add_columns({"jerk_score": EpisodeJerk(), "act_lag": ActLag(), "goal_emb": GoalEmb()})
episodes.backfill(["jerk_score", "act_lag", "goal_emb"], concurrency=8)
episodes.add_columns({"goal_dist": GoalDist(), "quality_flag": QualityFlag()})
episodes.backfill(["goal_dist", "quality_flag"])

keep = episodes.search().where("quality_flag = 'ok'").select(["episode_index"])

The filter returns 1,199 episode IDs for LeRobot's --dataset.episodes. No frames or video are rewritten or copied. A few rows after detection, from the published run:

Episode jerk_score act_lag goal_dist quality_flag
0 0.100 {lag: 1, gain: 0.006} 0.017 ok
12 0.617 {lag: 1, gain: 0.010} 0.006 noise
2 0.110 {lag: 6, gain: 0.185} 0.027 misaligned
20 0.065 {lag: 1, gain: 0.002} 0.105 label_mismatch

We trained SmolVLA four times for 40,000 steps and evaluated each model on 400 simulator rollouts. The last row is a ceiling: perfect removal using our hidden list of corrupted episodes, which is not available on real data.

Training data Episodes Spatial Object Goal LIBERO-10 Overall
Clean 1,693 91 87 88 78 86.0
Messy 1,693 74 71 48 58 62.7
Detector-cleaned 1,199 79 75 85 69 77.0
Perfect removal 1,186 84 86 80 71 80.2

Corruption dropped overall success from 86.0% to 62.7%. Filtering the flagged episodes brought it back to 77.0%, close to the 80.2% ceiling. On the Goal suite it recovered from 48% to 85%. The checks ran at 96% precision and 93% recall. Most misses were wrong labels on tasks whose final frames look nearly identical.

In the video below, the model trained on messy data fails to put the bowl on top of the cabinet in all 10 attempts. The detector-cleaned model succeeds 8 out of 10 times.

Inspect episodes without moving the data

Curation is easier when you can inspect the episodes behind the scores. Because the video remains in the same LanceDB table, you can open any episode directly from object storage without downloading the dataset or converting it to another format.

lerobot-dataset-viz --repo-id ... --root <object-store-uri>/data.lance --display-mode foxglove

Episode 7 of DROID played from object storage: three camera streams plus the proprioceptive state and action traces, with no local copy of the 369 GB dataset. This capture pulled 132.7 MB over 54 seconds of playback, which is 0.036% of the dataset.

Try it yourself

pip install "lerobot[lancedb]" "lerobot-lancedb>=0.3.1"       # native reader + converter

lerobot-lance-convert --repo-id lerobot/pusht --out ./pusht-lance

hf buckets sync ./pusht-lance hf://buckets/<namespace>/<bucket>/pusht-lance

lerobot-train \
  --dataset.repo_id=<namespace>/<bucket>/pusht-lance \
  --dataset.repo_type=bucket \
  --policy.type=act

Convert any LeRobot v3.0 dataset, push it to the Hugging Face Hub, a Storage Bucket, or any object store, and point your existing training job at it. Datasets without storage_format: "lance" keep using LeRobot's default reader, so nothing else changes.

The LeRobot-LanceDB integration docs cover conversion, loading, training, and benchmarks. Ready-made datasets are under the lance-format organization on Hugging Face, the curated DROID table is live in the robot data curation demo, and the benchmark scripts, recorded batch orders and full measurement log are in the companion repo.

Community

Sign up or log in to comment