FAQ
Common questions
A LeRobot dataset is episodic robot learning data built for the Hugging Face LeRobot library. Each episode holds camera frames, actions, and the metadata you need for imitation learning. Teams use it so loading and training stay consistent across tasks. We collect the demos and package them so they drop into LeRobot workflows without a custom converter.
Robot learning data is recorded experience used to train robots: camera frames, actions or teleop signals, depth when you need it, and labels for the task. It powers imitation learning, demo-based RL, and vision-language-action models. We collect, annotate, and check this data for teams building embodied AI.
Most teams use egocentric human recording, teleoperation, kinesthetic teaching, or multi-camera task scripts, then annotate and QA the results. Plenty of groups outsource because sync, labeling, and validation eat engineering time. We run that work end to end and deliver in LeRobot, RLDS, or HDF5.
LeRobot expects episodic trajectories with observations, actions, and metadata its loaders can read. The on-disk layout follows current LeRobot conventions and whatever backend you use. We package to your LeRobot version, and we can ship RLDS or HDF5 alongside when your stack needs them.
It depends on task complexity, how many episodes you need, sensors (RGB vs RGB-D, multi-camera), how deep the annotation goes, exclusivity, and format. Pilots are usually a fixed package. Larger programs run on milestones or retainers. We quote after a short discovery call, so request pricing when you want a written estimate.
A focused pilot can often be scoped and collected in a few weeks. Bigger multi-task or multi-site jobs take longer, especially with depth calibration and richer labels. After discovery we share a schedule for capture, annotation, validation, and delivery.
Yes. You can set tasks, objects, environments, cameras, depth, action schemas, language labels, splits, formats, and licensing. We write a dataset specification before capture starts. Exclusive rights are available. Book a discovery call or ask for a custom quote to lock scope.
Yes. We label action segments, objects and scenes, optional language instructions, and quality flags. Taxonomies are versioned. Human review plus automated checks ship with QA metrics. Bring your ontology, or we propose one that fits imitation learning and VLA training.
Yes. We record synchronized RGB and depth with calibration and timestamp alignment. That can be egocentric, wrist, or scene-camera setups. Tell us your preferred sensors, resolution, and frame rate during scoping so capture matches what your models expect.
Industrial automation, warehouse logistics, household and service robots, humanoids, agriculture, and university labs. Buyers are usually startups, embodied AI companies, foundation model teams, OEMs, and research groups training on real demonstrations rather than simulation alone.
Yes. We sign NDAs before confidential task details or site access. Private deliveries and restricted raw footage access are standard for commercial work. Send your NDA on the contact form, or ask for ours. Security handling is summarized on our Security page and fixed in the statement of work.
Yes. Exclusive or privately commissioned datasets are not resold or released publicly. Terms cover usage rights, retention, and whether related task variants can be collected for others. Pricing reflects exclusivity and volume. Call out exclusivity when you request a custom quote.