KUAVO-VLA-1.0

A domain-specialized VLA model for industrial humanoids — turning new tasks from relearning into rapid adaptation.

Leju Robotics introduces KUAVO-VLA-1.0, a domain-specialized model built for real industrial deployment. Starting from a general-purpose VLA foundation model, it retrains on high-quality same-embodiment Kuavo real-robot data, pre-embedding body adaptation, basic manipulation, and industrial common capabilities into the model.

For developers, this means a new skill no longer starts from scratch — resources can focus on scene-specific adaptation, significantly shortening the path from foundation model to skill deployment.

Domain training corpus
680hours
Industrial tasks
100
Demonstration frames
24.5M
Benchmark tasks
25
Collection
Real robotteleoperated

Abstract #

Vision-Language-Action (VLA) foundation models have recently demonstrated impressive open-domain manipulation skills, yet most public benchmarks and deployments remain centered on household and tabletop settings. Industrial manipulation imposes a substantially different set of requirements: narrow tolerances, long-horizon multi-stage assembly, heterogeneous tooling, and a hard expectation of high, repeatable success rates under fixed line-side conditions.

We present KUAVO-VLA-1.0, a domain-specialized vision-language-action foundation model (DSFM) purpose-built for dual-arm humanoid manipulation in industrial scenarios on the Kuavo humanoid platform. KUAVO-VLA-1.0 is obtained through a two-stage pipeline on top of LingBot-VLA-2.0, a large-scale pretrained VLA foundation model. The first stage, domain-specific retraining, trains on a self-collected corpus of 100 industrial manipulation tasks totaling 680 hours of dual-arm teleoperated demonstrations, instilling broad industrial-domain competence and cross-task transfer — this is what turns a general-purpose backbone into a DSFM. The second stage, task-specific post-training, specializes the DSFM on individual tasks to push per-task reliability toward production-grade success rates.

We evaluate KUAVO-VLA-1.0 on a 25-task industrial manipulation benchmark — 20 industrial tasks plus 5 unseen in-domain tasks (GM100) — against a panel of state-of-the-art VLA baselines (π0.5, GR00T N1.7, and LingBot-VLA-2.0 itself, which as our own backbone isolates domain specialization with architecture and pretraining held fixed). To support reproducible research on DSFMs for real industrial deployment, we release the technical report, this project webpage, training code built on our open-source LeTools framework, model checkpoints, and the underlying industrial manipulation dataset, released on the OpenLET platform.

Method #

We do not propose a new backbone. We treat industrial deployment as a domain-adaptation problem and design a two-stage pipeline on top of a general-purpose VLA foundation model.

KUAVO-VLA-1.0 framework: configuration space on Kuavo4pro across Kuavo4pro/5/5W, the data collection and processing pipeline, model architecture with vision-language backbone and Mixture-of-Experts action head, and the two-stage training pipeline.
Configuration space on Kuavo4pro (a), the data collection and processing pipeline (b), model architecture (c), and the two-stage training pipeline (d). End-effector and active-arm selection are deployment-time configuration, not hard-coded — which_arm determines whether the action head predicts an 8-dimensional single-arm or 16-dimensional dual-arm chunk.
BASE
Pretrained backbone: LingBot-VLA-2.0

A large-scale pretrained vision-language-action foundation model: a vision-language understanding backbone pretrained on internet-scale image-text and video data, coupled with a Mixture-of-Experts action head that refines noised action tokens into a denoised action chunk via flow matching. The policy consumes one head camera, two wrist cameras, and proprioceptive arm state. LingBot-VLA-2.0 is also evaluated as a baseline in Results, isolating what domain-specific retraining adds on top of an identical backbone.

STAGE 1
Domain-specific retraining

Retraining on 100 distinct industrial tasks (680 hours) turns the general-purpose backbone into a domain-specialized foundation model (DSFM): it shifts the visual and behavioral distribution toward line-side fixtures, connectors, packaging, and tooling — all underrepresented in general-purpose corpora. Spanning pick-and-place, sorting, assembly and insertion, packaging, and tool-based operations, this stage is designed to encourage positive skill transfer across tasks.

STAGE 2
Task-specific post-training

A lightweight per-task stage starting from the DSFM checkpoint produced by Stage 1. Because the model already holds broad industrial competence, this typically needs far fewer demonstrations and steps than specializing from the general backbone directly, while reaching higher final success rates.

Dataset #

The corpus comprises 100 industrial manipulation tasks, 58,780 demonstration episodes and 24.5M synchronized frames — 680 hours at the native 10 Hz recording rate. All training data is collected by teleoperation on physical Kuavo dual-arm humanoids across real industrial scenarios. Raw ROS bags are screened by operator review, converted into standardized episodic trajectories with synchronized multi-view RGB, proprioceptive state, and language annotations, then filtered again automatically before entering the corpus. The resulting dataset is released on the OpenLET platform.

Benchmark #

Twenty-five industrial manipulation tasks on the Kuavo dual-arm humanoid, spanning the same skill categories as the domain-specialization corpus. Twenty of them also appear in the 100-task Stage 1 corpus, so those measure task-specific specialization on already-seen tasks. The remaining five are unseen in-domain tasks from the GM100 collection, exercising the same Kuavo4pro dual-arm manipulation skills but held out of the Stage 1 corpus. Each task is evaluated over multiple real-robot rollouts with fixed initial-condition randomization.

Full 25-task list filter by category
# Category Task prompt Stage 1

Each benchmark task is initialized from the Stage 1 checkpoint and task-specifically post-trained for 20 epochs.

Demo footage

Real-robot rollouts of KUAVO-VLA-1.0 on six of the benchmark's key scenarios. SR — success rate, the share of rollouts that complete the task end-to-end; PS — progress score, the mean fraction of a task's steps completed.

PCB board
SR 60.0% · PS 83.3%
SMT tray handling
SR 73.3% · PS 95.6%
FMCG unloading
SR 73.3% · PS 82.2%
Line-side switch sorting
SR 80.0% · PS 95.0%
Canned-goods loading
SR 46.7% · PS 66.7%
FMCG loading
SR 60.0% · PS 85.6%

Unseen in-domain tasks (GM100)

Five tasks from the public GM100 collection, held out of the Stage 1 corpus and adapted only in Stage 2.

BM-107 · QR code scanning
SR 80.0% · PS 88.0%
BM-96 · Squeaky chicken
SR 0.0% · PS 65.6%
BM-47 · Egg packing
SR 60.0% · PS 82.2%
BM-45 · Snack sorting
SR 73.3% · PS 93.3%
BM-39 · Paper-towel roll
SR 0.0% · PS 51.7%

Results #

Full 25-task benchmark — 20 industrial tasks (also in the 100-task Stage 1 corpus) plus 5 unseen in-domain tasks — each task evaluated over multiple real-robot rollouts with fixed initial-condition randomization. KUAVO-VLA-1.0 (task-specific) is the DSFM after Stage 2 per-task post-training. GR00T N1.7 and π0.5 are evaluated as external baselines under the same protocol; LingBot-VLA-2.0 is evaluated too — it is also the backbone KUAVO-VLA-1.0 is built from, so this comparison isolates what domain-specific retraining adds with architecture and pretraining held fixed. See Seen vs. unseen tasks below for the seen/unseen split.

Success rate & progress score

SR — share of rollouts that complete the task end-to-end; PS — mean fraction of each task's steps completed, credited even on rollouts that fall short. Mean over all 25 tasks; all four systems fully evaluated (25/25).

Success rate
0% 25% 50% 75% 100% 6.7% GR00T N1.7 12.8% π0.5 16.3% LingBot-VLA-2.0 48.3% KUAVO-VLA-1.0
Progress score
0% 25% 50% 75% 100% 21.2% GR00T N1.7 32.1% π0.5 42.1% LingBot-VLA-2.0 74.5% KUAVO-VLA-1.0

All four systems above are evaluated on the complete 25-task benchmark under a matched real-robot protocol.

Seen vs. unseen tasks #

The 25-task benchmark pairs 20 industrial tasks, all present in the 100-task Stage 1 corpus, with 5 unseen in-domain tasks from the GM100 collection, which exercise the same Kuavo4pro dual-arm manipulation skills but were held out of Stage 1. On the unseen group, after the same 20-epoch task-specific post-training, KUAVO-VLA-1.0 still leads at 42.7% SR against 25.3% for LingBot-VLA-2.0, its own backbone — the 1.7× gap indicates that domain-specific retraining instills transferable manipulation skill rather than rote memory of the training tasks. Two caveats bound this: these are public tasks with published demonstrations, adapted under the same 20-epoch budget, not zero-shot evaluations; and they differ from the seen tasks in scene rather than in embodiment or skill, so this is evidence of skill transfer to unseen in-domain tasks, not of zero-shot generalization.

Model SR PS
Seen 20 Unseen 5 Seen 20 Unseen 5
GR00T N1.7 8.3%0.0% 22.5%16.2%
π0.5 12.3%14.7% 32.9%28.8%
LingBot-VLA-2.0 14.0%25.3% 40.1%49.7%
KUAVO-VLA-1.0 (task-specific) 49.7%42.7% 74.1%76.2%
Per-task breakdown 25 tasks × SR / PS

Success rate (SR) and progress score (PS) by task, shown as two separate rows per system — on many of the harder tasks SR alone is 0% for most systems, but PS shows how far they actually got. Task codes are our internal task IDs; titles are our own short-form translations and are independent of the task-prompt wording used in the Benchmark table above. The 5 cards tagged Unseen are unseen in-domain tasks from the GM100 collection — see Seen vs. unseen tasks above for the aggregate comparison.

GR00T N1.7 / π0.5 / LingBot-VLA-2.0 KUAVO-VLA-1.0 Unseen = unseen in-domain task
056-044
Desktop parts pick-and-place
GR00T N1.7
SR
0%
PS
11%
π0.5
SR
0%
PS
20%
LingBot-VLA-2.0
SR
0%
PS
6%
KUAVO-VLA-1.0
SR
0%
PS
31%
052-036
Shock-absorber unloading
GR00T N1.7
SR
0%
PS
27%
π0.5
SR
20%
PS
53%
LingBot-VLA-2.0
SR
0%
PS
31%
KUAVO-VLA-1.0
SR
100%
PS
100%
030-077
Part-size sorting
GR00T N1.7
SR
0%
PS
16%
π0.5
SR
0%
PS
3%
LingBot-VLA-2.0
SR
0%
PS
9%
KUAVO-VLA-1.0
SR
0%
PS
19%
022-035
Automated mechanical-parts kitting
GR00T N1.7
SR
0%
PS
9%
π0.5
SR
20%
PS
57%
LingBot-VLA-2.0
SR
0%
PS
41%
KUAVO-VLA-1.0
SR
60%
PS
84%
036-047
Canned-goods boxing
GR00T N1.7
SR
0%
PS
2%
π0.5
SR
0%
PS
22%
LingBot-VLA-2.0
SR
0%
PS
28%
KUAVO-VLA-1.0
SR
47%
PS
67%
013-052
Bucket-liquid cap sealing
GR00T N1.7
SR
0%
PS
0%
π0.5
SR
0%
PS
8%
LingBot-VLA-2.0
SR
0%
PS
38%
KUAVO-VLA-1.0
SR
13%
PS
55%
005-006
FMCG two-hand coordinated loading
GR00T N1.7
SR
20%
PS
60%
π0.5
SR
80%
PS
94%
LingBot-VLA-2.0
SR
33%
PS
80%
KUAVO-VLA-1.0
SR
60%
PS
86%
002-001
FMCG unloading
GR00T N1.7
SR
33%
PS
76%
π0.5
SR
20%
PS
61%
LingBot-VLA-2.0
SR
13%
PS
66%
KUAVO-VLA-1.0
SR
73%
PS
82%
004-007
Small-part flipping
GR00T N1.7
SR
0%
PS
2%
π0.5
SR
0%
PS
33%
LingBot-VLA-2.0
SR
0%
PS
29%
KUAVO-VLA-1.0
SR
47%
PS
76%
038-047
PCB inspection
GR00T N1.7
SR
0%
PS
13%
π0.5
SR
0%
PS
14%
LingBot-VLA-2.0
SR
0%
PS
39%
KUAVO-VLA-1.0
SR
60%
PS
83%
041-022
Test-tube loading
GR00T N1.7
SR
60%
PS
78%
π0.5
SR
33%
PS
67%
LingBot-VLA-2.0
SR
60%
PS
80%
KUAVO-VLA-1.0
SR
60%
PS
86%
029-065
Seal-flap flattening
GR00T N1.7
SR
0%
PS
2%
π0.5
SR
53%
PS
62%
LingBot-VLA-2.0
SR
0%
PS
2%
KUAVO-VLA-1.0
SR
87%
PS
88%
024-048
Push-to-fixed-position
GR00T N1.7
SR
40%
PS
73%
π0.5
SR
13%
PS
62%
LingBot-VLA-2.0
SR
67%
PS
90%
KUAVO-VLA-1.0
SR
87%
PS
97%
019-002
Schaeffler small-parts loading
GR00T N1.7
SR
13%
PS
29%
π0.5
SR
7%
PS
21%
LingBot-VLA-2.0
SR
27%
PS
52%
KUAVO-VLA-1.0
SR
60%
PS
88%
090-015
Switch-part sorting
GR00T N1.7
SR
0%
PS
9%
π0.5
SR
0%
PS
23%
LingBot-VLA-2.0
SR
0%
PS
28%
KUAVO-VLA-1.0
SR
60%
PS
88%
025-055
Assembly
GR00T N1.7
SR
0%
PS
4%
π0.5
SR
0%
PS
4%
LingBot-VLA-2.0
SR
0%
PS
5%
KUAVO-VLA-1.0
SR
0%
PS
51%
045-082
Steel-bar extraction
GR00T N1.7
SR
0%
PS
0%
π0.5
SR
0%
PS
2%
LingBot-VLA-2.0
SR
7%
PS
31%
KUAVO-VLA-1.0
SR
0%
PS
56%
104-005
SMT tray handling
GR00T N1.7
SR
0%
PS
1%
π0.5
SR
0%
PS
0%
LingBot-VLA-2.0
SR
47%
PS
78%
KUAVO-VLA-1.0
SR
73%
PS
96%
057-025
Cosmetic-box loading
GR00T N1.7
SR
0%
PS
6%
π0.5
SR
0%
PS
17%
LingBot-VLA-2.0
SR
0%
PS
8%
KUAVO-VLA-1.0
SR
27%
PS
56%
073-042
Line-side switch sorting
GR00T N1.7
SR
0%
PS
32%
π0.5
SR
0%
PS
35%
LingBot-VLA-2.0
SR
27%
PS
63%
KUAVO-VLA-1.0
SR
80%
PS
95%
BM-107 Unseen
Medicine QR-code scanning
GR00T N1.7
SR
0%
PS
8%
π0.5
SR
0%
PS
4%
LingBot-VLA-2.0
SR
53%
PS
65%
KUAVO-VLA-1.0
SR
80%
PS
88%
BM-96 Unseen
Squawking-toy squeeze
GR00T N1.7
SR
0%
PS
3%
π0.5
SR
13%
PS
49%
LingBot-VLA-2.0
SR
7%
PS
54%
KUAVO-VLA-1.0
SR
0%
PS
66%
BM-47 Unseen
Egg carton filling
GR00T N1.7
SR
0%
PS
4%
π0.5
SR
0%
PS
7%
LingBot-VLA-2.0
SR
67%
PS
81%
KUAVO-VLA-1.0
SR
60%
PS
82%
BM-45 Unseen
Snack sorting
GR00T N1.7
SR
0%
PS
40%
π0.5
SR
60%
PS
84%
LingBot-VLA-2.0
SR
0%
PS
34%
KUAVO-VLA-1.0
SR
73%
PS
93%
BM-39 Unseen
Toilet-paper roll replacement
GR00T N1.7
SR
0%
PS
25%
π0.5
SR
0%
PS
0%
LingBot-VLA-2.0
SR
0%
PS
13%
KUAVO-VLA-1.0
SR
0%
PS
52%

BibTeX #

@techreport{kuavovla2026,
  title  = {KUAVO-VLA-1.0: A Domain-Specialized Vision-Language-Action
            Foundation Model for Industrial Dual-Arm Humanoid Manipulation},
  author = {Leju Robotics},
  year   = {2026},
  institution = {Leju Robotics}
}