Skip to content
RobotWorld
Back to Papers

PAPER DEEP DIVE

数据采集双臂开源

MEVION: Low-Cost Open-Source Data Collection System for Powerful and High-Speed Dual-Arm Manipulation

The global competition for developing robotic foundation models is intensifying. Among the data collection systems used for dual-arm robots, ALOHA is representative of being low-cost and open-source, and is widely adopted by researchers as a de facto standard. However, due to its limited ability to generate high forces and speeds, it is difficult to handle heavy objects or perform fast manipulations. To address this, we developed MEVION, a low-cost and open-source dual-arm robot data collection system capable of generating greater force and speed. All parts of this robot can be sourced through e-commerce, and by extensively utilizing sheet metal welding, its large body structure is constructed with a small number of components at low cost, while also simplifying assembly. MEVION is equipped with four 6-DoF arms with parallel grippers. Each arm weighs 7.0 kg and has a maximum torque of 60 Nm, and the entire system can be constructed for about USD 14,000. The elbow joint adopts a closed-link mechanism similar to those used in quadruped robots, which reduces the distal mass and enables higher force and speed output at the end-effector. We demonstrate that MEVION enables data collection for object manipulation tasks not previously possible and supports imitation learning-based motion generation. All hardware and software of this work are included in the Supplementary Materials or https://github.com/haraduka/mevion.

Kento Kawaharazuka, Yoshiki Obinata, Hirokazu Ishida, Jihoon Oh, Temma Suzuki, Shintaro Inoue, Keita Yoneda, Ayumu Iwata, Kei OkadaJuly 20, 20267 min read
中文

Paper: MEVION: Low-Cost Open-Source Data Collection System for Powerful and High-Speed Dual-Arm Manipulation
Authors: Kento Kawaharazuka, Yoshiki Obinata, Hirokazu Ishida, Jihoon Oh, Temma Suzuki, Shintaro Inoue, Keita Yoneda, Ayumu Iwata, Kei Okada
Affiliation: Department of Mechano-Informatics, The University of Tokyo
Link: arXiv:2607.17970v1 [cs.RO], July 2026
Code: ✅ Open-source at github.com/haraduka/mevion (hardware + software + STEP files)

1. Abstract

The global competition for developing robotic foundation models is intensifying. Among dual-arm robot data collection systems, ALOHA is representative for being low-cost and open-source, widely adopted as a de facto standard. However, its limited force and speed make it difficult to handle heavy objects or perform fast manipulations. This paper develops MEVION — a low-cost, open-source dual-arm robot data collection system capable of generating greater force and speed. All parts can be sourced through e-commerce, with extensive sheet metal welding enabling a large body structure from few components at low cost while simplifying assembly. MEVION is equipped with four 6-DOF arms with parallel grippers; each arm weighs 7.0 kg with 60 Nm maximum torque, and the entire system costs approximately USD 14,000. The elbow joint adopts a closed-link mechanism similar to quadruped robots, reducing distal mass and enabling higher force and speed output at the end-effector. Experiments demonstrate MEVION enables data collection for previously impossible manipulation tasks and supports imitation learning-based motion generation.

2. Background and Motivation

Imitation learning and its extension to VLA models have spurred active research on dual-arm manipulation. Existing data collection systems fall into teleoperation and autonomous types. ALOHA is the representative low-cost open-source teleoperation system, but its limited force and speed mean related research focuses on slow manipulation of lightweight objects, making diverse data collection challenging.

Key limitations of existing systems:

  • SO-101: All-plastic structure, max torque 1.9 Nm, speed 5.5 rad/s, $120/arm — extremely low force and speed
  • ALOHA: Metal arms only, max torque 21.2 Nm, speed 3.1 rad/s, $4,900/arm — cannot handle heavy or fast manipulation
  • OpenArm: Standing-type all-metal, max torque 40 Nm, speed 16.8 rad/s, $3,200/arm — higher performance but 31 metal components, heavier end-effector

Figure 1: MEVION — low-cost open-source data collection system for powerful and high-speed dual-arm manipulation.

3. Core Method

3.1 Design Overview

The design is divided into arm and hand components, with 17 essential metal components (11 arm + 6 hand), or 21 including joint limiters and covers. All machined parts use A7075 aluminum alloy; sheet metal uses SUS304 stainless steel (Shoulder-Link, Hand-Slider-Cover) or A5052 aluminum alloy. Some limiters and cable covers use flexible TPU 3D printing; fingertips use PLA 3D printing.

Figure 2: Design overview. 17 essential metal components (11 arm + 6 hand), 3 of which are integrated via sheet metal welding into single complex large parts.

3.2 Joint Configuration and Closed-Link Mechanism

The arm's joint configuration matches ALOHA: 6 DOF (Shoulder-Yaw, Shoulder-Pitch, Elbow-Pitch, Elbow-Yaw, Wrist-Pitch, Wrist-Yaw). Key innovation: the elbow joint uses a parallel-link mechanism similar to quadruped robots, driving the joint from a proximal motor to reduce distal mass.

The joint torque relates to motor torque as:

$$\tau_{joint}=N \cdot \tau_{motor} - \tau_{friction}$$

The hand uses a slider-crank parallel gripper mechanism driven by a single motor. Motor configuration:

JointMotorMax Torque
Shoulder-Yaw, Shoulder-Pitch, Elbow-PitchRobStride0360 Nm
Elbow-Yaw, Wrist-Pitch, Wrist-YawRobStride0217 Nm
Hand (gripper)RobStride0117 Nm

Figure 3: Sheet metal welding details for Shoulder-Link, Lower2-Link, and Hand-Slider-Base. Welded sections highlighted in red.

3.3 Sheet Metal Welding: Key to Low Cost and Few Parts

The Shoulder-Link and Lower2-Link (arm) and Hand-Slider-Base (hand) integrate large complex shapes into single parts via sheet metal welding. This maintains structural strength while allowing large-shape integration, achieving both low cost and easy assembly with few parts. All metal components are fully compatible with MISUMI's meviy online machining service, which auto-generates quotes and orders from 3D CAD files.

3.4 Comparison with Existing Systems

NameMorphologyDOFsWeightLengthMaterialsMax TorqueMax SpeedPrice/arm
SO-101Tabletop4+10.7kg0.45mPlastic1.9Nm5.5rad/s$120
ALOHATabletop6+14.3kg0.75mMetal arms21.2Nm3.1rad/s$4,900
OpenArmStanding7+15.5kg0.63mAll metal40Nm16.8rad/s$3,200
MEVIONTabletop6+17.0kg0.83mAll metal60Nm20.4rad/s$3,500

Table 1: Comparison of open-source leader-follower bimanual data collection systems with MEVION.

Per-arm total cost is the sum of metal, motors, and sensors:

$$C_{arm}=C_{metal}+C_{motors}+C_{sensors}=1465+1645+422=3532 \text{ USD}$$

vs ALOHA: MEVION is ~1.6× heavier but achieves ~3× max torque and ~7× max speed. Peak payload at max extension: MEVION 7.6kg vs OpenArm 4.5kg. Metal components: MEVION 17 vs OpenArm 31 — sheet metal welding dramatically reduces part count.

3.5 Software: Unified Control and Gravity Compensation

MEVION is controlled by a unified Python script, with MuJoCo simulation and real-world control sharing the same code. Supports ROS/RViz real-time visualization and Scikit-Robot interactive visualization. Uses Intel RealSense D405 cameras.

Motor power relates to torque as:

$$P=\tau \cdot \omega=\tau_{max} \cdot \omega_{max}$$

Power: 24V or 48V; communication: CAN. Four CAN-USB interfaces connect to individual arms then to PC. Control scheme: PD controller with gravity compensation torque:

$$\tau^{cmd}_{i}=K_{p,i}(q_{i}^{ref}-q_{i})+K_{d,i}(\dot{q}_{i}^{ref}-\dot{q}_{i})+\tau^{ff}_{i}$$

where feedforward term $\tau^{ff}$ is computed using Pinocchio. Limits on $q_{i}^{ref}$, $\dot{q}_{i}^{ref}$, $\tau^{ff}_{i}$ for joint angle, velocity, and torque. Control loop at 200 Hz.

Unlike ALOHA, MEVION uses software-side gravity compensation:

$$\tau^{ff} = \tau_{\text{gravity}}(q) + \tau_{\text{Coriolis}}(q, \dot{q})$$

eliminating additional frames or mechanisms, with safety via hardware (wireless e-stop) and software (torque/velocity limits).

Figure 4: Software architecture. Unified Python script controls both MuJoCo simulation and real robot. Supports RViz (ROS) and Scikit-Robot visualization.

flowchart TD
    A[Python Control Script] --> B{MuJoCo Sim or Real Control}
    B -->|Simulation| C[MuJoCo XML Model]
    B -->|Real| D[CAN-USB x4]
    D --> E[4x 6-DOF Arms]
    E --> F[RobStride03/02/01 Motors]
    F --> G[PD Controller + Gravity Comp]
    G --> H[τ_cmd = Kp(q_ref-q) + Kd(q_dot_ref-q_dot) + τ_ff]
    H --> I[200Hz Control Loop]
    A --> J[RViz/ROS Visualization]
    A --> K[Scikit-Robot Interactive]
    A --> L[RealSense D405 Cameras]

4. Key Experiments

4.1 Teleoperation Experiments

Five tasks validating diverse manipulation:

  • (a) Bottle cap opening: bimanual cooperation
  • (b) Object packing: transporting slender heavy objects (wrenches, 1.5kg bottles)
  • (c) Frying pan operation: high-speed manipulation
  • (d) 3.6kg dumbbell manipulation: heavy object handling
  • (e) Daruma Otoshi: traditional Japanese game requiring fast striking

Figure 5: Joint torque transitions during dumbbell manipulation. Confirms MEVION generates sufficient force for 3.6kg objects.

4.2 Imitation Learning Experiments

Trained imitation learning on teleoperated data, validating autonomous motion generation:

  • (a) Towel manipulation
  • (b) Dumbbell packing

Confirms MEVION supports the complete pipeline from data collection to autonomous motion generation.

5. Limitations and Future Work

  • Sheet metal welding limitations: Unlike metal machining, difficult to independently modify or repair. Re-ordering is often simplest due to low cost, but trade-offs with machining should be considered. Welding quality varies by vendor, relevant for large-scale data collection.
  • Large-scale data collection: The ultimate goal. Software gravity compensation reduces physical footprint, allowing more robots to operate simultaneously. Plans to build multiple MEVION systems with collaborators. Scalability introduces mass-production challenges requiring design simplification.
  • Safety concerns: No hardware gravity compensation mechanism increases risk, mitigated by hardware e-stop and software limits.

6. Conclusion

MEVION achieves far superior force (60Nm vs 21.2Nm) and speed (20.4rad/s vs 3.1rad/s) output over ALOHA at approximately USD 14,000 total cost, through sheet metal welding and closed-link elbow design. 17 metal components (vs OpenArm 31) demonstrate sheet metal welding's advantage in integrating large complex shapes. The proximal-motor parallel-link elbow reduces distal mass, achieving 7.6kg peak payload (vs OpenArm 4.5kg). Software-side gravity compensation (Pinocchio) eliminates additional hardware. Teleoperation experiments (3.6kg dumbbell, Daruma Otoshi) and imitation learning (towel, dumbbell packing) validate the complete pipeline from data collection to autonomous motion generation. All hardware (STEP files, parts list) and software (MuJoCo/ROS/Scikit-Robot) are open-source, with all metal parts orderable via MISUMI meviy e-commerce.

Sheet metal welding lets 17 parts replace 31; a closed-link mechanism lets a 7kg arm produce 60Nm — open-source doesn't mean low-performance. Engineering ingenuity turns $14,000 into four powerful, fast robotic arms.

Related Papers

HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone

HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone

Learning deployable manipulation policies is bottlenecked by the scarcity of data that is both high-fidelity and scalable. Real-robot teleoperation is accurate but costly to scale; robot-free UMI capture scales readily, and current practice uses the resulting data mainly for pre-training, adding a small real-robot "anchor" at post-training. We ask whether raising the fidelity of robot-free UMI data, rather than shrinking the real-robot fraction, can remove that anchor. We present HiFi-UMI, a portable UMI data-production system co-designed for trajectory accuracy, inter-gripper relative pose, synchronization, and field of view: head-mounted offline stereo-inertial SLAM, native rather than reconstructed relative pose, a shared microsecond GPIO trigger, and two wide-angle cameras per hand covering ~200 degrees. It reaches 3 mm workspace-local end-effector accuracy without external tracking infrastructure. Using this corpus, we demonstrate zero-robot post-training: a policy post-trained solely on HiFi-UMI demonstrations deploys directly on a real robot and matches in-domain teleoperation across three backbones spanning the vision-language-action and world-action-model families, with success-rate differences of -2.5, +3.1, and -0.6 percentage points on StarVLA-QwenPI, OpenPI-pi_0.5, and LingBot-VA; the strongest policy reaches 85% on a precision insertion task, even though the teleoperation baseline is collected in the evaluation scene and no HiFi-UMI trajectory is. Pre-training on 4,000 hours from the same corpus lowers action error on ten unseen tasks by 41% and, on StarVLA-QwenPI, raises real-robot success by a further 18.1 percentage points. We open-source HiFi-UMI-2K, 2,000 hours of microsecond-synchronized, ultra-wide-FoV demonstrations, each automatically reconstructed and validated through simulation replay, as a large-scale, high-fidelity resource for the robot-learning community.

操作策略UMI数据采集Jul 28, 2026
Scale Up Strategically: Learning Compositional Generalization via Bias-Aware Evaluation and Data Collection for Robotic Manipulation

Scale Up Strategically: Learning Compositional Generalization via Bias-Aware Evaluation and Data Collection for Robotic Manipulation

Compositional generalization is essential for robot to follow diverse instructions. However, pretrained policies are known to take shortcuts, deferring to salient cues rather than grounding language. We introduce a diagnostic framework that localizes this failure to individual \textit{instruction factors}, \textit{e.g.,} reusable semantic components such as color, verb, object, size, and spatial attribute. Our framework formalizes instruction factor bias, the tendency of fine-tuned policies to over-rely on dominant factors as shortcuts, and quantifies it through two metrics: Factor Dominance Rate (FDR), capturing pairwise bias between factors, and Factor Dominance Hierarchy (FDH), aggregating these into a global ranking. Evaluation on six foundation policies reveals broadly consistent ordering, \textit{i.e.}, color $\geq$ object $\geq$ spatial $\geq$ verb $\geq$ size, with color dominant, and verb and size most under-grounded. We further show the diagnosis is actionable: a bias-aware data collection strategy that reallocates a fixed budget toward under-grounded factors outperforms baselines in simulation and on a real robot using half the demonstrations, thereby enabling more sample-efficient and generalizable policy learning.

机器人操作组合泛化数据采集Jul 23, 2026