ModPack

An Extensible Teleoperation Interface for Bimanual Mobile Manipulation

Joshua Citron, Renee Zbizika, Zeyi Liu, Shuran Song

Stanford University

ModPack teaser
Existing teleoperation systems are often tailored to specific robot hardware and task domains, limiting their scalability and adaptability. We present ModPack, a modular and extensible teleoperation system designed to support diverse robot embodiments and task requirements within a unified framework. At the core of ModPack is a self-contained wearable "backpack" that integrates onboard computation, power, communication, and data storage. Built on top of this shared interface, the system supports plug-and-play capability modules including joint-level teleoperation with haptic feedback, mobile manipulation, and active perception. Experiments across two distinct robot platforms and real-world mobile manipulation tasks demonstrate that ModPack provides a flexible and reusable framework for data collection and policy learning.

Inside the Backpack

The ModPack core is a wearable backpack base unit that moves with the operator while providing a configurable mounting and housing structure for modular hardware components. The main backpack assembly consists of 3D-printed front and rear panels secured with toggle latches, allowing convenient access to the internal components. The internal volume is organized into five removable shelves that house the onboard compute, power, and optional accessories. Toggle the exploded view to separate the assembly, and drag the model to orbit and inspect the components.
drag to orbit ยท ctrl+drag to pan ยท scroll to zoom
3D-Printed Hardware

The backpack is fabricated from 3D-printed parts, improving accessibility and enabling rapid modification by the research community.

The internal volume is organized into five removable shelves that house the onboard compute, power, and optional accessories. Following ergonomic observations from prior backpack-based systems, which suggest that vertically distributed loads near the thoracic region can reduce carrying strain, the PC and power banks are arranged along the operator's spine. To further reduce weight and improve thermal ventilation, excess material is removed from both the structural panels and internal shelves.

Print Settings
Material
PLA
Infill
15% ยท Gyroid
Layer height
0.2 mm
Walls
2 ยท 0.4 mm nozzle
Supports
Tree
Printers
Bambu H2D ยท X1 Carbon

The choice for the backpack to be fabricated in two parts is such that it can be easily 3D printed.

Hardware & software will be open-sourced.


Leader Arms

We design and 3D-print two leader-arm variants: a pair of 6-DoF ARX5 leader arms for a customized mobile manipulation platform, and a pair of 7-DoF leader arms for the RB-Y1m robot. Given the CAD model of a target robot, each leader arm is constructed to be kinematically equivalent to its corresponding follower arm, following GELLO, enabling direct joint-space teleoperation. The leader arms are actuated with off-the-shelf Dynamixel servo motors. Drag to orbit ยท scroll to zoom ยท toggle between variants below.
drag to orbit ยท ctrl+drag to pan ยท scroll to zoom
ARX5 Leader Arm

A 6-DoF leader arm kinematically equivalent to the ARX5 follower arm, enabling direct joint-space teleoperation. Motors were selected to provide the minimum torque necessary to hold the arm, with gravity compensation to alleviate operator fatigue during extended teleoperation.

Specifications
DoF
6
Actuators
Dynamixel XM540 ยท XM430
Structure
3D-printed PLA

Hardware & software will be open-sourced.


Teleoperation

Cloth Placement
Take Elevator
Box Transfer (Top Shelf)
Box Transfer (Bottom Shelf)
Season Food
Recycle

For the box transfer with haptic feedback task, we show both the robot and operator as well as the POV of the robot/operator and haptic feedback to the leader arms.


Capability Experiments

(a) Cloth Placement with Active Perception
A cloth is placed in front of the robot on a counter, where a basket is either to the left or to the right of the robot. The robot must move forward to grasp the cloth and then reverse, at which point it should locate the basket on either side and traverse towards it, finishing the task once it places the cloth in the basket. The main challenge is that the robot must use its active perception module to search for the basket, which may not be in view at the start of the task.

POLICY ROLLOUTS

All cameras โ€” success
Wrist camera โ€” success
Head camera only โ€” success
Head camera only โ€” success

FAILURE CASES

Wrist cam only: loses sight of cloth during navigation
Head cam only: backs up without cloth
Wrist cam only: incorrect placement โ€” cloth dropped on basket edge
All cameras: bad grasp of cloth
Policy Successes Success Rate
Head Camera Only 22 / 25 88%
All Cameras 20 / 25 80%
Wrist Camera Only 3 / 25 12%
Key Findings

(b) Box Transfer with Haptic Feedback
The robot is required to pick up a box from a cabinet, move to a rack, and place the box onto the unoccupied shelf (either top or bottom). This task requires whole-system coordination, bimanual coordination, active perception to keep the box in view, and navigation to the placement location. The haptic feedback module prevents excessive force and facilitates bimanual coordination during grasping and transport.

ALL CAMERAS + TORQUE

Top shelf
Bottom shelf

ALL CAMERAS

Top shelf
Bottom shelf

HEAD CAMERA ONLY

Top shelf
Bottom shelf

FAILURE CASES

All cameras: base overshoot โ€” robot base pushes cabinet forward
Head cam only: near collision โ€” arm nearly collides during placement
All cameras: bad grasp of box
All cameras + torque: drops box during transport
Policy Successes Success Rate
All Cameras + Torque 12 / 20 60%
Head Camera Only 11 / 20 55%
All Cameras 6 / 20 30%
Key Findings

More Results

Attention Pattern Analysis
Visual attention visualization

We provide a qualitative analysis of the model's attention patterns. As shown in the above figure, we visualize the attention weights over image patches from the final layer of the ViT encoder for both the [All Cam] and [All Cam + Torque] policies. The torque-conditioned policy appears to attend more strongly to task-relevant regions, such as the edges of the box that are informative for grasping, whereas the vision-only policy attends to less relevant regions. We also visualize the aggregated attention weights assigned to the torque tokens in the DiT model for the [All Cam + Torque] policy. The attention weights for both the left- and right-arm torque tokens increase as the robot approaches the box and initiates grasping, suggesting that the policy increasingly relies on torque information during contact-critical phases of the task.