Work / Manipulation, reinforcement learning Simulation

Dual-arm manipulation research

A bimanual platform for reinforcement learning on contact-rich tasks

Two simulated arms stacking coloured blocks, with per-arm camera views.Simulation

This platform uses two AgileX PiPER 6-DOF arms with two-finger grippers, arranged for bimanual manipulation. A three-camera RGB-D rig (overhead, side and wrist-level views) removes occlusion blind spots during grasping and handovers.

Challenge

Reinforcement learning on two arms is unstable: sparse rewards, collisions between the arms and infeasible policy outputs all stall training.

Approach

Custom reward functions combine sparse task-completion rewards with dense shaping terms such as distance to goal, grasp stability, collision penalties and bimanual coordination bonuses. Observations fuse joint states, end-effector poses and RGB-D features; actions are defined in Cartesian space with safety clamps. A trajectory-planning layer provides collision-aware fallback paths when the learned policy outputs infeasible commands.

Result

A simulation research platform, also used as the base for vision-language-action experiments including a stack-and-sort task that combines vision and RL to detect, classify and manipulate objects.

In brief

  • Dual PiPER arms with a three-camera RGB-D perception rig
  • Custom dense and sparse reward design
  • Curriculum-style task progression
  • Planning fallback for safe policy execution
  • Stack-and-sort task combining vision and RL