ACT + M3 for bimanual YAM: PutPotOnCooktop

An ACT policy for the simulated bimanual YAM robot, trained on kabilanKB/yam_put_pot with M3 modality masking, adapted from "Robust Bimanual Vision-Language-Action Models via Embarrassingly Simple Modality Masking" (arXiv 2608.22419).

The task is PutPotOnCooktop in Isaac Sim 6 / Isaac Lab 3: both arms grasp a pot by its handles, lift it and place it on a cooktop.

Code: kabilankb/Isaac-Lab-3-and-add-SO-101-leader-teleoperation (see ACT_M3.md).

This is an early research checkpoint. It completes the task in roughly one episode in four, and M3 masking showed no measurable benefit over plain ACT in the evaluation below.

Model

Inputs Three 320×240 RGB cameras (top_rgb, left_rgb, right_rgb) and the 14-value joint state
Output Chunk of 100 actions × 14 joint-position targets (6 joints + gripper per arm)
Backbone ImageNet-pretrained ResNet-18, shared across cameras
Transformer 4 encoder layers, 4 decoder layers, width 512, 8 heads, VAE latent 32
Size 68M parameters

M3 is applied during training only; this checkpoint is an ordinary ACT checkpoint.

  • Wrist-camera masking: for 30% of training samples, both wrist cameras are hidden from attention together. The top camera is never hidden.
  • Query masking: each of the 100 action queries is hidden from the others with probability 0.1, and the visible ones are rescaled by 1/0.9.
  • The paper's language masking does not apply, since ACT has no language input.

Training

Data kabilanKB/yam_put_pot: 20 MimicGen episodes, 7,779 frames at 30 FPS, domain randomization on
Steps / batch size 20,000 / 32 (about 82 passes over the data)
Optimizer AdamW, learning rate 1e-5
Seed 1000 (single run)
Final training loss 0.041

Evaluation

PutPotOnCooktop-v0 with pot_000 / cooktop_000, plain scene (no domain randomization), 25 actions executed per prediction, 900-step limit, 50 episodes, 10 parallel environments.

Model Full task succeeded Pot lifted
ACT + M3 (this checkpoint) 11 of 50 (22%) 13 of 50 (26%)
Plain ACT, same settings 12 of 50 (24%) 15 of 50 (30%)
  • The two models are indistinguishable on this test.
  • The rates are slightly optimistic: the run stopped at the first 50 finished episodes, and successes finish sooner than timeouts.
  • Not tested: randomized or cluttered scenes, which is where the paper claims M3 helps.
  • The dataset is small (two demonstrations per pot and cooktop pair); more data is the most likely way to raise the success rate.

Usage

Requires the LeRobot v2.0 fork used by YAMLab (RogerDAI1217/lerobot, branch lerobotv2.0). Its config.json has no type key, so pass the config explicitly:

import draccus
from huggingface_hub import snapshot_download
from lerobot.common.policies.act.configuration_act import ACTConfig
from lerobot.common.policies.act.modeling_act import ACTPolicy

ckpt = snapshot_download("kabilanKB/yam_put_pot_act_m3")
with draccus.config_type("json"):
    config = draccus.parse(ACTConfig, f"{ckpt}/config.json", args=[])
config.n_action_steps = 25
policy = ACTPolicy.from_pretrained(ckpt, config=config)

To run it in the simulator, download the checkpoint and use scripts/eval/eval_act.py from the code repository:

D=yamlab_datasets/tasks_data/PutPotOnCooktop/objects
python scripts/eval/eval_act.py --checkpoint <downloaded folder> --n_action_steps 25 \
    --task PutPotOnCooktop-v0 --asset pot=$D/Pot/pot_000 --asset cooktop=$D/Cooktop/cooktop_000 \
    --num_episodes 5 --enable_gripper_clamp --enable_cameras --viz kit
Downloads last month
13
Safetensors
Model size
67.8M params
Tensor type
F32
·
Video Preview
loading

Dataset used to train kabilanKB/yam_put_pot_act_m3

Papers for kabilanKB/yam_put_pot_act_m3