Instructions to use kabilanKB/yam_put_pot_act_m3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LeRobot
How to use kabilanKB/yam_put_pot_act_m3 with LeRobot:
- Notebooks
- Google Colab
- Kaggle
ACT + M3 for bimanual YAM: PutPotOnCooktop
An ACT policy for the simulated bimanual YAM robot, trained
on kabilanKB/yam_put_pot with M3
modality masking, adapted from "Robust Bimanual Vision-Language-Action Models via Embarrassingly
Simple Modality Masking" (arXiv 2608.22419).
The task is PutPotOnCooktop in Isaac Sim 6 / Isaac Lab 3: both arms grasp a pot by its handles, lift it and place it on a cooktop.
Code: kabilankb/Isaac-Lab-3-and-add-SO-101-leader-teleoperation
(see ACT_M3.md).
This is an early research checkpoint. It completes the task in roughly one episode in four, and M3 masking showed no measurable benefit over plain ACT in the evaluation below.
Model
| Inputs | Three 320×240 RGB cameras (top_rgb, left_rgb, right_rgb) and the 14-value joint state |
| Output | Chunk of 100 actions × 14 joint-position targets (6 joints + gripper per arm) |
| Backbone | ImageNet-pretrained ResNet-18, shared across cameras |
| Transformer | 4 encoder layers, 4 decoder layers, width 512, 8 heads, VAE latent 32 |
| Size | 68M parameters |
M3 is applied during training only; this checkpoint is an ordinary ACT checkpoint.
- Wrist-camera masking: for 30% of training samples, both wrist cameras are hidden from attention together. The top camera is never hidden.
- Query masking: each of the 100 action queries is hidden from the others with probability 0.1, and the visible ones are rescaled by 1/0.9.
- The paper's language masking does not apply, since ACT has no language input.
Training
| Data | kabilanKB/yam_put_pot: 20 MimicGen episodes, 7,779 frames at 30 FPS, domain randomization on |
| Steps / batch size | 20,000 / 32 (about 82 passes over the data) |
| Optimizer | AdamW, learning rate 1e-5 |
| Seed | 1000 (single run) |
| Final training loss | 0.041 |
Evaluation
PutPotOnCooktop-v0 with pot_000 / cooktop_000, plain scene (no domain randomization),
25 actions executed per prediction, 900-step limit, 50 episodes, 10 parallel environments.
| Model | Full task succeeded | Pot lifted |
|---|---|---|
| ACT + M3 (this checkpoint) | 11 of 50 (22%) | 13 of 50 (26%) |
| Plain ACT, same settings | 12 of 50 (24%) | 15 of 50 (30%) |
- The two models are indistinguishable on this test.
- The rates are slightly optimistic: the run stopped at the first 50 finished episodes, and successes finish sooner than timeouts.
- Not tested: randomized or cluttered scenes, which is where the paper claims M3 helps.
- The dataset is small (two demonstrations per pot and cooktop pair); more data is the most likely way to raise the success rate.
Usage
Requires the LeRobot v2.0 fork used by YAMLab (RogerDAI1217/lerobot, branch lerobotv2.0).
Its config.json has no type key, so pass the config explicitly:
import draccus
from huggingface_hub import snapshot_download
from lerobot.common.policies.act.configuration_act import ACTConfig
from lerobot.common.policies.act.modeling_act import ACTPolicy
ckpt = snapshot_download("kabilanKB/yam_put_pot_act_m3")
with draccus.config_type("json"):
config = draccus.parse(ACTConfig, f"{ckpt}/config.json", args=[])
config.n_action_steps = 25
policy = ACTPolicy.from_pretrained(ckpt, config=config)
To run it in the simulator, download the checkpoint and use scripts/eval/eval_act.py from the
code repository:
D=yamlab_datasets/tasks_data/PutPotOnCooktop/objects
python scripts/eval/eval_act.py --checkpoint <downloaded folder> --n_action_steps 25 \
--task PutPotOnCooktop-v0 --asset pot=$D/Pot/pot_000 --asset cooktop=$D/Cooktop/cooktop_000 \
--num_episodes 5 --enable_gripper_clamp --enable_cameras --viz kit
- Downloads last month
- 13