Skip to content

Data auto-generation part using a simple state machine - #127

Merged
EverNorif merged 56 commits into
LightwheelAI:mainfrom
Papaercold:auto-data-generation
Mar 10, 2026
Merged

EverNorif merged 56 commits into
LightwheelAI:mainfrom
Papaercold:auto-data-generation

Conversation

@Papaercold

@Papaercold Papaercold commented Jan 28, 2026 •

Copy link
Copy Markdown
Contributor

Summary

This PR introduces a state-machine–based data generation pipeline for the SO101 pick-orange task in LeIsaac.

The main implementation lives under scripts/environments/state_machine, where a scripted finite state machine is used to generate deterministic pick-and-place demonstrations. To integrate this pipeline with the existing teleoperation and replay infrastructure, a new teleoperation device type named so101_state_machine is added, with corresponding changes under source/leisaac.

The implementation intentionally follows the coding style and structure of
scripts/environments/teleoperation/teleop_se3_agent.py, and includes detailed comments for readability and maintainability.


Key Features

  • State-machine–based data generation

    • Implemented under scripts/environments/state_machine
    • Designed to be deterministic, readable, and easy to extend
  • New teleoperation device: so101_state_machine

    • Integrated into the existing teleoperation interface
    • Allows scripted policies to drive the environment using the same action pipeline as keyboard/gamepad devices
  • Replay compatibility

    • Recorded datasets can be replayed using the existing
      scripts/environments/teleoperation/replay.py
    • No custom replay logic is required

Usage Examples

Generate data

python scripts/environments/state_machine/pick_orange.py \
  --dataset_file=./datasets/dataset_test.hdf5 \
  --task=LeIsaac-SO101-PickOrange-v0 \
  --num_envs=1 \
  --device=cuda \
  --enable_cameras \
  --record \
  --num_demos=1

Replay recorded data

python scripts/environments/teleoperation/replay.py \
  --dataset_file=./datasets/dataset_test.hdf5 \
  --task=LeIsaac-SO101-PickOrange-v0 \
  --num_envs=1 \
  --device=cuda \
  --enable_cameras \
  --select_episodes 1 \
  --replay_mode=action \
  --task_type=so101_state_machine

Observed Issue / Open Question

When running the state-machine–based data generation script, the robot gripper exhibits a brief downward drop at the beginning of each episode before stabilizing.

At first glance, this behavior appears to be gravity-related. However, gravity is explicitly disabled at spawn time for all relevant teleoperation devices, including the newly introduced state-machine device:

def use_teleop_device(self, teleop_device) -> None:
    self.task_type = teleop_device
    if teleop_device in ["keyboard", "gamepad", "so101_state_machine"]:
        self.scene.robot.spawn.rigid_props.disable_gravity = True

Given this configuration, it is unclear whether the observed initial drop is truly caused by gravity. Other potential factors may include controller or drive initialization behavior, insufficient drive stiffness during the first few simulation steps, or the absence of an explicit action warm-start when the episode begins.

I would appreciate feedback on the following questions:

  • Is this type of initial transient expected when introducing a scripted state machine as a teleoperation device?
  • Could this behavior be related to controller initialization, drive parameter settings, or action preprocessing rather than gravity itself?
  • Are there recommended best practices in Isaac Lab / LeIsaac for avoiding such initial transients when using scripted or non-interactive teleoperation devices?

Any insights, suggestions, or references to similar patterns in existing tasks would be greatly appreciated.


Notes

  • This PR focuses on clarity, determinism, and minimal intrusion into existing teleoperation and replay logic.
  • The issue described above does not affect replay correctness once the episode is running, but it does impact the visual behavior at the very beginning of an episode.

@EverNorif

Copy link
Copy Markdown
Collaborator

@Papaercold Thank you for your PR. Due to a busy schedule recently, I may not be able to review the code right away, I'll check it later.

Regarding the brief drop of the gripper that you mentioned, I haven’t observed this behavior so far. I’ll check for it again using the latest code.

@EverNorif EverNorif linked an issue Jan 28, 2026 that may be closed by this pull request
@EverNorif
EverNorif self-requested a review January 28, 2026 07:54
@Papaercold

Copy link
Copy Markdown
Contributor Author

The fold_cloth state machine is still under development.

At the moment, it is very difficult to complete the cloth-folding task using the current state machine.

I believe there are still some issues with the collision configuration. In some cases, the gripper penetrates the cloth mesh, which prevents it from being properly grasped and lifted.

@EverNorif

Copy link
Copy Markdown
Collaborator

@Papaercold Thank you! I will review it next week.

@Papaercold

Copy link
Copy Markdown
Contributor Author

I’ve started working on the reinforcement learning module based on the rsl_rl library, but it is still under development.

For now, please ignore the files under the rl folder as well as the new rl_config settings during your review.

The directory structure and corresponding file descriptions are also documented in the Markdown file.

If you prefer, I can submit a separate PR for the RL part after you finish reviewing the state machine implementation.

@EverNorif

Copy link
Copy Markdown
Collaborator

@Papaercold Yes, I think it would be better to keep the PR focused on a single feature. You can save your progress locally or in another branch first. This PR should only track the state machine–related functionality.

@Papaercold

Copy link
Copy Markdown
Contributor Author

@Papaercold Yes, I think it would be better to keep the PR focused on a single feature. You can save your progress locally or in another branch first. This PR should only track the state machine–related functionality.

Sure, I’ll make the changes now. I’ll keep this PR focused on the state machine functionality.

@Papaercold

Papaercold commented Mar 3, 2026 •

Copy link
Copy Markdown
Contributor Author

@EverNorif The modifications have been completed. In this commit, the RL module has been fully removed. I have also removed fold_cloth.py, as its current performance is not yet satisfactory.

The remaining code consists only of the components that I have tested and verified to be functioning correctly. You may refer to the Markdown documentation for guidance during your review.

@EverNorif EverNorif left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I tested this PR locally and it runs successfully.

There are still some issues with the structure and implementation that need to be adjusted.

In terms of structure, I think this organization would be better:

scripts/datagen/state_machine/
-- generate.py(or any other name you think is appropriate.)      # Unified runner script: run the state machine based on the provided task, rather than being limited to specific one.
-- replay.py           # Replay script for state-machine demonstrations

source/leisaac/leisaac/datagen/state_machine/
-- base.py             # StateMachineBase abstract class
-- pick_orange.py      # PickOrangeStateMachine
-- ... # other state machine implement

Comment thread dependencies/IsaacLab
Comment thread scripts/environments/state_machine/pick_orange.py Outdated
Comment thread source/leisaac/leisaac/state_machine/__init__.py Outdated
Comment thread source/leisaac/leisaac/state_machine/pick_orange.py Outdated
Comment thread record_pick_orange.sh Outdated
Comment thread replay_pick_orange.sh Outdated
Comment thread STATE_MACHINE_README.md Outdated
Comment thread STATE_MACHINE_README.md Outdated
Comment thread STATE_MACHINE_README.md Outdated
@Papaercold

Copy link
Copy Markdown
Contributor Author

@EverNorif Thanks for the review. I’ll make the requested changes shortly and follow up here if anything comes up.

@Papaercold

Papaercold commented Mar 6, 2026 •

Copy link
Copy Markdown
Contributor Author

I’ve finished the updates. Here’s a summary of the changes.

Summary of Changes

  1. Refactored the code according to the provided feedback.
  2. Aligned the coding style with the teleoperation module and added keyboard input signal handling (e.g., Ctrl+C).
  3. Removed outdated Markdown and bash files. Added state_machine.md to the official documentation. Updated js files and introduction.md. The project should now be directly deployable to the website.
  4. Fixed the incorrect numbering issue in the recorded data. All data numbering rules are now consistent with the teleoperation module.
  5. Removed inappropriate code comments.
  6. The Isaac Lab version is now aligned with the main project.

@EverNorif Thank you for your review and support.

@EverNorif EverNorif left a comment •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall there are no major issues; there may only be two remaining bugs and the documentation location needs to be updated. And we need to ensure that the linters run successfully.

Comment thread docs/docs/docs/features/state_machine.md
Comment thread scripts/datagen/state_machine/generate.py
Comment thread docs/docs/docs/getting_started/state_machine.md Outdated
@Papaercold

Copy link
Copy Markdown
Contributor Author

@EverNorif You can test it again now. All the known issues should be resolved.

@EverNorif

Copy link
Copy Markdown
Collaborator

@Papaercold I think overall it is already fine, but it still needs to pass the linters. You can refer to the configuration in .pre-commit-config.yaml to run the pre-commit checks.

@Papaercold
Papaercold requested a review from EverNorif March 6, 2026 22:49
@Papaercold

Copy link
Copy Markdown
Contributor Author

@EverNorif Finished.
By the way, when I tested the entire project, following pre-existing files fail flake8 E226 (missing whitespace around arithmetic operator), but they are not part of my changes:

  • scripts/convert/isaaclab2lerobot.py:167 — hdf5_id+1
  • scripts/convert/isaaclab2lerobotv3.py:168 — hdf5_id+1

These issues already exist on the main branch prior to this PR. Just as a heads-up — I’ll leave it to you to decide whether to fix it.

@EverNorif EverNorif left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@EverNorif
EverNorif merged commit 41696e4 into LightwheelAI:main Mar 10, 2026
2 checks passed
harsh-mulodhia pushed a commit to harsh-mulodhia/leisaac that referenced this pull request Jun 24, 2026
)

* Add data auto-generate module.

* Add data auto-generate module.

* Add data auto-generate module.

* Add data auto-generate module.

* Add data auto-generate module.

* Add auto_terminate.

* Add auto_terminate.

* Add auto_terminate.

* Add Description.

* State Machinecode refactoring.

* State Machinecode refactoring.

* State Machinecode refactoring.

* State Machinecode refactoring.

* State Machinecode refactoring.

* State Machinecode refactoring.

* Add State Machine code.

* Apply pre-commit fixes (black/isort/pyupgrade) for several files

* Apply pre-commit fixes (black/isort/pyupgrade) for pick_orange.py

* Change structure..

* Create StateMacchine Class.

* Refactor code.

* Fix bugs.

* Delete redundant files

* Delete redundant files.

* Change PickOrangeStateMachine

* Change PickOrangeStateMachine

* Change PickOrangeStateMachine

* Change PickOrangeStateMachine

* Change PickOrangeStateMachine

* Change PickOrangeStateMachine

* Add state_machine/fold_cloth.py

* Fix bugs

* Add state_machine/replay.py

* Add readme

* fix bugs

* fix bugs

* fix bugs

* fix bugs

* fix bugs

* Change documents

* Change bi_arm_cfg

* Add RL module - 1st version.

* Change documents.

* Change bash.

* Delete RL part.

* Refactor

* Change format

* Refactor

* Change Isaaclab version==2.3.2

* Change Isaaclab version==2.3.0

* Change documents.

* Fix bugs.

* Change format.

* Change format.

---------

Co-authored-by: Zihan Gao <zg137@duke.edu>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Feature Request] Automatic Data Generation Pipeline

2 participants