A shared, growing interface between planning and control.

InterEvolve

Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation.

Zhuo Lin† Sirui Xu† Liuyu Bian Yu-Xiong Wang‡ Liang-Yan Gui‡

University of Illinois Urbana-Champaign †Equal contribution  ‡Equal advising

Explore the evolution
Scroll to explore

01 / Method

Evolve the task, not the controller.

An LLM agent writes a staged reward program, a frozen controller runs it in parallel simulation, and a fixed verifier's feedback drives the next revision.

02 / Evolve

Evolve reward programs at test time.

03 / Discover

Discover unseen skills.

Novel use of skills from OMOMO.

04 / Adapt

Adapt fast from the skill library.

Large-box programs stored in the skill library evolve into the same skills for a plastic box, a small box, and a suitcase.

05 / Generalize

Generalize to novel goals and initializations.

Each selected program runs unchanged as the target moves or the box starts elsewhere.

06 / Compose

Compose strategies for complex scenarios.

07 / Arrange

Long-horizon box arrangement.

08 / Grasp

Grasp smaller objects.

Inspire hands.

09 / Deploy

Deploy autonomously on a real G1.

Onboard perception only, except where noted.

Kick after kick.

On a real G1, the same kicking program kicks the box in succession. Wherever one kick sends the box, the next kick follows it. Box pose from motion capture.