All inputs in this section are real-world IMU captures. Meeting Room and Apartment GT captions are annotated from the RGB videos with Qwen3-VL and condensed to one short sentence.
The IMU signals in this section are synthesized from motion sequences.
The IMU signals in this section are synthesized from motion sequences.
To show out-of-domain generalization, we test our model on HumanML3D and LINGO, two datasets without object annotations. Our model is never trained on these two datasets for the object generation task, yet it still predicts plausible objects from the IMU signals alone.
Our model supports sampling diverse text descriptions. Given the same IMU input (HumanML3D, 5 sensors), we sample 32 captions from our model. Each bar shows how often a caption was sampled, with its perplexity (PPL); the greedy-decoded caption is tagged.
GT A person is kicking a ball in a manner similar to soccer.
7 distinct captions in 32 samples
GT A person ascends four steps with their left hand on the railing.
13 distinct captions in 32 samples
GT A person stands up and then bounces up and down.
15 distinct captions in 32 samples
GT A person slowly lowers themselves and places their right hand on the ground.
13 distinct captions in 32 samples
Our model also supports sampling diverse objects. The body is decoded once and shared by every sample; each sample then draws its own caption, object categories, object instances (meshes) and first-frame layout.
For each clip, we keep the model's predicted motion and text, and replace only the object name in the text. Given the same IMU input and the edited text, the model outputs a different object that fits the new text and still follows the motion.
IMU signals only encode motion relative to the first frame. After fine-tuning on ParaHome, where every sequence is captured in the same fixed scene, our model predicts the person's initial global position and orientation in that scene from new IMU signals alone. This relies on a scene-specific prior learned during fine-tuning.