IMU4D: 4D Human-Object Understanding from Wearable IMUs

IMU4D teaser: IMU time series from earbuds, watch and smartphone are turned into 4D human-object recovery (SMPL-X motion, text description, 3D objects)

We propose IMU4D, a unified multimodal model that jointly predicts human motion, activity descriptions, and the 3D objects the person interacts with purely from wearable IMU signals.

Real-world IMU → Motion / Text


All inputs in this section are real-world IMU captures. Meeting Room and Apartment GT captions are annotated from the RGB videos with Qwen3-VL and condensed to one short sentence.

Use ‹ / › on the video or the dots to browse samples.
Basic preview mode. Open through a local HTTP server for synchronized playback and sample browsing.
DIP-IMU (3pt, head + left wrist + right thigh)
DIP-IMU (3pt, head + left wrist + right thigh)
DIP-IMU (3pt, head + left wrist + right thigh)
DIP-IMU (3pt, head + left wrist + right thigh)
DIP-IMU (3pt, head + left wrist + right thigh)
DIP-IMU (3pt, head + left wrist + right thigh)
DIP-IMU (3pt, head + left wrist + right thigh)
DIP-IMU (3pt, head + left wrist + right thigh)
DIP-IMU (3pt, head + left wrist + right thigh)
DIP-IMU (3pt, head + left wrist + right thigh)
DIP-IMU (3pt, head + left wrist + right thigh)
DIP-IMU (3pt, head + left wrist + right thigh)
DIP-IMU (3pt, head + left wrist + right thigh)
DIP-IMU (3pt, head + left wrist + right thigh)
DIP-IMU (3pt, head + left wrist + right thigh)
DIP-IMU (3pt, head + left wrist + right thigh)
DIP-IMU (3pt, head + left wrist + right thigh)
DIP-IMU (3pt, head + left wrist + right thigh)
DIP-IMU (3pt, head + left wrist + right thigh)
DIP-IMU (3pt, head + left wrist + right thigh)
IMUPoser (3pt, head + left wrist + right pocket)
IMUPoser (3pt, head + left wrist + right pocket)
IMUPoser (3pt, head + left wrist + right pocket)
IMUPoser (3pt, head + left wrist + right pocket)
IMUPoser (3pt, head + left wrist + right pocket)
IMUPoser (3pt, head + left wrist + right pocket)
IMUPoser (3pt, head + left wrist + right pocket)
IMUPoser (3pt, head + left wrist + right pocket)
IMUPoser (3pt, head + left wrist + right pocket)
IMUPoser (3pt, head + left wrist + right pocket)
IMUPoser (3pt, head + left wrist + right pocket)
IMUPoser (3pt, head + left wrist + right pocket)
IMUPoser (3pt, head + left wrist + right pocket)
IMUPoser (3pt, head + left wrist + right pocket)
IMUPoser (3pt, head + left wrist + right pocket)
IMUPoser (3pt, head + left wrist + right pocket)
IMUPoser (3pt, head + left wrist + right pocket)
IMUPoser (3pt, head + left wrist + right pocket)
IMUPoser (3pt, head + left wrist + right pocket)
IMUPoser (3pt, head + left wrist + right pocket)

Synthetic IMU → Motion / Text


The IMU signals in this section are synthesized from motion sequences.

Use ‹ / › on the video or the dots to browse samples.
Basic preview mode. Open through a local HTTP server for synchronized playback and sample browsing.
HumanML3D (5pt, head + both wrists + both pockets)
HumanML3D (5pt, head + both wrists + both pockets)
HumanML3D (5pt, head + both wrists + both pockets)
HumanML3D (5pt, head + both wrists + both pockets)
HumanML3D (5pt, head + both wrists + both pockets)
HumanML3D (5pt, head + both wrists + both pockets)
HumanML3D (5pt, head + both wrists + both pockets)
HumanML3D (5pt, head + both wrists + both pockets)
HumanML3D (5pt, head + both wrists + both pockets)
HumanML3D (5pt, head + both wrists + both pockets)
HumanML3D (5pt, head + both wrists + both pockets)
HumanML3D (5pt, head + both wrists + both pockets)
HumanML3D (5pt, head + both wrists + both pockets)
HumanML3D (5pt, head + both wrists + both pockets)
HumanML3D (5pt, head + both wrists + both pockets)
HumanML3D (5pt, head + both wrists + both pockets)
HumanML3D (5pt, head + both wrists + both pockets)
HumanML3D (5pt, head + both wrists + both pockets)
HumanML3D (5pt, head + both wrists + both pockets)
HumanML3D (5pt, head + both wrists + both pockets)

Synthetic IMU → Motion / Text / Object


The IMU signals in this section are synthesized from motion sequences.

Use ‹ / › on the video or the dots to browse samples.
Basic preview mode. Open through a local HTTP server for synchronized playback and sample browsing.

Out-Of-Distribution Object Generation


To show out-of-domain generalization, we test our model on HumanML3D and LINGO, two datasets without object annotations. Our model is never trained on these two datasets for the object generation task, yet it still predicts plausible objects from the IMU signals alone.

Use ‹ / › on the video or the dots to browse samples.
Basic preview mode. Open through a local HTTP server for synchronized playback and sample browsing.

Text Sampling Diversity


Our model supports sampling diverse text descriptions. Given the same IMU input (HumanML3D, 5 sensors), we sample 32 captions from our model. Each bar shows how often a caption was sampled, with its perplexity (PPL); the greedy-decoded caption is tagged.

GT A person is kicking a ball in a manner similar to soccer.

samplessampled captionPPL
  1. 9/32A person repeatedly kicks forward and walks back.1.16
  2. 6/32A person steps left, kicks with the right foot, steps back, and delivers two left-foot kicks.1.08
  3. 6/32A person is mimicking a soccer kick.1.26
  4. 5/32A person is kicking a ball in the style of football/soccer. greedy1.12

7 distinct captions in 32 samples

GT A person ascends four steps with their left hand on the railing.

samplessampled captionPPL
  1. 10/32A person ascends the stairs, holding onto the handrail.1.12
  2. 4/32A person ascends the stairs with an outstretched arm for support. greedy1.13
  3. 4/32A person climbs the stairs with their left hand on the railing.1.26
  4. 3/32A person climbs the stairs twice while holding the railing with his left hand.1.19

13 distinct captions in 32 samples

GT A person stands up and then bounces up and down.

samplessampled captionPPL
  1. 6/32A person is jumping up and down with both legs. greedy1.18
  2. 5/32A person is bouncing in one spot.1.41
  3. 4/32A person is hopping with a light, effortless motion.1.17
  4. 3/32A person hops with a light, airy step.1.20

15 distinct captions in 32 samples

GT A person slowly lowers themselves and places their right hand on the ground.

samplessampled captionPPL
  1. 9/32A person bends down, picks up an object with their right hand, and comes back to standing. greedy1.08
  2. 5/32A person bends down, picks up something with their right hand, and comes back to standing.1.11
  3. 3/32A person bends down to grasp an object with his right hand.1.15
  4. 3/32A person slowly bends down and places their right hand on the floor.1.21

13 distinct captions in 32 samples

Object Sampling Diversity


Our model also supports sampling diverse objects. The body is decoded once and shared by every sample; each sample then draws its own caption, object categories, object instances (meshes) and first-frame layout.

Use ‹ / › on the video or the dots to browse samples.
Basic preview mode. Open through a local HTTP server for synchronized playback and sample browsing.

Text-Conditioned Object Editing


For each clip, we keep the model's predicted motion and text, and replace only the object name in the text. Given the same IMU input and the edited text, the model outputs a different object that fits the new text and still follows the motion.

Use ‹ / › on the video or the dots to browse samples.
Basic preview mode. Open through a local HTTP server for synchronized playback and sample browsing.

Other Applications


Relocalization (IMU → Motion + Absolute Location)

IMU signals only encode motion relative to the first frame. After fine-tuning on ParaHome, where every sequence is captured in the same fixed scene, our model predicts the person's initial global position and orientation in that scene from new IMU signals alone. This relies on a scene-specific prior learned during fine-tuning.

Ground Truth Ours
Sample 1
Sample 2