MoMask logo

MoMask

Paid

Generative Masked Modeling of 3D Human Motions

5.0
Inputs: text
Type
Saas

About MoMask

MoMask is a novel masked modeling framework for text-driven 3D human motion generation, developed by researchers at the University of Alberta and Google Research and published at CVPR 2024. It employs a hierarchical quantization scheme to represent human motion as multi-layer discrete motion tokens with high-fidelity details. A Masked Transformer predicts randomly masked base-layer motion tokens conditioned on text input during training and iteratively fills in missing tokens during inference. A Residual Transformer progressively predicts next-layer tokens based on current layer results. MoMask achieves state-of-the-art performance on text-to-motion generation, with an FID of 0.045 on HumanML3D (vs 0.141 for T2M-GPT) and 0.228 on KIT-ML (vs 0.514). It also supports text-guided temporal inpainting (e.g., inbetweening, prefix, suffix) without requiring additional fine-tuning.

Key Features

Hierarchical quantization scheme (RVQ-VAE) for multi-layer discrete motion tokens
Masked Transformer for base-layer token prediction conditioned on text
Residual Transformer for progressive next-layer token prediction
State-of-the-art text-to-motion generation (FID 0.045 on HumanML3D, 0.228 on KIT-ML)
Seamless application to text-guided temporal inpainting without fine-tuning
High-fidelity motion reconstruction with residual quantization layers

Pros & Cons

Pros
  • Outperforms existing methods on text-to-motion benchmarks
  • Supports temporal inpainting without additional model fine-tuning
  • Hierarchical quantization enables high-fidelity motion detail

Best For

Text-driven 3D human motion generationText-guided temporal inpainting (inbetweening, prefix, suffix) of motion clipsMotion reconstruction with varying numbers of residual quantization layers

Alternatives to MoMask