About MoMask
MoMask is a novel masked modeling framework for text-driven 3D human motion generation, developed by researchers at the University of Alberta and Google Research and published at CVPR 2024. It employs a hierarchical quantization scheme to represent human motion as multi-layer discrete motion tokens with high-fidelity details. A Masked Transformer predicts randomly masked base-layer motion tokens conditioned on text input during training and iteratively fills in missing tokens during inference. A Residual Transformer progressively predicts next-layer tokens based on current layer results. MoMask achieves state-of-the-art performance on text-to-motion generation, with an FID of 0.045 on HumanML3D (vs 0.141 for T2M-GPT) and 0.228 on KIT-ML (vs 0.514). It also supports text-guided temporal inpainting (e.g., inbetweening, prefix, suffix) without requiring additional fine-tuning.
Key Features
Pros & Cons
- Outperforms existing methods on text-to-motion benchmarks
- Supports temporal inpainting without additional model fine-tuning
- Hierarchical quantization enables high-fidelity motion detail
Best For
Alternatives to MoMask
PlugSugar
Automate conversations, answer questions with Web Search plugin, and customize ChatGPT experience using powerful AI plugins.
100DaysOfAI Challenge
Respage
Automate lead acquisition, interact with potential leads, and capture lead information and preferences.
Travel Plan AI
Your personal AI guide for unforgettable journeys.
3D Avataaars Generator
Create custom avatars for storytelling, game development, and marketing campaigns with ease.
AnimateDiff