Logo image
M3T: Discrete Multi-Modal Motion Tokens for Sign Language Production
Conference proceeding   Peer reviewed

M3T: Discrete Multi-Modal Motion Tokens for Sign Language Production

Proceedings of BMVC 2026
BMVC 2026, 37 (Lancaster, United Kingdom, 23/11/2026–23/11/2026)
07/08/2026

Abstract

Sign Language Production Motion Modeling

Sign language production requires more than hand motion generation. Non-manual features, including mouthings, eyebrow raises, gaze, and head movements, are grammatically obligatory and cannot be recovered from manual articulators alone. Existing 3D production systems face two barriers to integrating them: the body-model fittings they train on retain a facial space too low-dimensional to encode these articulations, and, if richer representations are adopted, standard discrete tokenization suffers from codebook collapse, leaving most of the expression space largely unreachable. We propose SMPL-FX, which couples FLAME's rich expression space with the SMPL-X body parameterisation, and tokenize the resulting representation with modality-specific Finite Scalar Quantization VAEs for body, hands, and face, raising face codebook utilization from 78.8% to 99.0%. M³T is an autoregressive transformer trained on this multi-modal motion vocabulary, with an auxiliary sign-to-text translation objective that encourages semantically grounded embeddings. Across three standard datasets, M³T achieves state-of-the-art sign language production quality, and on NMFs-CSL, where signs are distinguishable only by non-manual features, our model reaches 58.3% accuracy against 49.0% for the strongest comparable baseline, without large-scale sign-language pre-training. Project page: https://cogvis-cvssp.github.io/papers/m3t/

pdf
390_M3T_Discrete_Multi_Modal_M5.16 MB
Author's Accepted Manuscript Embargoed Access, Embargo ends: 23/11/2026

Metrics

1 File views/ downloads
2 Record Views

Details

Logo image

Usage Policy