M2M-HMR
Monocular Models are Strong Learners for Multi-View Human Mesh Recovery
Applied Scientist II · Amazon
Scholar· GitHub· LinkedIn· X· Email
I'm an Applied Scientist II at Amazon, where I develop large-scale multimodal AI systems for real-world operational and customer experiences. I received my Ph.D. in Computer Science from the University of North Carolina at Charlotte in 2026, advised by Prof. Pu Wang in the GENIUS Lab. I was previously a Research Scientist Intern at Google's Extended Reality (AR/VR) team, and my work includes a U.S. Patent on 3D human pose and movement estimation from monocular images.
My research builds multimodal foundation models that unify real-time perception with high-fidelity synthesis: 3D human pose estimation and mesh reconstruction via generative masked modeling, and multimodal motion synthesis frameworks for controllable, high-quality 3D human animation in real time. Ultimately, I aim to create AI systems that both understand human behavior in the physical world and synthesize interactive digital counterparts within immersive XR environments.
Research interests
3D Human Understanding·Generative Motion·Multimodal AI·Interactive XR
Open to research collaborations and opportunities — reach me at usama.saleem7977@gmail.com.




Monocular Models are Strong Learners for Multi-View Human Mesh Recovery
Streamable Co-Speech Gesture Generation Model
High-fidelity and Editable Dance Synthesis via Generative Masked Motion Prior
Generative Masked Modeling for Robust Hand Mesh Reconstruction in the Wild
Spatio-Temporal Control for Masked Motion Synthesis

Generative Human Mesh Recovery
Biomechanically-accurate 3D Pose Estimation from Monocular Videos
Bidirectional Autoregressive Motion Model