Applied Scientist II · Amazon

Muhammad Usama Saleem

I'm an Applied Scientist II at Amazon, where I develop large-scale multimodal AI systems for real-world operational and customer experiences. I received my Ph.D. in Computer Science from the University of North Carolina at Charlotte in 2026, advised by Prof. Pu Wang in the GENIUS Lab. I was previously a Research Scientist Intern at Google's Extended Reality (AR/VR) team, and my work includes a U.S. Patent on 3D human pose and movement estimation from monocular images.

My research builds multimodal foundation models that unify real-time perception with high-fidelity synthesis: 3D human pose estimation and mesh reconstruction via generative masked modeling, and multimodal motion synthesis frameworks for controllable, high-quality 3D human animation in real time. Ultimately, I aim to create AI systems that both understand human behavior in the physical world and synthesize interactive digital counterparts within immersive XR environments.

Research interests

3D Human Understanding·Generative Motion·Multimodal AI·Interactive XR

Open to research collaborations and opportunities — reach me at usama.saleem7977@gmail.com.

Muhammad Usama Saleem

Experience

2026 — Present
Amazon
Applied Scientist II · Seattle, WA
2025 — 2026
Google
Research Scientist Intern · Extended Reality (AR/VR) · San Francisco, CA
Summer 2025
Amazon
Research Scientist · Multimodal GenAI · Boston, MA
2023 — 2025
Lowe's
Research Lead · Charlotte, NC

News

Selected Research

ECCV 2026 Oral

M2M-HMR

Monocular Models are Strong Learners for Multi-View Human Mesh Recovery

Publications

2026

U.S. Patent — 3D human pose from monocular images
MAGE
Google
MAGE: Modality-Agnostic Music Generation and Target-Source Extraction
Muhammad Usama Saleem, Ravi Tejasvi, Tianyu Xu, Rajeev Nongpiur, Ishan Chatterjee, Mayur Jagdishbhai Patel, Pu Wang
In collaboration with Google
arXiv, 2026
Work done during research internship at Google.
Real-Time Neural Musculoskeletal Pose Estimation
Real-Time Neural Musculoskeletal Pose Estimation
Shengkai Xu, Farnoosh Koleini, Muhammad Usama Saleem, Abbey Thomas, Ahmed Helmy, Pu Wang
IEEE/ACM CHASE, 2026
LiveGesture
LiveGesture: Streamable Co-Speech Gesture Generation Model
Muhammad Usama Saleem, Mayur Jagdishbhai Patel, Ekkasit Pinyoanuntapong, Zhongxing Qin, Li Yang, Hongfei Xue, Ahmed Helmy, Chen Chen, Pu Wang
CVPR, 2026
M2M-HMR
Walk Before You Dance: High-fidelity and Editable Dance Synthesis via Generative Masked Motion Prior
Foram Shah*, Parshwa Shah*, Muhammad Usama Saleem, Ekkasit Pinyoanuntapong, Pu Wang, Hongfei Xue, Ahmed Helmy
AAAI, 2026
*Equal Contribution

2025

MaskHand: Generative Masked Modeling for Robust Hand Mesh Reconstruction in the Wild
Muhammad Usama Saleem, Ekkasit Pinyoanuntapong, Mayur Jagdishbhai Patel, Hongfei Xue, Ahmed Helmy, Srijan Das, Pu Wang
ICCV, 2025
Snapchat
MaskControl: Spatio-Temporal Control for Masked Motion Synthesis
Ekkasit Pinyoanuntapong, Muhammad Usama Saleem, Korrawe Karunratanakul, Pu Wang, Hongfei Xue, Chen Chen, Chuan Guo, Junli Cao, Jian Ren, Sergey Tulyakov
In collaboration with Snapchat
ICCV, 2025Oral · Best Paper Award Nominee
GenHMR: Generative Human Mesh Recovery
Muhammad Usama Saleem, Ekkasit Pinyoanuntapong, Pu Wang, Hongfei Xue, Srijan Das, Chen Chen
AAAI, 2025
BioPose: Biomechanically-accurate 3D Pose Estimation from Monocular Videos
Muhammad Usama Saleem*, Farnoosh Koleini*, Pu Wang, Hongfei Xue, Ahmed Helmy, Abbey Fenwick
WACV, 2025
*Equal Contribution

2024

BAMM: Bidirectional Autoregressive Motion Model
Ekkasit Pinyoanuntapong, Muhammad Usama Saleem, Pu Wang, Minwoo Lee, Srijan Das, Chen Chen
ECCV, 2024

2023

DPFedProxGAN

2022

Privacy Enhancement
Privacy Enhancement for Cloud-Based Few-Shot Learning
A. Parnami, Muhammad Usama Saleem, L. Fan, M. Lee
IJCNN, 2022
DP-Shield: Face Obfuscation with Differential Privacy
Muhammad Usama Saleem, D. Reilly, L. Fan
EDBT, 2022