Applied Scientist II · Amazon Music

Muhammad Usama Saleem

I'm an Applied Scientist II at Amazon Music, where I work on multimodal content understanding for music, including track and artist understanding across all genres, as well as sonics, sequencing, and recommendation, building LLM-based systems with LoRA fine-tuning and automated evaluators. I received my Ph.D. in Computer Science from the University of North Carolina at Charlotte in 2026, advised by Prof. Pu Wang in the GENIUS Lab. I was previously a Research Scientist Intern at Google's Extended Reality (AR/VR) team, and my work includes a U.S. Patent on 3D human pose and movement estimation from monocular images.

Research interests

3D Computer Vision·Digital Humans·Generative Motion·Diffusion Models·Multimodal Learning·Audio & Music Representation·Interactive XR

Open to research collaborations and opportunities. Reach me at usama.saleem7977@gmail.com.

Muhammad Usama Saleem

Experience

2026 - Present
Amazon
Applied Scientist II · Amazon Music · Seattle, WA
2025 - 2026
Google
Research Scientist Intern · Extended Reality (AR/VR) · San Francisco, CA
Summer 2025
Amazon
Research Scientist · Multimodal GenAI · Boston, MA
2023 - 2025
Lowe's
Research Lead · Charlotte, NC

News

Selected Research

ECCV 2026 Oral

M2M-HMR

Monocular Models are Strong Learners for Multi-View Human Mesh Recovery

Publications

2026

U.S. Patent: 3D human pose from monocular images
MAGE
Google
MAGE: Modality-Agnostic Music Generation and Target-Source Extraction
Muhammad Usama Saleem, Ravi Tejasvi, Tianyu Xu, Rajeev Nongpiur, Ishan Chatterjee, Mayur Jagdishbhai Patel, Pu Wang
In collaboration with Google
arXiv, 2026
Work done during research internship at Google.
Real-Time Neural Musculoskeletal Pose Estimation
Real-Time Neural Musculoskeletal Pose Estimation
Shengkai Xu, Farnoosh Koleini, Muhammad Usama Saleem, Abbey Thomas, Ahmed Helmy, Pu Wang
IEEE/ACM CHASE, 2026
LiveGesture
LiveGesture: Streamable Co-Speech Gesture Generation Model
Muhammad Usama Saleem, Mayur Jagdishbhai Patel, Ekkasit Pinyoanuntapong, Zhongxing Qin, Li Yang, Hongfei Xue, Ahmed Helmy, Chen Chen, Pu Wang
CVPR, 2026
M2M-HMR
Walk Before You Dance: High-fidelity and Editable Dance Synthesis via Generative Masked Motion Prior
Foram Shah*, Parshwa Shah*, Muhammad Usama Saleem, Ekkasit Pinyoanuntapong, Pu Wang, Hongfei Xue, Ahmed Helmy
AAAI, 2026
*Equal Contribution

2025

MaskHand: Generative Masked Modeling for Robust Hand Mesh Reconstruction in the Wild
Muhammad Usama Saleem, Ekkasit Pinyoanuntapong, Mayur Jagdishbhai Patel, Hongfei Xue, Ahmed Helmy, Srijan Das, Pu Wang
ICCV, 2025
Snapchat
MaskControl: Spatio-Temporal Control for Masked Motion Synthesis
Ekkasit Pinyoanuntapong, Muhammad Usama Saleem, Korrawe Karunratanakul, Pu Wang, Hongfei Xue, Chen Chen, Chuan Guo, Junli Cao, Jian Ren, Sergey Tulyakov
In collaboration with Snapchat
ICCV, 2025Oral · Best Paper Award Nominee
GenHMR: Generative Human Mesh Recovery
Muhammad Usama Saleem, Ekkasit Pinyoanuntapong, Pu Wang, Hongfei Xue, Srijan Das, Chen Chen
AAAI, 2025
BioPose: Biomechanically-accurate 3D Pose Estimation from Monocular Videos
Muhammad Usama Saleem*, Farnoosh Koleini*, Pu Wang, Hongfei Xue, Ahmed Helmy, Abbey Fenwick
WACV, 2025
*Equal Contribution

2024

BAMM: Bidirectional Autoregressive Motion Model
Ekkasit Pinyoanuntapong, Muhammad Usama Saleem, Pu Wang, Minwoo Lee, Srijan Das, Chen Chen
ECCV, 2024

2023

DPFedProxGAN

2022

Privacy Enhancement
Privacy Enhancement for Cloud-Based Few-Shot Learning
A. Parnami, Muhammad Usama Saleem, L. Fan, M. Lee
IJCNN, 2022
DP-Shield: Face Obfuscation with Differential Privacy
Muhammad Usama Saleem, D. Reilly, L. Fan
EDBT, 2022

Research Focus

My research spans three areas: generative 3D human reconstruction, controllable motion synthesis, and multimodal audio generation and separation, with collaborations across Google DeepMind/XR, Amazon, Snapchat, and Lowe's.

Generative 3D Human Reconstruction

3D Pose Estimation·Mesh Recovery

Generative models that recover accurate 3D human pose and mesh directly from monocular images and video, replacing single-prediction regressors with masked and diffusion-based generation to handle depth ambiguity and occlusion. This spans full-body mesh recovery (GenHMR, M2M-HMR), hand mesh reconstruction in the wild (MaskHand), and biomechanically accurate pose estimation grounded in anatomical constraints (BioPose), alongside a U.S. patent on 3D pose and movement estimation from monocular images.

Controllable Motion Synthesis

Masked Motion Modeling·Co-Speech Gesture·Dance Synthesis

Generative frameworks for controllable, high-fidelity 3D human motion: streamable co-speech gesture generation with zero look-ahead (LiveGesture), editable dance synthesis via a generative masked motion prior (Walk Before You Dance), spatio-temporal control over masked motion synthesis (MaskControl, with Snapchat), and bidirectional autoregressive motion modeling (BAMM). I'm interested in how masking, tokenization, and diffusion jointly enable real-time, controllable generation.

Multimodal Audio Generation and Separation

Music Generation·Source Separation·Music Understanding·LLMs

Extending generative modeling beyond vision to audio and cross-modal settings, including a modality-agnostic framework for conditional music generation and mixture-grounded target-source extraction (MAGE), developed during a research internship at Google. At Amazon Music, I work on multimodal content understanding for music, including track and artist understanding across all genres, sonics, sequencing, and recommendation, using LLMs with LoRA fine-tuning and automated evaluators. I'm broadly interested in how the same generative principles (masking, tokenization, diffusion) transfer across modalities to build unified multimodal systems.