Item

Monocular Markerless Motion Capture for 3D Whole-Body Biomechanical Motion Analysis

Al Kamali, Abdulaziz Ahmed Yahya
Citations
Altmetric:
Department
Robotics
Embargo End Date
Type
Thesis
Date
2026
License
Language
English
Collections
Research Projects
Organizational Units
Journal Issue
Abstract
Motion capture (MoCap) is widely used in biomechanics, sports science, rehabilitation, animation, and robotics to estimate human movement over time. Classical laboratory motion capture systems usually rely on multiple calibrated cameras and reflective markers, which can provide accurate kinematic measurements but remain expensive, intrusive, and difficult to deploy outside controlled environments. Monocular markerless motion capture is attractive because it reduces the hardware requirement to a single RGB camera, making motion analysis more accessible in real-world settings. However, monocular reconstruction remains fundamentally challenging because a single view does not directly provide depth, and because camera uncertainty, self-occlusion, and fast motion can destabilize the recovered body over time. These difficulties become more severe for children and other subjects whose morphology is weakly represented by adult-dominated body priors. This thesis presents a practical pipeline for monocular, markerless 4D human reconstruction, where ``4D'' refers to a 3D body representation evolving over time. The pipeline builds on SAM-Body4D, a training-free video framework that generates temporally consistent person masks across a sequence and uses them to guide frame-by-frame body reconstruction. For mesh recovery, the pipeline uses \emph{SAM 3D Body}, a full-body human mesh recovery model that predicts a parametric body representation called the Momentum Human Rig (MHR). This representation is suitable for the present work because it separates identity-dependent body shape from articulated state parameters, which enables later shape correction and motion retargeting. In the fitting stage used in this thesis, the recovered articulated state is converted to a canonical body model by resetting the joint-angle subset to zero giving an A-pose; anthropometric optimization is then performed over the 45-dimensional identity coefficients together with a scale subset of the MHR model parameters. To address the mismatch between generic priors and subject-specific geometry, the thesis formulates anthropometric refinement as an inverse problem: given a set of target body measurements, estimate body-model parameters such that the forward-generated mesh matches those measurements. Three fitting strategies are investigated: 1) a learned inverse mapping trained on synthetic data, 2) direct black-box optimization using Covariance Matrix Adaptation Evolution Strategy (CMA-ES), and 3) a multi-stage coarse-to-fine fitting strategy. CMA-ES is well suited to this problem because the measurement process includes non-differentiable operations such as mesh slicing, contour selection, and feasibility checks. In practice, the multi-stage optimization strategy produces the most stable results because it first fits heights, then torso and limb shape, and finally refines head shape and standing height. In addition, temporally stabilizing the estimated camera intrinsics centering the principal point, reusing first-frame intrinsics across the sequence at fixed resolution, and lightly smoothing camera-related outputs after inference reduces visible temporal drift in the reconstructed global trajectory. Experiments on in-the-wild monocular sequences, including a child athlete performing a balance beam routine, show that multi-stage optimization produces more plausible retargeted meshes and lower anthropometric error than learned inversion and single-shot optimization.In addition, a quantitative comparison on an in-place jumping-jacks sequence against Human3R, a recent strong monocular 4D reconstruction baseline, shows that the proposed stabilized reconstruction yields lower centroid jitter, lower depth-axis jitter, and lower in-place drift, while Human3R remains smoother in pose-only dynamics. These results support the claim that the proposed reconstruction stage improves global temporal stability, especially in the depthwise direction, before later anthropometric refinement is applied. These results support anthropometric refinement as a useful step toward subject-specific biomechanics from monocular video. With stable reconstruction and corrected body geometry established, the next step is to connect the recovered kinematics to biomechanical dynamics and contact-related inference.
Citation
Al Kamali, Abdulaziz Ahmed Yahya, "Monocular Markerless Motion Capture for 3D Whole-Body Biomechanical Motion Analysis," M.S. Thesis, Robotics, MBZUAI, 2026.
Source
Conference
Keywords
Temporal camera stabilization, Anthropometrics, Momentum Human Rig, Biomechanics, SAM-Body4D, Human mesh recovery
Subjects
Source
Publisher
DOI
Full-text link