CAST: Real-Time Motion Capture for Any Skeleton Topology

Anonymous Authors

Under review · research draft

Abstract

This paper studies recovering motion from a monocular video, even when the target skeleton topology differs from the subject observed in the video. Although previous work has made efforts to resolve this problem, they remain vulnerable to joint-rotation errors that are coupled across the skeleton and can propagate along its kinematic hierarchy. To mitigate these errors, we present CAST, a topology-aware motion capture framework that combines learned local rotations with geometric correction guided by predicted global bone directions.

Our observability analysis characterizes which rotational components can be constrained by outgoing bone directions. Guided by this analysis, GRACE uses reliability-gated alignment to correct composed orientations in parent-to-child order, leaving unconstrained components to the learned rotation predictor. The motion predictor and correction module are trained jointly with motion-derived supervision, using a frozen image encoder.

Experiments under the TopoCap evaluation protocol show that the proposed framework improves motion recovery across different target topologies, including unseen structures. Reference-based cross-skeleton evaluation on a curated Mixamo benchmark and a user preference study further assess transfer to different target structures. A causal CAST-B variant achieves over 51 FPS on a single NVIDIA A100 GPU, with timing covering image encoding and motion prediction.

One input, four target rigs.

Choose an input to see the same action expressed through different skeletons.

CAST-L
IN Input RGB
01 Horse
02 Crab
03 Ostrich
04 Human
0.0 / 2.7 s

Learn the motion. Let geometry guide the correction.

Visual evidence meets the structure of the target skeleton.

CAST architecture: DINOv3 video features and target-skeleton queries feed an encoder and rotation decoder. A bone direction head provides global directions and confidence to GRACE, which refines the rotations. The lower panels explain rotational observability and geometric alignment.
CAST jointly trains its motion predictor and GRACE. The target’s outgoing bone geometry determines which rotational components can be constrained.
01

Condition on the target

Skeleton queries gather image evidence, exchange information across joints, and aggregate motion over time.

02

Predict global directions

The network estimates local rotations, global bone directions, and directional reliability from the learned features.

03

Correct along the hierarchy

GRACE aligns bone directions and gates each update before passing the corrected orientation to child joints.

Across familiar and unseen structures.

Matched-rig reconstruction and reference-based cross-skeleton transfer.

Matched-rig motion capture

Compare the captured motion with the input across ten Zoo and ten MObj sequences.

Zoo10 sequences
MObj10 sequences

Cross-skeleton motion capture

Transfer motion across different source and target skeletons. Each row pairs the input clip with the motion CAST transfers to the target rig.

Obj → Obj15 sequences
Lower is better ↓
Zoo · Matched-rig motion capture
MethodMPJPE (mm)PA-MPJPE (mm)MPJVE (mm/frame)CD
Puppeteer454.70236.3819.740.245
TopoCap————
MCA V2*105.5267.0410.350.077
CAST-B Ours42.6833.235.480.036
CAST-L Ours39.5931.055.140.034

Results from the current manuscript. MCA V2* is our reimplementation trained on the same Zoo and MObj data. Metrics use root-centered predictions and the TopoCap normalization convention; CD is in normalized coordinates. A dash denotes an unreported result.

Unseen skeletons

CAST-L reaches 74.63 mm MPJPE on MObj Unseen, compared with 103.04 mm for MCA V2*.

Cross-skeleton transfer

CAST-L reduces MPJPE by 54.7% relative to MCA V2* across 300 ordered source–target pairs on Mixamo.

Perceptual quality

CAST-L receives 63.3% of first-choice votes from 18 participants across 360 rankings.

Motion, on your terms.

Beyond dense video: streaming inference and key-frame completion.

Causal inference

Incoming frames. Immediate motion.

A causal CAST-B variant processes incoming frames without waiting for future observations.

51.3–53.9 FPS

Single NVIDIA A100 · 30–89 joints

Measured with bf16 and torch.compile. Includes the frozen image backbone and motion prediction; excludes video acquisition and mesh rendering.

Key-frame completion

Fewer observations. A complete motion sequence.

With one observed image every four frames, the completion variant increases MPJPE by 5.7–8.5% relative to dense input.

Sketch-guided completion examples showing two observed sketches and the completed poses in between.

Sketch-guided examples at observation interval k = 4. Open the figure to view all 35 examples.

What remains challenging

Bone directions cannot resolve every rotation: axial twist and leaf-joint orientation still depend on the learned predictor. CAST also assumes fixed parent-local bone offsets. Motion that changes these offsets lies outside its rotation-only representation.