GVHMR Online

Fine-tuned causal GVHMR for one person and a static camera. Uses the current frame and cached past information to predict the current body pose; no future frames or revision of earlier outputs.

Code and installation · Download checkpoint

The code is available on the online-gvhmr branch of the Otago repository (repository access required). The checkpoint download is public and verified without authentication.

Use

After installing the code and frontend/body-model assets described in the repository README:

python tools/online/download_checkpoint.py --repo zzwalala/gvhmr-online
python tools/online/demo.py --video /path/to/video.mp4 \
  --target-fps 10 --no-flip-test --fast --frontend-precision fp32 \
  --save-meshes --output outputs/my_video
# Or show a live camera beside the predicted mesh:
python tools/online/preview.py --camera 0 --fps 10 --output outputs/my_camera

Run from Otago/online-gvhmr. No Hugging Face login is required to download. The checkpoint is inference-only; optimizer state, datasets, body-model assets and frontend weights are not included as separate assets.

Checkpoint

  • Epoch 50, global step 264,150; fine-tuned from the original GVHMR release.
  • Fine-tuning data: prepared BEDLAM static-camera sequences and Human3.6M S1/S5/S6/S7; development uses disjoint BEDLAM recordings and H36M S8.
  • 10 Hz observations, 120-observation local attention cache; internal velocity units retain the 30 Hz model convention.
  • Shape averaging uses a running prefix; world reconstruction uses forward-only velocity/contact updates and a fixed initial floor.
  • SHA-256: 87f5782cb36b2d4df1d806db0d8d085ca6496e8870529dd07b489ab22651e1cd.

Evaluation

Matched H36M S9/S11 subset, 12 clips, six actions, 10 Hz. All models use identical detected RGB observations. 2,746/2,880 selected frames scored; 134 common detection failures excluded.

Model MPJPE (mm) ↓ PA-MPJPE (mm) ↓
Original offline 61.29 43.05
Original weights, causal 65.50 46.81
Fine-tuned causal 61.97 42.89

Fine-tuning lowers causal MPJPE by 5.4%, while acceleration error increases 6.0%. This is not the full official benchmark. Root-relative pose accuracy does not measure global trajectory quality. Floor/world drift, occlusions, incorrect initial floor estimates, and moving cameras remain limitations.

Attribution and terms

Adapted from GVHMR. Retains the upstream educational/research/non-commercial terms in LICENSE. Use the original authors' citation from the repository. Separate frontend and body-model assets are subject to their own terms.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support