GVHMR Online
Fine-tuned causal GVHMR for one person and a static camera. Uses the current frame and cached past information to predict the current body pose; no future frames or revision of earlier outputs.
Code and installation · Download checkpoint
The code is available on the online-gvhmr branch of the Otago repository (repository access required). The checkpoint download is public and verified without authentication.
Use
After installing the code and frontend/body-model assets described in the repository README:
python tools/online/download_checkpoint.py --repo zzwalala/gvhmr-online
python tools/online/demo.py --video /path/to/video.mp4 \
--target-fps 10 --no-flip-test --fast --frontend-precision fp32 \
--save-meshes --output outputs/my_video
# Or show a live camera beside the predicted mesh:
python tools/online/preview.py --camera 0 --fps 10 --output outputs/my_camera
Run from Otago/online-gvhmr. No Hugging Face login is required to download. The checkpoint is inference-only; optimizer state, datasets, body-model assets and frontend weights are not included as separate assets.
Checkpoint
- Epoch 50, global step 264,150; fine-tuned from the original GVHMR release.
- Fine-tuning data: prepared BEDLAM static-camera sequences and Human3.6M S1/S5/S6/S7; development uses disjoint BEDLAM recordings and H36M S8.
- 10 Hz observations, 120-observation local attention cache; internal velocity units retain the 30 Hz model convention.
- Shape averaging uses a running prefix; world reconstruction uses forward-only velocity/contact updates and a fixed initial floor.
- SHA-256:
87f5782cb36b2d4df1d806db0d8d085ca6496e8870529dd07b489ab22651e1cd.
Evaluation
Matched H36M S9/S11 subset, 12 clips, six actions, 10 Hz. All models use identical detected RGB observations. 2,746/2,880 selected frames scored; 134 common detection failures excluded.
| Model | MPJPE (mm) ↓ | PA-MPJPE (mm) ↓ |
|---|---|---|
| Original offline | 61.29 | 43.05 |
| Original weights, causal | 65.50 | 46.81 |
| Fine-tuned causal | 61.97 | 42.89 |
Fine-tuning lowers causal MPJPE by 5.4%, while acceleration error increases 6.0%. This is not the full official benchmark. Root-relative pose accuracy does not measure global trajectory quality. Floor/world drift, occlusions, incorrect initial floor estimates, and moving cameras remain limitations.
Attribution and terms
Adapted from GVHMR. Retains the upstream educational/research/non-commercial terms in LICENSE. Use the original authors' citation from the repository. Separate frontend and body-model assets are subject to their own terms.