People as rulers

A posed body's metric torso against its 2D keypoints gives a closed-form scale label, σ = (f·e₃/e₂)/z̃. Agreeing labels pretrain a Scale Readout that then predicts σ from backbone tokens alone.
arXiv 2026 · 2609.29106
Vision and Image Processing Lab, University of Waterloo, Canada
Drag to orbit, scroll or pinch to zoom, right-drag to pan. Every scene is WildHSR's own prediction at its own metric scale, on a 1 m grid. Switch to Human3R to compare where it puts the same people.
3D foundation models recover video cameras and geometry in one forward pass, but some of the strongest are up to scale. Joint people-scene reconstruction then requires two missing outputs: metric scale and persistent person identity. We ask whether one up-to-scale foundation representation can support both through lightweight adaptation. Exact metric labels are scarce, but unlabeled in-the-wild video is abundant. We use people in curated web video to initialise the solution: a posed metric body and 2D keypoints give an approximate, closed-form scale pseudo-label. These pseudo-labels pretrain a Scale Readout, which is then fine-tuned together with a lightweight adapter using exact metric supervision from standard real-video training splits. At inference the Scale Readout predicts metric scale from foundation-model tokens, without the ruler or its teachers. For person identity, we probe the pretrained foundation model alone and find evidence that its intermediate query-key features encode person correspondence across frames. In most evaluated moving-person clips, a mid-layer token prefers that person over the vacated location and other people. A tiny projection reads this correspondence; together with metric pelvis motion and proposal confidence, it drives dustbin-aware Sinkhorn association of per-frame bodies. WildHSR combines both readouts to reconstruct metric cameras, scene and people from monocular video.
A frozen 3D foundation model (VGGT-Ω) gives up-to-scale cameras, depth and scene tokens. A mesh branch proposes people. Two light readouts add what the backbone lacks, and every window is predicted in one forward pass.


A posed body's metric torso against its 2D keypoints gives a closed-form scale label, σ = (f·e₃/e₂)/z̃. Agreeing labels pretrain a Scale Readout that then predicts σ from backbone tokens alone.

At layer 13, query-key features on a person in one frame match the same person later, even though the backbone was never trained to match moving people. A small projection turns this into an association cue.

When a person moves, a pelvis token's attention lands on their new position (5.0× uniform), not the place they left (1.2×).
Input on the left, WildHSR's metric reconstruction on the right. Crowds, running, stairs and long walks, all from a single moving camera.
@article{bright2026wildhsr,
title = {WildHSR: Metric Feed-Forward 4D People-Scene Reconstruction
from a 3D Foundation Model},
author = {Bright, Jerrin and Zelek, John},
journal = {arXiv preprint arXiv:2609.29106},
year = {2026}
}