arXiv 2026 · 2609.29106

WildHSRMetric Feed-Forward 4D People-Scene Reconstruction from a 3D Foundation Model

Jerrin Bright  ·  John Zelek

Vision and Image Processing Lab, University of Waterloo, Canada

One moving camera in, one metric world out. Inset: a hand-held video of a crowded street (3DPW). Around it: WildHSR's reconstruction of the same 10 seconds, with 14 persistent people and their recent paths, the person-free scene, and the recording camera, on a 1 m grid. Everything is predicted feed-forward.

Explore the reconstructions

Drag to orbit, scroll or pinch to zoom, right-drag to pan. Every scene is WildHSR's own prediction at its own metric scale, on a 1 m grid. Switch to Human3R to compare where it puts the same people.

Loading…camera to nearest person: n/a
drag to orbit · 1 m grid
Loading scene…
0.0 s

Abstract

3D foundation models recover video cameras and geometry in one forward pass, but some of the strongest are up to scale. Joint people-scene reconstruction then requires two missing outputs: metric scale and persistent person identity. We ask whether one up-to-scale foundation representation can support both through lightweight adaptation. Exact metric labels are scarce, but unlabeled in-the-wild video is abundant. We use people in curated web video to initialise the solution: a posed metric body and 2D keypoints give an approximate, closed-form scale pseudo-label. These pseudo-labels pretrain a Scale Readout, which is then fine-tuned together with a lightweight adapter using exact metric supervision from standard real-video training splits. At inference the Scale Readout predicts metric scale from foundation-model tokens, without the ruler or its teachers. For person identity, we probe the pretrained foundation model alone and find evidence that its intermediate query-key features encode person correspondence across frames. In most evaluated moving-person clips, a mid-layer token prefers that person over the vacated location and other people. A tiny projection reads this correspondence; together with metric pelvis motion and proposal confidence, it drives dustbin-aware Sinkhorn association of per-frame bodies. WildHSR combines both readouts to reconstruct metric cameras, scene and people from monocular video.

How it works

A frozen 3D foundation model (VGGT-Ω) gives up-to-scale cameras, depth and scene tokens. A mesh branch proposes people. Two light readouts add what the backbone lacks, and every window is predicted in one forward pass.

WildHSR architecture: 3D foundation model with Scale Readout, mesh branch with Pelvis Readout and cross-attention, Sim(3) composition into one metric world
Metric scale

People as rulers

Body ruler: backbone depth, 2D torso length in pixels and metric torso length give a closed-form scale

A posed body's metric torso against its 2D keypoints gives a closed-form scale label, σ = (f·e₃/e₂)/z̃. Agreeing labels pretrain a Scale Readout that then predicts σ from backbone tokens alone.

Identity

Correspondence already inside

Person-patch matching matrix at layer 13 showing same-person blocks

At layer 13, query-key features on a person in one frame match the same person later, even though the backbone was never trained to match moving people. A small projection turns this into an association cue.

Probe

It follows the person, not the pixel

Attention of a pelvis token at the query frame and 20 frames later, following the moving person

When a person moves, a pelvis token's attention lands on their new position (5.0× uniform), not the place they left (1.2×).

Results on real video

Input on the left, WildHSR's metric reconstruction on the right. Crowds, running, stairs and long walks, all from a single moving camera.

BibTeX

@article{bright2026wildhsr,
  title   = {WildHSR: Metric Feed-Forward 4D People-Scene Reconstruction
             from a 3D Foundation Model},
  author  = {Bright, Jerrin and Zelek, John},
  journal = {arXiv preprint arXiv:2609.29106},
  year    = {2026}
}