3D dog pose estimation is hindered by two limitations in ex- isting benchmarks: distributional bias toward canonical viewpoints, sta- ble poses, and narrow morphologies, and alignment degradation between images and 3D annotations. We instead synthesize a large-scale bal- anced benchmark (4.2M images with exact 3D ground truth) by decou- pling pose, shape, and texture into independent libraries (70K retargeted poses, 500 breed shapes, 1,000 appearance maps) and rendering directly to guarantee pixel-perfect alignment and broad view coverage. However, this balanced distribution exposes a challenge previously masked by bi- ased data: non-canonical combinations of pose, shape, and viewpoint create multimodal 3D ambiguities that deterministic regression cannot resolve. We therefore propose a flow-based probabilistic framework that models the conditional distribution of 3D poses given an image, pro- ducing multiple plausible hypotheses under ambiguity while converging to precise estimates for clear views. Experiments on StanfordExtra, An- imal3D, and our synthetic benchmark demonstrate that our method, trained solely on synthetic data, outperforms state-of-the-art approaches (79.9 vs. 33.2 PCK on the synthetic benchmark). Dataset available here.
Joo Young Choi*, Wonkwang Lee*, Ju-hyeong Seon, Gunhee Kim
Abstract
Data-centric AI
ECCV 2026