Logo image
Sketch-a-Pose: Bridging Abstract Drawings and 6-DoF Camera Estimation
Preprint   Open access

Sketch-a-Pose: Bridging Abstract Drawings and 6-DoF Camera Estimation

Benjamin Paul Canini, Richard Bowden Prof and Yi-Zhe Song
07/08/2026

Abstract

Pose Estimation Gaussian Splatting Sketch Embodied AI Robotics

Camera pose estimation has traditionally relied on RGB images, assuming access to rich visual detail and low-level features. In this paper, we introduce a novel framework for estimating 6-DoF camera pose directly from sketch inputs. This new setting is both highly challenging—sketches are abstract, sparse, and stylistically diverse—and uniquely valuable, enabling pose estimation when photos are unavailable or when structural abstraction is preferable (e.g., design, ideation, cultural heritage).

We present Sketch-a-Pose (SAP), a lightweight framework that bridges the modality gap between drawings and geometry through three components: (i) a sketch adapter that maps CLIP embeddings into the pose-sensitive DINO space for robust initialization, (ii) a sparse Gaussian edge representation that captures sketch-aligned structure and supports efficient rendering, and (iii) an analysis-by-synthesis refinement loop guided by a structural LPIPS loss.

Across the Tanks and Temples and MIP-NeRF 360 datasets, SAP demonstrates that sketches alone can yield accurate camera pose recovery, often rivaling or even surpassing RGB-based baselines such as iNeRF, NeMO+VoGE, and 6DGS. Beyond raw numbers, SAP reveals a surprising insight: structural alignment from sketches can outperform photorealistic cues, suggesting a new paradigm of semantic-guided geometric vision.

pdf
Sketch_a_Pose_cam_ready23.46 MBDownloadView
Preprint (Author's original) Open Access CC BY V4.0

Metrics

2 File views/ downloads
1 Record Views

Details

Logo image

Usage Policy