Abstract
The goal of portrait animation is to dynamically change the appearance of a source photo using a motion sequence while preserving the identity. Existing approaches suffer from limited generalisation across subjects, often resorting to ever-expanding video training datasets that yield marginal gains in motion fidelity. To address this problem, we propose a landmark-centric framework, leveraging the facial structure as an intrinsic, compact, and interpretable prior. Our solution exploits spatial relationships encoded in landmarks to reduce video training data requirements and enhance cross-identity generalisation. Built on StyleGAN3, our model seamlessly integrates attribute editing for flexible control. Furthermore, we introduce a latent refinement algorithm performing an inference-time optimisation guided by the proportional–derivative control theory. This algorithm refines latent codes to achieve precise facial motions for unseen identities, circumventing subject specific training. Evaluated on HDTF and MEAD, our approach, trained on limited data, matches the performance of state-of-the-art models reliant on large-scale video corpora. Our study draws on the potential of the structural prior conveyed by landmarks to achieve superior data efficiency and performance generalisation, demonstrating the merit of the proposed approach in portrait animation practice.
•Landmark-driven portrait animation with 468 facial landmarks.•Facial structural priors enable data-efficient portrait animation.•PD-guided latent refinement ensures precise facial motion.•StyleGAN3-based latent editing enables controllable animation.