Abstract
Ranging from augmented reality and visual effects to digital art, graphics applications increasingly demand photorealistic image editing that maintains physical consistency within a scene. Achieving such realism requires physically plausible modifications that respect the underlying structure and illumination of the visual world. This thesis focuses on intrinsic image decomposition, the task of recovering the physical factors that dominate scene appearance, such as reflectance and shading, from a single image.This decomposition problem is motivated by human perceptual abilities such as colour constancy: reflectance is considered an intrinsic and illumination-invariant property of surface materials, while shading accounts for the variations caused by lighting and geometry. Decomposing an image into multiple intrinsic layers supports a wide range of applications, including relighting, material editing, and, as proposed in this thesis, realistic object insertion.
This thesis introduces both theoretical advances and practical applications in scene-level decomposition, with particular emphasis on perceptual quality, generalisability, and practical applicability. Three primary contributions are presented. First, CRefNet, a hybrid transformer-convolutional architecture, improves spatial reflectance consistency under complex illumination by incorporating long-range dependencies and reflectance-specific training strategies. Second, IntrinsicDiffusion repurposes large-scale latent diffusion models for intrinsic image decomposition, offering a multimodal learning framework that supports training with heterogeneous annotations and improves generalisation to diverse real-world scenes. Finally, ProjectiveShading presents a method for shadow-aware object insertion, using dense shading information to simulate physically correct occlusion and illumination effects in virtual-real interactions.
Through these contributions, this thesis not only improves intrinsic image decomposition methods, but also demonstrates their broader potential for future work in physical and perceptual scene understanding, vision-based rendering, and AI-driven visual content creation.
| Date of Award | 22 Apr 2026 |
|---|---|
| Original language | English |
| Awarding Institution |
|
| Supervisor | Wenbin Li (Supervisor), Christian Richardt (Supervisor), Nanxuan Zhao (Supervisor) & Yongliang Yang (Supervisor) |
Cite this
- Standard