Skip to main navigation Skip to search Skip to main content

3D Reconstruction and View Synthesis of 360-degree Images for Virtual Reality

  • Manuel Rey-Area

Student thesis: Doctoral ThesisPhD

Abstract

As humans, we perceive the world in three dimensions. When moving our heads around, how we see the world changes accordingly. Objects change in appearance based on their distance to our eyes: they appear larger as we move closer and smaller as we move away. We can even look behind and around corners, revealing occluded parts of the world that were previously unseen. These human-world interactions are crucial and must be present in virtual reality to achieve a feeling of immersion. 360° images are a popular format to bring immersive content into digital displays, encapsulating the viewer within a captured environment. However, a 360° image represents the world as seen from a fixed viewpoint and cannot be looked at from a different perspective, reducing the feeling of immersion. In this dissertation, I hence investigate methods to transform 360° images into 360° 3D photos that can be experienced interactively with six degrees of freedom, by both rotating and translating the viewpoint.

For an image to react when looked at from novel viewpoints, it needs to encode three-dimensional information. I propose a framework to estimate per-pixel depth from a single 360° image that leverages the power of monocular depth estimators trained on perspective images. The system operates as a three-step pipeline that projects the 360° image onto a series of perspective images for which depth is estimated, brings the individual depth maps into agreement using deformable alignment, and finally blends them seamlessly into a final 360° depth map.

A 360° image augmented with depth reacts accordingly to changes in viewpoint. Looking at it from new viewpoints reveals regions of the scene which are empty. Therefore, I propose two approaches to fill these regions plausibly with content that agrees both in appearance and geometry with the initial 360° image. The first approach combines a layered representation with an inpainting technique that directly operates on the sphere surface. This setting naturally handles wrap around and spherical distortions typical of 360° images in equirectangular format.

The second approach iteratively fills empty regions in space by choosing views inside an action sphere, inpainting the empty areas with a latent diffusion model, and finally incorporating the new content into the underlying representation to enforce multi-view consistency. The method uses a representation based on Gaussian point clouds that dynamically refines occluded areas as new views become available, allowing high-quality visual renderings at real-time speeds close to virtual reality requirements.
Date of Award25 Jun 2025
Original languageEnglish
Awarding Institution
  • University of Bath
SupervisorWenbin Li (Supervisor), Neill Campbell (Supervisor) & Christian Richardt (Supervisor)

Cite this

'