Skip to main navigation Skip to search Skip to main content

Autonomous Robots for Professional Filming

  • Philip Lorimer

Student thesis: Doctoral ThesisDoctor of Engineering (EngD)

Abstract

The filmmaking industry continually seeks advanced tools to enhance creative expression, with dynamic camera movement being a cornerstone of cinematography. While innovations like free-roaming ground dollies provide flexibility, they often rely on manual operation or pre-programmed paths, limiting their adaptability and efficiency.

This thesis investigates the integration of machine learning and robotics to develop an autonomous cinematography pipeline for mobile ground robots. The autonomous pipeline is designed for a mobile ground-robot to autonomously plan moving trajectories and perform the capture, focusing on the traditional dolly-in shot as a proof of concept. Data-driven methods are leveraged for their adaptability and minimal reliance on explicit task modelling, offering scalability for broader applications.

The first contribution introduces a novel reinforcement learning (RL) application for executing the dolly-in shot, using a ground-based robot. This involves: (1) manually engineering image-based features representing the subject’s position \& size within the image frame, robot distance, orientation, and joint configurations; (2) training an RL agent on these features in simulation, achieving performance comparable to a traditional proportional-derivative (PD) controller; and (3) validating the RL agent through sim-to-real transfer, achieving strong predictive accuracy quantified by a Sim-vs-Real Correlation Coefficient (SRCC).

The work is extended to address the challenge of crafting effective reward functions for complex tasks. Learning from Demonstration (LfD) methods are applied to teach the dolly-in shot directly from expert demonstrations, enabling non-technical users to guide the robot. Benchmarks show that LfD improves learning efficiency, and real-world experiments confirm its practical applicability, maintaining strong positive SRCC values.

Finally, the perceptual pipeline transitions from manually engineered features to learned representations using Variational Autoencoders (VAEs). VAEs encode visual information into compact latent spaces, facilitating efficient feature extraction. Comparisons of learned versus manual features across RL and LfD frameworks reveal that VAEs enhance task success and learning efficiency, underscoring their scalability for real-world deployment.
Date of Award8 Oct 2025
Original languageEnglish
Awarding Institution
  • University of Bath
SupervisorWenbin Li (Supervisor) & Alan Hunter (Supervisor)

Keywords

  • autonomous intelligent system
  • autonomous cinematography
  • Robotic filmmaking
  • Mobile ground robots
  • Robotics
  • reinforcement learning
  • Imitation learning
  • Learning from Demonsration
  • learning (artificial intelligence)
  • artificial intelligence
  • Variational Autoencoders
  • Sim-to-Real Transfer
  • Visual Feature Extraction
  • Trajectory planning

Cite this

'