Skip to main navigation Skip to search Skip to main content

HD-EPIC: A Highly-Detailed Egocentric Video Dataset

  • Toby Perrett
  • , Ahmad Darkhalil
  • , Saptarshi Sinha
  • , Omar Emara
  • , Sam Pollard
  • , Kranti Parida
  • , Kaiting Liu
  • , Prajwal Gatti
  • , Siddhant Bansal
  • , Kevin Flanagan
  • , Jacob Chalk
  • , Zhifan Zhu
  • , Rhodri Guerrier
  • , Fahd Abdelazim
  • , Bin Zhu
  • , Davide Moltisanti
  • , Michael Wray
  • , Hazel Doughty
  • , Dima Damen
  • University of Bristol
  • Leiden University
  • Singapore Management University

Research output: Chapter or section in a book/report/conference proceedingChapter in a published conference proceeding

9   Link opens in a new tab Citations (SciVal)
229 Downloads (Pure)

Abstract

We present a validation dataset of newly-collected kitchen-based egocentric videos, manually annotated with highly detailed and interconnected ground-truth labels covering: recipe steps, fine-grained actions, ingredients with nutritional values, moving objects, and audio annotations. Importantly, all annotations are grounded in 3D through digital twinning of the scene, fixtures, object locations, and primed with gaze. Footage is collected from unscripted recordings in diverse home environments, making HD-EPIC the first dataset collected in-the-wild but with detailed annotations matching those in controlled lab environments. We show the potential of our highly-detailed annotations through a challenging VQA benchmark of 26K questions assessing the capability to recognise recipes, ingredients, nutrition, fine-grained actions, 3D perception, object motion, and gaze direction. The powerful long-context Gemini Pro only achieves 37.6% on this benchmark, showcasing its difficulty and highlighting shortcomings in current VLMs. We additionally assess action recognition, sound recognition, and long-term video-object segmentation on HD-EPIC. HD-EPIC is 41 hours of video in 9 kitchens with digital twins of 413 kitchen fixtures, capturing 69 recipes, 59K fine-grained actions, 51K audio events, 20K object movements and 37K object masks lifted to 3D. On average, we have 263 annotations per minute of our unscripted videos.

Original languageEnglish
Title of host publication2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
Place of PublicationU. S. A.
PublisherIEEE
Pages23901-23913
Number of pages13
ISBN (Electronic)9798331543648
ISBN (Print)9798331543655
DOIs
Publication statusPublished - 13 Aug 2025
EventThe IEEE/CVF Conference on Computer Vision and Pattern Recognition 2025 - Music City Center, Nashville, USA United States
Duration: 11 Jun 202515 Jun 2025
https://cvpr.thecvf.com/

Publication series

NameProceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition
PublisherIEEE
ISSN (Print)1063-6919

Conference

ConferenceThe IEEE/CVF Conference on Computer Vision and Pattern Recognition 2025
Abbreviated titleCVPR Nashville
Country/TerritoryUSA United States
CityNashville
Period11/06/2515/06/25
Internet address

Bibliographical note

Publisher Copyright:
©2025 IEEE.

ASJC Scopus subject areas

  • Software
  • Computer Vision and Pattern Recognition

Fingerprint

Dive into the research topics of 'HD-EPIC: A Highly-Detailed Egocentric Video Dataset'. Together they form a unique fingerprint.

Cite this