DeepMaven: Deep question answering on long-distance movie/TV show videos with multimedia knowledge extraction and synthesis

Yi Fung; Han Wang; Tong Wang; Ali Kebarighotbi; Mohit Bansal; Heng Ji; Prem Natarajan

Publication

DeepMaven: Deep question answering on long-distance movie/TV show videos with multimedia knowledge extraction and synthesis

By Yi Fung, Han Wang, Tong Wang, Ali Kebarighotbi, Mohit Bansal, Heng Ji, Prem Natarajan

2023

Download Copy BibTeX

Share

Download

Copy BibTeX

Share

Long video content understanding poses a challenging set of research questions as it involves long-distance, cross-media reasoning and knowledge awareness. In this paper, we present a new benchmark for this problem domain, targeting the task of deep movie/TV question answering (QA) beyond previous work’s focus on simple plot summary and short video moment settings. We define several baselines based on direct retrieval of relevant context for long-distance movie QA. Observing that real-world QAs may require higher-order multi-hop inferences, we further propose a novel framework, called the DEEPMAVEN, which extracts events, entities, and relations from the rich multimedia content in long videos to pre-construct movie knowledge graphs (movieKGs), and at the time of QA inference, complements general semantics with structured knowledge for more effective information retrieval and knowledge reasoning. We also introduce our recently collected DeepMovieQA dataset, including 1,000 long-form QA pairs from 41 hours of videos, to serve as a new and useful resource for future work. Empirical results show the DeepMaven performs competitively for both the new DeepMovieQA and the pre-existing MovieQA dataset.

DeepMaven: Deep question answering on long-distance movie/TV show videos with multimedia knowledge extraction and synthesis

Latest news

Work with us