NarrativeBridge: Enhancing Video Captioning with Causal-Temporal Narrative

Asmar Nadeem; Faegheh Sardari; Robert Dawes; Syed Sameed Husain; Adrian Hilton; Armin Mustafa

doi:10.48550/arxiv.2406.06499

Back

NarrativeBridge: Enhancing Video Captioning with Causal-Temporal Narrative

Preprint

Open access

NarrativeBridge: Enhancing Video Captioning with Causal-Temporal Narrative

Asmar Nadeem, Faegheh Sardari, Robert Dawes, Syed Sameed Husain, Adrian Hilton and Armin Mustafa

Arxiv

10/06/2024

DOI: https://doi.org/10.48550/arxiv.2406.06499

Abstract

Existing video captioning benchmarks and models lack causal-temporal narrative, which is sequences of events linked through cause and effect, unfolding over time and driven by characters or agents. This lack of narrative restricts models' ability to generate text descriptions that capture the causal and temporal dynamics inherent in video content. To address this gap, we propose NarrativeBridge, an approach comprising of: (1) a novel Causal-Temporal Narrative (CTN) captions benchmark generated using a large language model and few-shot prompting, explicitly encoding cause-effect temporal relationships in video descriptions; and (2) a Cause-Effect Network (CEN) with separate encoders for capturing cause and effect dynamics, enabling effective learning and generation of captions with causal-temporal narrative. Extensive experiments demonstrate that CEN significantly outperforms state-of-the-art models in articulating the causal and temporal aspects of video content: 17.88 and 17.44 CIDEr on the MSVD-CTN and MSRVTT-CTN datasets, respectively. Cross-dataset evaluations further showcase CEN's strong generalization capabilities. The proposed framework understands and generates nuanced text descriptions with intricate causal-temporal narrative structures present in videos, addressing a critical limitation in video captioning. For project details, visit https://narrativebridge.github.io/.

Files and links (1)

url

https://doi.org/10.48550/arXiv.2406.06499View

Preprint (Author's original) Open CC BY-NC-ND V4.0

Metrics

1 Record Views

Details

Title: NarrativeBridge: Enhancing Video Captioning with Causal-Temporal Narrative
Creators: Asmar Nadeem (Corresponding Author) - University of Surrey, School of Computer Science & Electronic Engineering
Faegheh Sardari (Author) - University of Surrey, School of Computer Science & Electronic Engineering
Robert Dawes (Author) - BBC Research and Development, United Kingdom
Syed Sameed Husain (Author) - University of Surrey, Surrey Institute of Education
Adrian Hilton (Author) - University of Surrey, School of Computer Science & Electronic Engineering
Armin Mustafa (Corresponding Author) - University of Surrey, School of Computer Science & Electronic Engineering
Publisher: Arxiv
Grants: BBC Prosperity Partnership: Future Personalised Object-Based Media Experiences Delivered at Scale Anywhere, EP/V038087/1, Engineering and Physical Sciences Research Council (United Kingdom, Swindon) - EPSRC
Grant note: This research was partly supported by the British Broadcasting Corporation Research and Development (BBC R&D), Engineering and Physical Sciences Research Council (EPSRC) Grant EP/V038087/1 “BBC Prosperity Partnership: Future Personalised Object-Based Media Experiences Delivered at Scale Anywhere”.
Identifiers: 991123794802346
Academic Unit: Surrey Institute of Education; School of Computer Science & Electronic Engineering
Language: English
Resource Type: Preprint

NarrativeBridge: Enhancing Video Captioning with Causal-Temporal Narrative

Abstract

Files and links (1)

Metrics

Details

Usage Policy