A Pipeline for Creative Visual Storytelling

Lukin, Stephanie M.; Hobbs, Reginald L.; Voss, Clare R.

doi:10.18653/v1/w18-1503

Cited by 20 publications

(26 citation statements)

References 12 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…Moving from objects to actions, several tasks have been proposed to mimic more realistic settings where a higher degree of integration between modalities is required. One is visual storytelling (Huang et al, 2016;Gonzalez-Rico and Pineda, 2018;Lukin et al, 2018), where models have to understand the action depicted in each photo and their relations to generate a story. Similar abilities are required in the task of generating non-grounded, human-like questions about an image (Mostafazadeh et al, 2016;Jain et al, 2017), and in that of asking discriminative questions over pairs of similar scenes .…”

Section: Related Workmentioning

confidence: 99%

Be Different to Be Better! A Benchmark to Leverage the Complementarity of Language and Vision

Pezzelle¹,

Greco²,

Gandolfi³

et al. 2020

Findings of the Association for Computational Linguistics: EMNLP 2020

View full text Add to dashboard Cite

This paper introduces BD2BB, a novel language and vision benchmark that requires multimodal models combine complementary information from the two modalities. Recently, impressive progress has been made to develop universal multimodal encoders suitable for virtually any language and vision tasks. However, current approaches often require them to combine redundant information provided by language and vision. Inspired by real-life communicative contexts, we propose a novel task where either modality is necessary but not sufficient to make a correct prediction. To do so, we first build a dataset of images and corresponding sentences provided by human participants. Second, we evaluate state-of-the-art models and compare their performance against human speakers. We show that, while the task is relatively easy for humans, best-performing models struggle to achieve similar results.

show abstract

Section: Related Workmentioning

confidence: 99%

Be Different to Be Better! A Benchmark to Leverage the Complementarity of Language and Vision

Pezzelle¹,

Greco²,

Gandolfi³

et al. 2020

Findings of the Association for Computational Linguistics: EMNLP 2020

View full text Add to dashboard Cite

show abstract

“…First, all were publicly available and well-documented, ensuring easy replicability. Other existing visual storytelling models (Huang et al, 2016;Yu et al, 2017;Hsu et al, 2018;Lukin et al, 2018) would have required reimplementation. Doing so introduces the possibility of unintentionally crippling performance (e.g., when setting required but unreported parameters), which we wished to avoid.…”

Section: Methodsmentioning

confidence: 99%

The Steep Road to Happily Ever after: an Analysis of Current Visual Storytelling Models

Modi¹,

Parde²

2019

Proceedings of the Second Workshop on Shortcomings in Vision and Language

View full text Add to dashboard Cite

Visual storytelling is an intriguing and complex task that only recently entered the research arena. In this work, we survey relevant work to date, and conduct a thorough error analysis of three very recent approaches to visual storytelling. We categorize and provide examples of common types of errors, and identify key shortcomings in current work. Finally, we make recommendations for addressing these limitations in the future.

show abstract

“…First, all were publicly available and well-documented, ensuring easy replicability. Other existing visual storytelling models Yu et al, 2017;Hsu et al, 2018;Lukin et al, 2018) would have required reimplementation. Doing so introduces the possibility of unintentionally crippling performance (e.g., when setting required but unreported parameters), which we wished to avoid.…”

Section: Methodsmentioning

confidence: 99%

Proceedings of the Second Workshop on Shortcomings in Vision and Language

2019

View full text Add to dashboard Cite

Visual question answering (VQA) models have been shown to over-rely on linguistic biases in VQA datasets, answering questions "blindly" without considering visual context. Adversarial regularization (AdvReg) aims to address this issue via an adversary subnetwork that encourages the main model to learn a bias-free representation of the question. In this work, we investigate the strengths and shortcomings of AdvReg with the goal of better understanding how it affects inference in VQA models. Despite achieving a new stateof-the-art on VQA-CP, we find that AdvReg yields several undesirable side-effects, including unstable gradients and sharply reduced performance on in-domain examples. We demonstrate that gradual introduction of regularization during training helps to alleviate, but not completely solve, these issues. Through error analyses, we observe that AdvReg improves generalization to binary questions, but impairs performance on questions with heterogeneous answer distributions. Qualitatively, we also find that regularized models tend to over-rely on visual features, while ignoring important linguistic cues in the question. Our results suggest that AdvReg requires further refinement before it can be considered a viable bias mitigation technique for VQA.

show abstract

A Pipeline for Creative Visual Storytelling

Cited by 20 publications

References 12 publications

Be Different to Be Better! A Benchmark to Leverage the Complementarity of Language and Vision

Be Different to Be Better! A Benchmark to Leverage the Complementarity of Language and Vision

The Steep Road to Happily Ever after: an Analysis of Current Visual Storytelling Models

Proceedings of the Second Workshop on Shortcomings in Vision and Language

Contact Info

Product

Resources

About