Bhavan Jasani scite author profile

We present DocFormer -a multi-modal transformer based architecture for the task of Visual Document Understanding (VDU). VDU is a challenging problem which aims to understand documents in their varied formats (forms, receipts etc.) and layouts. In addition, DocFormer is pre-trained in an unsupervised fashion using carefully designed tasks which encourage multi-modal interaction. DocFormer uses text, vision and spatial features and combines them using a novel multi-modal self-attention layer. DocFormer also shares learned spatial embeddings across modalities which makes it easy for the model to correlate text to visual tokens and vice versa. DocFormer is evaluated on 4 different datasets each with strong baselines. DocFormer achieves state-of-the-art results on all of them, sometimes beating models 4x its size (in no. of parameters).

show abstract

Are we Asking the Right Questions in MovieQA?

Jasani

Girdhar

Ramanan

2019

View full text Add to dashboard Cite

Word2VecQuestion OUR MODEL Uses a single layer perceptron + well tuned Word2Vec trained on movie plots and doesn't use any other information like videos or subtitles Plot Subtitle Video Good Word2Vec is able to capture enough semantics to answer half the questions in dataset Answer choices Figure 1: Answering questions about movies without watching any movies. The MovieQA task is: Given a question and multiple answer choices, find the correct answer by using the context provided in the corresponding videos and subtitles. Prior works use deep networks to incorporate information from videos and subtitles to do this task. We show a much simpler model that achieves state of the art performance, without using any video or subtitles context. Our model uses a well-tuned word embedding trained in an unsupervised manner on Wikipedia movie plots (movie summaries), and is able to answer about half of the questions in the dataset by just looking at the questions and choices. AbstractJoint vision and language tasks like visual question answering are fascinating because they explore high-level understanding, but at the same time, can be more prone to language biases. In this paper, we explore the biases in the MovieQA dataset and propose a strikingly simple model which can exploit them. We find that using the right word embedding is of utmost importance. By using an appropriately-trained word embedding, about half the Question-Answers (QAs) can be answered by looking at the questions and answers alone, completely ignoring narrative context from video clips, subtitles, and movie scripts. Com-pared to the best published papers on the leaderboard, our simple question+answer only model improves accuracy by 5% for video + subtitle category, 5% for subtitle, 15% for DVS and 6% higher for scripts.

show abstract

Data-path unrolling with logic folding for area-time-efficient FPGA-based FAST corner detector

Lam

Lim

et al. 2017

J Real-Time Image Proc

View full text Add to dashboard Cite

Stream-Based ORB Feature Extractor with Dynamic Power Optimization

Tran

Pham

Lam

et al. 2018

View full text Add to dashboard Cite

Threshold-Guided Design and Optimization for Harris Corner Detector Architecture

Jasani

Lam

Meher

et al. 2018

IEEE Trans. Circuits Syst. Video Technol.

View full text Add to dashboard Cite

YORO - Lightweight End to End Visual Grounding

Appalaraju²,

Jasani³

et al. 2023

View full text Add to dashboard Cite

Area-Time Efficient FAST Corner Detector Using Data-Path Transposition

Lam

Lim

et al. 2018

IEEE Trans. Circuits Syst. II

View full text Add to dashboard Cite

scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.

Contact Info

hi@scite.ai

10624 S. Eastern Ave., Ste. A-614

Henderson, NV 89052, USA

Blog Terms and Conditions API Terms Privacy Policy Contact Cookie Preferences Do Not Sell or Share My Personal Information

Made with 💙 for researchers

Part of the Research Solutions Family.

Bhavan Jasani

DocFormer: End-to-End Transformer for Document Understanding

DocFormer: End-to-End Transformer for Document Understanding

Are we Asking the Right Questions in MovieQA?

Data-path unrolling with logic folding for area-time-efficient FPGA-based FAST corner detector

Stream-Based ORB Feature Extractor with Dynamic Power Optimization

Threshold-Guided Design and Optimization for Harris Corner Detector Architecture

YORO - Lightweight End to End Visual Grounding

Area-Time Efficient FAST Corner Detector Using Data-Path Transposition

Contact Info

Product

Resources

About