Decoupled Spatial Temporal Graphs for Generic Visual Grounding

Feng, Qianyu; Wei, Yunchao; Cheng, Ming–Ming; Yang, Yi

doi:10.48550/arxiv.2103.10191

Cited by 1 publication

(2 citation statements)

References 60 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…The objective of video REC is to localize the spatial-temporal tube according to the natural language query. Most of the previous works [19,22,43,45,62] can be divided into two categories, i.e., two-stage methods and one-stage methods. However, both kinds of methods require time-consuming post-processing steps, which hinders their practical applications.…”

Section: Related Workmentioning

confidence: 99%

“…Current video REC methods can be classified into two major categories: two-stage, proposal-driven methods and one-stage, proposal-free methods. For the two-stage methods [19,20,26,62], they extract potential spatio-temporal tubes and then align these candidates to the sentence to find the best matching one. The other stream of one-stage methods [7,13,43,45,59] fuses visual-text features and directly predicts bounding boxes densely at all spatial locations.…”

Section: Introductionmentioning

confidence: 99%

See 1 more Smart Citation

Video Referring Expression Comprehension via Transformer with Content-aware Query

Jiang¹,

Cao²,

Song³

et al. 2022

Preprint

View full text Add to dashboard Cite

Video Referring Expression Comprehension (REC) aims to localize a target object in video frames referred by the natural language expression. Recently, the Transformerbased methods have greatly boosted the performance limit. However, we argue that the current query design is suboptima and suffers from two drawbacks: 1) the slow training convergence process; 2) the lack of fine-grained alignment. To alleviate this, we aim to couple the pure learnable queries with the content information. Specifically, we set up a fixed number of learnable bounding boxes across the frame and the aligned region features are employed to provide fruitful clues. Besides, we explicitly link certain phrases in the sentence to the semantically relevant visual areas. To this end, we introduce two new datasets (i.e., VID-Entity and VidSTG-Entity) by augmenting the VID-Sentence and VidSTG datasets with the explicitly referred words in the whole sentence, respectively. Benefiting from this, we conduct the fine-grained cross-modal alignment at the region-phrase level, which ensures more detailed feature representations. Incorporating these two designs, our proposed model (dubbed as ContFormer) achieves the stateof-the-art performance on widely benchmarked datasets. For example on VID-Entity dataset, compared to the previous SOTA, ContFormer achieves 8.75% absolute improvement on Accu.@0.6. The dataset, code and models are available at https://github.com/mengcaopku/ ContFormer.

show abstract