2024
DOI: 10.1145/3654796
|View full text |Cite
|
Sign up to set email alerts
|

A Novel Pretrained General-purpose Vision Language Model for the Vietnamese Language

Dinh Anh Vu,
Quang Nhat Minh Pham,
Giang Son Tran

Abstract: Lying in the cross-section of computer vision and natural language processing, vision language models are capable of processing images and text at once. These models are helpful in various tasks: text generation from image and vice versa, image-text retrieval, or visual navigation. Besides building a model trained on a dataset for a task, people also study general-purpose models to utilize many datasets for multitasks. Their two primary applications are image captioning and visual question answering. For Engli… Show more

Help me understand this report

Search citation statements

Order By: Relevance

Paper Sections

Select...

Citation Types

0
0
0

Publication Types

Select...

Relationship

0
0

Authors

Journals

citations
Cited by 0 publications
references
References 33 publications
0
0
0
Order By: Relevance

No citations

Set email alert for when this publication receives citations?