Weight-based multi-stream model for Multi-Modal Video Question Answering

There has been a tremendous success in individual domains of Computer Vision, Natural Language Processing, and Knowledge Representation. Videos are a rich source of information with the multi-modal data forms of images, audio, and optionally subtitles blended. Current research is going on in combini...

Full description

Saved in:

Bibliographic Details
Main Authors:	Mohith Rajesh, Sanjiv Sridhar, Chinmay Kulkarni, Aaditya Shah, Natarajan S
Format:	Article
Language:	English
Published:	LibraryPress@UF 2023-05-01
Series:	Proceedings of the International Florida Artificial Intelligence Research Society Conference
Subjects:	video question answering attention mechanism computer vision natural language processing neural networks pretrained models transfer learning weight-based multi-stream model tvqa dataset clip vision transformers deberta multimedia multi-modal
Online Access:	https://journals.flvc.org/FLAIRS/article/view/133306
Tags:	Add Tag No Tags, Be the first to tag this record!

Internet

https://journals.flvc.org/FLAIRS/article/view/133306

Weight-based multi-stream model for Multi-Modal Video Question Answering

Internet

Similar Items