ended3월 14일· 1 sources

VIOLET : End-to-End Video-Language Transformers with Masked Visual-tokenModeling

VIOLET: 마스크드 비주얼 토큰 모델링을 활용한 엔드투엔드 비디오-언어 Transformer

Why it matters

VIOLET proposes an end-to-end Video-Language Transformer that introduces Masked Visual-token Modeling (MVM) as a pre-training objective, enabling the model to learn joint video-text representations without relying on pre-extracted sparse features. The approach unifies video-language tasks under a single transformer architecture, achieving strong performance on downstream tasks such as video question answering and text-to-video retrieval.

1
Sources
+0
24h
Growth
191d
Active
VIOLETVideo-Language TransformerMasked Visual-token Modelingmultimodal learningend-to-end

Sources

Related Issues