WangJ. x 3

bookRéférences 3

Vx2text: End-to-end learning of video-based text generation from multimodal inputs

We present Vx2Text, a framework for text generation from multimodal inputs consisting of video plus ...

2026-01-20 00:00:00

WangJ.LinX.BertasiusG.ChangS.F.

A cry for help: Early detection of brain injury in newborns

… base model (encoder) is a 76M-parameter convolutional neural network that has demonstra...

2023-11-06 20:02:46

OnuC.C.LatremouilleS.GorinA.WangJ.

Emu: Enhancing image generation models using photogenic needles in a haystack

Training text-to-image models with web scale image-text pairs enables the generation of a wide range...

WangJ.DaiX.HouJ.MaC.Y.TsaiS.WangR.

Mots-clés associés