Rendered at 20:40:19 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
jerlendds 1 days ago [-]
VLM != LLM. Vision language models basically treat text tokens and image tokens the same. Post-training an LLM on images+text can improve its capabilities. Id recommend searching around the keyword VLM to find more resources on how multi-modal AI works.
transformers were first an image understanding technique, the text processing came later, it's all about the training data and gradient descent, and allegedly attention
laruss5 11 hours ago [-]
It's actually the other way round - the Transformer architecture was introduced for text (machine translation) in "Attention Is All You Need" (2017). Vision Transformers, which apply it to images, came three years later in 2020: https://arxiv.org/abs/2010.11929
verdverm 5 hours ago [-]
right, it was not text generation per-se (completion/contemporary understanding) that came first, vision was before that, translation before that
- https://huggingface.co/blog/vlms
- https://en.wikipedia.org/wiki/Multimodal_learning