Sitemap

Review — Flamingo: A Visual Language Model for Few-Shot Learning

6 min readMar 14, 2023

--

Press enter or click to view image in full size
A Family of Flamingo Models (Figure from National Geographics, Photograph by Klaus Nigge, Nat Geo Image Collection)

Flamingo: A Visual Language Model for Few-Shot Learning,
Flamingo, by DeepMind
2022 NeurIPS, Over 250 Citations (

@ Medium)
Image-Text Foundation Model, Vision Language Model, Visual Language Model, VLM

3.1. Visual/Vision/Video Language Model (VLM)20172021 [CLIP] [VinVL] [ALIGN] [VirTex] [ALBEF] [Conceptual 12M (CC12M) 2022 [FILIP] [Wukong] [LiT]
My Other Previous Paper Readings Are Also Over Here

Press enter or click to view image in full size
Selected examples of inputs and outputs obtained from Flamingo-80B. Flamingo can rapidly adapt to various image/video understanding tasks with few-shot prompting (top). Out of the box, Flamingo is also capable of multi-image visual dialogue (bottom).
Press enter or click to view image in full size
Flamingo architecture overview. Flamingo is a family of visual language models (VLMs) that take as input visual data interleaved with text and produce free-form text as output.
Press enter or click to view image in full size
The Perceiver Resampler Module.
Press enter or click to view image in full size
GATED XATTN-DENSE layers.
Press enter or click to view image in full size
Flamingo results overview.
Press enter or click to view image in full size
Comparison to the state of the art.

Left Figure & Table: A single Flamingo model reaches the state of the art on a wide array of image (I) and video (V) understanding tasks with few-shot learning, significantly outperforming previous best zero- and few-shot methods with as few as four examples.

Right Figure: The larger the model, the better the few-shot performance, similar to GPT-3. The performance also improves with the number of shots.

Press enter or click to view image in full size
Comparison to SotA when fine-tuning Flamingo.

By fine-tuning the model on a short schedule with a small learning rate by additionally unfreezing the vision backbone to accommodate a higher input resolution. Results are improved over the previously presented in-context few-shot learning results, setting a new state of the art on five additional tasks: VQAv2, VATEX, VizWiz, MSRVTTQA, and HatefulMemes.

Press enter or click to view image in full size
Ablation studies. Each row should be compared to the baseline Flamingo run (top row).

--

--

Sik-Ho Tsang
Sik-Ho Tsang

Written by Sik-Ho Tsang

PhD, Researcher. I share what I learn. :) Linktree: https://linktr.ee/shtsang for Twitter, LinkedIn, etc.