PIXELBANKv9.1.0
Menu
Loading...
VLM 1: Vision-Language Models cover

VLM 1: Vision-Language Models — all 50 problems

Fifteen hands-on exercises covering the arithmetic that actually runs a vision-language model: patch tokenisation and image-token budgets, the CLIP and SigLIP contrastive objectives, retrieval and zero-shot transfer, bridge modules from linear projectors to the Perceiver Resampler, and the metrics for efficient adaptation and hallucination. Every problem takes its arrays as input - no vision encoder is ever run.