Multimodal architectures that ground language in vision — contrastive pretraining, cross-attention fusion, and instruction tuning.