Multimodal AI
Models that see, hear, read, and reason across modalities.
Цена: 9 644 ₽
Длительность: 21 ч
Автор: John Jackson
Программа курса
- Vision Transformers and the Patch-Token Primitive
- CLIP and Contrastive Vision-Language Pretraining
- From CLIP to BLIP-2 — Q-Former as Modality Bridge
- Flamingo and Gated Cross-Attention for Few-Shot VLMs
- LLaVA and Visual Instruction Tuning
- Any-Resolution Vision: Patch-n'-Pack and NaFlex
- Open-Weight VLM Recipes: What Actually Matters
- LLaVA-OneVision: Single-Image, Multi-Image, Video in One Model
- Qwen-VL Family and Dynamic-FPS Video
- InternVL3: Native Multimodal Pretraining
- Chameleon and Early-Fusion Token-Only Multimodal Models
- Emu3: Next-Token Prediction for Image and Video Generation
- Transfusion: Autoregressive Text + Diffusion Image in One Transformer
- Show-o and Discrete-Diffusion Unified Models
- Janus-Pro: Decoupled Encoders for Unified Multimodal Models
- MIO and Any-to-Any Streaming Multimodal Models
- Video-Language Models: Temporal Tokens and Grounding
- Long-Video Understanding at Million-Token Context
- Audio-Language Models: the Whisper to Audio Flamingo 3 Arc
- Omni Models: Qwen2.5-Omni and the Thinker-Talker Split
- Embodied VLAs: RT-2, OpenVLA, π0, GR00T
- Document and Diagram Understanding
- ColPali and Vision-Native Document RAG
- Multimodal RAG and Cross-Modal Retrieval
- Multimodal Agents and Computer-Use (Capstone)
- Итоговое задание