Transformers Deep Dive

The architecture that changed everything. Understand every layer.

Цена: 6 918 ₽

Длительность: 14 ч

Автор: John Jackson

Программа курса

  1. Why Transformers — The Problems with RNNs
  2. Self-Attention from Scratch
  3. Multi-Head Attention
  4. Positional Encoding — Sinusoidal, RoPE, ALiBi
  5. The Full Transformer — Encoder + Decoder
  6. BERT — Masked Language Modeling
  7. GPT — Causal Language Modeling
  8. T5, BART — Encoder-Decoder Models
  9. Vision Transformers (ViT)
  10. Audio Transformers — Whisper Architecture
  11. Mixture of Experts (MoE)
  12. KV Cache, Flash Attention & Inference Optimization
  13. Scaling Laws
  14. Build a Transformer from Scratch — The Capstone
  15. Attention Variants — Sliding Window, Sparse, Differential
  16. Speculative Decoding — Draft, Verify, Repeat
  17. Итоговое задание