Reinforcement Learning

Agents that learn by doing. The foundation of RLHF.

Цена: 8 781 ₽

Длительность: 11 ч

Автор: John Jackson

Программа курса

  1. MDPs, States, Actions & Rewards
  2. Dynamic Programming — Policy Iteration & Value Iteration
  3. Monte Carlo Methods — Learning from Complete Episodes
  4. Temporal Difference — Q-Learning & SARSA
  5. Deep Q-Networks (DQN)
  6. Policy Gradient — REINFORCE from Scratch
  7. Actor-Critic — A2C and A3C
  8. Proximal Policy Optimization (PPO)
  9. Reward Modeling & RLHF
  10. Multi-Agent RL
  11. Sim-to-Real Transfer
  12. RL for Games — AlphaZero, MuZero, and the LLM-Reasoning Era
  13. Итоговое задание