Reinforcement Learning
Agents that learn by doing. The foundation of RLHF.
Цена: 8 781 ₽
Длительность: 11 ч
Автор: John Jackson
Программа курса
- MDPs, States, Actions & Rewards
- Dynamic Programming — Policy Iteration & Value Iteration
- Monte Carlo Methods — Learning from Complete Episodes
- Temporal Difference — Q-Learning & SARSA
- Deep Q-Networks (DQN)
- Policy Gradient — REINFORCE from Scratch
- Actor-Critic — A2C and A3C
- Proximal Policy Optimization (PPO)
- Reward Modeling & RLHF
- Multi-Agent RL
- Sim-to-Real Transfer
- RL for Games — AlphaZero, MuZero, and the LLM-Reasoning Era
- Итоговое задание