Reinforcement Learning
Train agents that learn by doing — from Markov decision processes to deep RL playing real environments.
About this course
Reinforcement learning is how AI learns behaviour: game-playing systems, robotics, recommendation, and the RLHF techniques behind modern LLMs. This course builds RL from first principles — states, actions, rewards, value functions — then moves into deep RL: DQN, policy gradients and actor-critic methods, trained in simulated environments you can watch learn. It closes with how RLHF aligns large language models.
Who it's for: Practitioners comfortable with deep learning basics who want the most distinctive specialisation in ML.
What you'll learn
Syllabus
Foundations
Weeks 1–2- MDPs, rewards and returns
- Value functions
- Q-learning from scratch
Deep RL
Weeks 3–4- DQN
- Experience replay
- Watching agents learn
Policy methods
Weeks 5–6- Policy gradients
- Actor-critic
- PPO in practice
Capstone
Weeks 7–8- Train an agent in a chosen environment
- RLHF and LLMs
- Results defence
Tools you'll use
Certificate
Finish the course and its capstone project to earn a FuturAIse Academy Certificate of Completion — verifiable online, with the projects to back it up.
