From Autocomplete to Assistant: RL for Beginners
Published 8/2026Created by Shashank Jain
MP4 | Video: h264, 1920x1080 | Audio: AAC, 44.1 KHz, 2 Ch
Level: Intermediate | Genre: eLearning | Language: English | Duration: 15 Lectures ( 2h 11m ) | Size: 1.7 GB
Learn how ᑕᕼᗩTGᑭT is really trained: RLHF, reward models, PPO, GRPO, DPO and reasoning models. No maths required.
What you'll learn
Explain how modern AI assistants are really trained, from pretraining to RLHF, and why a model that read the internet still can't answer your question
Describe what a reward model is and how it learns human taste from simple A/B comparisons, without anyone ever writing down a single score
Understand reinforcement learning from zero: policy, reward, baseline and advantage, explained with no maths and no prior RL background
Tell PPO, GRPO, RLOO and DPO apart, and explain in plain English what these algorithms actually disagree about
Explain why verifiable rewards (RLVR) produced today's reasoning models, and how thinking for longer makes answers better
Diagnose why AI models turn sycophantic, verbose or over-cautious, and trace each behaviour back to its real cause in the training data
Read published benchmark scores critically, understanding contamination, format sensitivity and why cross-lab comparisons mislead
Hold a fluent, informed conversation about LLM post-training, alignment, and where this fast-moving field is heading nextRequirements
No programming, maths or machine learning experience needed. If you have used an AI chatbot and wondered how it was built, you have everything required
No prior knowledge of reinforcement learning is assumed. We build it from scratch using everyday examples like learning to ride a bicycle
Just curiosity and about two hours. Every concept is introduced with a real-world analogy before any technical term appearsDescription
A machine has read almost everything humans have ever written. Ask it a simple question, and it cannot help you. Understanding why — and what we do about it — is the story of this course.
Every AI assistant you have ever used went through a second, far less understood stage of training after it finished reading the internet. That stage is where a raw text predictor becomes something you can actually talk to. It is where helpfulness, tone, refusals, reasoning and personality are decided. It is also where things go wrong: it is the reason models flatter you, pad their answers, and refuse perfectly harmless requests.
This course explains that stage from absolute zero, in plain English, with no mathematics and no prior knowledge of machine learning or reinforcement learning.
WHAT MAKES THIS COURSE DIFFERENT
Most material on this subject opens with equations. This one opens with a cook who has read every cookbook and never tasted food. Reinforcement learning is built up from learning to ride a bicycle. Reward models are explained through chess ratings. Over-optimisation is explained through schools teaching to the test. Every technical idea arrives only after you already feel the problem it solves.
The fifteen episodes tell one continuous story rather than fifteen disconnected lectures, and each one ends where the next begins.
WHAT YOU WILL UNDERSTAND
Why a model that read the internet still cannot answer your question
How instruction tuning turns a text predictor into an assistant, and where it hits a wall
Why comparing two answers beats writing one, and how that single insight lifted the ceiling on AI quality
What a reward model is, and how it learns human taste from comparisons alone
Reinforcement learning explained from nothing: policy, reward, baseline, advantage
What PPO, GRPO, RLOO and DPO actually disagree about
Why models must be held on a leash, and what happens when they are not
How verifiable rewards produced today's reasoning models
Where preference data really comes from, and how human biases become permanent model behaviour
Why AI turns sycophantic, verbose and over-cautious, and how to trace each behaviour to its cause
How to read published benchmark scores with appropriate scepticismWHO THIS IS FOR
Anyone who wants to understand modern AI properly rather than superficially: curious beginners, product managers and designers working alongside AI teams, engineers new to machine learning, students, writers and analysts.
By the end you will not just recognise these terms. You will be able to explain them, argue about them, and follow any conversation in this field with confidence.
Who this course is for
Anyone curious about how ᑕᕼᗩTGᑭT and similar AI assistants are actually built, with no technical background required
Product managers, designers, founders and analysts working alongside AI teams who want to follow the conversation properly
Software engineers new to machine learning who want the concepts behind RLHF before diving into papers or code
Students and career changers exploring AI who want a solid conceptual foundation rather than a maths-heavy course
Writers, journalists and researchers who need to discuss AI alignment and model behaviour accurately and confidently
Anyone who has noticed AI being sycophantic, wordy or needlessly cautious and wants to understand exactly why it happensHomepage
You do not have permission to view the full content of this post. Log in or register now.
You do not have permission to view the full content of this post. Log in or register now.