Intermediate · $99
RLHF and Model Evaluation
RLHF and evaluation work pays more than labeling because the judgement is harder to verify. This course covers how preference data is produced, how comparisons are justified, and how safety and harm categories are applied in practice.
Join the enrolment list
Enrolment is not open yet. Leave your details and we will email you when it opens, at $99.
What to expect
- Around 95 exercises on real model output, each with the accepted judgement and the reasoning behind it
- Written justification drills where you rewrite a weak rationale into one a reviewer accepts
- A safety taxonomy you apply to borderline items rather than read about
- The heaviest writing load of the Intermediate courses, since the writing is what gets audited
What is included
- 95 exercises with annotated answers, across comparison, safety, and failure modes
- A justification template and the phrasing that gets rationales flagged
- A harm and safety taxonomy reference you can work from
- A calibration set for checking your own consistency before a shift
Before you start: Comfortable reading a rubric and applying it consistently. Able to write clear English prose, since most of the assessed work is written.
Who this is for
- Annotators already passing assessments who want the better paid evaluation work
- People on labeling projects who keep getting turned down for preference work
- Anyone whose written rationales are being flagged without a clear reason
Skip it if
- Complete beginners, who should start with the free foundations course
- Anyone looking for machine learning theory or how to train a reward model themselves
What you will be able to do
- Produce preference comparisons that survive reviewer audit
- Write justifications that explain a ranking rather than restate it
- Apply harm and safety taxonomies to borderline model output
- Spot reward hacking, sycophancy, and confident hallucination in responses
- Hold your own standard steady across a batch of several hundred items
- Defend a judgement when a reviewer disagrees, and concede when they are right
Syllabus
7 modules, 32 lessons, about 9 hours of work in total.
- 01
What RLHF is, from the annotator's side
1 hrWhere your comparisons go, how they train a reward model, and why consistency matters more than taste.
- The pipeline from your comparison to a trained reward model
- Why an inconsistent annotator is worse than a strict one
- Preference data, demonstration data, and critique data, and who pays most for each
- The vocabulary used in evaluation project briefs
- 02
Ranking model outputs
1 hr 20 minComparison across several axes at once: helpfulness, correctness, and instruction following.
- The standard axes and the order to apply them in
- Correctness against helpfulness, when the two conflict
- Following the instruction versus answering the better question
- Length, formatting, and the presentation bias to resist
- Exercise set: 30 comparisons with annotated answers
- 03
Writing usable justifications
1 hr 25 minThe difference between a rationale a reviewer accepts and one that gets flagged.
- What an auditor checks a rationale against
- Naming the deciding property, with evidence quoted from the response
- Phrasing that gets a rationale flagged as unsupported
- Rewriting drill: 20 weak rationales into accepted ones
- Writing to a limit when the deciding reason is complicated
- 04
Safety, harm, and refusals
1 hr 30 minApplying policy taxonomies, and judging a model that refuses too much as carefully as one that refuses too little.
- How a harm taxonomy is structured and how to apply one you did not write
- Judging a refusal: appropriate, excessive, or missing
- Dual use requests and the reasoning that separates them
- Borderline set: 25 items where reasonable annotators disagree
- Escalating an item instead of forcing a judgement
- 05
Failure modes in model output
1 hr 20 minHallucination, sycophancy, and reward hacking, with annotated examples.
- Confident hallucination and how to check a claim quickly
- Sycophancy: agreeing with the user against the evidence
- Reward hacking, and why the output that looks best often is not
- Formatting that performs well and says nothing
- Exercise set: 20 responses with the failure identified
- 06
Calibration and disagreement
1 hr 15 minWorking with agreement scores between annotators and handling reviewer pushback.
- How agreement between annotators is measured and reported to you
- Detecting your own drift across a long batch
- Handling reviewer pushback: when to defend and when to concede
- Recalibrating after a guideline change without redoing everything
- 07
Working a real evaluation queue
1 hr 10 minPace, note keeping, and the habits that keep quality scores stable over weeks.
- Structuring a shift so quality does not fall in the last hour
- Keeping notes that make you consistent with yourself next week
- Reading a quality report and choosing the one thing to change
- Moving from a general queue onto a specialist project
Roles this prepares you for
Free reading first
These guides are free and cover some of the same ground. Read them before paying for anything.
Questions about this course
Do I need a machine learning background?
+
No. The course explains where your judgements go so the work makes sense, but the skills being taught are reading, judgement, and writing. There is no mathematics and no code.
Does evaluation work really pay more than labeling?
+
Usually, because the judgement is harder to verify and harder to replace. The salary guide on this site has current ranges by task type, and it is free to read before you buy.
How much writing is involved?
+
A lot. Written justifications are what reviewers audit, so roughly a third of the course is writing drills. If you dislike writing prose in English, this is the wrong specialism.
Before you buy
This is training you work through at your own pace, written by DataAnnotationJobs.org. It does not guarantee a job, an assessment pass, or any level of earnings, and it is not affiliated with or endorsed by any employer listed on this site. See terms and refunds.