Overview
Alignment is not a philosophy seminar here. It is measurement. You will build a benchmark that can actually fail, write a rubric a model can apply and then check it against independent graders, push a model toward and away from a set of values by prompt and by fine-tune, and prove the change was real rather than a story you told yourself afterwards. The discussion threads — reward misspecification, scalable oversight, hallucination and sleeper agents, and whose values we are aligning to in the first place — run alongside the work rather than instead of it.
Where This Class Leads
What you leave able to do
- configure inference hyperparameters
- parameterise a prompt template
- constrain model output to a schema
- characterise a models failure modes
- falsify a prompt with a benchmark
- verify a rubric against independent graders
- measure whether a fine tune changed behaviour
How you show it. The training set, a held-out set built from different sources, and base-versus-tuned scores on both.
Join Us
- Your seat in the class
- Lifetime access to materials
- Sponsors 1 scholarship seat for someone else
A minimum price keeps out spam. Need to pay less? Apply for a scholarship — no one is turned away for lack of funds.
🔒 Secure payment via Stripe
🎓 0 scholarship seats available. Not enrolling, but want to put someone through this class?
Apply for a scholarshipWant to schedule this class for your team?
Contact us: liz@themultiverse.school