Overview

Alignment is not a philosophy seminar here. It is measurement. You will build a benchmark that can actually fail, write a rubric a model can apply and then check it against independent graders, push a model toward and away from a set of values by prompt and by fine-tune, and prove the change was real rather than a story you told yourself afterwards. The discussion threads — reward misspecification, scalable oversight, hallucination and sleeper agents, and whose values we are aligning to in the first place — run alongside the work rather than instead of it.

Where This Class Leads

What you leave able to do

  • configure inference hyperparameters
  • parameterise a prompt template
  • constrain model output to a schema
  • characterise a models failure modes
  • falsify a prompt with a benchmark
  • verify a rubric against independent graders
  • measure whether a fine tune changed behaviour

How you show it. The training set, a held-out set built from different sources, and base-versus-tuned scores on both.

Join Us

Want to schedule this class for your team?

Contact us: liz@themultiverse.school