Autoraters 101

A short course on evaluating generative AI with LLMs as judges — so you can stop guessing whether a model is good and start measuring it. Three videos, a hands-on demo, and the Python templates to run it yourself.

Course 4 is coming

Running the judge over a real test set, catching drift, and reporting agreement you can defend in a review. Get it in your inbox when it lands.

Subscribe — it's free