AI evaluation & LLM judges
I test whether automated evaluators distinguish correct answers from incorrect ones. My commit-first judging study found that asking a judge to solve a task before scoring candidates removed incorrect acceptances on one coding task, but increased them on another when the judge’s own answer was wrong. A census of eight frameworks found that none of 24 applicable default configurations implemented commit-first judging.
At Evaluator Integrity, we also develop methods that can detect evaluator failures without knowing the correct answers, by finding contradictions between accepted outputs. Our latest MASK experiments in Inspect Evals, developed with AISI contributions, reproduced cases where the recorded score contradicted the judge’s own conclusion across GPT-4.1 and GPT-4o. These controlled results have undergone model-assisted review; human review is pending.