LLM-as-a-Judge
LLM-as-a-judge is an evaluation method in which a language model grades the output of another model, or of itself, against a rubric, a reference answer, or a competing output. The judge receives the input, the response under test, and grading instructions, then returns a score, a label, or a preference, often with a short written rationale. Unlike exact-match metrics or string overlap scores, it can assess open-ended qualities such as helpfulness, faithfulness to sources, or tone, but it inherits the biases of the judging model. In agentic systems, LLM judges score final answers, individual tool-use steps, and full trajectories, both in offline test suites and on sampled production traffic.
Published research on the method has documented systematic judge biases, including a preference for the answer shown first, for longer answers, and for outputs from the judge's own model family. Common mitigations are pairwise comparison with swapped order, narrow rubrics with one criterion per question, and a small human-labeled set used to measure how often the judge agrees with people before its scores are trusted. A judge verdict is a measurement with error, so it is tracked as a trend and disagreements are reviewed, rather than a single score being treated as ground truth.
Related terms
Related services: LLM development, Agentic AI consulting.
Ready to put agentic AI to work?
Book a free 45-minute consultation. We'll map one real process worth automating with production-grade AI.