Judge prompts.
Inspect the VEX-Bench judge prompts for GPT-5.2, Gemini 3.1 Pro, and Claude Opus 4.7, alongside their agreement with human annotators.
Judge agreement.
Pairwise ordinal α between human annotators and model judges.
Scroll sideways to compare all five raters
Tentative agreement ≥ .667
Reliable agreement ≥ .800
How agreement is measured
- Observed Do
- Two raters assign different scores to the same article.
- Random De
- A chance baseline created by randomly matching scores from different articles.
- Agreement α
- Closer to 1 means stronger agreement between two raters.
Krippendorff’s α is widely used in content analysis, human coding, media research, and communication research (Krippendorff, 2004; 2018).