Judge prompts.

Inspect the VEX-Bench judge prompts for GPT-5.2, Gemini 3.1 Pro, and Claude Opus 4.7, alongside their agreement with human annotators.

Judge agreement.

Pairwise ordinal α between human annotators and model judges.

Scroll sideways to compare all five raters

Loading pairwise agreement…

Tentative agreement ≥ .667 Reliable agreement ≥ .800
How agreement is measured
αord =1− DoDe
Observed Do
Two raters assign different scores to the same article.
Random De
A chance baseline created by randomly matching scores from different articles.
Agreement α
Closer to 1 means stronger agreement between two raters.

Krippendorff’s α is widely used in content analysis, human coding, media research, and communication research (Krippendorff, 2004; 2018).

Evaluation prompts.

Loading checked-in VEX-Bench prompts…