Benchmark results.
5,880 articles7 methods × 7 models2 tasks
Key Finding 1: Different tasks induce distinct verification-complexity profiles.
Rewrite starts from a real article; fabrication starts without one.
Fabrication and rewrite
D1–D5 scores by method, averaged over valid, non-refused outputs.
Key Finding 2: Success-based evaluations underestimate verification complexity.
Both tasks pooled (840 conditions per method); D1–D5 use scorable outputs.
Bars show VEX (%). SR = elicitation success; NR = non-refusal.
Key Finding 3: LLMs, methods, and their interactions exhibit distinct integrated risk.
- ISC leads method averages; DeepSeek V4 Pro leads model averages.
- Grok × DisinfoCap leads the pairwise mean rank.
Ranks across 62 settings
31 dimension combinations × two tasks. Lower rank means higher risk. “#1” counts first-place finishes.
- ISC ranks first among methods in 61/62 settings.
- Grok × DisinfoCap has the strongest pairwise mean.
- Grok × PoisonedRAG finishes first in 15/62.
Full results by model
Fabrication
Rewrite
Key Finding 4: Fabrication combines real institutions with fabricated individuals.
- Fabrication often names real institutions but invents people and contacts.
- These are agent verdicts about entity mentions, not verified facts about whole articles.
Entity results · fabrication
Share of each entity type among real or fabricated mentions.
Percent judged real. Counts show real / total mentions.
Paper entity results. “Real” and “fabricated” are agent verdicts; unverifiable mentions are excluded.
Key Finding 5: Generation is cheaper than verification.
- Generating an article costs 1.45¢ on average; checking it with an agent costs 17.86¢—12.4× as much.
- For Grok, the ratio reaches 169×.
Costs in US cents per article. Ratio = verification / generation.
Paper cost estimates. US cents per valid, non-refused article; LLM and search costs are estimated.
Domain diagnostics.
Verification complexity, model yield and entity realism across the six benchmark domains.