Benchmark results.

5,880 articles7 methods × 7 models2 tasks

Loading benchmark results…

Key Finding 1: Different tasks induce distinct verification-complexity profiles.

Rewrite starts from a real article; fabrication starts without one.

Fabrication and rewrite

D1–D5 scores by method, averaged over valid, non-refused outputs.

    Key Finding 2: Success-based evaluations underestimate verification complexity.

    Both tasks pooled (840 conditions per method); D1–D5 use scorable outputs.

    Bars show VEX (%). SR = elicitation success; NR = non-refusal.

      Key Finding 3: LLMs, methods, and their interactions exhibit distinct integrated risk.

      • ISC leads method averages; DeepSeek V4 Pro leads model averages.
      • Grok × DisinfoCap leads the pairwise mean rank.

      Ranks across 62 settings

      31 dimension combinations × two tasks. Lower rank means higher risk. “#1” counts first-place finishes.

      Mean rank by method
      Mean rank by model
      Mean rank by model × method
      • ISC ranks first among methods in 61/62 settings.
      • Grok × DisinfoCap has the strongest pairwise mean.
      • Grok × PoisonedRAG finishes first in 15/62.
      Full results by model

      Fabrication

        Rewrite

          Key Finding 4: Fabrication combines real institutions with fabricated individuals.

          • Fabrication often names real institutions but invents people and contacts.
          • These are agent verdicts about entity mentions, not verified facts about whole articles.

          Entity results · fabrication

          Entity composition

          Share of each entity type among real or fabricated mentions.

          Recurring real institutions
          Recurring fabricated people
          Recurring real people
          Entity realism by model

          Percent judged real. Counts show real / total mentions.

          Recurring fabricated contacts
          Recurring real contacts

          Paper entity results. “Real” and “fabricated” are agent verdicts; unverifiable mentions are excluded.

          Key Finding 5: Generation is cheaper than verification.

          • Generating an article costs 1.45¢ on average; checking it with an agent costs 17.86¢—12.4× as much.
          • For Grok, the ratio reaches 169×.
          • Generation
          • Agent LLM
          • Web search

          Costs in US cents per article. Ratio = verification / generation.

          Paper cost estimates. US cents per valid, non-refused article; LLM and search costs are estimated.

          Domain diagnostics.

          Verification complexity, model yield and entity realism across the six benchmark domains.

          Verification complexity by domain
          Elicitation success by domain and model
          Entity realism by domain