Factuality Evaluation

A hop-stratified and random factuality audit of the large-scale GPT-5-mini run. Each extracted claim is checked against Wikipedia (when the subject exists there) and against curated web sources (the "frontier" check used when the subject is not on Wikipedia). Only subjects with a completed evaluation record are listed here.

2,010
evaluated subjects
1,277
on Wikipedia
733
frontier (web-only)
20,092
claims checked
17,224
supported (wiki+web)
254
refuted (wiki+web)
15,390
unverifiable (wiki+web)
Supported Refuted Unverifiable
Wikipedia verdicts
Web verdicts
Supported rate by hop
WikipediaWeb
Coverage
Mode:
SubjectHopMode Wikipedia (S/R/U)Web (S/R/U)