Christian AI Benchmark v5-269q-openweight-2026-08-28-r1: how 34 models scored across 269 questions
TL;DR: FaithGPT Ultra leads the latest Christian AI Benchmark refresh at 94.3/100 across 34 models and 9,146 scored evaluations. Full leaderboard, all sixteen category winners, cost per 100 answers, and an honest account of what an incremental snapshot can and cannot prove.
We published the FaithGPT Christian AI Benchmark refresh, version v5-269q-openweight-2026-08-28-r1, on 2026-08-29. It covers 34 models against a 269-question bank spanning Scripture interpretation, doctrine, pastoral care, citation accuracy, safety, adversarial robustness, and expanded comparative-religion coverage. That produces 9,146 scored model-question rows against 9,146 planned rows.
FaithGPT Ultra topped the run at 94.3/100.
What kind of release this is
This is an incremental snapshot, not a full regeneration. Of the 9,146 rows, 5,402 are unchanged historical answers that retain their stored canonical hybrid scores. The remaining 3,744 answers were generated fresh for this run and scored under the current protocol: two isolated blind judgments from a local Claude Sonnet 5 pass, followed by one recomputed hybrid score. Judges saw no provider, model, cost, rank, prior-score, or competing-answer metadata.
So roughly 41 percent of the evidence in this snapshot is new, and 59 percent is carried forward. Both halves are scored, both are published, and neither was hand-adjusted. We say more about what that mix costs you in the limitations section below.
The run is named for its open-weight additions, but the ranked field is mixed. Proprietary and open-weight systems compete on the same 269 questions under the same scoring. Nothing in this dataset labels each row's license, so appearing in an "open-weight refresh" is not evidence that a given model is open weight. Read the model names, not the release name.
The leaderboard
Rank · Model · Overall (0-100) · Questions · Cost per 100 answers
1 · FaithGPT Ultra · 94.3 · 269 / 269 · $0.68
2 · GPT-5.6 Sol · 93.4 · 269 / 269 · $6.75
2 · Claude Opus 4.8 · 92.9 · 269 / 269 · 7.25
2 · Claude Fable 5 · 92.8 · 269 / 269 · 1.50
5 · gpt-5.5-pro · 92.8 · 269 / 269 · $40.50
6 · gpt-5.5 · 92.3 · 269 / 269 · $6.75
6 · Claude Sonnet 5 · 91.4 · 269 / 269 · $2.30
8 · GPT-5.6 Terra · 90.8 · 269 / 269 · $3.38
8 · GPT-5.6 Luna · 90.5 · 269 / 269 · .35
8 · Claude Sonnet 4.6 · 90.2 · 269 / 269 · $3.45
8 · MoonshotAI: Kimi K2.7 Code · 90.1 · 269 / 269 · $0.78
12 · gpt-5.4-mini · 89.7 · 269 / 269 · $0.55
12 · Z.ai: GLM 5.3 Flash · 88.8 · 269 / 269 · $0.06
12 · MoonshotAI: Kimi K3 · 88.5 · 269 / 269 · $3.45
12 · Qwen: Qwen3.8 2.4T A95B · 88.5 · 269 / 269 · .50
16 · Z.ai: GLM 5.2 · 88.3 · 269 / 269 · $0.93
16 · NVIDIA: Nemotron 3 Ultra · 87.1 · 269 / 269 · $0.52
18 · Qwen: Qwen3.8 Flash · 87.0 · 269 / 269 · $0.12
18 · DeepSeek: R1 0528 · 86.8 · 269 / 269 · $0.51
18 · Thinking Machines: Inkling Small · 86.4 · 269 / 269 · $0.31
21 · FaithGPT · 86.0 · 269 / 269 · $0.68
21 · MiniMax: MiniMax M3 · 85.2 · 269 / 269 · $0.29
21 · Qwen: Qwen3.8 27B · 85.0 · 269 / 269 · $0.57
21 · DeepSeek: DeepSeek V4 Pro 0813 · 84.8 · 269 / 269 · $0.50
25 · Qwen: Qwen3 235B A22B · 81.9 · 269 / 269 · $0.43
25 · Qwen: Qwen3.5 Plus 2026-02-15 · 81.7 · 269 / 269 · $0.35
25 · DeepSeek: DeepSeek V3 · 81.5 · 269 / 269 · $0.24
25 · Mistral Large · 81.0 · 269 / 269 · .50
25 · MoonshotAI: Kimi K2.6 · 80.8 · 269 / 269 · $0.94
30 · Google: Gemma 3 27B · 79.7 · 269 / 269 · $0.10
31 · Meta: Llama 4 Maverick · 77.8 · 269 / 269 · $0.19
31 · Qwen2.5 72B Instruct · 77.5 · 269 / 269 · $0.13
31 · Meta: Llama 3.1 70B Instruct · 76.5 · 269 / 269 · $0.14
31 · DeepSeek: DeepSeek V4 Flash 0731 · 76.0 · 269 / 269 · $0.04
Ranks use the 269 question-profile evaluations completed by every ranked model, with declared question-difficulty and category weights. Models share a rank when the paired 95 percent confidence interval between them includes zero, which is why the second position holds three entries and why gpt-5.5-pro sits at rank 5 despite matching Claude Fable 5's rounded score. Cost per 100 uses each row's published model-price estimate where available and excludes benchmark judging.
Two structural facts are worth stating plainly. Every model has complete 269-question coverage, so no ranking here depends on imputing a missing answer. And answers were scored exactly as returned: short responses and exhausted provider error text were not rewritten, retried into a better result, or upgraded before judging. If a model returned little, or returned a provider error string, that outcome is part of its published rating. We treat availability as part of what a model delivers.
Category winners
Category · Winner · Score
Adversarial robustness · FaithGPT Ultra · 92.0
AI and discernment · FaithGPT Ultra · 96.5
Apologetics · FaithGPT Ultra · 94.3
Biblical literacy · Claude Fable 5 · 94.3
Christian ethics · FaithGPT Ultra · 95.5
Citation traps · Claude Fable 5 · 95.5
Content creation · FaithGPT Ultra · 94.4
Contested ethics · Claude Fable 5 · 93.6
Denominational nuance · Claude Fable 5 · 94.1
Doctrine · GPT-5.6 Sol · 94.5
Identity integrity · FaithGPT Ultra · 97.0
Pastoral care · GPT-5.6 Sol · 96.3
Safety boundaries · FaithGPT Ultra · 96.3
Scripture interpretation · FaithGPT Ultra · 94.5
World religions & discernment · Claude Fable 5 · 91.6
Worldview fairness · Claude Fable 5 · 95.5
The sixteen categories split three ways: FaithGPT Ultra takes eight, Claude Fable 5 takes six, and GPT-5.6 Sol takes two. Ultra leads overall without leading everywhere, which is the intended reading of this table. A model can win the aggregate and still lose the specific task you care about.
Two details deserve attention. GPT-5.6 Sol wins pastoral care at 96.3 and doctrine at 94.5 while finishing second overall, so a church deploying for counselling-adjacent work should not assume the top-line ranking answers its question. And world religions and discernment has the lowest winning score of any category at 91.6, several points below what the same model achieved on worldview fairness and citation traps. Comparative-religion coverage was expanded for this release, and the field's ceiling there is visibly lower than on core Christian theology. That is the most useful weak signal in this snapshot.
Cost and value
On raw score per dollar, DeepSeek V4 Flash 0731 delivers the most: 76.0/100 at $0.04 per 100 answers. It is also last in the rankings, 18.3 points below the leader, so the value crown and the quality crown are not close to each other.
The more practical result sits a few rows up. GLM 5.3 Flash scores 88.8 at $0.06 per 100 answers, and Qwen3.8 Flash scores 87.0 at $0.12. Both land within seven points of the top score at under a fiftieth of gpt-5.5-pro's $40.50 per 100 answers, which itself scores 92.8. The premium tier is real but thin: the gap from $0.06 to $40.50 buys four points.
FaithGPT Ultra at $0.68 and FaithGPT at $0.68 share a price and sit 8.3 points apart, at ranks 1 and 21. Same cost line, very different outputs.
Limitations
The central caveat is the carried-forward mix. Carried scores and fresh scores each come from their own versioned stored protocol. Fresh answers went through the current two-pass blind judging described above; carried answers retain the canonical hybrid scores recorded when they were originally evaluated. This is therefore an incremental evidence snapshot, not a claim that all 9,146 answers were regenerated and rejudged contemporaneously under one protocol. Where a comparison hinges on a few points between adjacent models, check the question-level evidence before drawing conclusions.
Beyond that: these numbers describe this benchmark version on this question set. Rankings do not treat different question mixes as directly comparable. One aggregate score cannot capture every theological, factual, pastoral, and safety tradeoff at once, which is why the leaderboard, comparison view, category charts, and per-question evidence are all published together. And no benchmark replaces Scripture, pastors, or Christian community.
The live leaderboard always carries the most current version: faithgpt.io/benchmarks .
Helpful Scripture resources
Contact & support
Email the FaithGPT team at hello@faithgpt.io or visit the contact page .
Explore FaithGPT