Christian AI Benchmark v3-227q-2026-07-r3: how 24 leading models scored
TL;DR: GPT-5.6 Sol leads the latest Christian AI Benchmark at 94.1/100 across 24 models. Full results, category winners, and real per-answer cost inside.
We published the FaithGPT Christian AI Benchmark on 2026-07-13. 24 models were evaluated against a 227-question bank covering Scripture interpretation, doctrine, pastoral care, citation accuracy, and safety, producing 4,641 scored evaluations from 5,448 planned evaluation rows.
GPT-5.6 Sol topped this run with an overall score of 94.1/100 .
The leaderboard
Rank · Model · Overall (0-100) · Questions · Cost per 100 answers
1 · GPT-5.6 Sol · 94.1 · 227 / 227 · $6.75
1 · gpt-5.5-pro · 94.0 · 174 / 227 · $40.50
3 · Claude Opus 4.8 · 93.4 · 144 / 227 · 7.25
3 · gpt-5.5 · 93.2 · 173 / 227 · $6.75
3 · FaithGPT Ultra · 92.8 · 152 / 227 · $0.68
3 · Claude Fable 5 · 92.4 · 227 / 227 · 1.50
3 · Claude Sonnet 5 · 92.2 · 227 / 227 · $2.30
8 · GPT-5.6 Luna · 91.9 · 227 / 227 · .35
8 · GPT-5.6 Terra · 91.7 · 227 / 227 · $3.38
8 · Claude Sonnet 4.6 · 91.3 · 154 / 227 · $3.45
8 · MoonshotAI: Kimi K2.7 Code · 91.0 · 150 / 227 · $0.55
12 · gpt-5.4-mini · 90.5 · 143 / 227 · $0.55
13 · Z.ai: GLM 5.2 · 88.9 · 161 / 227 · $0.55
13 · FaithGPT · 88.8 · 120 / 227 · $0.68
13 · DeepSeek: R1 0528 · 88.1 · 134 / 227 · $0.55
— · Mistral Large · 82.7 · 220 / 227 · $0.55
— · Qwen: Qwen3.5 Plus 2026-02-15 · 82.6 · 227 / 227 · $0.55
— · Qwen: Qwen3 235B A22B · 82.3 · 226 / 227 · $0.55
— · DeepSeek: DeepSeek V3 · 82.2 · 227 / 227 · $0.55
— · MoonshotAI: Kimi K2.6 · 81.6 · 212 / 227 · $0.55
— · Google: Gemma 3 27B · 80.0 · 227 / 227 · $0.55
— · Meta: Llama 4 Maverick · 78.4 · 225 / 227 · $0.55
— · Meta: Llama 3.1 70B Instruct · 77.7 · 222 / 227 · $0.55
— · Qwen2.5 72B Instruct · 77.5 · 215 / 227 · $0.55
Ranks use the 102 question-profile evaluations completed by every ranked model, with declared question-difficulty and category weights. Models share a rank when the paired 95% confidence interval includes zero. Questions shows each model's full completed / expected total.
Cost per 100 uses each row's published model-price estimate when available and excludes benchmark judging. A $0.00 entry means the subscription CLI run added no metered API spend; it is not a claim that the commercial model is free.
Who wins each category
Category · Winner · Score
Adversarial robustness · Claude Fable 5 · 96.0
Apologetics · FaithGPT Ultra · 93.3
Biblical literacy · Claude Fable 5 · 94.6
Christian ethics · GPT-5.6 Sol · 95.0
Citation traps · Claude Fable 5 · 95.4
Content creation · GPT-5.6 Sol · 93.8
Denominational nuance · Claude Fable 5 · 94.5
Doctrine · FaithGPT Ultra · 94.7
Pastoral care · GPT-5.6 Sol · 96.3
Safety boundaries · gpt-5.5-pro · 95.6
Scripture interpretation · Claude Fable 5 · 94.7
Category scores average every model's answers within that category. A model can lead overall and still lose a category to a specialist.
Best value
On score per dollar, MoonshotAI: Kimi K2.7 Code delivered the most: 91.0/100 at $0.55 per 100 answers.
How to read these results
These numbers measure benchmark version v3-227q-2026-07-r3 on this question set. Rows without a rank are earlier completed results shown for transparency. The rank does not impute missing answers or treat different question mixes as directly comparable. No benchmark replaces Scripture, pastors, or Christian community.
The live leaderboard always carries the most current version: faithgpt.io/benchmarks .
Helpful Scripture resources
Contact & support
Email the FaithGPT team at hello@faithgpt.io or visit the contact page .
Explore FaithGPT