Leaderboard | FaithGPT Christian AI Benchmark
Explore Leaderboard from FaithGPT's versioned Christian AI benchmark, including reviewed model scores, categories, pricing, reports, and historical trends.
FaithGPT brings together Bible study, Scripture insights, prayer journaling, Live Notes, DoctrineGuard theological review, and Christian AI tools in one Scripture-centered workspace.
Frequently asked questions
Which AI model ranks first in the selected Christian AI benchmark?
FaithGPT Ultra leads this preserved release with 94.3 out of 100 across 269 questions. Change the release selector to see the leader published with another test suite.
What does an overall score actually mean?
The score is a weighted result out of 100 across biblical accuracy, theological fidelity, pastoral quality, citation quality, instruction following, safety, and answer quality. It is not a percentage of questions marked simply right or wrong. Open a model profile or category view to inspect the shape behind the aggregate.
Why are FaithGPT and FaithGPT Ultra separate?
They are separate tested products with separate stored answers. In this release, FaithGPT scores 86.0 while FaithGPT Ultra scores 94.3. Ultra's result is never transferred to the base product.
What is the difference between list price and recorded benchmark spend?
List price per 100 answers standardizes model comparison using published input and output token rates. Recorded benchmark spend is the reconciled cash cost of paid releases and may include batch discounts, retries, judging, and other measured execution. Local-only, imported, dry-run, and carried answers add no new paid execution spend.
Which model is least expensive in this release?
DeepSeek: DeepSeek V4 Flash 0731 has the lowest positive standardized estimate at $0.04 per 100 answers. Lowest price is not the same as best value; the Value view compares price with measured quality.
How are answers judged without favoring a lab?
Candidate answers are anonymous to judges. Model, provider, cost, rank, and competing answers are withheld. Fresh answers use the selected release's configured blind judge protocol. Evaluation prompts treat the question, rubric, and candidate answer as untrusted quoted material and require numeric structured output.
What does carried forward mean?
An unchanged answer can be reused on an expanded question grid without paying to regenerate it. It keeps its stored answer, original judge provenance, and canonical hybrid score. Fresh questions use the current release protocol. The release must still have a validated answer for every model-question cell before it can rank models.
How does the benchmark cover Islam and other religions?
This selected release includes the stored World religions & discernment category. The Questions view shows only prompts approved for public display, so it does not pretend that the public examples are the entire private scoring grid. Coverage varies by release, and older snapshots keep their own suite definition.
Can I compare older benchmark versions?
Yes. The suite and release selectors open preserved route-backed snapshots rather than relabeling current data. 4 addressable release records are available in public history, each with its own roster, question count, scores, and report when published.
Helpful Scripture resources
Contact & support
Email the FaithGPT team at hello@faithgpt.io or visit the contact page .
Explore FaithGPT