ExploringEvidence: Medium50/100

Pushpay: Agentic AI search on Amazon Bedrock with custom GenAI evaluation, reaching 95% accuracy and reducing time-to-insight to <4 seconds

Pushpay built an agentic AI search feature for church and ministry staff to ask natural-language questions over community data and to get actionable insights in real time. The solution was built on Amazon Bedrock and paired with a custom generative AI evaluation framework, including semantic selection of filters, a golden dataset, LLM-as-a-judge evaluation, and domain-level dashboards for continuous improvement.

Organization
Pushpay
Industry
Other
Location
New Zealand
Published
January 2026

Reported outcomes

15 x faster

time-to-insightTime & speed

+95%accuracy60-70%accuracy

Strategic outcomes

Better decisions & insightFaster access to community insightsBetter decisions & insightRapid data-driven iteration
Why do we believe this?Outcome claims, sources, and evidence checks

Normalized claim

Time-to-insight: 15 x faster increase

AWS Machine Learning BlogJan 27, 2026Blog postExplicit claimMedium evidence strength

reduced time-to-insight from approximately 120 seconds to under 4 seconds

Normalized claim

Accuracy: 95% increase

AWS Machine Learning BlogJan 27, 2026Blog postExplicit claimMedium evidence strength

achieved 95% overall accuracy

Normalized claim

Accuracy: 60-70% increase

AWS Machine Learning BlogJan 27, 2026Blog postExplicit claimMedium evidence strength

while the team reached an accuracy plateau at 60-70%

Why do we believe this deployment?Customer identity, provider attribution, maturity, and source checks
Customer
Pushpay
Provider
AWS
Maturity
Exploring

The solution was built on Amazon Bedrock and paired with a custom generative AI evaluation framework, including semantic selection of filters, a golden dataset, LLM-as-a-judge evaluation, and domain-level dashboards for continuous improvement

Customer identity supportedSource describes one deploymentMaturity supported

Primary read

Use case focus

Showing 2 of 2

  • 1Prompt optimization
  • 2Content discovery assistant
  • Ministry staff needed fast, plain-English access to community insights from complex church data and more than 100 configurable filters.
  • Initial agent accuracy plateaued at 60-70%, and manual evaluation slowed improvements.
  • Pushpay built an agentic AI search feature that accepts natural-language queries in the existing application interface.
  • The team used Amazon Bedrock with prompt caching, Claude Sonnet 4.5, semantic search to select relevant filters, a dynamic prompt constructor, a golden dataset of over 300 queries, LLM-as-a-judge evaluation, and dashboards to support continuous optimization and rollout decisions.
  • The solution reduced time-to-insight from about 120 seconds to under 4 seconds.
  • Pushpay improved accuracy from 60-70% to more than 95% for high-performance domains.
  • The new evaluation workflow enabled rapid, data-driven iteration and targeted optimization by domain.
Architecture

An AI search agent embedded in the existing Pushpay application uses Amazon Bedrock and prompt caching to process natural-language queries, a dynamic prompt constructor to tailor prompts using query content, user persona, and tenant-specific requirements, semantic search to select relevant filters, and a custom generative AI evaluation framework with a golden dataset, LLM-as-a-judge comparison, domain categorization, dashboards, and staged rollout controls.

Sources & evidence1
Evidence: Medium50/100Evidence strength
  • Customer explicitly identified
  • Deployment status explicitly supported
  • Quantified outcome available
  • Technical implementation details available
Type: Blog PostPublished: Jan 27, 2026Publisher: AWSEvidence: VendorConfidence: Medium

AI-generated summary. Verify important details with the linked sources before relying on this case.

Explore related AI use cases

Was this useful?

Community

Comments

No published comments yet.