Scaled productionEvidence: Medium50/100

Forethought uses Amazon SageMaker multi-model endpoints to reduce generative AI inference costs

Use case typeAI platformUpdated Jun 13, 2026

Forethought Technologies’ SupportGPT platform serves more than 30 million customer interactions annually with personalized embeddings, intent detection, autocomplete, and response generation models. The company migrated embeddings and autocomplete models from Amazon EKS to Amazon SageMaker multi-model endpoints to improve GPU utilization, reduce OOM errors, and lower inference hosting costs.

Industry
Tech & Comms
Published
June 2023
Why do we believe this?Outcome claims, sources, and evidence checks

Normalized claim

Cost savings: 80% decrease

AWS Machine Learning BlogJun 13, 2023Blog postExplicit claimMedium evidence strength

we realized a cost saving of up to 80%

Normalized claim

Cost savings: 66-74% decrease

AWS Machine Learning BlogJun 13, 2023Blog postExplicit claimMedium evidence strength

we’ve now realized cost savings in the 66–74% range

Normalized claim

Embeddings models per instance: 12 models per instance increase

AWS Machine Learning BlogJun 13, 2023Blog postExplicit claimMedium evidence strength

we discovered it could house 12 embeddings models per instance while achieving 80% GPU memory utilization

Why do we believe this deployment?Customer identity, provider attribution, maturity, and source checks
Customer
Forethought Technologies
Provider
AWS
Maturity
Scaled Production

Run many hyper-personalized generative AI and LLM models at scale while keeping GPU inference cost, reliability, and deployment overhead under control

Customer identity supportedSource describes one deploymentMaturity supported

Primary read

Use case focus

Showing 3 of 3

  • 1Model serving optimization
  • 2Generative AI infrastructure
  • 3Inference cost reduction
  • Use Amazon SageMaker multi-model endpoints for real-time inference.
  • Package models and custom inference logic with NVIDIA Triton Inference Server, store artifacts in Amazon S3, and automate endpoint creation and model loading/unloading.
  • Apply autoscaling policies based on GPU memory utilization and traffic patterns, and warm models periodically to avoid cold starts.
  • Reported hosting cost reductions of 66% to 74%, with up to 80% savings in some scenarios.
  • Achieved higher model packing density, including 12 embeddings models per instance at about 80% GPU utilization.
Architecture

Forethought migrated customer-specific embeddings and autocomplete models from Amazon EKS to Amazon SageMaker multi-model endpoints. The implementation uses a shared Triton serving container, custom Python and PyTorch/TorchScript inference logic, Amazon S3 model artifacts, Boto3 automation, and autoscaling policies based on traffic and GPU memory utilization. The team also warms models on a schedule to reduce cold-start latency.

Sources & evidence1
Evidence: Medium50/100Evidence strength
  • Customer explicitly identified
  • Deployment status explicitly supported
  • Quantified outcome available
  • Technical implementation details available
Type: Blog PostPublished: Jun 13, 2023Publisher: AWSEvidence: VendorConfidence: High

AI-generated summary. Verify important details with the linked sources before relying on this case.

Explore related AI use cases

Was this useful?

Community

Comments

No published comments yet.