Normalized claim
Cost savings: 80% decrease
we realized a cost saving of up to 80%
Forethought Technologies’ SupportGPT platform serves more than 30 million customer interactions annually with personalized embeddings, intent detection, autocomplete, and response generation models. The company migrated embeddings and autocomplete models from Amazon EKS to Amazon SageMaker multi-model endpoints to improve GPU utilization, reduce OOM errors, and lower inference hosting costs.
Normalized claim
Cost savings: 80% decrease
we realized a cost saving of up to 80%
Normalized claim
Cost savings: 66-74% decrease
we’ve now realized cost savings in the 66–74% range
Normalized claim
Embeddings models per instance: 12 models per instance increase
we discovered it could house 12 embeddings models per instance while achieving 80% GPU memory utilization
Run many hyper-personalized generative AI and LLM models at scale while keeping GPU inference cost, reliability, and deployment overhead under control
Primary read
Showing 3 of 3
Forethought migrated customer-specific embeddings and autocomplete models from Amazon EKS to Amazon SageMaker multi-model endpoints. The implementation uses a shared Triton serving container, custom Python and PyTorch/TorchScript inference logic, Amazon S3 model artifacts, Boto3 automation, and autoscaling policies based on traffic and GPU memory utilization. The team also warms models on a schedule to reduce cold-start latency.
AI-generated summary. Verify important details with the linked sources before relying on this case.
Was this useful?
Community
No published comments yet.