Normalized claim
Cost: 50% decrease
Cisco says the collaboration helped reduce foundation model deployment costs by 50% on average and latency by 20% on average.
Cisco's Webex AI team migrated LLM workloads from application containers to Amazon SageMaker Inference to improve speed, scalability, and price-performance. The solution powers Webex Contact Center topic analytics and other AI experiences, including transcript analysis, conversational assistants, and meeting summaries.
Reported outcomes
Cost: −50%
Cost savings
Catalog median for cost savings deployments: −40% across 171 reported metrics. Compare benchmarks →
Normalized claim
Cost: 50% decrease
Cisco says the collaboration helped reduce foundation model deployment costs by 50% on average and latency by 20% on average.
Normalized claim
Cost: 20% decrease
Cisco says the collaboration helped reduce foundation model deployment costs by 50% on average and latency by 20% on average.
No explicit deployment-stage evidence found.
Primary read
Showing 3 of 4
Cisco's WxAI architecture decouples model hosting from Webex applications by moving LLMs to Amazon SageMaker Inference. An LLM proxy runs on Amazon EKS and routes requests to multiple models. Webex Contact Center Topic Analytics uses a three-model pipeline: FLAN-T5 for call-driver extraction, clustering over embeddings, and a generative topic-labeling model. The architecture uses SageMaker auto scaling, Amazon Bedrock as an optional provider, and AWS services for ingress, logging, metrics, and data storage.
AI-generated summary. Verify important details with the linked sources before relying on this case.
Was this useful?
Community
No published comments yet.