Evidence: Low35/100

Cisco Webex Contact Center Topic Analytics: migrating LLM workloads to SageMaker Inference

Cisco's Webex AI team migrated LLM workloads from application containers to Amazon SageMaker Inference to improve speed, scalability, and price-performance. The solution powers Webex Contact Center topic analytics and other AI experiences, including transcript analysis, conversational assistants, and meeting summaries.

Organization
Cisco
Industry
Tech & Comms
Published
August 2024

Reported outcomes

Cost: −50%

Cost savings

Cost: −20%

Catalog median for cost savings deployments: −40% across 171 reported metrics. Compare benchmarks →

Why do we believe this?Outcome claims, sources, and evidence checks

Normalized claim

Cost: 50% decrease

AWS BlogAug 8, 2024Blog postInferred claimLow evidence strength

Cisco says the collaboration helped reduce foundation model deployment costs by 50% on average and latency by 20% on average.

Normalized claim

Cost: 20% decrease

AWS BlogAug 8, 2024Blog postInferred claimLow evidence strength

Cisco says the collaboration helped reduce foundation model deployment costs by 50% on average and latency by 20% on average.

Why do we believe this deployment?Customer identity, provider attribution, maturity, and source checks
Customer
Cisco, Webex, Webex Contact Center
Provider
AWS
Maturity
Unknown
Linked source
AWS Blog

No explicit deployment-stage evidence found.

Customer identity supportedSource describes one deploymentMaturity evidence evaluated

Primary read

Use case focus

Showing 3 of 4

  • 1Contact Center Analytics
  • 2Conversational AI
  • 3Model Inference Optimization
  • Cisco decoupled model hosting from applications and moved LLM inference to Amazon SageMaker Inference.
  • The architecture uses an LLM proxy on Amazon EKS, auto scaling, and a three-model pipeline for topic analytics with FLAN-T5 call-driver extraction, clustering, and topic labeling.
  • The stack also uses Amazon Bedrock as an optional model provider and AWS infrastructure services for ingress and observability.
  • The platform can handle hundreds of inferences per minute at peak and overnight batch jobs reliably.
  • Cisco says the collaboration helped reduce foundation model deployment costs by 50% on average and latency by 20% on average.
Architecture

Cisco's WxAI architecture decouples model hosting from Webex applications by moving LLMs to Amazon SageMaker Inference. An LLM proxy runs on Amazon EKS and routes requests to multiple models. Webex Contact Center Topic Analytics uses a three-model pipeline: FLAN-T5 for call-driver extraction, clustering over embeddings, and a generative topic-labeling model. The architecture uses SageMaker auto scaling, Amazon Bedrock as an optional provider, and AWS services for ingress, logging, metrics, and data storage.

Sources & evidence1
Evidence: Low35/100Evidence strength
  • Customer explicitly identified
  • Quantified outcome available
  • Technical implementation details available
Type: Blog PostPublished: Aug 8, 2024Publisher: AWSEvidence: VendorConfidence: Medium

AI-generated summary. Verify important details with the linked sources before relying on this case.

Explore related AI use cases

Was this useful?

Community

Comments

No published comments yet.