ProductionEvidence: Medium55/100

TwelveLabs builds multimodal video understanding using Amazon Bedrock and SageMaker HyperPod

TwelveLabs developed multimodal video foundation models and deployed them on AWS to support training and high-volume inference at global production scale. Its Marengo and Pegasus models power semantic video search, summarization, and structured outputs over massive archives.

Organization
TwelveLabs
Industry
Tech & Comms
Published
July 2026

Reported outcomes

Strategic outcomes

Scale & capacityProduction deployment at petabyte scaleBetter decisions & insightMade video archives searchable and actionableCost efficiencyImproved unit economics
Why do we believe this deployment?Customer identity, provider attribution, maturity, and source checks
Customer
TwelveLabs
Provider
AWS
Maturity
Production

TwelveLabs developed multimodal video foundation models and deployed them on AWS to support training and high-volume inference at global production scale

Customer identity supportedSource describes one deploymentMaturity supported

Primary read

Use case focus

Showing 3 of 3

  • 1Content discovery assistant
  • 2Multimodal analytics
  • 3Search modernization
  • Scaling video AI from research to global production required infrastructure for training multimodal models and running inference on petabytes of video data.
  • Traditional search methods fail on video archives with millions of hours because the data volume and complexity are too massive.
  • TwelveLabs built a unified AWS production stack with EKS, EC2 GPU instances, S3, and S3 Vectors to store video assets and embeddings together.
  • The company trained models across thousands of GPUs with Amazon SageMaker HyperPod and used vector search for sub-second retrieval across billions of embeddings.
  • The architecture enabled production deployment and unit-economics improvements at petabyte scale.
  • With TwelveLabs models available in Amazon Bedrock, developers can build video AI applications while maintaining control over their data.
Architecture

TwelveLabs trains multimodal video foundation models on Amazon SageMaker HyperPod, runs production workloads on Amazon EKS and Amazon EC2 GPU instances, stores raw video and processed multimodal data in Amazon S3, and uses Amazon S3 Vectors to co-locate searchable embeddings with video assets for sub-second semantic retrieval.

Sources & evidence1
Evidence: Medium55/100Evidence strength
  • Customer explicitly identified
  • Deployment status explicitly supported
  • Primary source available
  • Technical implementation details available
Type: Customer StoryPublished: Jul 15, 2026Publisher: AWSEvidence: PrimaryConfidence: High

AI-generated summary. Verify important details with the linked sources before relying on this case.

Explore related AI use cases

Was this useful?

Community

Comments

No published comments yet.