TwelveLabs builds multimodal video understanding using Amazon Bedrock and SageMaker HyperPod
TwelveLabs developed multimodal video foundation models and deployed them on AWS to support training and high-volume inference at global production scale. Its Marengo and Pegasus models power semantic video search, summarization, and structured outputs over massive archives.
- Organization
- TwelveLabs
- Industry
- Tech & Comms
- Location
- United States
- Published
- July 2026
Reported outcomes
Strategic outcomes
Why do we believe this deployment?Customer identity, provider attribution, maturity, and source checks
- Customer
- TwelveLabs
- Provider
- AWS
- Maturity
- Production
- Linked source
- AWS Solutions Case Study
TwelveLabs developed multimodal video foundation models and deployed them on AWS to support training and high-volume inference at global production scale
Primary read
Use case focus
Showing 3 of 3
- 1Content discovery assistant
- 2Multimodal analytics
- 3Search modernization
- Scaling video AI from research to global production required infrastructure for training multimodal models and running inference on petabytes of video data.
- Traditional search methods fail on video archives with millions of hours because the data volume and complexity are too massive.
- TwelveLabs built a unified AWS production stack with EKS, EC2 GPU instances, S3, and S3 Vectors to store video assets and embeddings together.
- The company trained models across thousands of GPUs with Amazon SageMaker HyperPod and used vector search for sub-second retrieval across billions of embeddings.
- The architecture enabled production deployment and unit-economics improvements at petabyte scale.
- With TwelveLabs models available in Amazon Bedrock, developers can build video AI applications while maintaining control over their data.
Architecture
TwelveLabs trains multimodal video foundation models on Amazon SageMaker HyperPod, runs production workloads on Amazon EKS and Amazon EC2 GPU instances, stores raw video and processed multimodal data in Amazon S3, and uses Amazon S3 Vectors to co-locate searchable embeddings with video assets for sub-second semantic retrieval.
Sources & evidence1
- Customer explicitly identified
- Deployment status explicitly supported
- Primary source available
- Technical implementation details available
AI-generated summary. Verify important details with the linked sources before relying on this case.
Explore related AI use cases
Was this useful?
Community
Comments
No published comments yet.