ExploringEvidence: Medium50/100

Embed the world: Multimodal AI for searchable aerial imagery at scale

Use case typeSemantic searchUpdated Jun 22, 2026

Vexcel worked with AWS to evaluate multimodal embeddings, captioning, fusion strategies, and vector search for turning multi-view aerial imagery into a natural-language-searchable knowledge base. The system used Amazon Bedrock, Amazon OpenSearch Serverless, Amazon S3, AWS Secrets Manager, and automated evaluation against OpenStreetMap ground truth across about 100 configurations. The solution evolved into a preview product for searchable vector embeddings across Vexcel's global imagery library spanning 45+ countries.

Organization
Vexcel
Industry
Tech & Comms
Published
June 2026

Reported outcomes

+11%

F1 improvement for pools with captionsOther quantified impact

+13%F1 improvement for roads with captions

Strategic outcomes

Other strategic outcomeTurned aerial imagery into a searchable knowledge baseNew product / capabilityEvolved into a preview searchable product
Why do we believe this?Outcome claims, sources, and evidence checks

Normalized claim

F1 improvement for pools with captions: 11% increase

AWS Machine Learning BlogJun 22, 2026Blog postInferred claimMedium evidence strength

an 11% F1 score improvement for pools

Normalized claim

F1 improvement for roads with captions: 13% increase

AWS Machine Learning BlogJun 22, 2026Blog postInferred claimMedium evidence strength

13% for roads

Why do we believe this deployment?Customer identity, provider attribution, maturity, and source checks
Customer
Vexcel
Provider
AWS
Maturity
Exploring

The system used Amazon Bedrock, Amazon OpenSearch Serverless, Amazon S3, AWS Secrets Manager, and automated evaluation against OpenStreetMap ground truth across about 100 configurations

Customer identity supportedSource describes one deploymentMaturity supported

Primary read

Use case focus

Showing 2 of 2

  • 1Semantic search
  • 2Multimodal analytics
  • Turn billions of pixels across multi-view aerial imagery into a natural-language-searchable knowledge base without manual inspection or bespoke retraining for each query.
  • Evaluate which embedding model, fusion strategy, captioning approach, and search method work best for geospatial semantic search.
  • Built a five-stage modular pipeline for AOI selection, imagery ingestion, embedding and indexing, search, and evaluation.
  • Tested Amazon Nova Multimodal Embeddings, Amazon Titan Multimodal Embeddings, and Cohere embeddings with caption integration and metadata filtering.
  • Used OpenStreetMap ground truth and a configurable evaluation harness to compare search quality across many configurations.
  • Caption integration improved F1 scores by 11% for pools and 13% for roads.
  • Amazon Nova Multimodal Embeddings achieved the highest average F1 scores across benchmark queries.
  • The concepts evolved into a preview searchable product spanning Vexcel's global imagery library.
Architecture

A five-stage modular pipeline: AOI selection persisted to Amazon S3; imagery ingestion from Vexcel's API with Amazon S3 caching and AWS Secrets Manager for credentials; embedding and caption generation using Amazon Bedrock models; indexing in Amazon OpenSearch Serverless or Amazon S3 Vectors; natural-language search with multi-view fusion, image-caption fusion, and metadata filtering; and automated evaluation against OpenStreetMap ground truth.

Sources & evidence1
Evidence: Medium50/100Evidence strength
  • Customer explicitly identified
  • Deployment status explicitly supported
  • Quantified outcome available
  • Technical implementation details available
Type: Blog PostPublished: Jun 22, 2026Publisher: AWSEvidence: VendorConfidence: Medium

AI-generated summary. Verify important details with the linked sources before relying on this case.

Explore related AI use cases

Was this useful?

Community

Comments

No published comments yet.