GCPPilotEvidence: Medium65/100

Stanford Center for Genomics and Personalized Medicine: Building a mega-scale genetic variation analysis pipeline

Stanford Center for Genomics and Personalized Medicine (SCGPM) at Stanford University built a mega-scale genetic variation analysis pipeline on Google Cloud Platform using Google Genomics and BigQuery to analyze large DNA sequencing datasets faster than on-premises clusters and enable secure sharing of genomic data. The team processed hundreds of whole genomes for the Million Veteran Program pilot and established security best practices for storing and sharing genomic data in the cloud.

Industry
Healthcare
Published
June 2016

Reported outcomes

10 seconds

quantified impactOther quantified impact

Strategic outcomes

New product / capabilityBuilt mega-scale variant analysis pipelineRisk & complianceEstablished secure genomic data sharing practicesScale & capacityExpanded access to cloud genomics resourcesBetter decisions & insightEnabled rapid variant-analysis queries
Why do we believe this?Outcome claims, sources, and evidence checks

Normalized claim

Quantified impact: 10 seconds

Google Cloud Customer StoryJun 1, 2016Customer storyInferred claimMedium evidence strength

Returned variant-analysis query results in less than 10 seconds even with 500 genomes and millions of variants.

Why do we believe this deployment?Customer identity, provider attribution, maturity, and source checks
Customer
Stanford Center for Genomics and Personalized Medicine, Stanford University
Provider
GCP
Maturity
Pilot

The team processed hundreds of whole genomes for the Million Veteran Program pilot and established security best practices for storing and sharing genomic data in the cloud

Customer identity supportedSource describes one deploymentMaturity supported

Primary read

Use case focus

Showing 2 of 2

  • 1Genomics Analytics
  • 2Data Platform Modernization
  • Process and analyze extremely large DNA sequencing datasets faster than on-premises clusters.
  • Enable secure storage and sharing of sensitive genomic data.
  • Scale variant analysis to hundreds of genomes without long local-cluster turnaround times.
  • Built a genetic variation, or variant, analysis pipeline on Google Cloud Platform using Google Genomics.
  • Used Google BigQuery to mine variants and compare them against publicly accessible databases.
  • Worked with the Google Genomics team and implemented security best practices such as encryption, user permissions, and limiting data access.
  • Processed 500 genomes from raw data to variant calls for the Million Veteran Program pilot data in days.
  • Returned variant-analysis query results in less than 10 seconds even with 500 genomes and millions of variants.
  • Established security best practices to help genomics labs confidently store and share data in the cloud.
  • Expanded access to cloud resources for Stanford genomics researchers, with over a petabyte of genomic data stored and analyzed on Google Cloud Platform.
Architecture

SCGPM built a genetic variation analysis pipeline on Google Cloud Platform with Google Genomics to process large DNA sequencing datasets, using BigQuery for fast variant queries and cloud-based security controls for data protection and sharing.

Implementation partners1
Sources & evidence1
Evidence: Medium65/100Evidence strength
  • Customer explicitly identified
  • Deployment status explicitly supported
  • Primary source available
  • Quantified outcome available
  • Technical implementation details available
Type: Customer StoryPublished: Jun 1, 2016Publisher: Google CloudEvidence: VendorConfidence: Medium

AI-generated summary. Verify important details with the linked sources before relying on this case.

Explore related AI use cases

Was this useful?

Community

Comments

No published comments yet.