Case Studies | Healthcare & Life Sciences

Turning scattered biomedical literature into an analytics-ready corpus

Biomedical Data Lake

About

A Healthcare & Life Sciences business working with Cloud Combinator on AWS. The client is anonymised at their request.

Challenge

The challenge had three focus areas, each a prerequisite for the AI work that would follow.

Ingesting messy, heterogeneous sources

The corpus arrives from biomedical literature APIs such as PubMed and from parsed PDF documents, with text, metadata, tables, and figures all mixed together. The platform had to land every source reliably and without transformation first, so nothing was lost and everything remained traceable back to its origin.

Cleaning and conforming to a canonical model

Raw literature is full of duplicates, inconsistent fields, and missing values, and the same paper can appear more than once under different identifiers. The lake needed to deduplicate by DOI and PMID, enforce a schema, and normalise fields into a single canonical data model that downstream consumers could rely on.

Shaping data for downstream AI, cost-effectively

The final layer had to be shaped specifically for a knowledge graph and vector-embedding pipelines, not for general reporting. At the same time, at tens of millions of documents, storage and processing costs had to be controlled through versioning, lifecycle policies, and right-sized compute.

Solution

The lake follows a medallion architecture with three progressively refined layers held in versioned Amazon S3 storage. The bronze layer lands raw API responses and PDF extraction output with zero transformation, partitioned by source and ingestion date. The silver layer holds deduplicated, schema-validated documents in a canonical model, written as Parquet for analytics alongside structured JSON payloads. The gold layer reshapes that data specifically for downstream consumers, with its schema co-designed with the knowledge-graph and vector-store workstreams.

3

Medallion layers: bronze, silver, and gold, progressively refining the corpus

Tens of

Documents the lake is designed to hold at full corpus scale (projected)

100%

AWS resources defined as infrastructure as code (Terraform or CDK)

By the numbers:

  • 3 - Medallion layers: bronze, silver, and gold, progressively refining the corpus
  • Tens of - Documents the lake is designed to hold at full corpus scale (projected)
  • 100% - AWS resources defined as infrastructure as code (Terraform or CDK)
Changes

The engagement delivered a complete, self-contained data platform: a versioned medallion lake, ingestion pipelines, a catalogued schema registry, transformation jobs, orchestration, a validation query library, and full infrastructure as code with runbooks and a data dictionary. It gives the client a curated corpus that the knowledge-graph and AI-reporting projects can build on directly. Figures below reflect the design targets set in the scope of work.

  • Reliable, traceable ingestionLambda-based ingest functions land every API source and parsed PDF into the bronze layer unchanged, with S3 object versioning and Glacier archival keeping the raw record complete and cost-controlled.
  • A single canonical modelGlue PySpark jobs deduplicate by DOI and PMID, enforce schema, and normalise fields, turning inconsistent source material into conformed silver-layer datasets.
  • Ready for AI from day oneThe gold layer is shaped for knowledge-graph ingestion and vector-embedding pipelines, so the downstream projects start from clean, purpose-built data.
  • Built-in quality assuranceA saved library of Amazon Athena queries validates record counts, schema compliance, and data profiles after every stage, making data loss and mapping errors visible early.
  • HandoverCloud Combinator delivered the platform as infrastructure as code with architecture decision records, runbooks, and a data dictionary, leaving the client with a lake it can operate and extend.

With the data foundation in place, the client is positioned to build its biomedical knowledge graph and AI reporting on top of a corpus that is clean, versioned, and analytics-ready, and to add new literature sources over time without re-architecting the platform.

AWS Stack

Amazon S3

For versioned object storage across the bronze, silver, and gold medallion layers.

Amazon Athena

For SQL-based validation, auditing, and data profiling across all layers.

AWS Step Functions

For orchestrating the ingest, crawl, transform, and validate pipeline.

AWS Lambda

And Amazon EventBridge for scheduled, event-driven ingestion from literature APIs.

YOU MIGHT LIKE

Related success stories

View all case studies

Case Studies | Insights

Utilising Language Recognition, Speed, and Enhanced Security to Make Social Media a Force for Good

  • Here, we take a detailed look at how the Cloud Combinator team collaborated with another cutting-edge AI service provider that provides intelligent systems to “make social media more social” for brands and users alike.
  • Arwen AI is a UK-based startup specialising in AI solutions to manage and enhance brands’ social media interactions. Founded in 2020 by Matt McGrory, Dr. David Cole, and Joel Bailey, Arwen. AI focuses on using AI to automatically detect and remove spam, toxic comments, and other unwanted content from social media platforms.
  • The team at Arwen have three core products. ‘Moderate’ is focused on identifying and removing toxic content from social media channels. ‘Engage’ helps brands identify and engage with meaningful conversations on social media, and ‘Customize’ allows brands to apply bespoke algorithms to their channels - creating an even more effective moderation and engagement.
Read more
CONTACT US

Ready to turn AI into impact?

We'll help you spot the highest-value opportunities, reduce risk around your first AI initiative, and define a clear path to results from day one.

Why talk to us:

Outcome-driven recommendations

AWS-recognised delivery expertise

Risk-aware AI adoption

Clear next step, not a sales pitch

Start with a focused 20-minute conversation about your goals — no pressure, no commitment.

This website uses cookies to enhance user experience and to analyze performance and traffic on our website.

See our Privacy Policy for details.