Case Studies | Healthcare & Life Sciences

From curated literature to a queryable knowledge graph

Biomedical Knowledge Graph

About

A Healthcare & Life Sciences business working with Cloud Combinator on AWS. The client is anonymised at their request.

Challenge

The challenge had three focus areas.

Extracting structure from unstructured science

Scientific papers describe entities such as microbes, diseases, dietary factors, medications, and metabolites, and the relationships between them, but they do so in prose. The platform had to identify those entities and relationships reliably and turn them into graph nodes and edges, each linked back to its source paper as evidence, without a hand-built extraction pipeline.

Combining graph and semantic search in one place

Downstream reporting needs both graph traversal, following the links between a microbe and a condition, and semantic similarity over the underlying literature. Running a separate graph database and vector store would have added complexity and cost. The design needed one engine that could do both.

Controlling the cost of an always-available graph

A graph database sized for query performance is expensive to leave running continuously, especially when ingestion happens only weekly or monthly. The platform had to stay affordable while remaining ready to serve queries when needed.

Solution

At the core is Amazon Bedrock Knowledge Bases with Amazon Neptune Analytics as the graph store. When documents from the data lake gold layer are synced from Amazon S3, Bedrock automatically chunks them, uses a large language model to extract entities and relationships, generates vector embeddings with Titan Embeddings V2, and constructs the graph nodes, edges, and vector index in Neptune Analytics. This removes the need for a custom natural-language extraction pipeline, a manual schema framework, or hand-loading of the graph. Because Neptune Analytics is an unified graph and vector engine, downstream consumers can run graph traversal and semantic similarity in the same place, queried with openCypher.

80

New papers per day the graph is designed to ingest (projected)

Up to 500K

Graph nodes anticipated at twelve-month literature scale (projected)

10%

Neptune Analytics cost while paused, versus its active running cost

By the numbers:

  • 80 - New papers per day the graph is designed to ingest (projected)
  • Up to 500K - Graph nodes anticipated at twelve-month literature scale (projected)
  • 10% - Neptune Analytics cost while paused, versus its active running cost
Changes

The engagement delivered an automated knowledge-graph platform: a configured Bedrock Knowledge Base backed by Neptune Analytics, a Step Functions orchestration handling scheduled ingestion, validation, and the pause and resume lifecycle, full infrastructure as code, and documentation. It gives the client a graph that grows with the literature and answers both relationship and similarity queries from a single engine. Figures below reflect the design targets set in the scope of work.

  • Automated extraction, no custom pipelineBedrock Knowledge Bases reads each document and builds the graph directly, extracting entities and relationships with a LLM rather than a hand-maintained NLP stack.
  • Evidence-linked relationshipsEvery extracted relationship is tied back to its source paper, so the graph doubles as an evidence trail for the AI reporting that consumes it.
  • Graph and vector in one engineNeptune Analytics serves both graph traversal and semantic similarity through openCypher, removing the need for a separate vector store.
  • Pause-and-resume economicsA scheduled Step Functions lifecycle keeps Neptune Analytics paused between ingestion runs, cutting its cost to a tenth of the active rate while keeping the graph ready to serve.
  • HandoverCloud Combinator delivered the platform as infrastructure as code with architecture decision records, runbooks, and a data dictionary, including an optional entity-standardisation path if naming variants fragment the graph.

With a queryable, evidence-linked knowledge graph in place, the client is ready to build its AI reporting prototype on top, and to grow the graph continuously as new literature is published without changing the underlying architecture.

AWS Stack

Amazon Bedrock Knowledge Bases

For automated entity and relationship extraction, embedding generation, and graph construction from documents.

Amazon Neptune Analytics

As an unified knowledge-graph store with native vector search, queried with openCypher.

AWS Step Functions

And Amazon EventBridge for scheduled ingestion and the Neptune Analytics pause and resume lifecycle.

AWS Lambda

For ingestion triggers and pause and resume automation, with Amazon CloudWatch for monitoring and alerting.

AWS Secrets Manager

And AWS KMS for credential security and encryption at rest.

YOU MIGHT LIKE

Related success stories

View all case studies

Case Studies | Insights

Utilising Language Recognition, Speed, and Enhanced Security to Make Social Media a Force for Good

  • Here, we take a detailed look at how the Cloud Combinator team collaborated with another cutting-edge AI service provider that provides intelligent systems to “make social media more social” for brands and users alike.
  • Arwen AI is a UK-based startup specialising in AI solutions to manage and enhance brands’ social media interactions. Founded in 2020 by Matt McGrory, Dr. David Cole, and Joel Bailey, Arwen. AI focuses on using AI to automatically detect and remove spam, toxic comments, and other unwanted content from social media platforms.
  • The team at Arwen have three core products. ‘Moderate’ is focused on identifying and removing toxic content from social media channels. ‘Engage’ helps brands identify and engage with meaningful conversations on social media, and ‘Customize’ allows brands to apply bespoke algorithms to their channels - creating an even more effective moderation and engagement.
Read more
CONTACT US

Ready to turn AI into impact?

We'll help you spot the highest-value opportunities, reduce risk around your first AI initiative, and define a clear path to results from day one.

Why talk to us:

Outcome-driven recommendations

AWS-recognised delivery expertise

Risk-aware AI adoption

Clear next step, not a sales pitch

Start with a focused 20-minute conversation about your goals — no pressure, no commitment.

This website uses cookies to enhance user experience and to analyze performance and traffic on our website.

See our Privacy Policy for details.