Case Studies | Healthcare & Life Sciences

Training generative imaging models at scale across GPU and AWS Trainium

SageMaker HyperPod Model Training

About

A Healthcare & Life Sciences business working with Cloud Combinator on AWS. The client is anonymised at their request.

Challenge

The challenge had three focus areas.

Compute at scale

Training the Sana model, a deep-compression autoencoder combined with a linear diffusion transformer, requires large clusters of accelerators working in concert. The client needed an environment that could provision anywhere from eight to well over a hundred accelerators, keep them fed with data, and recover automatically when long-running jobs are interrupted.

Portability across accelerators

Most training code is written for NVIDIA GPUs and CUDA. To take advantage of the price and performance of AWS Trainium, the Sana codebase had to be ported to the AWS Neuron SDK and PyTorch XLA, replacing CUDA-specific operations and adopting Trainium-native distributed training and BF16 precision.

Cost and performance visibility

The client wanted more than a single working run. The team needed a like-for-like comparison of throughput, scaling efficiency and cost between GPU and Trainium, so future training decisions could rest on evidence rather than assumption.

Solution

Training data and model artefacts live in Amazon S3. A SageMaker HyperPod cluster provisions the accelerators, whether NVIDIA GPUs or AWS Trainium chips, and the training job runs across them in parallel. On Trainium, the Sana model runs through PyTorch XLA on the AWS Neuron SDK, with gradients synchronised across every worker so each replica stays in step.

8-128

Accelerators per HyperPod cluster, GPU or Trainium (planned scale)

2

Accelerator platforms benchmarked side by side (GPU and Trainium)

4

Delivery phases from environment build to production handover

By the numbers:

  • 8-128 - Accelerators per HyperPod cluster, GPU or Trainium (planned scale)
  • 2 - Accelerator platforms benchmarked side by side (GPU and Trainium)
  • 4 - Delivery phases from environment build to production handover
Changes

The platform is designed to take the client from constrained, single-machine experiments to elastic, multi-accelerator training that scales on demand. Acceptance is defined against successful distributed training of the Sana model on both GPU and Trainium, a documented performance comparison and production-ready pipelines for future runs. The figures below describe the designed scale of the platform and are engagement targets rather than measured production results.

  • Elastic trainingA SageMaker HyperPod environment that provisions accelerators on demand and recovers long-running jobs automatically, removing the ceiling imposed by local hardware.
  • Accelerator choiceThe same Sana workload runs on both NVIDIA GPUs and AWS Trainium, giving the client a genuine price and performance choice for every future training run.
  • Evidence-based decisionsA documented cost-performance comparison across the two platforms, so scaling and budget decisions rest on measured throughput and efficiency.
  • Production-ready pipelinesContainerised training, checkpointing and BF16 optimisation packaged into repeatable pipelines the team can reuse.
  • HandoverRunbooks, technical documentation and knowledge-transfer sessions so the client's engineers can operate and extend the platform independently.

With a scalable training foundation in place across GPU and Trainium, the client is positioned to train larger models and iterate faster, no longer limited by local compute and ready to scale whatever it takes on next.

AWS Stack

Amazon SageMaker HyperPod

For managed, resilient clusters that scale distributed training across many accelerators.

AWS Trainium

(Trn1 instances) for purpose-built, cost-efficient machine learning training silicon.

Amazon EC2 P4d

And P5 GPU instances for high-performance GPU training and baseline comparison.

Amazon S3

For durable storage of datasets, checkpoints and model artefacts.

Amazon CloudWatch

For centralised metrics, monitoring and logging of training jobs.

YOU MIGHT LIKE

Related success stories

View all case studies

Case Studies | Insights

Utilising Language Recognition, Speed, and Enhanced Security to Make Social Media a Force for Good

  • Here, we take a detailed look at how the Cloud Combinator team collaborated with another cutting-edge AI service provider that provides intelligent systems to “make social media more social” for brands and users alike.
  • Arwen AI is a UK-based startup specialising in AI solutions to manage and enhance brands’ social media interactions. Founded in 2020 by Matt McGrory, Dr. David Cole, and Joel Bailey, Arwen. AI focuses on using AI to automatically detect and remove spam, toxic comments, and other unwanted content from social media platforms.
  • The team at Arwen have three core products. ‘Moderate’ is focused on identifying and removing toxic content from social media channels. ‘Engage’ helps brands identify and engage with meaningful conversations on social media, and ‘Customize’ allows brands to apply bespoke algorithms to their channels - creating an even more effective moderation and engagement.
Read more
CONTACT US

Ready to turn AI into impact?

We'll help you spot the highest-value opportunities, reduce risk around your first AI initiative, and define a clear path to results from day one.

Why talk to us:

Outcome-driven recommendations

AWS-recognised delivery expertise

Risk-aware AI adoption

Clear next step, not a sales pitch

Start with a focused 20-minute conversation about your goals — no pressure, no commitment.

This website uses cookies to enhance user experience and to analyze performance and traffic on our website.

See our Privacy Policy for details.