Case Studies | SaaS B2B

From one GPU to an elastic ML platform

Migration to Amazon SageMaker

About

A SaaS B2B business working with Cloud Combinator on AWS. The client is anonymised at their request.

Challenge

The challenge had three focus areas.

Breaking the single-GPU ceiling

Training a 30-million-parameter transformer on around 100GB of data per run, over 48 to 72 hours, on one local GPU limited how much the client could train and how fast. They needed elastic GPU capacity that spins up for a job and shuts down afterwards, so throughput was no longer capped by one machine.

Serving developers and customers on one platform, safely

The client wanted internal developers to have full access while customers could only train on their own data and see only their own models. That demanded strict tenant isolation, per-user roles, and access mediated through the client's own application, with no direct customer access to AWS.

Keeping GPU training affordable and visible

GPU training is expensive, so the client needed to cap what each customer could consume, stop jobs that exceeded their quota, and see per-user usage clearly, all without surprises on the monthly bill.

Solution

We ran the engagement as advisory and enablement, prototyping and validating the approach and then equipping the client's team with the architecture and step-by-step guidance to build and run it themselves.

Up to 70%

Projected training-cost saving using SageMaker managed spot training

Multi-user

Isolated developer and customer access via ABAC and per-user roles

30M

Parameter transformer moved from one on-prem GPU to elastic AWS GPUs

By the numbers:

  • Up to 70% - Projected training-cost saving using SageMaker managed spot training
  • Multi-user - Isolated developer and customer access via ABAC and per-user roles
  • 30M - Parameter transformer moved from one on-prem GPU to elastic AWS GPUs
Changes

The engagement delivered a validated architecture and the hands-on guidance for the client to stand up a multi-user SageMaker platform, meeting the success criteria of an operational domain, validated access controls, an end-to-end pipeline design, model versioning, per-user monitoring, and a completed knowledge transfer.

  • Scale on demandEphemeral GPU training jobs replace a single fixed GPU, so training throughput is no longer capped by one machine and the client only pays for compute in use.
  • Safe multi-tenancyA single domain with VPC isolation, per-user execution roles, and attribute-based access control lets developers and customers share the platform without seeing each other's data or models.
  • Cost under controlManaged spot training with checkpointing, per-user GPU quotas enforced automatically, and budget alerts keep expensive GPU training predictable.
  • Enterprise-grade traceabilityCloudWatch dashboards plus CloudTrail with SourceIdentity give the client per-user audit and failure visibility.
  • HandoverIncluded the architecture, guidance documents, and recorded walkthroughs so the client's team can build, run, and extend the platform independently.

With a clear SageMaker blueprint and a cost model that works at scale, the client is positioned to onboard more customers onto self-service training and to grow its AI-driven manufacturing platform, no longer bound to a single GPU but ready to support whatever the business takes on next.

AWS Stack

Amazon SageMaker

For managed, multi-user model training, pipelines, and the Model Registry.

Amazon S3

For training data, checkpoints, and model artifacts.

AWS IAM

And IAM Identity Center for per-user roles and attribute-based access control.

Amazon CloudWatch

And AWS CloudTrail for per-user monitoring, dashboards, and audit trails.

AWS Lambda

And Amazon EventBridge for automatic enforcement of per-user GPU quotas.

AWS Budgets

For cost alerts, and Amazon EC2 G5 GPU instances for training compute.

YOU MIGHT LIKE

Related success stories

View all case studies

Case Studies | Insights

Utilising Language Recognition, Speed, and Enhanced Security to Make Social Media a Force for Good

  • Here, we take a detailed look at how the Cloud Combinator team collaborated with another cutting-edge AI service provider that provides intelligent systems to “make social media more social” for brands and users alike.
  • Arwen AI is a UK-based startup specialising in AI solutions to manage and enhance brands’ social media interactions. Founded in 2020 by Matt McGrory, Dr. David Cole, and Joel Bailey, Arwen. AI focuses on using AI to automatically detect and remove spam, toxic comments, and other unwanted content from social media platforms.
  • The team at Arwen have three core products. ‘Moderate’ is focused on identifying and removing toxic content from social media channels. ‘Engage’ helps brands identify and engage with meaningful conversations on social media, and ‘Customize’ allows brands to apply bespoke algorithms to their channels - creating an even more effective moderation and engagement.
Read more
CONTACT US

Ready to turn AI into impact?

We'll help you spot the highest-value opportunities, reduce risk around your first AI initiative, and define a clear path to results from day one.

Why talk to us:

Outcome-driven recommendations

AWS-recognised delivery expertise

Risk-aware AI adoption

Clear next step, not a sales pitch

Start with a focused 20-minute conversation about your goals — no pressure, no commitment.

This website uses cookies to enhance user experience and to analyze performance and traffic on our website.

See our Privacy Policy for details.