Case Studies | SaaS B2B

Benchmarking ad-labelling AI against AWS-native models

AI Ad-Labelling Infrastructure

About

A SaaS B2B business working with Cloud Combinator on AWS. The client is anonymised at their request.

Challenge

The challenge had three focus areas.

Proving AWS-native models could match the baseline

The existing OpenAI-based workflow set the bar for accuracy and consistency. Any AWS alternative had to be measured against that baseline on accuracy, consistency, latency, throughput, and cost before it could be trusted

Controlling cost at rapidly growing volume

Labelling is performed per taxonomy, so a single ad can generate five to seven requests today and ten or more as customers add bespoke taxonomies. With volumes projected to grow from around 3,000 requests a day toward hundreds of thousands, the economics of the inference approach mattered enormously.

Building for a closed, enterprise-ready environment

Future enterprise customers need assurances about where models run and how data is handled. The client needed the labelling to run inside its own AWS account, with a fine-tuning pipeline it could re-run for future retraining and reproducibility, rather than a black-box external service.

Solution

Cloud Combinator structured the work in four phases with clear decision points, so the client only progressed to fine-tuning and deployment once the evidence justified it.

10M+

Ad-labelling requests per month at the 12-month high-volume target (projected)

50%

Inference cost reduction from batch pricing versus on-demand

$799

One-time cost to fine-tune a candidate model on 100M tokens

By the numbers:

  • 10M+ - Ad-labelling requests per month at the 12-month high-volume target (projected)
  • 50% - Inference cost reduction from batch pricing versus on-demand
  • $799 - One-time cost to fine-tune a candidate model on 100M tokens
Changes

The engagement gives the client an evidence-based path off its external API dependency: a documented benchmark against the OpenAI baseline, an AWS model recommendation with confirmed fine-tuning support, a validated training-data schema, and a fine-tuning pipeline configured for repeatable retraining, all running within the client's own AWS environment. Figures below reflect the volume projections and pricing assumptions set out in the scope of work.

  • A measured baseline, not a guessThe client's OpenAI workflow is benchmarked for accuracy, consistency, latency, throughput, and cost, giving an objective bar for every AWS option.
  • Model selection on evidenceBedrock and SageMaker candidates are evaluated with an AWS-based framework, and only a model that clears the benchmark and supports fine-tuning proceeds.
  • Fine-tuned to the client's taxonomiesThe chosen model is adapted to the client's own labelled data, improving accuracy and consistency on both global and bespoke taxonomies.
  • Built to scale affordablyA batch-heavy, provisioned-throughput approach is designed to hold cost down as volumes climb from thousands to millions of requests a month.
  • HandoverCloud Combinator delivered the fine-tuning pipeline and configuration documentation for repeatable retraining, leaving the client able to run and improve the model in its own AWS account.

With a benchmarked, fine-tuned AWS labelling capability, the client is positioned to scale into its LinkedIn Marketing API migration and later Meta and Google Ads expansion, offering enterprise customers a closed, controllable environment as request volumes grow.

AWS Stack

Amazon Bedrock

For evaluating, fine-tuning, and serving foundation models such as Meta Llama 3.1 Instruct (70B) for ad topic classification.

Amazon SageMaker

For model training, fine-tuning jobs, and ML workflow support where the selected model requires it.

Amazon Bedrock Evaluations

And AWS-based benchmarking frameworks for measuring model accuracy, consistency, and cost against the baseline.

Amazon S3

For structured storage of training, validation, and benchmark datasets and model artefacts, with versioning for retraining.

AWS Identity

And Access Management for least-privilege access within the client's own AWS account.

YOU MIGHT LIKE

Related success stories

View all case studies

Case Studies | Insights

Utilising Language Recognition, Speed, and Enhanced Security to Make Social Media a Force for Good

  • Here, we take a detailed look at how the Cloud Combinator team collaborated with another cutting-edge AI service provider that provides intelligent systems to “make social media more social” for brands and users alike.
  • Arwen AI is a UK-based startup specialising in AI solutions to manage and enhance brands’ social media interactions. Founded in 2020 by Matt McGrory, Dr. David Cole, and Joel Bailey, Arwen. AI focuses on using AI to automatically detect and remove spam, toxic comments, and other unwanted content from social media platforms.
  • The team at Arwen have three core products. ‘Moderate’ is focused on identifying and removing toxic content from social media channels. ‘Engage’ helps brands identify and engage with meaningful conversations on social media, and ‘Customize’ allows brands to apply bespoke algorithms to their channels - creating an even more effective moderation and engagement.
Read more
CONTACT US

Ready to turn AI into impact?

We'll help you spot the highest-value opportunities, reduce risk around your first AI initiative, and define a clear path to results from day one.

Why talk to us:

Outcome-driven recommendations

AWS-recognised delivery expertise

Risk-aware AI adoption

Clear next step, not a sales pitch

Start with a focused 20-minute conversation about your goals — no pressure, no commitment.

This website uses cookies to enhance user experience and to analyze performance and traffic on our website.

See our Privacy Policy for details.