Case Studies | PropTech

Scaling computer vision inference on AWS

AWS Batch YOLO Inference

About

A PropTech business working with Cloud Combinator on AWS. The client is anonymised at their request.

Challenge

The evaluation focused on three questions that decide whether a machine learning pipeline can scale sustainably.

Is AWS Batch the right engine for YOLO inference

The client needed a clear, practical answer on whether AWS Batch could handle batch inference jobs well: job management, performance, ease of integration, and operational overhead, rather than a theoretical recommendation.

Which compute gives the best value

GPU-enabled EC2 instances such as g6.2xlarge deliver acceleration but start slower, while AWS Fargate provisions in around thirty seconds but offers no GPU. The engagement had to benchmark both paths and quantify the trade-off for the client's workload.

Does cost track demand

With inference volumes set to grow year on year, the environment needed auto scaling driven by job-queue demand, so resources scale up under load and back down when idle, avoiding over-provisioning.

Solution

Video is uploaded to an Amazon S3 input bucket. That upload triggers an AWS Lambda function, which submits a new job to an AWS Batch job queue. The job runs the YOLO model inference on the AWS Batch compute environment using a container image held in Amazon ECR, and writes its results, the bounding boxes and predictions, to a S3 output bucket.

4,200

Videos per day at the 2026 baseline (projected)

2

Compute paths benchmarked, Fargate CPU and EC2 GPU

7

G6.2xlarge instances projected at 2028 scale

By the numbers:

  • 4,200 - Videos per day at the 2026 baseline (projected)
  • 2 - Compute paths benchmarked, Fargate CPU and EC2 GPU
  • 7 - G6.2xlarge instances projected at 2028 scale
Changes

The project was accepted against its success criteria: a feasibility assessment of AWS Batch for YOLO inference, validated instance selection, a cost-aware environment with auto scaling, and documentation and knowledge transfer for a clean handover, all deployed into the client's AWS environment via CloudFormation. The figures below are drawn from the Statement of Work's scaling model and are projections.

  • A proven scaling engineAWS Batch was validated as a practical way to manage and scale YOLO inference jobs, with job queues, retry strategies, and managed compute removing server-management overhead.
  • An evidence-based compute choiceBenchmarking CPU Fargate against GPU EC2 gave the client a quantified basis for choosing instance types on performance, cost, and resource utilisation rather than guesswork.
  • Cost that follows demandAuto scaling on job-queue depth means capacity rises under load and falls when idle, keeping the solution cost-effective as volumes grow toward tens of thousands of videos per day.
  • Event-driven and hands-offA S3 upload automatically triggers a Lambda that submits the Batch job, so inference runs without manual intervention.
  • HandoverCalls, Loom walkthroughs, and documentation were provided so the client can operate the pipeline and take it toward production.

The pipeline delivers the start of the client's ML inference capability on AWS, a foundation that is no longer a proof of concept but a scalable base to build on as demand grows and the service moves toward production.

AWS Stack

AWS Batch

For managing and scaling YOLO inference jobs across a managed compute environment.

Amazon EC2 GPU

Instances (g6.2xlarge) and AWS Fargate for the two benchmarked execution paths.

Amazon ECR

For storing the YOLO inference container image with its dependencies.

Amazon S3

For input video and inference output storage.

AWS Lambda

For automatically submitting Batch jobs when new video arrives.

Amazon CloudWatch

For monitoring, logging, and performance comparison.

Amazon VPC

And AWS CloudFormation for secure, repeatable deployment.

YOU MIGHT LIKE

Related success stories

View all case studies

Case Studies | Insights

Utilising Language Recognition, Speed, and Enhanced Security to Make Social Media a Force for Good

  • Here, we take a detailed look at how the Cloud Combinator team collaborated with another cutting-edge AI service provider that provides intelligent systems to “make social media more social” for brands and users alike.
  • Arwen AI is a UK-based startup specialising in AI solutions to manage and enhance brands’ social media interactions. Founded in 2020 by Matt McGrory, Dr. David Cole, and Joel Bailey, Arwen. AI focuses on using AI to automatically detect and remove spam, toxic comments, and other unwanted content from social media platforms.
  • The team at Arwen have three core products. ‘Moderate’ is focused on identifying and removing toxic content from social media channels. ‘Engage’ helps brands identify and engage with meaningful conversations on social media, and ‘Customize’ allows brands to apply bespoke algorithms to their channels - creating an even more effective moderation and engagement.
Read more
CONTACT US

Ready to turn AI into impact?

We'll help you spot the highest-value opportunities, reduce risk around your first AI initiative, and define a clear path to results from day one.

Why talk to us:

Outcome-driven recommendations

AWS-recognised delivery expertise

Risk-aware AI adoption

Clear next step, not a sales pitch

Start with a focused 20-minute conversation about your goals — no pressure, no commitment.

This website uses cookies to enhance user experience and to analyze performance and traffic on our website.

See our Privacy Policy for details.