Case Studies | FinTech

An AI invoice and receipt extraction pipeline on AWS

Receipt Extraction

About

A FinTech business working with Cloud Combinator on AWS. The client is anonymised at their request.

Challenge

The challenge had four focus areas.

Extracting data from difficult documents

The system had to accurately extract critical fields, buyer, seller, dates, line items with quantity, unit price and total, and the full VAT breakdown, from documents that were often handwritten, crinkled or image-based. Getting this right on challenging inputs, not just clean ones, was the core requirement.

Cost efficiency

Extraction quality is only useful if it is affordable at scale. The pipeline had to establish a realistic per-user, per-document cost and minimise expensive model calls, on the principle that well-formatted documents can be handled by cheaper models while difficult ones justify more capable ones.

Scalability and reliability under spiky load

Real usage is uneven, with peaks on Fridays and around end-of-month VAT deadlines. The system had to prove it could absorb bursts of uploads reliably, and to show the trade-offs between on-demand and batch processing.

Clean integration and compliance

The pipeline had to fit the client's Go-based backend patterns for a lightweight integration, output a fixed JSON schema, and handle sensitive financial data in line with GDPR.

Solution

Documents are uploaded securely to Amazon S3 using pre-signed URLs generated on demand by AWS Lambda, so uploads are simple and access is time-limited. AWS Step Functions then orchestrates the flow, passing each document from S3 to a large language model in Amazon Bedrock while controlling concurrency so that usage limits are respected even under load. Amazon CloudWatch and Amazon DynamoDB track every document from upload to completion, recording timestamps, file size, status and extraction results.

$0.075

Indicative processing cost per document (projected)

50,000

Documents per month the pipeline is sized for (projected)

1 JSON

Consistent structured schema produced for every document

By the numbers:

  • $0.075 - Indicative processing cost per document (projected)
  • 50,000 - Documents per month the pipeline is sized for (projected)
  • 1 JSON - Consistent structured schema produced for every document
Changes

The engagement delivered the extraction pipeline defined in the project success criteria: accurate data extraction across difficult document types into the client's JSON schema, a clear view of per-document cost, and demonstrated behaviour under simulated peak load. Acceptance was based on meeting the success criteria and confirming the AWS environment is suitable for continued development of the client's service. The figures below are indicative estimates from project sizing rather than measured production results.

  • Accurate extraction from hard inputsHandwritten, crinkled and image-based documents are turned into structured fields including full VAT breakdowns, in the client's specified JSON schema.
  • Cost optimisation built inRouting well-formatted documents to cheaper models while reserving capable models for difficult ones keeps the per-document cost realistic, with on-demand and batch options assessed.
  • Proven under loadSimulated continuous uploads validated that the pipeline handles spiky, high-volume periods reliably.
  • Lightweight integrationThe solution follows the client's Go-based backend patterns and outputs directly to their schema for a clean fit.
  • HandoverA fully documented GitHub repository with Terraform infrastructure as code, Lambda and Step Functions code and test scripts was delivered, alongside a knowledge-transfer workshop covering the AWS services and prompt engineering.

With an accurate, cost-aware extraction pipeline proven, the client is positioned to move it towards production and integrate it into the live product, giving users faster, more reliable data capture from every receipt and invoice.

AWS Stack

Amazon Bedrock

For AI extraction of structured data from invoices and receipts.

Amazon S3

For secure document storage with pre-signed URL uploads.

AWS Step Functions

For orchestrating processing and simulating load for scalability testing.

AWS Lambda

For generating pre-signed URLs and processing documents.

Amazon DynamoDB

For document metadata and status tracking.

Amazon CloudWatch

For real-time monitoring, and Amazon API Gateway for the upload endpoint.

YOU MIGHT LIKE

Related success stories

View all case studies

Case Studies | Insights

Utilising Language Recognition, Speed, and Enhanced Security to Make Social Media a Force for Good

  • Here, we take a detailed look at how the Cloud Combinator team collaborated with another cutting-edge AI service provider that provides intelligent systems to “make social media more social” for brands and users alike.
  • Arwen AI is a UK-based startup specialising in AI solutions to manage and enhance brands’ social media interactions. Founded in 2020 by Matt McGrory, Dr. David Cole, and Joel Bailey, Arwen. AI focuses on using AI to automatically detect and remove spam, toxic comments, and other unwanted content from social media platforms.
  • The team at Arwen have three core products. ‘Moderate’ is focused on identifying and removing toxic content from social media channels. ‘Engage’ helps brands identify and engage with meaningful conversations on social media, and ‘Customize’ allows brands to apply bespoke algorithms to their channels - creating an even more effective moderation and engagement.
Read more
CONTACT US

Ready to turn AI into impact?

We'll help you spot the highest-value opportunities, reduce risk around your first AI initiative, and define a clear path to results from day one.

Why talk to us:

Outcome-driven recommendations

AWS-recognised delivery expertise

Risk-aware AI adoption

Clear next step, not a sales pitch

Start with a focused 20-minute conversation about your goals — no pressure, no commitment.

This website uses cookies to enhance user experience and to analyze performance and traffic on our website.

See our Privacy Policy for details.