Case Studies | FinTech

Reading 100-page loan documents in minutes, not hours

Intelligent Document Processing

About

A FinTech business working with Cloud Combinator on AWS. The client is anonymised at their request.

Challenge

The client brought three connected challenges to the engagement, each rooted in the volume and complexity of the documents flowing through origination.

Manual review of very large documents

Origination decisions rely on reading long documents, frequently running to 100 pages or more. Doing this manually is slow, hard to scale during busy

Unstructured, variable content

The documents arrive as PDFs containing a mix of text and images, in a range of formats and layouts. Extracting the relevant content reliably, and then distilling it into something a decision-maker can act on, is difficult to do by hand at any meaningful scale.

The need for a scalable, cost-efficient foundation

The client needed more than an one-off script. It wanted a solution that scales automatically with demand, keeps costs proportional to usage, and provides a clean foundation for future enhancements such as notifications, human review, or a front-end interface, all without over-engineering the first release.

Solution

Cloud Combinator structured the work so that value was proven early and then hardened into a deployable service. The proof of concept validated the core pipeline against the client's real documents; the initial-works build then added the secure networking, orchestration, monitoring, and documentation needed to run it reliably. The build itself progressed through four phases.

100-page

Scale of documents the pipeline processes automatically

~$2,025

Projected monthly AWS running cost (USD, indicative)

$0

Upfront cost; pay-as-you-go serverless model

By the numbers:

  • 100-page - Scale of documents the pipeline processes automatically
  • ~$2,025 - Projected monthly AWS running cost (USD, indicative)
  • $0 - Upfront cost; pay-as-you-go serverless model
Changes

Cloud Combinator delivered a fully functional, serverless document-processing pipeline that meets the project's success criteria: automated ingestion, Textract-based extraction, and Bedrock-driven analysis and origination decision support, all on a scalable, cost-efficient architecture, and accepted on sign-off against those criteria by the client's technical and business stakeholders.

  • Automated end-to-end processingDocuments move from upload to extracted text to AI analysis with no manual intervention, removing the manual reading bottleneck at the front of origination.
  • AI-assisted origination decisionsAmazon Bedrock with Claude (3.5/3.7) turns dense, unstructured PDFs into concise page-level analysis that supports faster, more consistent decisions.
  • Scalable and cost-efficient by designA fully serverless architecture scales automatically with demand, with no upfront cost and a projected 12-month AWS spend of approximately $24,302 (USD, indicative).
  • A foundation for future growthThe architecture is ready to accommodate optional enhancements such as SNS notifications, human review via Amazon A2I, or a front-end interface, without rework.
  • HandoverThe client received source code, CloudFormation templates for reproducible infrastructure, detailed technical documentation and user manuals, and recorded demonstration sessions for a clean transfer of ownership.

With the core pipeline proven, accepted, and running in the client's own AWS account, the business now has a scalable platform to build on. The natural next steps, scoped separately, include stakeholder notifications, human-in-the-loop review, a front-end upload and review interface, and a Well-Architected review to take the service to full production standard.

AWS Stack

Amazon S3

For secure, encrypted document storage across raw uploads, extracted data, and analysis outputs, with event notifications that trigger the pipeline.

AWS Lambda

For the serverless functions that orchestrate ingestion, text extraction, and document analysis.

Amazon Textract

For extracting text and images from large, multi-page PDF documents.

Amazon Bedrock

(Claude 3.5/3.7) for AI-driven analysis of each page and support for origination decisions.

Amazon API Gateway

For RESTful endpoints that trigger the pipeline and enable future front-end integration.

Amazon CloudWatch

For logging, metrics, and alarms across the processing pipeline.

Amazon VPC

And IAM for secure networking, private connectivity, and least-privilege access control.

YOU MIGHT LIKE

Related success stories

View all case studies

Case Studies | Insights

Utilising Language Recognition, Speed, and Enhanced Security to Make Social Media a Force for Good

  • Here, we take a detailed look at how the Cloud Combinator team collaborated with another cutting-edge AI service provider that provides intelligent systems to “make social media more social” for brands and users alike.
  • Arwen AI is a UK-based startup specialising in AI solutions to manage and enhance brands’ social media interactions. Founded in 2020 by Matt McGrory, Dr. David Cole, and Joel Bailey, Arwen. AI focuses on using AI to automatically detect and remove spam, toxic comments, and other unwanted content from social media platforms.
  • The team at Arwen have three core products. ‘Moderate’ is focused on identifying and removing toxic content from social media channels. ‘Engage’ helps brands identify and engage with meaningful conversations on social media, and ‘Customize’ allows brands to apply bespoke algorithms to their channels - creating an even more effective moderation and engagement.
Read more
CONTACT US

Ready to turn AI into impact?

We'll help you spot the highest-value opportunities, reduce risk around your first AI initiative, and define a clear path to results from day one.

Why talk to us:

Outcome-driven recommendations

AWS-recognised delivery expertise

Risk-aware AI adoption

Clear next step, not a sales pitch

Start with a focused 20-minute conversation about your goals — no pressure, no commitment.

This website uses cookies to enhance user experience and to analyze performance and traffic on our website.

See our Privacy Policy for details.