Case Studies | FinTech

Proving trustworthy AI for due diligence answers

LLM Document Processing

About

A FinTech business working with Cloud Combinator on AWS. The client is anonymised at their request.

Challenge

The PoC concentrated on four areas: measuring answer accuracy, guaranteeing source traceability, eliminating hallucination, and comparing models fairly across varied datasets.

Accuracy and controlled generation

The system needed to answer based on exact matches with the client's bank of answers, make reasoned assumptions from previous client data where

Reference traceability

Every generated answer had to include precise source citations mapped back to the original Q&A documents, so a reviewer could always see where an answer came from.

Hallucination suppression

The framework had to validate that the model does not invent answers when no reference exists, assessing its ability to suppress generation when ground truth is absent, the single most important safeguard for due diligence use.

Fair, repeatable model comparison across diverse data

The system needed to interchange different Amazon Bedrock models and score them on a like-for-like basis, and to perform across datasets of varying completeness and age, from clients with rich five-year corpuses to those with limited histories, in a reliable, repeatable pipeline.

Solution

Cloud Combinator structured the PoC across weekly milestones, from infrastructure configuration through pipeline build, evaluation, multi-model testing and handover, with a final week of support for implementing in the client's own AWS environment. Delivery was collaborative throughout, with pair programming and code walkthroughs to upskill the client's team.

3

Configurable output styles evaluated, from definite answers only to fully generated content

~5,000

Model requests per year modelled in the PoC cost basis (approx. 100 questions per document)

$2,306

Projected monthly AWS run cost (MRR) for the evaluation pipeline

By the numbers:

  • 3 - Configurable output styles evaluated, from definite answers only to fully generated content
  • ~5,000 - Model requests per year modelled in the PoC cost basis (approx. 100 questions per document)
  • $2,306 - Projected monthly AWS run cost (MRR) for the evaluation pipeline
Changes

Acceptance was based on signing off delivery against the project success criteria, confirming the technical competence of the client's internal tech lead, and demonstrating that the AWS environment is suitable for continued development of the client's service. The PoC delivered a working, repeatable framework for measuring how trustworthy LLM-generated DDQ answers are, along with a comparative report across models.

  • Trust made measurableThe framework scores model answers on correctness, traceability and hallucination rate, using RAG evaluation and Bedrock's LLM-as-a-judge with custom metrics.
  • Hallucination guardrail provenThe PoC specifically validated the model's ability to suppress generation when no ground truth exists, the key safeguard for due diligence.
  • Model choice de-riskedThe same the client's dataset was run across multiple Bedrock models, including Claude 3.7 Sonnet and Claude Sonnet 3.0, producing a like-for-like comparison of accuracy, traceability, hallucination and latency.
  • Team capability liftedCollaborative pair programming, code walkthroughs and prompt documentation strengthened the client's internal LLM integration and AI workflow skills.
  • HandoverFinal CloudFormation, documented Lambda code, prompt templates and test harnesses were delivered in a shared repository, with a knowledge transfer session and a week of post-project support for implementation in the client's AWS environment.

With an evidence-based framework in hand, the client can decide with confidence which Bedrock model and evaluation mode best fit each client's data, and has a repeatable, Infrastructure-as-Code foundation to build on as it moves from proof of concept towards a production DDQ answering capability in its own AWS environment.

AWS Stack

Amazon Bedrock

For LLM inference and, via its Evaluate API and LLM-as-a-judge, for scoring answer quality across models.

Amazon OpenSearch Service

For indexing the client's Q&A documents and JSON test sets and for exact-match and chunked retrieval.

AWS Lambda

For orchestrating retrieval, Bedrock invocation, response serialisation and post-processing.

Amazon S3

For secure storage of raw test data, evaluation artifacts, logs and structured JSON outputs.

AWS Step Functions

For orchestrating the evaluation workflows across each mode.

AWS CloudFormation

For delivering the whole environment as repeatable, auditable Infrastructure as Code.

YOU MIGHT LIKE

Related success stories

View all case studies

Case Studies | Insights

Utilising Language Recognition, Speed, and Enhanced Security to Make Social Media a Force for Good

  • Here, we take a detailed look at how the Cloud Combinator team collaborated with another cutting-edge AI service provider that provides intelligent systems to “make social media more social” for brands and users alike.
  • Arwen AI is a UK-based startup specialising in AI solutions to manage and enhance brands’ social media interactions. Founded in 2020 by Matt McGrory, Dr. David Cole, and Joel Bailey, Arwen. AI focuses on using AI to automatically detect and remove spam, toxic comments, and other unwanted content from social media platforms.
  • The team at Arwen have three core products. ‘Moderate’ is focused on identifying and removing toxic content from social media channels. ‘Engage’ helps brands identify and engage with meaningful conversations on social media, and ‘Customize’ allows brands to apply bespoke algorithms to their channels - creating an even more effective moderation and engagement.
Read more
CONTACT US

Ready to turn AI into impact?

We'll help you spot the highest-value opportunities, reduce risk around your first AI initiative, and define a clear path to results from day one.

Why talk to us:

Outcome-driven recommendations

AWS-recognised delivery expertise

Risk-aware AI adoption

Clear next step, not a sales pitch

Start with a focused 20-minute conversation about your goals — no pressure, no commitment.

This website uses cookies to enhance user experience and to analyze performance and traffic on our website.

See our Privacy Policy for details.