Case Studies | FinTech

One governed inference path

Bedrock AI Enablement and Agent Platform

About

A FinTech business working with Cloud Combinator on AWS. The client is anonymised at their request.

Challenge

The challenge had three focus areas, each essential to putting the client's engineering AI on a governed footing.

Unused credits and ungoverned spend

Paying Anthropic and Microsoft directly meant the client's AWS promotional credits went unused and inference spend never landed on the

Governance and data residency

As a regulated firm, the client needed inference to run in-region with data kept inside the AWS boundary, no static long-lived keys in the request path, and the ability to block a team or user that exceeds its cap, demonstrated live rather than asserted.

A reusable path for agents

Beyond coding assistants, the client wanted to prove that a real agent could run on AWS under the same controls. An existing compliance app needed to move onto a governed agent runtime as a blueprint for the agents that follow, without a rebuild.

Solution

Cloud Combinator delivered the work as one fixed-scope phase over six weeks, built on the client's existing Control Tower and Sentinel logging foundation, with all infrastructure delivered as Terraform to match the client's own practice.

100-120

Engineers targeted onto Claude Code on Amazon Bedrock (target)

Zero

Static long-lived keys used for inference; IAM role only

6 weeks

End-to-end delivery, with production scoped separately

By the numbers:

  • 100-120 - Engineers targeted onto Claude Code on Amazon Bedrock (target)
  • Zero - Static long-lived keys used for inference; IAM role only
  • 6 weeks - End-to-end delivery, with production scoped separately
Changes

The engagement is built to deliver a single, governed Bedrock inference path for the client's engineering AI, accepted against a short, testable set of conditions measured from the gateway and the AWS bill rather than asserted. Acceptance covers adoption, credit drawdown, governance, observability, the agent and a full handover.

  • Credits put to workIn-scope engineer inference is served by Amazon Bedrock and billed to the credit-holding AWS organisation, with no direct Anthropic or Microsoft inference billing for that usage.
  • Governed spendThe LiteLLM gateway issues virtual keys and enforces per-team and per-user budgets, caps and model routing, blocking a team or user that reaches its cap, demonstrated live.
  • Full observabilityAll in-scope Claude Code and agent traffic is visible in Langfuse with per-team and per-model cost attribution, on a single pane kept out of the request path.
  • An agent blueprintThe compliance app runs on Amazon Bedrock AgentCore across dev, test and prod with Guardrails and Entra identity, delivered with a reusable Terraform blueprint for the agents that follow.
  • HandoverAn adoption and spend report, the architecture, the reusable blueprint and a production and scale recommendation are handed over, alongside participation in the AWS Hands-On Labs.

With one governed inference path in place and a reusable agent blueprint proven, the client is positioned to harden the platform for production, roll AI out to the wider workforce, and add further agents on the same foundation, building on a setup that is ready for whatever it takes on next.

AWS Stack

Amazon Bedrock

For in-region inference on Claude models, with token spend on the AWS bill drawing AWS credits.

Amazon Bedrock AgentCore

For the isolated, governed runtime hosting the compliance agent.

Amazon Bedrock Guardrails

For policy controls applied to agent and inference traffic.

AWS Fargate

For running the LiteLLM gateway as a serverless container in a private subnet.

Amazon Aurora PostgreSQL

(Serverless v2) for the gateway's virtual key, budget and cap state.

AWS IAM Identity Center

For Entra-federated single sign-on governing human access.

AWS Control Tower

For baseline governance, guardrails and centralised log-archive and audit accounts.

AWS PrivateLink

For private VPC endpoints so inference traffic never leaves the AWS network.

YOU MIGHT LIKE

Related success stories

View all case studies

Case Studies | Insights

Utilising Language Recognition, Speed, and Enhanced Security to Make Social Media a Force for Good

  • Here, we take a detailed look at how the Cloud Combinator team collaborated with another cutting-edge AI service provider that provides intelligent systems to “make social media more social” for brands and users alike.
  • Arwen AI is a UK-based startup specialising in AI solutions to manage and enhance brands’ social media interactions. Founded in 2020 by Matt McGrory, Dr. David Cole, and Joel Bailey, Arwen. AI focuses on using AI to automatically detect and remove spam, toxic comments, and other unwanted content from social media platforms.
  • The team at Arwen have three core products. ‘Moderate’ is focused on identifying and removing toxic content from social media channels. ‘Engage’ helps brands identify and engage with meaningful conversations on social media, and ‘Customize’ allows brands to apply bespoke algorithms to their channels - creating an even more effective moderation and engagement.
Read more
CONTACT US

Ready to turn AI into impact?

We'll help you spot the highest-value opportunities, reduce risk around your first AI initiative, and define a clear path to results from day one.

Why talk to us:

Outcome-driven recommendations

AWS-recognised delivery expertise

Risk-aware AI adoption

Clear next step, not a sales pitch

Start with a focused 20-minute conversation about your goals — no pressure, no commitment.

This website uses cookies to enhance user experience and to analyze performance and traffic on our website.

See our Privacy Policy for details.