Case Studies | GovTech

A real-time speech-to-speech AI service on AWS

Speech-to-Speech AI Service

About

A GovTech business working with Cloud Combinator on AWS. The client is anonymised at their request.

Challenge

The challenge had three focus areas, each mapping to one link in the voice chain, plus the underlying need for the whole thing to be reproducible.

High-accuracy speech-to-text

Authenticated callers had to be able to send audio and receive an accurate transcript. If the transcription is wrong, everything downstream is wrong, so the quality of this first step sets the ceiling for the entire experience.

Context-aware answer generation

A transcript on its own is not an answer. The service needed to interpret what was said and produce a coherent, relevant response for the client's specific domain, with room to shape those answers using prompts and, in future, caller metadata such as age or location.

Human-quality audio output

The reply had to come back as natural, human-like speech rather than a flat, robotic read, because in a voice product the quality of the voice is the product.

Solution

Cloud Combinator delivered the service as four clear milestones. Each stage was wired up and validated on its own before being chained into the end-to-end flow, and every piece of infrastructure was defined in AWS CloudFormation so the whole environment could be stood up, torn down, and rebuilt consistently in the client's account.

3-stage

Speech-to-speech pipeline: transcribe, reason, synthesise

100%

Infrastructure deployed as code with AWS CloudFormation

Serverless

Event-driven Lambda compute with nothing to manage

By the numbers:

  • 3-stage - Speech-to-speech pipeline: transcribe, reason, synthesise
  • 100% - Infrastructure deployed as code with AWS CloudFormation
  • Serverless - Event-driven Lambda compute with nothing to manage
Changes

The Proof of Concept met its success criteria across all three focus areas: audio in, accurate transcript, a context-aware answer, and natural speech back out, delivered as a single serverless pipeline that the client now owns in its own AWS account.

  • Accurate transcriptionAuthenticated callers send audio and receive a reliable transcript from Nova Sonic speech-to-text, setting a dependable foundation for everything downstream.
  • Relevant answersAn Amazon Bedrock large language model turns each transcript into a coherent, domain-relevant response, with a prompt template and response schema in place ready for further tuning.
  • Natural voice outputNova Sonic text-to-speech returns human-like audio rather than robotic playback, so the experience matches the product ambition.
  • Repeatable environmentBecause the whole stack is defined in CloudFormation, the service can be rebuilt consistently and handed over cleanly to the client's account.
  • HandoverA runbook, handover deck, and recorded demo left the client's team able to operate and extend the service without dependence on Cloud Combinator.

The Proof of Concept is the start of the client's AI capability, not the finish. With a proven serverless pattern and a clean handover in place, the natural next steps are hardening the pipeline towards production, fine-tuning the language model on the client's own material, and using caller metadata to tailor responses further. The client now has a working, ownable foundation to take that product forward on AWS.

AWS Stack

Amazon API Gateway

For the authenticated audio entry point into the service.

AWS Lambda

For serverless compute that runs each pipeline stage and scales with demand.

Amazon Bedrock

For the large language model that generates context-aware answers.

Amazon Nova Sonic

For high-accuracy speech-to-text and natural, human-like text-to-speech.

Amazon S3

For optional persistence of generated audio artefacts.

AWS Secrets Manager

For secure storage of credentials used by the service.

AWS CloudFormation

For deploying the entire environment as repeatable infrastructure as code.

YOU MIGHT LIKE

Related success stories

View all case studies

Case Studies | Insights

Utilising Language Recognition, Speed, and Enhanced Security to Make Social Media a Force for Good

  • Here, we take a detailed look at how the Cloud Combinator team collaborated with another cutting-edge AI service provider that provides intelligent systems to “make social media more social” for brands and users alike.
  • Arwen AI is a UK-based startup specialising in AI solutions to manage and enhance brands’ social media interactions. Founded in 2020 by Matt McGrory, Dr. David Cole, and Joel Bailey, Arwen. AI focuses on using AI to automatically detect and remove spam, toxic comments, and other unwanted content from social media platforms.
  • The team at Arwen have three core products. ‘Moderate’ is focused on identifying and removing toxic content from social media channels. ‘Engage’ helps brands identify and engage with meaningful conversations on social media, and ‘Customize’ allows brands to apply bespoke algorithms to their channels - creating an even more effective moderation and engagement.
Read more
CONTACT US

Ready to turn AI into impact?

We'll help you spot the highest-value opportunities, reduce risk around your first AI initiative, and define a clear path to results from day one.

Why talk to us:

Outcome-driven recommendations

AWS-recognised delivery expertise

Risk-aware AI adoption

Clear next step, not a sales pitch

Start with a focused 20-minute conversation about your goals — no pressure, no commitment.

This website uses cookies to enhance user experience and to analyze performance and traffic on our website.

See our Privacy Policy for details.