Case Studies | Healthcare & Life Sciences

A voice agent that answers every patient call

Voice AI Agent

About

A Healthcare & Life Sciences business working with Cloud Combinator on AWS. The client is anonymised at their request.

Challenge

The challenge had three focus areas, each of which had to hold up before a live patient pilot could be considered.

Answering in real time over the phone

A voice agent on a live call cannot feel laggy. A chained pipeline of separate speech, reasoning, and voice services typically adds three to five seconds of round-trip delay, which is the difference between a natural conversation and one a patient

Grounding answers in the clinic's own procedures

Procedural answers had to come from the client's SOPs, not from the model's pre-training. That required a retrieval layer that could turn a spoken question into the right SOP passage reliably, and a way to keep that knowledge current as SOPs change day to day without redeploying code.

Knowing when to escalate

The agent had to recognise the limits of its confidence and hand off to a human when clinical judgment was required, carrying the full conversation context across so the patient never has to repeat themselves. Handoff had to be treated as a first-class behaviour, measured as carefully as deflection.

Solution

Cloud Combinator ran the build in four phases across parallel delivery tracks, compressing the retrieval groundwork into week one to free time for the higher-risk voice work later in the engagement.

Sub-1s

Target voice response latency, versus 3 to 5 seconds for a chained pipeline (projected)

3 tiers

Answer, retrieve from SOPs, and escalate with context

4-6 wks

Proof of concept build on AWS in the London region

By the numbers:

  • Sub-1s - Target voice response latency, versus 3 to 5 seconds for a chained pipeline (projected)
  • 3 tiers - Answer, retrieve from SOPs, and escalate with context
  • 4-6 wks - Proof of concept build on AWS in the London region
Changes

The proof of concept was scoped to prove all three focus areas end to end and to leave the client with a working agent, a retrieval pipeline, clinic-controlled SOP management, and the metrics needed to make a confident decision on a live pilot. Acceptance is met when the agent answers basic questions directly, retrieves procedural answers from the SOPs with measured relevance, escalates to a human with context preserved, lets clinic staff upload SOPs via presigned URLs, and produces latency, deflection and handoff metrics for decision-making.

  • Real-time voice on an unified modelPersonaPlex with Nemotron 3 VoiceChat handles full-duplex speech-to-speech on a single L40S GPU, so callers can interrupt and be understood naturally.
  • Answers grounded in the clinic's SOPsA retrieval layer over an OpenSearch vector index, with a Lambda fallback, surfaces the right SOP passage for procedural questions rather than relying on model pre-training.
  • Handoff as a first-class behaviourThe agent escalates on uncertainty and carries the full conversation context to the human, so patients never repeat themselves, and handoff rate is measured alongside deflection.
  • Clinic-controlled knowledgeAPI Gateway presigned URLs let staff upload and update SOPs to S3 without AWS access, and scheduled re-ingestion syncs daily changes automatically without a redeploy.
  • HandoverThe client owns and operates the deployed infrastructure independently, with a staff handoff guide, a SOP management guide, a monitoring dashboard, and a go or no-go results summary.

With the highest-risk components validated first and the full stack handed over, the client is positioned to move from proof of concept to a live clinic pilot, and from a single

Clinic to a multi-clinic rollout, once the confirmed success metrics and any compliance gates are cleared.

AWS Stack

Amazon EC2

(G6e, NVIDIA L40S GPU) for the compute running PersonaPlex with Nemotron 3 VoiceChat.

Amazon OpenSearch Serverless

For the vector index that retrieves relevant SOP passages.

Amazon Bedrock

For embedding the SOPs into the retrieval index.

AWS Lambda

For the retrieval fallback tool and the scheduled SOP re-ingestion job.

Amazon API Gateway

For presigned URL uploads so clinic staff manage SOPs without AWS console access.

Amazon S3

And Amazon CloudWatch for SOP and transcript storage and for audit, monitoring and compliance logging.

YOU MIGHT LIKE

Related success stories

View all case studies

Case Studies | Insights

Utilising Language Recognition, Speed, and Enhanced Security to Make Social Media a Force for Good

  • Here, we take a detailed look at how the Cloud Combinator team collaborated with another cutting-edge AI service provider that provides intelligent systems to “make social media more social” for brands and users alike.
  • Arwen AI is a UK-based startup specialising in AI solutions to manage and enhance brands’ social media interactions. Founded in 2020 by Matt McGrory, Dr. David Cole, and Joel Bailey, Arwen. AI focuses on using AI to automatically detect and remove spam, toxic comments, and other unwanted content from social media platforms.
  • The team at Arwen have three core products. ‘Moderate’ is focused on identifying and removing toxic content from social media channels. ‘Engage’ helps brands identify and engage with meaningful conversations on social media, and ‘Customize’ allows brands to apply bespoke algorithms to their channels - creating an even more effective moderation and engagement.
Read more
CONTACT US

Ready to turn AI into impact?

We'll help you spot the highest-value opportunities, reduce risk around your first AI initiative, and define a clear path to results from day one.

Why talk to us:

Outcome-driven recommendations

AWS-recognised delivery expertise

Risk-aware AI adoption

Clear next step, not a sales pitch

Start with a focused 20-minute conversation about your goals — no pressure, no commitment.

This website uses cookies to enhance user experience and to analyze performance and traffic on our website.

See our Privacy Policy for details.