Case Studies | FinTech

Turning futures market data into predictive trading signals

TFT Model Training Pipeline

About

A FinTech business working with Cloud Combinator on AWS. The client is anonymised at their request.

Challenge

The challenge had three focus areas, each a prerequisite for a trustworthy signal.

Ingesting and preparing heterogeneous market data

The model depends on large volumes of historical futures and equities data drawn from multiple providers, in inconsistent formats. That data had to be ingested reliably from Amazon S3, cleaned, normalised and partitioned into a consistent shape before any feature work could begin. Get this wrong and every downstream signal inherits the error.

Engineering predictive features at scale

Raw ticks are not signals. The pipeline needed to compute the specific features that express mispricing, including Basis Z-Score as a valuation signal, Tick-Delta as a latency signal and EWMA Volatility, and turn them into a clean time series ready for training. This had to run over tens of gigabytes efficiently and repeatably.

Training and validating to an accuracy bar

The TFT model had to be trained on GPU-accelerated compute and then held to an explicit standard: an AUC above 0.70 on validation data, tested against separate stress periods. Without a defined bar and a disciplined validation and stress-testing regime, there is no basis for trusting the signal in a live product.

Solution

We built the pipeline in clear stages so each layer could be validated before the next was added, and so the client's engineers could follow every design decision.

AUC > 0.70

Target model accuracy bar on validation data

12-18 hrs

GPU training run on EC2 G5.xlarge (A10G)

~$886/mo

Projected AWS running cost for the pipeline

By the numbers:

  • AUC > 0.70 - Target model accuracy bar on validation data
  • 12-18 hrs - GPU training run on EC2 G5.xlarge (A10G)
  • ~$886/mo - Projected AWS running cost for the pipeline
Changes

The engagement delivered a complete, documented training pipeline in the client's own AWS environment, meeting the acceptance criteria set out at the start: successful data ingestion and feature extraction, a trained TFT predictive model, and a validated handover to the client's technical team.

  • Reliable data foundationA S3-to-Glue batch ingestion path cleans, normalises and partitions raw market data into a consistent, training-ready shape.
  • Predictive features on demandManaged Flink computes Basis Z-Score, Tick-Delta and EWMA Volatility features into a reusable time series, ready for repeated training runs.
  • A model held to a standardThe TFT model is trained on GPU compute and validated and stress tested against defined 2023 to 2025 periods, with an explicit AUC above 0.70 acceptance goal.
  • Signals ready to consumeThe pipeline produces mock trading signals in JSON, structured for interpretation by NinjaTrader.
  • HandoverA runbook, reference architecture and knowledge-transfer session left the client able to run, monitor and iterate on the pipeline independently.

With the pipeline proven, the client's roadmap points to scaling through 2026, moving from batch validation towards live data feeds and larger datasets on additional GPU instances, to support the third-party provider subscription signal service. These are the client's forward-looking projections rather than delivered outcomes.

AWS Stack

Amazon S3

For scalable storage of raw, enriched and output data across the pipeline.

AWS Glue

For batch ingestion, cleaning and partitioning of historical market data.

Amazon Managed Service

For Apache Flink for feature engineering over the prepared data.

Amazon EC2

(G5.xlarge, NVIDIA A10G GPU) for GPU-accelerated TFT model training.

Amazon EBS

For high-performance storage attached to the training instance.

YOU MIGHT LIKE

Related success stories

View all case studies

Case Studies | Insights

Utilising Language Recognition, Speed, and Enhanced Security to Make Social Media a Force for Good

  • Here, we take a detailed look at how the Cloud Combinator team collaborated with another cutting-edge AI service provider that provides intelligent systems to “make social media more social” for brands and users alike.
  • Arwen AI is a UK-based startup specialising in AI solutions to manage and enhance brands’ social media interactions. Founded in 2020 by Matt McGrory, Dr. David Cole, and Joel Bailey, Arwen. AI focuses on using AI to automatically detect and remove spam, toxic comments, and other unwanted content from social media platforms.
  • The team at Arwen have three core products. ‘Moderate’ is focused on identifying and removing toxic content from social media channels. ‘Engage’ helps brands identify and engage with meaningful conversations on social media, and ‘Customize’ allows brands to apply bespoke algorithms to their channels - creating an even more effective moderation and engagement.
Read more
CONTACT US

Ready to turn AI into impact?

We'll help you spot the highest-value opportunities, reduce risk around your first AI initiative, and define a clear path to results from day one.

Why talk to us:

Outcome-driven recommendations

AWS-recognised delivery expertise

Risk-aware AI adoption

Clear next step, not a sales pitch

Start with a focused 20-minute conversation about your goals — no pressure, no commitment.

This website uses cookies to enhance user experience and to analyze performance and traffic on our website.

See our Privacy Policy for details.