github.com/cloudcruncher Open Source & Enterprise Architecture

Systems & Architectures

A curated showcase of verified public open-source engineering systems and tier-1 enterprise architectures delivering mission-critical scale in UK banking.

Verified Public Code MIT / Open Source

Public Open-Source Repositories

Production-grade open-source implementations authored and maintained on GitHub.

Enterprise Scale Regulated UK Financial Services

Enterprise Architecture Case Studies

High-impact modernization and AI deployments delivered across NatWest Group and Lloyds Banking Group.

Enterprise Agentic AI NatWest Group (Enterprise Innovation)
Aug 2025 – Present

Data Foundation for Enterprise Agentic Investigation & Decision Support

KEY IMPACT: >90% reduction in manual case-handling time across 300+ daily onboarding applications

The data engineering side of enterprise agentic workflows: connecting agents to data sources, giving them trustworthy datasets and context, and helping take proof-of-concept agents to production. Delivered through reusable Model Context Protocol (MCP) servers, Pydantic data contracts, quality gates and a resilient Snowflake writer.

DELIVERY & ARCHITECTURAL HIGHLIGHTS
  • Packaged a reusable, PII-guardrailed MCP server (bounded JSON slices, role-based field allow-lists, full audit trail) consumed by downstream agents with zero ad-hoc data access code, and reused end-to-end by DAVE
  • Built the data engineering behind DAVE (Decision Assist for complaints) and the LLM-as-a-judge evaluation microservice, including a near-real-time Kafka evaluation pipeline with a batch REST path writing metrics to Snowflake
  • Built PySpark model-monitoring pipelines on AWS EMR Serverless (PostgreSQL and on-prem sources to S3 Parquet and Snowflake) computing accuracy, precision, recall, F1, LLM token usage and latency, orchestrated by Airflow
  • On the Colleague Assist / CAI contact-centre programme, modelled AWS Connect call transcripts and contact-trace records in Snowflake (turn-timing and silence analytics) for Data Science
  • Complaints investigation delivered in June, helping reimagine the complaints process with AI (see NatWest’s “Transforming Complaints with GenAI”)
  • Designed Pydantic schema validation gates and audit trails meeting UK banking and financial-crime standards
  • Engineered Splunk observability pipelines covering agent execution, tool-call performance, and decision quality
  • Built a resilient Snowflake writer for the agent evaluation service: jittered, batched writes so the warehouse is not woken for every insert, reducing deployment issues to near zero
PythonMCP ProtocolKiro CLISnowflakeAirflowPySparkAWS EMR ServerlessKafkaPostgreSQLSplunkPydantic
Lakehouse Modernisation NatWest Group (Climate Analytics)
Jul 2023 – Aug 2025

1TB+/Month Snowflake ESG & Climate Analytics Platform

KEY IMPACT: 40% faster batch processing and 30% reduction in data quality incidents

Led the engineering team delivering NatWest’s ESG and climate data platform on Snowflake, processing 1TB+ of data monthly from 50+ internal and external sources. Built with Snowflake, Airflow and dbt in a layered, medallion-style design, governed by a Data Guardian framework (contracts and data quality), enriched with many third-party datasets, and published as emissions data products on a data marketplace.

DELIVERY & ARCHITECTURAL HIGHLIGHTS
  • Layered, medallion-style Snowflake design (RAW → INT → PRS → presentation), orchestrated with Airflow on AWS MWAA and transformed with dbt
  • Data Guardian framework: data contracts and data-quality checks so each source is trusted before it is used
  • Ingested many third-party data sources to enrich the ESG lakehouse alongside internal data
  • Published emissions data products to a data marketplace for downstream consumers
  • Cut batch processing by 40% (PySpark and dbt optimisation) and data quality incidents by 30% (metadata-driven ingestion where new sources are onboarded by config not code, idempotent duplicate-load guards, dbt snapshots and automated dbt tests)
  • Cut Snowflake compute cost with per-model warehouse routing and transient tables; automated dbt deployment to MWAA via a build-and-promote pipeline (Artifactory → S3) with CloudWatch and SNS failure alerting
  • Worked with Data Architects and Data Scientists to deliver production climate risk models; set team engineering standards and mentored junior engineers
SnowflakeAirflow (AWS MWAA)dbtPySparkPythonSQLGitLab CI/CDDocker
Retail Analytics NatWest Group (Retail Data & Analytics Decisioning)
Sep 2021 – Jul 2023

Retail Data Marts for Decision-Making

KEY IMPACT: Trusted customer, mortgage and deposit data for retail leadership decisions

Built enterprise batch pipelines and Snowflake data models for the Retail Data & Analytics Decisioning team. Customer, mortgage and deposit data marts gave the business a dependable foundation for better decisions by retail leadership.

DELIVERY & ARCHITECTURAL HIGHLIGHTS
  • Built customer, mortgage and deposit data marts for retail analytics and decisioning
  • Engineered enterprise batch pipelines in Python, SQL, Snowflake, PySpark and StreamSets, including PySpark pipelines for Customer Lifetime Value (CLV) analytics
  • Designed Snowflake data models for high-volume analytical workloads, tuned with clustering, partitioning and materialised views
SnowflakePySparkPythonSQLStreamSets