Public Open-Source Repositories
Production-grade open-source implementations authored and maintained on GitHub.
Fiduciary Agent - Autonomous Personal Finance & Intelligence
Autonomous personal fiduciary intelligence harness. 100% on-device local AI (Apple Silicon Metal GPU), UK Open Banking (TrueLayer/Wise), deterministic math core & tax optimization.
London Pulse - Daily Open-Data Tracker for London Food & Drink
Live PWA tracking what is opening, closing and changing across London's 33 boroughs, built in public from open data (FSA hygiene registers, Companies House, TfL, police.uk). Day-over-day diffs, a brand tracker for roasters, bakeries and chains, an Area guide for movers (nearby venues, stations, crime vs London average, compare two areas), and an in-browser SQL lab with charts.
Open Lakehouse - Enterprise Streaming & Multi-Tenant Platform
Production-shaped open lakehouse: Apache Iceberg + Polaris Catalog + Trino + Open Policy Agent (OPA), CDC streaming, and a governed AI assistant. Self-healing, chaos-tested, SLO-driven.
Lakehouse Markets Data - Markets & Payments Intelligence
Markets & Payments Intelligence: Data engineering tenant of open-lakehouse handling Coinbase crypto streams, card authorisations, FX rates, sanctions screening; Kappa architecture and medallion data products.
Enterprise Architecture Case Studies
High-impact modernization and AI deployments delivered across NatWest Group and Lloyds Banking Group.
Data Foundation for Enterprise Agentic Investigation & Decision Support
The data engineering side of enterprise agentic workflows: connecting agents to data sources, giving them trustworthy datasets and context, and helping take proof-of-concept agents to production. Delivered through reusable Model Context Protocol (MCP) servers, Pydantic data contracts, quality gates and a resilient Snowflake writer.
- Packaged a reusable, PII-guardrailed MCP server (bounded JSON slices, role-based field allow-lists, full audit trail) consumed by downstream agents with zero ad-hoc data access code, and reused end-to-end by DAVE
- Built the data engineering behind DAVE (Decision Assist for complaints) and the LLM-as-a-judge evaluation microservice, including a near-real-time Kafka evaluation pipeline with a batch REST path writing metrics to Snowflake
- Built PySpark model-monitoring pipelines on AWS EMR Serverless (PostgreSQL and on-prem sources to S3 Parquet and Snowflake) computing accuracy, precision, recall, F1, LLM token usage and latency, orchestrated by Airflow
- On the Colleague Assist / CAI contact-centre programme, modelled AWS Connect call transcripts and contact-trace records in Snowflake (turn-timing and silence analytics) for Data Science
- Complaints investigation delivered in June, helping reimagine the complaints process with AI (see NatWest’s “Transforming Complaints with GenAI”)
- Designed Pydantic schema validation gates and audit trails meeting UK banking and financial-crime standards
- Engineered Splunk observability pipelines covering agent execution, tool-call performance, and decision quality
- Built a resilient Snowflake writer for the agent evaluation service: jittered, batched writes so the warehouse is not woken for every insert, reducing deployment issues to near zero
1TB+/Month Snowflake ESG & Climate Analytics Platform
Led the engineering team delivering NatWest’s ESG and climate data platform on Snowflake, processing 1TB+ of data monthly from 50+ internal and external sources. Built with Snowflake, Airflow and dbt in a layered, medallion-style design, governed by a Data Guardian framework (contracts and data quality), enriched with many third-party datasets, and published as emissions data products on a data marketplace.
- Layered, medallion-style Snowflake design (RAW → INT → PRS → presentation), orchestrated with Airflow on AWS MWAA and transformed with dbt
- Data Guardian framework: data contracts and data-quality checks so each source is trusted before it is used
- Ingested many third-party data sources to enrich the ESG lakehouse alongside internal data
- Published emissions data products to a data marketplace for downstream consumers
- Cut batch processing by 40% (PySpark and dbt optimisation) and data quality incidents by 30% (metadata-driven ingestion where new sources are onboarded by config not code, idempotent duplicate-load guards, dbt snapshots and automated dbt tests)
- Cut Snowflake compute cost with per-model warehouse routing and transient tables; automated dbt deployment to MWAA via a build-and-promote pipeline (Artifactory → S3) with CloudWatch and SNS failure alerting
- Worked with Data Architects and Data Scientists to deliver production climate risk models; set team engineering standards and mentored junior engineers
Retail Data Marts for Decision-Making
Built enterprise batch pipelines and Snowflake data models for the Retail Data & Analytics Decisioning team. Customer, mortgage and deposit data marts gave the business a dependable foundation for better decisions by retail leadership.
- Built customer, mortgage and deposit data marts for retail analytics and decisioning
- Engineered enterprise batch pipelines in Python, SQL, Snowflake, PySpark and StreamSets, including PySpark pipelines for Customer Lifetime Value (CLV) analytics
- Designed Snowflake data models for high-volume analytical workloads, tuned with clustering, partitioning and materialised views