Data platforms
AI agents can trust,
and precision-extracted pour-overs.
I'm Robin Saini. 15+ years in software & data engineering (10+ in data engineering), currently the data engineer on agentic AI projects at NatWest Group, connecting AI agents to data sources and giving them trustworthy datasets and context. Before that I led a 50+ source Snowflake platform. I build with AI coding tools (Kiro CLI at work, Claude Code and Google Antigravity on my own projects), and my personal projects are public on GitHub.
Working with AI Agents: Trustworthy Data & Context
Agents are only as good as the data and context behind them. At NatWest I'm the data engineer on agentic projects: connecting agents to the right data sources, giving them datasets they can trust, and building the context that leads to good outcomes.
Observability runs across all of it: Splunk dashboards for agent execution, tool-call performance and decision quality.
At NatWest: the data layer for agents
Data contracts, schema validation and quality gates so agent inputs meet banking and financial-crime standards. Reusable MCP servers consumed by several agents. Retrieval and context pipelines debugged at the source. A resilient Snowflake writer with jittered back-off, and Snowflake CI/CD with automated tests.
How I build: AI coding harnesses
I use AI coding harnesses to build faster: Kiro CLI at work, and Claude Code and Google Antigravity on my own projects. I use them to write and test pipelines, MCP servers and tests. They are tools I build with, not the systems I deliver.
Personal projects: other use cases
Different problems, all public on GitHub: a governed lakehouse that feeds an assistant on live customer calls (open-lakehouse), a local-first personal finance agent (fiduciary-agent), and markets and payments pipelines (lakehouse-markets-data).
Featured GitHub Projects
Personal projects on GitHub, each a different use case: governed data for AI assistants, a local-first finance agent, and markets data pipelines.
Fiduciary Agent - Autonomous Personal Finance & Intelligence
Autonomous personal fiduciary intelligence harness. 100% on-device local AI (Apple Silicon Metal GPU), UK Open Banking (TrueLayer/Wise), deterministic math core & tax optimization.
London Pulse - Daily Open-Data Tracker for London Food & Drink
Live PWA tracking what is opening, closing and changing across London's 33 boroughs, built in public from open data (FSA hygiene registers, Companies House, TfL, police.uk). Day-over-day diffs, a brand tracker for roasters, bakeries and chains, an Area guide for movers (nearby venues, stations, crime vs London average, compare two areas), and an in-browser SQL lab with charts.
Open Lakehouse - Enterprise Streaming & Multi-Tenant Platform
Production-shaped open lakehouse: Apache Iceberg + Polaris Catalog + Trino + Open Policy Agent (OPA), CDC streaming, and a governed AI assistant. Self-healing, chaos-tested, SLO-driven.
Lakehouse Markets Data - Markets & Payments Intelligence
Markets & Payments Intelligence: Data engineering tenant of open-lakehouse handling Coinbase crypto streams, card authorisations, FX rates, sanctions screening; Kappa architecture and medallion data products.
Latest Field Notes & Articles
Reflections on lakehouse performance, data for AI agents, and pour-over mechanics.
Lessons from Building a Local-First Fiduciary Agent: Deterministic Maths, Grounded AI
What I learned building fiduciary-agent, a personal open-source finance harness that runs on-device: keep arithmetic out of the LLM, ground every answer, guard against injection, and choose boring storage.
Production Agentic Systems Are a Data Engineering Problem: MCP Servers, Contracts and Observability
What a data engineer actually does to take AI agents from proof of concept to production: reusable MCP servers, data contracts, resilient writers and observability.
Building a 1TB+/Month ESG Data Platform on Snowflake: Medallion Layers, Data Guardian and Data Products
How a data engineering team built NatWest's 50+ source ESG and climate platform on Snowflake, Airflow and dbt, with contracts and data quality built in, third-party enrichment, and emissions products published to a data marketplace.