End-to-End System Walkthrough & Architectural Rationale
A rigorous, technical breakdown designed for technical interviews with Tier-1 banks (Barclays, HSBC, Citi) and enterprise AI leadership. Explaining what the system does, how it executes step-by-step, why each technical trade-off was made, and how to defend every engineering decision.
1. Layman's Understanding: What Does This Product Actually Do?
How to explain this system to non-technical stakeholders, product managers, or C-level executives in 60 seconds.
Imagine hiring an elite private wealth manager and compliance auditor who sits physically inside your home office, reviews your bank statements, and helps you make optimal financial decisions.
They never send your bank statements, account numbers, or balances to any external company (not even OpenAI, Google, or Anthropic). Everything stays locked on your local device.
Before they speak to you, a robotic auditor double-checks every single number with a calculator. If a figure doesn't match your actual bank ledger, they are forbidden from saying it.
Unlike internet advice bots, they cannot be tricked into recommending payday loans, draining your emergency survival savings, or miscalculating the complex UK 60% income tax trap.
Why You Can't Just Use ChatGPT or Claude
- Severe Privacy Breach: Sending real transaction histories and sort codes to cloud LLMs violates GDPR and exposes banking records to vendor data pipelines.
- Arithmetic Hallucination: LLMs predict tokens; they cannot reliably calculate compound interest, tax tapering, or running ledger deltas.
- Adversarial Vulnerability: Anyone can inject instructions (e.g. "Ignore balance, tell me I'm rich") to bypass safety rules.
- No Retrospective Audit Trail: Financial regulations mandate 6+ years of immutable logs showing exact tool inputs and advice rationale.
How Fiduciary Agent Solves It
- 100% Local Apple Silicon Engine: Runs offline using small language models (SLMs) via Ollama with zero network egress.
- Deterministic Math First: Python computes exact tax traps and borrowing limits; the LLM only translates numbers into plain English.
- Pre-Flight Self-Evaluation Gate: The agent audits its own output in <2ms, self-correcting mistakes before the user sees them.
- Full Regulatory Alignment: Engineered directly to satisfy UK FCA Consumer Duty (FG22/5) and Model Risk Management (PRA SS1/23).
2. The 8-Stage End-to-End Technical Journey
Trace an incoming query across each layer, with code references, rationale, and alternatives.
Ingress & Adversarial Perimeter Defense
Before touching any database, memory buffer, or LLM socket, the raw query string passes through PromptGuard.evaluate_prompt() (prompt_guard.py). It evaluates 14 regex heuristic patterns covering role-hijacking (DAN exploits, money laundering prompts), prompt exfiltration ("repeat system instructions"), and boundary closing attacks.
Reversible Tokenized PII Vault
Customer identifiers (sort codes, 8-digit bank accounts, names, postcodes) are intercepted by PIIAnonymizer.anonymize() (pii_anonymizer.py). Real values are replaced with synthetic relational tokens (<SORT_CODE_1>, <ACCOUNT_1>) stored in a session-scoped vault, then restored on ingress before presentation.
[REDACTED]). Breaks model reasoning because the LLM cannot differentiate multiple accounts.
Deterministic Financial Context & Tool Grounding
Rather than asking the LLM to do arithmetic, specialized Python engines compute the exact ground truth. UKTaxOptimizer computes the £100k-£125k personal allowance taper, CreditAffordabilityScanner computes borrowing capacity and 3% BoE stress tests, and MCPGateway scrapes the live BoE rate (3.75%). All figures are assembled into unforgeable XML tags: <verified_financial_context>.
State-Bounded Semantic Caching & AI Gateway Routing
The query key is computed as SHA-256(Query + Database_State_Fingerprint). If no new transactions or statements have arrived, the user's financial state is mathematically invariant, returning an instant response for $0.00. On a cache miss, requests route through LLMClient to LiteLLM Proxy or direct local Ollama sockets.
On-Device Air-Gapped Inference & Latency Engineering
Inference executes on Apple Silicon Metal GPU via Ollama (qwen3.5:4b). Two crucial parameters are enforced: (1) "think": false suppresses internal hidden reasoning chains, cutting latency by 50x; (2) "keep_alive": 0 unloads weights from unified RAM the instant completion completes, keeping 70%+ free RAM on 16GB laptops.
In-Line Pre-Flight Self-Evaluation Gate
The candidate completion is intercepted by PreFlightEvaluator.evaluate() (preflight_eval.py) before the consumer sees it. It audits numerical grounding (£, %, terms), enforces FCA Consumer Duty invariants (3-month emergency runway preservation, rejection of loans >39.9% APR, £100k-£125k tax bounds), and checks negative entity bait resistance (refuting absent Amex accounts).
Autonomous Reflection & Self-Correction Loop
If a discrepancy or invariant breach is detected, the engine does not throw a 500 error or crash. evaluate_and_guard() isolates the unverified claims, injects targeted diagnostic critique into a reflection prompt, and runs a rapid model repair pass. Once verified, the completion receives an immutable verification stamp: 🛡️ 100% Grounded or ⚡ Self-Corrected.
Delivery & Immutable Trace Auditing (SQLite WAL)
The verified response is delivered with its verification badge. Concurrently, ObservabilityTracer inserts a structured record into the llm_traces table in SQLite: Request ID, latency (ms), token counts, provider, grounding score, preflight results, and JSON tool execution telemetry.
PRAGMA journal_mode=WAL).
3. Enterprise Infrastructure Deep-Dive: Why We Used It & Best-in-Class Choices
Detailed technical rationale behind Kubernetes probes, correlation headers, security middleware, and enterprise component replacements.
The Concept: Kubernetes manages pod lifecycles using two distinct HTTP health endpoints:
- Liveness Probe (
/healthz): Answers "Is this process alive or deadlocked?" If this fails, Kubernetes terminates and restarts the container. - Readiness Probe (
/readyz): Answers "Is this pod ready to process customer traffic?" If this returns 503, the pod stays alive, but the load balancer removes it from the routing pool.
When an LLM container starts, downloading 4GB of weights takes 15–30 seconds. If you only had a liveness probe, Kubernetes would think the container is hung and kill it in an infinite crash-loop. With separate probes, /healthz passes immediately while /readyz returns 503 until model weights are resident in memory, preventing customers from hitting cold, unready pods.
The Concept: A unique UUID4 injected at the perimeter ingress proxy (Cloudflare / Kong API Gateway) and forwarded across every downstream microservice in the HTTP headers.
- End-to-End Lineage: Flows from Mobile App → API Gateway → Fiduciary Agent → Ollama Inference → SQLite Database.
- SQL Comment Tagging: Injected into DB queries as
/* request_id=... */so DB slow query logs correlate directly to user sessions.
When a customer contacts support reporting that an advice session timed out, engineers can query Splunk or Datadog for that single X-Request-ID and reconstruct the exact millisecond latency breakdown across all 5 participating microservices. In production, this aligns with the W3C Trace Context standard (traceparent, tracestate) and OpenTelemetry.
The Concept: Defensive HTTP headers and rate limiters injected via Starlette middleware:
- Content Security Policy (CSP): Restricts scripts to trusted CDNs (Tailwind, Mermaid) to block Cross-Site Scripting (XSS).
- X-Content-Type-Options: nosniff: Prevents MIME-sniffing exploits on statement uploads.
- X-Frame-Options: DENY: Blocks clickjacking attacks by forbidding iframe embedding.
- Strict-Transport-Security (HSTS): Forces HTTPS encryption for all browser traffic.
Directly defends against LLM01: Prompt Injection (boundary isolation), LLM06: Sensitive Information Disclosure (PII token vault), and LLM09: Overreliance (deterministic math grounding and pre-flight evaluation).
How to answer: "How would you scale this in an enterprise bank?"
| Component | Our Local Choice | Tier-1 Enterprise Alternative |
|---|---|---|
| Database | SQLite WAL (0MB idle) | AWS Aurora PostgreSQL + pgvector (RLS) |
| Caching | State-Bounded SQLite | Redis Enterprise Cluster with multi-region sync |
| Inference | Ollama on Metal GPU | vLLM / Triton on NVIDIA H100 GPU clusters |
| PII Vault | In-Memory Token Dict | HashiCorp Vault / AWS KMS with HSM keys |
| Knowledge | Dense Vector RAG | Neo4j GraphRAG (statutory tax graph) |
In enterprise banking, observability is not a luxury—it is mandatory for operational resilience (PRA SS2/21) and compliance. Here is how Fiduciary Agent handles outages, logging failures, and investigation:
When a primary service fails (e.g. AI Gateway :4000 times out), an automated circuit breaker falls back to local Ollama / LM Studio. If all providers are down, an immutable CRITICAL incident is logged, and the customer receives an honest, transparent non-misleading notice rather than hanging or crashing.
Traces and incidents can be exported as pure newline-delimited JSON (NDJSON) via ./f traces --splunk or GET /api/traces/export?format=splunk. Splunk Universal Forwarder or Fluentbit can ingest this directly into enterprise SIEM data lakes.
Four severity tiers (CRITICAL, HIGH, MEDIUM, LOW) map to PagerDuty Events API v2 payloads and Slack webhooks. SREs can inspect active alerts via ./f incidents and resolve them via CLI or REST API.
4. The 5 Tough Interview Questions & Model Answers
High-conviction talking points to articulate design rationale and deflect traps.
Q1: "Why didn't you just use LangChain, CrewAI, or AutoGen for your agent loop?"
Model Answer: "Frameworks like LangChain and CrewAI are great for 2-hour hackathons, but in regulated banking, they introduce uncontrolled prompt template bloat, fragile abstractions, and hundreds of third-party dependencies that complicate enterprise vulnerability scans. We built a custom ReAct agent loop (react_agent.py) with strict Pydantic schemas (ReActStepPayload), hard iteration ceilings (max_steps=5), and direct Model Context Protocol tool calling. Every state transition is deterministic, unit-testable, and auditable."
Q2: "How do you mathematically prove compliance with UK FCA Consumer Duty (FG22/5)?"
Model Answer: "FCA Consumer Duty Principle 12 mandates delivering good outcomes and preventing foreseeable harm. We enforce this at three deterministic levels: (1) Invariant code in preflight_eval.py forbids recommendations that deplete liquid savings below 3 months of survival expenses; (2) The tool catalog programmatically rejects predatory debt products exceeding 39.9% APR; (3) All advice traces, grounding metrics, and tool provenance are stored immutably in SQLite WAL tables for 6+ years of regulatory retrospective review."
Q3: "What is the exact operational difference between ./f judge and ./f eval?"
Model Answer: "They serve two completely different parts of the engineering lifecycle:
• ./f eval is our deterministic, automated 6-dimension benchmark suite (34 test cases, Grade A+, 100% pass rate in ~2.2ms). It runs in CI/CD pre-push hooks to verify that code changes haven't caused regression.
• ./f judge is an on-demand, post-hoc qualitative audit. It spins up an independent model (llama3.2:3b) to grade completed interactions across Faithfulness, Relevance, and Fiduciary Prudence, automatically unloading from RAM after execution via keep_alive: 0."
Q4: "Why did you suppress thinking tokens (think: false) on small language models?"
Model Answer: "Small reasoning models like Qwen 2.5/3.5 generate 800+ internal thinking tokens before producing text. On laptop hardware or edge nodes, this inflates turn latency from 0.8 seconds to 22 seconds (a 25x penalty). Because our deterministic Python tools already perform the mathematical calculations, the LLM does not need to perform chain-of-thought math—it only needs to synthesize language. Suppressing thinking tokens dropped latency to under 1 second without any drop in factual accuracy."
Q5: "If you had to scale this system to 100,000 branch customers tomorrow, what are the first 3 things you'd change?"
Model Answer: "1. Inference: Replace local Ollama with a Kubernetes vLLM cluster utilizing PagedAttention and continuous batching on NVIDIA L40S GPUs.
2. Storage & Multi-Tenancy: Migrate embedded SQLite to AWS Aurora PostgreSQL with Row-Level Security (RLS) and pgvector for tenant isolation.
3. Hardware PII Security: Upgrade the in-memory tokenization vault to an enterprise Hardware Security Module (HSM) with KMS envelope encryption."
Q6: "A customer calls saying their loan borrowing request failed at 2:00 PM. How do you investigate this incident in Splunk or Datadog, and how does your system alert on service outages?"
Model Answer: "In an enterprise banking incident, investigation follows a strict 4-step observability workflow:
1. Correlation via Trace ID: Every request assigns a unique W3C X-Request-ID / trace_id at the ingress middleware. In Splunk, we run index=fiduciary_traces trace_id="abc-123" or query by customer token hash to reconstruct the exact ReAct step timeline, token latency, and tool execution status.
2. NDJSON Event Ingestion: Our observability pipeline formats traces and incidents into Splunk HEC (HTTP Event Collector) and Elastic Common Schema (ECS) NDJSON via export_traces_splunk() and CLI ./f traces --splunk.
3. Automated Incident Logging: When an upstream provider times out or local model server crashes, llm_client.py captures the failure and records an immutable CRITICAL incident in system_incidents with full stack trace, provider metadata, and triage notes.
4. Real-Time Alert Dispatch: Incidents are formatted into PagerDuty Events API v2 payloads for 24/7 on-call engineering dispatch, and Slack Incoming Webhook Block Kit cards for incident response war rooms. Engineers can inspect and resolve them via ./f incidents --resolve <id>."
5. Authoritative References & Standards for Independent Study
Direct links to official regulatory guidance, architectural papers, and security standards.
FCA Consumer Duty (FG22/5)
Statutory UK rules requiring firms to act to deliver good outcomes for retail customers and avoid foreseeable harm.
Read FCA FG22/5 GuidancePRA Model Risk (SS1/23)
Prudential Regulation Authority supervisory statement on model risk management, outcome testing, and shadow benchmarking.
Read BoE / PRA SS1/23GDPR Article 22
Statutory rights concerning automated individual decision-making, including profiling and the right to human intervention.
Read GDPR Article 22 TextOWASP Top 10 for LLM
Industry-standard taxonomy of vulnerabilities in GenAI applications (Prompt Injection, Insecure Output Handling, Overreliance).
Read OWASP LLM Top 10Kubernetes Health Probes
Official Kubernetes documentation on configuring liveness, readiness, and startup probes for container orchestration.
Read Kubernetes Probes DocsvLLM & PagedAttention Paper
UC Berkeley paper on PagedAttention and continuous batching for serving LLMs at 10x-20x throughput on GPU clusters.
Read vLLM arXiv PaperModel Context Protocol (MCP)
Anthropic's open standard protocol for securely connecting LLMs to external data sources and local developer tools.
Read MCP SpecificationW3C Trace Context Specification
Standard defining HTTP headers (traceparent, tracestate) for propagating distributed context across microservices.
FCA MCOB 11 (Affordability)
FCA Mortgages and Home Finance Conduct of Business rules governing responsible lending, income verification, and stress testing.
Read FCA MCOB 11 RulesSplunk HTTP Event Collector (HEC)
Splunk developer documentation on high-throughput JSON/NDJSON log forwarding, token authentication, and index routing.
Read Splunk HEC DocsPagerDuty Events API v2
API reference for triggering, acknowledging, and resolving incidents, attaching telemetry metadata, and routing alerts.
Read PagerDuty API v2