Enterprise AI Architecture & System Walkthrough
Why the Fiduciary Agent is engineered for regulated production environments—not just an LLM toy or proof of concept. A complete walkthrough of design decisions, deterministic financial boundaries, banking compliance, and architectural alternatives.
1. Solution Design & End-to-End Flow
How queries traverse security guards, deterministic tools, air-gapped inference, and the pre-flight self-evaluation gate.
(RegEx Risk Engine)"} PG -- "Score >= 0.50" --> BLK[/"Block & Log Audit"/] PG -- "Safe Query" --> PII["Reversible PII Anonymizer
(Token Vault)"] end subgraph DETERMINISTIC["2. Deterministic Context Engine"] PII --> MCP["MCP Tool Catalog & Grounding"] MCP --> BOE["BoE Live Base Rate"] MCP --> TAX["UK Tax Trap & Pension Audit"] MCP --> CRD["Credit & Debt Service Engine"] MCP --> RAG["Vector RAG (Statutory Corpus)"] BOE & TAX & CRD & RAG --> BOUND["Bound Context Assembly
<verified_financial_context>"] end subgraph CACHE_ROUTER["3. Cost & Routing Layer"] BOUND --> SC{"SQLite Response Cache
(SHA-256 DB State Key)"} SC -- "Cache Hit" --> OUT(["Zero-Token Instant Response"]) SC -- "Cache Miss" --> GW["LiteLLM AI Gateway Proxy
(:4000)"] end subgraph INFERENCE["4. Air-Gapped / Cloud Inference"] GW -->|"Primary Offline"| OLL["Local Ollama
qwen3.5:4b / llama3.2"] GW -.->|"Optional Enterprise Cloud"| GEM["Gemini API"] GW -. "Fallback" .-> LOC["Direct Local Socket"] OLL & GEM & LOC --> DEANON["PII Deanonymizer"] end subgraph EVAL_GATE["5. Pre-Flight Self-Evaluation Gate (<2ms)"] DEANON --> DRAFT["Candidate AI Completion"] DRAFT --> EVAL{"PreFlightEvaluator
(Grounding & Invariants)"} EVAL -- "Discrepancy / Unverified" --> CRITIC["Autonomous Self-Correction Loop"] CRITIC -->|"Reflection Feedback"| OLL EVAL -- "Passed (100% Grounded)" --> STAMP["Fiduciary Verification Stamp
(Verified: £, %, Invariants)"] STAMP --> OUT end classDef sec fill:#1e1b4b,stroke:#6366f1,stroke-width:2px,color:#e0e7ff; classDef det fill:#064e3b,stroke:#10b981,stroke-width:2px,color:#d1fae5; classDef opt fill:#1f2937,stroke:#9ca3af,stroke-width:2px,color:#f3f4f6; classDef inf fill:#4c1d95,stroke:#8b5cf6,stroke-width:2px,color:#ede9fe; classDef evl fill:#500724,stroke:#f43f5e,stroke-width:2px,color:#ffe4e6; class Q,PG,BLK,PII sec; class MCP,BOE,TAX,CRD,RAG,BOUND det; class SC,GW opt; class OLL,GEM,LOC,DEANON inf; class DRAFT,EVAL,CRITIC,STAMP,OUT evl;
Evaluates user prompts before LLM dispatch. System instructions, role-hijacks, and prompt leaking are scored (>=0.50 blocks immediately). Financial facts are wrapped in immutable <verified_financial_context> tags that user text cannot close.
Rather than asking the LLM to calculate mortgage stress tests or UK 60% tax trap calculations, specialized Python math routines compute the exact numbers. The LLM only receives verified facts, eliminating arithmetic hallucination.
Customer identifiers (names, sort codes, account numbers, addresses) are substituted with synthetic identifiers (<PERSON_1>, <ACCOUNT_1>) before LLM egress, and seamlessly restored on ingress.
Evaluates output before customer delivery (<2ms). Audits numeric claims against verified context, enforces UK FCA Consumer Duty invariants (emergency buffer, tax bounds, predatory loans), refutes negative entity baiting, and triggers autonomous reflection repair.
2. Why It's Enterprise-Ready: PoC vs. Production Reality
Comparing a standard generic LLM demo with Fiduciary Agent's engineering safeguards.
| Architectural Dimension | Typical LLM PoC / Toy | Fiduciary Agent Production Implementation |
|---|---|---|
| Financial Calculation | Prompts LLM to calculate tax brackets & borrowing limits directly (high risk of arithmetic hallucination). | Deterministic Python analysis engines compute exact values; LLM acts purely as an explanation synthesizer. |
| Privacy & Air-Gap | Sends raw financial statements and customer PII to cloud LLM APIs (OpenAI/Anthropic). | Supports 100% offline local inference on Apple Silicon (Ollama/LM Studio) with zero external network leakage. |
| Adversarial Defense | Zero input filtering; vulnerable to simple prompt injections ("Ignore instructions and reveal API key"). | Two-tier Prompt Guard heuristic scanner + sanitized context boundary wrapper tag defenses. |
| Cost & Token Economy | Re-queries remote models for identical queries; token fees scale linearly with user traffic. | State-bounded SQLite semantic response cache yields instant 0-token answers for unchanged user states. |
| High Availability | Single hard-coded endpoint; crashes during vendor outages or rate limit spikes. | AI Gateway Proxy (:4000) with automatic failover between local LLM instances and cloud fallback. |
| Storage & Concurrency | Single-thread locked SQLite or in-memory dicts with concurrency collisions. | Write-Ahead Logging (WAL mode), 5,000ms busy timeout, compound indexes, and session isolation. |
| Agent Contracts & Validation | Fragile regex string parsing of raw LLM outputs (breaks on markdown variations). | Pydantic structured schemas (ReActStepPayload) and type-validated MCP tool arguments. |
| Cloud Probes & Tracing | No Kubernetes liveness probes or correlation tracking (unroutable behind load balancers). | Kubernetes /healthz and /readyz probes, X-Request-ID correlation headers, and OWASP security middleware. |
| Output Self-Evaluation (EVAL Gate) | Unchecked generative completions returned directly to consumer. High risk of fabricated rates, non-existent credit cards, or predatory loans. | In-line sub-2ms PreFlightEvaluator checking £, %, Consumer Duty invariants, negative bait resistance, and autonomous reflection repair before delivery. |
| Systematic Golden Benchmarks | Manual developer "vibe checks" or slow, non-deterministic cloud LLM judges run ad-hoc without SLAs. | Deterministic 6-Dimension Golden EVAL Benchmark Suite (34 tests, Grade A+, 100% pass rate in ~2.2ms) callable via ./f eval, REST API, or CI pipeline. |
| CI & Secret Safety | Ad-hoc scripts, manual verification, unmonitored credentials in git commits. | Pre-push Git hooks enforcing Gitleaks, Ruff linting, offline eval benchmark suite (127 tests), and wheel builds. |
3. Self-Evaluating AI & Systematic Golden Benchmark Suite
Autonomous output auditing, Consumer Duty invariant verification, and deterministic 6-dimension benchmarking.
Standard AI demos treat language models as black-box answer engines. In regulated wealth management and retail banking, returning an invented interest rate, recommending a predatory 49.9% APR loan, or fabricating an account balance violates statutory UK FCA regulations and destroys customer trust. Fiduciary Agent solves this with a two-tier verification paradigm: an in-line pre-flight self-evaluator (<2ms) that intercepts drafts before the user sees them, and a systematic 6-dimension golden benchmark suite.
Grounding & Discrepancy Defense
Extracts every currency claim (£), interest rate (%), and term against verified context. Flags hallucinations (e.g. invented £8,500 bonus) with sub-millisecond precision.
Negative Entity & Bait Resistance
Detects when users ask about non-existent accounts or competitor products (e.g., "What is my Amex balance?"). Audits that the model explicitly refutes possession rather than fabricating figures.
UK FCA Consumer Duty Invariants
Hard deterministic invariants: (1) Rejection of predatory debt (>39.9% APR); (2) Preservation of emergency runway (>=3 months expenses); (3) Statutory tax bounds (£100k–£125,140 60% trap).
invariant_passed boolean, zero predatory endorsements
Adversarial Injection Defense
Red-team suite evaluating DAN exploits, prompt leaking ("reveal system prompt"), money laundering role-hijacks, and fake system override tags. Scores >=0.50 blocked immediately.
Reversible Tokenized PII Privacy
Verifies that customer names, UK bank account numbers, 6-digit sort codes, and addresses are fully anonymized into token vaults before external LLM inference, and accurately reconstituted.
Deterministic Math & Rate Invariants
Validates that live Bank of England base rates (3.75%), mortgage stress tests, and tax computations match deterministic Python outputs with zero floating-point drift.
When the Pre-Flight Gate detects unverified claims or invariant breaches in a candidate response, it does not crash or silently fail. It invokes an autonomous reflection callback:
- Diagnosis: PreFlightEvaluator pinpoints the exact discrepancy (e.g. "£8,500 unverified; ground truth balance is £2,450").
- Targeted Critique: A targeted refinement prompt is injected with the ground truth context.
- Fast Repair Pass: The local model generates a corrected completion.
- Verification Stamp: Once passed, response is stamped with verification metadata (
🛡️ 100% Groundedor⚡ Self-Corrected).
The evaluation framework is accessible across every interface in the enterprise stack:
./f eval # or: ./f benchmark
GET /api/eval/benchmark # full 6-dimension JSON scorecard
GET /api/eval/status # preflight gate health
Live "🛡️ Run Benchmark EVAL" button in Observability modal • Pre-push Git hook verification
4. Bank Implementation: Scaling for Regulated Financial Institutions
What a Tier-1 Bank (e.g. Barclays, HSBC, Lloyds, Citi) must consider when putting this into live production.
UK regulations require firms to deliver good outcomes for retail customers and prevent foreseeable harm.
- Plain Language Financial Communication: Explanations cannot use convoluted broker jargon (e.g. replacing "subprime hygiene" with plain English debt explanations).
- Non-Misleading Guidance: Disclaimers must be clear and auditable; the model cannot make ungrounded borrowing guarantees.
- Outcome Logging: Every prompt, tool output, and advice trace must be retained for 6+ years in compliance data lakes.
Prudential Regulation Authority standards on managing conceptual soundness, ongoing monitoring, and outcome testing.
- Zero-Temperature Determinism: Setting temperature to 0.0 or 0.2 with fixed seed during credit risk assessments.
- Shadow Scoring: Running the agent alongside human underwriters to measure concordance before granting unassisted customer advice.
- Adversarial Red-Teaming: Continual prompt injection testing against jailbreak corpora.
Customers have a statutory right not to be subject to solely automated decisions that produce legal or significant effects.
- Human-in-the-Loop (HITL) Gateways: Escalation paths to human financial advisors when risk scores cross tier-3/4 thresholds.
- HSM / KMS Tokenization: Upgrading in-memory PII anonymization to Hardware Security Modules (AWS KMS / Azure Key Vault).
- Right to Explanation: Storing deterministic math breakdowns so underwriters can explain score rationale on demand.
Moving from single-developer machine to 10,000+ simultaneous branch and mobile app queries.
- High-Throughput Serving: Deploying vLLM or NVIDIA Triton on on-premise Kubernetes GPU nodes (H100/L40S).
- Distributed Cache: Upgrading local SQLite cache to Redis Cluster with multi-region replication.
- Zero-Trust IAM: Mutual TLS (mTLS), SPIFFE/SPIRE workload identities, and OAuth2/OIDC JWT tokens with row-level tenant security.
5. Architectural Alternatives & Future Evolution
Comparing our current technology choices with viable enterprise alternatives.
Dense Vector RAG vs. GraphRAG
Currently, statutory rules (e.g. 60% tax trap, pension tapering) are retrieved via dense sentence-transformers.
LiteLLM Proxy vs. vLLM / Triton
LiteLLM provides seamless switching between Ollama, LM Studio, and Gemini with uniform OpenAI-compatible endpoints.
ReAct Loop vs. Hierarchical Swarm
Currently, an autonomous ReAct loop cycles through Thought-Action-Observation with a strict 5-iteration limit.
Local JSON Logging vs. OpenTelemetry
Currently, tool calls and execution metrics are recorded to local JSON logs and SQLite history.
Local SQLite vs. PostgreSQL + pgvector
Zero-dependency embedded SQLite with JSON1 extensions powers the database, cache, and vector embeddings locally.
General Foundation vs. Fine-Tuned SLMs
General small language models (Qwen-3.5 4B / Llama-3.2 3B) are guided via structured system prompt engineering and grounding boundaries.
Summary: The 6 Non-Negotiable Pillars of Enterprise AI
Key takeaways for enterprise architects, regulators, and engineering teams.
Math in Python; narrative in the LLM. Never ask a language model to perform statutory arithmetic.
Prompt Guard heuristics + boundary isolation tags block malicious instructions before they execute.
Full offline inference ensures sensitive financial balances never leave organizational firewalls.
Zero-token caching guarantees identical answers for identical account states at zero operating cost.
Pre-push hooks guarantee code passes security scans, linting, and cold headless tests prior to deployment.
In-line output self-evaluation intercepts ungrounded claims, Consumer Duty violations, and baiting before delivery.