FOUNDER · TRACELANE AI agent reliability Bengaluru, India

I keep critical systems alive.
Now I build the tools that keep AI agents alive.

Fifteen years running mission-critical financial systems — currently Second Vice President at Northern Trust. Founder of Tracelane, an open-core LLM gateway and AI-agent observability platform. I bring an operator’s paranoia to AI infrastructure: assume failure, trace everything, catch it before it executes.

Portrait of Sanjeev Kumar Singh
Sanjeev Kumar SinghFounder, Tracelane · Production reliability engineer
15yrs
Mission-critical production ops
99%
SLA uptime, institutional platforms
4.6ms
Tracelane p99 gateway overhead
375µs
p99 inline guardrail dispatch

// currently

FOUNDER Mar 2026 — Present

Founder & Principal Engineer

Tracelane ↗

Open-core LLM gateway + AI-agent reliability platform. Observability, guardrails, audit trails — catching agent failures before execution.

OPERATOR Oct 2021 — Present

Second VP, Application Support Consultant

Northern Trust

Senior escalation point for institutional financial platforms across EMEA/APAC. 99% SLA. SOC 2 Type II audits with KPMG. 12+ engineers mentored.

// founder work

Open-core · LLM gateway + AI-agent observability

tracelane.dev ↗

Reliability infrastructure for AI agents.

4.6 ms p99 gateway overhead
375 µs p99 guardrail dispatch
11.5k rows/sec ingest, single node

One platform for routing, tracing and governing production agents: an LLM gateway with provider failover, OpenTelemetry GenAI traces over ClickHouse, inline guardrails, MCP failure-signature detection, and cryptographic audit trails. Built solo, in Rust and TypeScript, while running a day job in institutional finance.

  • Multi-provider routing · failover · queue replay
  • OpenTelemetry GenAI traces — LLM, tool, browser, MCP
  • Inline guardrails + eval-gated CI
  • SSO/RBAC · BYOK · tenant quotas · audit trails

// principles

How I operate

Fifteen years of production incidents leave you with convictions. These are mine.

01

Assume failure

Every system fails; the question is whether you find out from a trace or from a customer. Design for the failure first — observability, guardrails and rollback are the product, not the afterthought.

02

Be customer zero

I run Tracelane on Tracelane. Production use exposes the failures that automated tests never catch — that feedback loop is worth more than any roadmap.

03

Boring reliability, novel systems

AI agents are new; the discipline of keeping them alive is not. SLOs, runbooks, DR drills and post-incident rigor transfer directly from institutional finance to agent infrastructure.

// track record

Fifteen years, zero abandoned systems.

2026 — now Bengaluru

Founder & Principal Engineer

Tracelane

Architected and launched an open-core AI-agent reliability platform: LLM gateway, OTel GenAI tracing, ClickHouse analytics, guardrails, MCP failure detection. Eval-gated CI, blue-green deploys, provider failover. Used it as customer zero to find failures tests never caught.

2021 — now London → Bengaluru

Second VP, Application Support Consultant

Northern Trust

L2 lead for mission-critical fund-servicing platforms, EMEA/APAC. 99% SLA. Led Bedrock monitoring initiative end-to-end. Annual EMEA/US disaster-recovery exercises to RTO/RPO. Hired and mentored 7 DevOps engineers through Kubernetes transformation. Cut recurring incidents ~30% and manual ops ~30%. KPMG / SOC 2 Type II audit partner.

2018 — 2021 London

Technology Lead, L2 Application Support

Northern Trust · via Infosys

Managed 12+ specialists across 20+ global fund-servicing applications. Zero-downtime production upgrades. ~30% recurring-incident reduction via trend analysis and automation.

2015 — 2017 Chennai

Technology Analyst, Application Support

NatWest Group · via Infosys

24/7 production support for trade affirmation and confirmation platforms. Live SQL fixes for customer-impacting incidents. ServiceNow / JIRA incident and change management.

2011 — 2014 Chennai

Software Engineer, ETL Support

AIG · via Mphasis

Oracle data-warehouse ETL, PL/SQL development, SQL performance tuning, business-critical BFSI reporting.

B.E. Electronics & Communication — Anna University, Chennai · 2007 — 2011 · First Class with Distinction  ·  ITIL V3 Foundation · Agile Kanban

// stack

AI infrastructure

LLM gateway architectureOpenTelemetry GenAIAgent observabilityGuardrailsEval pipelinesMCPClaude APIProvider routing

Engineering

Rust (Axum / tokio)TypeScriptPythonSQL / PL-SQLClickHousePostgreSQLKafkaNATS JetStream

Operations

KubernetesDockerAWSAzureCloudflare WorkersIncident response / RCADR · RTO / RPOSOC 2 · SSO/RBAC · BYOK

// also running