Senior Technical Lead
Technology, Data & Digital · Data, AI & Analytics · Machine Learning · Data Engineering · Software Engineering
In short
We are seeking a Senior Technical Lead specializing in AI Observability to oversee the monitoring, measurement, and trust of AI systems in production. This role bridges MLOps, platform engineering, and applied AI, focusing on building infrastructure for real-time visibility into model behavior, quality, cost, and risk.
Responsibilities
- Own the end-to-end observability strategy for AI systems in production, including tracing, logging, metrics, evaluation pipelines, and alerting.
- Design and build systems to monitor model/agent quality in production: accuracy drift, hallucination rate, latency, cost per request, token usage, and task success rate.
- Establish golden signals and SLOs for AI systems, distinct from traditional infra SLOs (e.g., output quality, safety, groundedness, factuality).
- Build or integrate tracing across multi-step/agentic workflows so failures can be root-caused across prompts, tool calls, retrieval steps, and model versions.
- Stand up automated evaluation frameworks (offline and online/production evals) to continuously score live traffic and catch regressions after model, prompt, or data updates.
- Partner with ML/platform engineering to instrument new models and features with observability hooks before they reach production.
- Define and drive incident response processes specific to AI failures (silent quality degradation, drift, prompt injection, unsafe outputs) — not just uptime.
- Build dashboards and reporting for engineering, product, and executive stakeholders on production AI health, cost, and risk posture.
- Lead, mentor, and grow a team of engineers focused on observability tooling, or act as the technical lead embedded across ML/platform teams.
- Evaluate, select, and manage the observability toolchain (build vs. buy) across tracing, evals, monitoring, and cost-tracking platforms.
- Partner with security, compliance, and legal on auditability, data retention, and responsible-AI monitoring requirements.
Requirements
- + years in software/ML engineering, with 3+ years focused on observability, monitoring, reliability, or MLOps.
- Hands-on experience running AI/ML systems in production, including at least one LLM-based or generative AI system at scale.
- Strong understanding of distributed tracing, structured logging, and metrics pipelines (e.g., OpenTelemetry-style concepts), applied to AI/agentic workflows.
- Experience designing evaluation frameworks for generative AI (offline benchmarks, online/production evals, human-in-the-loop review).
- Solid grasp of the unique failure modes of production AI: drift, hallucination, prompt injection, latency/cost blowups, silent quality regressions.
- Track record of building or leading a team, or serving as a technical lead across cross-functional engineering groups.
- Proficiency in at least one major programming language (Python, Go, or similar) and comfort working across the ML/platform stack.
- Excellent cross-functional communication — able to translate observability data into decisions for engineers, product managers, and executives.
Desired Qualifications
- Experience with vector databases, RAG pipelines, or agentic frameworks in production.
- Background in SRE/DevOps prior to moving into ML/AI observability.
- Familiarity with responsible AI, model risk management, or AI governance frameworks.
- Experience presenting observability/risk posture to executive or board-level audiences.
#AI Observability#MLOps#Platform Engineering#AI Systems#Monitoring#Tracing#Evaluation#Alerting#LLM#Generative AI#SRE#DevOps#Responsible AI