Principal AI Architect

Teknik, data och digitalt · Data, AI och analys · Datavetenskap · Maskininlärning · Artificiell intelligens

I korthet

We are seeking a Principal AI Architect to define, build, and scale evaluation systems for AI products, focusing on LLMs, RAG, and agentic systems. This hybrid role involves designing and integrating evaluation frameworks across the product lifecycle to ensure quality and measure performance, collaborating with engineering and science teams to drive rigorous AI practices.

Ansvarsområden

  • Define the technical vision and architecture for AI, LLM, RAG, and agent evaluation systems.
  • Design evaluation frameworks to assess product quality throughout the lifecycle: pre-implementation, development, launch, and post-ship.
  • Understand the behavior of nondeterministic agentic systems across various tasks, users, and contexts.
  • Identify behavioral gaps, failure modes, and product-quality risks before they impact customers.
  • Determine where issues should be addressed: model behavior, prompts, tools, retrieval, UX, safety, or code.
  • Partner with engineering, applied science, and data science teams to integrate evaluation methods into product codebases and workflows.
  • Integrate evaluations into build pipelines, release gates, and experimentation systems.
  • Make evaluation results accessible and interpretable through dashboards and reports.
  • Build systems connecting telemetry, offline evaluation, human judgment, and automated evals.
  • Understand and evaluate enterprise search, RAG, grounding, and relevance systems for products like Copilot.
  • Translate ambiguous product goals into measurable evaluation strategies and technical plans.
  • Drive architecture decisions across components, services, and data pipelines.
  • Work with product leaders to prioritize evaluation investments and align with product milestones.
  • Mentor senior engineers and applied scientists on building evaluation infrastructure.
  • Stay current with LLM evaluation methods, agentic systems, and responsible AI practices.
  • Help evaluate products before code completion, before launch, and after shipping.
  • Help teams understand AI system performance, failure modes, improvements, and responsible scaling.
  • Evaluate products before code is fully ready, before launch, and after they ship.
  • Work across product, engineering, applied science, and data science teams to bring rigorous AI evaluation practices into the product lifecycle.
  • Help IC3 and partner teams understand whether AI systems are working as intended, where they fail, how they improve, and what it takes to ship them responsibly at scale.

Krav

  • Bachelor's Degree in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or related field AND 6+ years related experience OR Master's Degree AND 4+ years related experience OR Doctorate AND 3+ years related experience OR equivalent experience.
  • Demonstrated experience evaluating LLMs, including designing eval datasets, defining quality metrics, analyzing model behavior, identifying failure modes, and using results to guide product or system improvements.
  • Understanding of modern AI systems, including LLMs, agentic systems, RAG, ML pipelines, evaluation methodology, experimentation, and product telemetry.
  • Ability to architect complex systems where multiple components, services, models, data flows, tools, retrieval systems, and product surfaces interact.
  • Experience evaluating or building nondeterministic AI systems where quality must be understood statistically, behaviorally, and through product impact.
  • Experience bringing ML, DS, LLM, or applied science concepts into production systems and product codebases.
  • Ability to integrate evaluation into engineering systems such as CI/CD, build pipelines, release gates, dashboards, and monitoring workflows.
  • Understanding of search, retrieval, grounding, relevance, ranking, and enterprise RAG concepts.
  • Ability to define technical strategy, product-quality metrics, milestones, and execution plans across teams.
  • Coding and technical design skills, with the ability to work directly in product codebases when needed.
  • Effective communication skills with the ability to influence engineers, scientists, product managers, and executives.
  • Track record of leading ambiguous, cross-functional technical initiatives from concept through delivery.
  • Experience evaluating LLM-powered products, agents, enterprise search, recommendation systems, or generative AI applications.
  • Experience with offline evals, online experimentation, human evaluation, red teaming, synthetic data, model monitoring, RAG evaluation, and agent behavior analysis.
  • Experience with evaluation dashboards, scorecards, quality reporting, product-health monitoring, or data visualization systems.
  • Familiarity with responsible AI, safety, reliability, privacy, security, permissions, compliance, and enterprise-readiness considerations for AI systems.
  • Experience building evaluation platforms, experimentation systems, model observability, agent evaluation infrastructure, or product-quality infrastructure.
  • Experience operating at principal, architect, or senior technical leadership level.

Önskade kvalifikationer

  • Master's Degree in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or related field AND 9+ years related experience OR Doctorate AND 6+ years related experience OR equivalent experience.
  • 5+ years experience creating publications (e.g., patents, libraries, peer-reviewed academic papers).
  • 2+ years experience presenting at conferences or other events in the outside research/industry community as an invited speaker.
  • 5+ years experience conducting research as part of a research program (in academic or industry settings).
  • 3+ years experience developing and deploying live production systems, as part of a product team.
  • 3+ years experience developing and deploying products or systems at multiple points in the product cycle from ideation to shipping.

Förmåner

  • Certain roles may be eligible for benefits and other compensation.
#LLM#Architect#EngineerScientist
Microsoft Logo

Företag

Microsoft

Publicerade jobb

för 1 månad sedan

Anställningstyp

Heltid

Arbetsform

Hybrid

Erfarenhetsnivå

Mellannivå

Platser

United States

Kvalifikation

Kandidatexamen, Masterexamen, Doktorand

Sökande

Ansök tidigt