Senior Technical Architect

Technology, Data & Digital · Software & Web Development · Software Engineering · DevOps · Data Science

In short

The AI Observability Principal Architect is responsible for defining and delivering end-to-end observability and AIOps capabilities across enterprise applications and platforms, enabling proactive detection, intelligent event correlation, and automated incident response. This role will lead observability strategy, standardization, and implementation across teams, ensuring clear visibility into system health, business workflows, and performance outcomes.

Responsibilities

  • Define and implement observability strategy, standards, and roadmap
  • Establish telemetry framework (logs, metrics, traces) and instrumentation standards
  • Define Critical User Journeys (CUJs) and map Service Level Indicators (SLIs) / SLOs
  • Enable actionable alerting aligned to user/business impact
  • Implement alert-to-incident automation with correct routing and ownership
  • Drive AIOps capabilities including event correlation, alert noise reduction, root-cause-based incident generation, and predictive detection/anomaly identification
  • Build and optimize observability dashboards for operations and leadership visibility
  • Enable automation and self-healing playbooks for recurring incidents
  • Lead major incident support and post-incident improvements
  • Ensure continuous improvement through observability lifecycle management

Requirements

  • 18+ years experience in IT Operations / SRE / Observability / Platform Engineering
  • Strong expertise in Observability (logs, metrics, traces, distributed tracing), SRE practices (SLIs, SLOs, error budgets), and Incident management and automation
  • Experience in AIOps / Event Management (Event correlation, alert deduplication, noise reduction)
  • Hands-on experience with observability platforms and ITSM integrations
  • Strong understanding of distributed systems and cloud environments
  • Strong understanding of Telemetry pipelines and data integration
  • Experience in designing automation workflows, runbooks, and self-healing mechanisms
  • Strong stakeholder management and ability to work across application, infra, and operations teams

Desired Qualifications

  • Experience with Azure observability stack / OpenTelemetry
  • Experience with ServiceNow ITOM / AIOps
  • Exposure to AI-driven observability (anomaly detection, predictive analytics)
  • Experience in building executive dashboards and reporting frameworks
#AI Observability#AIOps#SRE#IT Operations#Platform Engineering#Incident Management
HCLTech Logo

Company

HCLTech

Job Posted

2 weeks ago

Employment Type

Full Time

WorkMode

On Site

Experience Level

Executive

Locations

Bangalore, India

Applicants

Be an early applicant