Senior Engineer - Monitoring Tools, Event Monitoring

Technology, Data & Digital

In short

The Senior Engineer - Monitoring Tools, Event Monitoring role in Noida, India, is responsible for proactively monitoring IT infrastructure, applications, and endpoints to identify anomalies, analyze trends, and automate remediation. The position focuses on enhancing user experience, reducing disruptions, and improving operational efficiency through advanced monitoring and analytics platforms, requiring expertise in tools like Grafana, Datadog, Splunk, and scripting languages such as PowerShell and Python.

Responsibilities

  • Monitor infrastructure, applications, endpoints, network services, and business-critical systems.
  • Configure and maintain monitoring dashboards, alerts, thresholds, and performance baselines.
  • Identify service degradations and potential incidents before they impact business operations.
  • Perform root cause identification through proactive monitoring and correlation of events.
  • Analyze operational data, alerts, incidents, tickets, and endpoint health metrics.
  • Develop trend analysis, capacity reports, service health dashboards, and predictive insights.
  • Generate executive-level reports highlighting service performance, recurring issues, and improvement opportunities.
  • Leverage AI/ML driven analytics to identify patterns and anomalies.
  • Design and implement automated remediation workflows for common incidents and endpoint issues.
  • Develop scripts and automation runbooks to reduce manual support efforts.
  • Drive self-healing capabilities across infrastructure and digital workplace environments.
  • Collaborate with engineering teams to continuously improve automation effectiveness.
  • Act as the operational bridge between Monitoring, Service Desk, Infrastructure, and Application teams.
  • Support major incident investigations by providing monitoring insights and analytics.
  • Identify recurring service issues and contribute to Problem Management initiatives.
  • Recommend preventive actions based on operational data.
  • Improve monitoring coverage and alert accuracy.
  • Reduce false positives and optimize event management processes.
  • Identify opportunities for operational excellence and service automation.
  • Contribute to Experience Management (XMO) and User Experience improvement initiatives.

Requirements

  • Monitoring & Observability Tools: Grafana, Datadog, Dynatrace, Splunk, Azure Monitor, AppDynamics, Elastic Stack (ELK), Prometheus, SolarWinds, LogicMonitor
  • Endpoint Management & Digital Workplace: Microsoft Intune, BigFix, SCCM/MECM, JAMF, Nexthink, Lakeside SysTrack
  • Analytics & Reporting: Power BI, Tableau, Excel
  • Advanced Analytics: SQL, Azure Data Explorer, Kusto Query Language (KQL)
  • Automation & Scripting: PowerShell, Python, Bash Scripting, ServiceNow Workflows, Power Automate, Azure Automation
  • ITSM Platforms: ServiceNow, BMC Helix, Jira Service Management
  • Experience with AI-driven Operations (AIOps)
  • Understanding of ITIL processes
  • Knowledge of Event Management and Correlation Engines
  • Exposure to Cloud Monitoring (Azure, AWS, GCP)
  • Experience with Digital Employee Experience (DEX) platforms
  • Strong analytical and problem-solving capabilities
  • Bachelor's Degree in Computer Science, Information Technology, Engineering, or related field
  • ITIL Foundation Certification preferred
  • Relevant vendor certifications in Monitoring, Cloud, Automation, or Analytics are an advantage
  • Proactive Monitoring
  • Data Analytics
  • Automation & Self-Healing
  • Problem Solving
  • Incident Analysis
  • Stakeholder Management
  • Continuous Improvement
  • Business Reporting
  • Service Reliability Engineering
  • Operational Excellence

Desired Qualifications

  • Business Outcomes Expected: Reduced Mean Time to Detect (MTTD), Reduced Mean Time to Resolve (MTTR), Increased Service Availability, Improved Digital Employee Experience (DEX), Reduction in Repeat Incidents, Increased Automation and Self-Healing Coverage, Enhanced Operational Visibility and Predictive Insights, Improved Customer Satisfaction and Service Reliability.

Benefits

  • At HCLTech, you'll supercharge your potential. You'll find your career. And you'll find your spark. All at a place that knows that helping its customers stay on top starts by putting its people first.
#Monitoring Tools#Event Monitoring#IT Infrastructure#Application Monitoring#Endpoint Monitoring#Service Performance#User Experience#Operational Efficiency#Data-driven Decision-making#Alerting#Root Cause Analysis#Trend Analysis#Capacity Planning#Predictive Insights#Automation#Remediation Workflows#Self-healing#ITSM#AIOps#ITIL#Cloud Monitoring#Digital Employee Experience#Service Reliability Engineering
HCLTech Logo

Company

HCLTech

Job Posted

2 weeks ago

Employment Type

Full Time

WorkMode

On Site

Experience Level

Senior

Locations

Noida, India

Qualification

Bachelor

Applicants

Be an early applicant