Senior Administrator - Monitoring Tools, Event Monitoring
Teknik, data och digitalt · IT-infrastruktur och säkerhet · DevOps · Dataanalys · Mjukvaruutveckling
I korthet
The Senior Administrator - Monitoring Tools, Event Monitoring is responsible for proactive monitoring of IT infrastructure and applications to identify anomalies, analyze trends, and automate remediation actions. This role focuses on improving user experience and operational efficiency using advanced monitoring platforms like Grafana, Datadog, and Splunk.
Ansvarsområden
- Monitor infrastructure, applications, endpoints, network services, and business-critical systems.
- Configure and maintain monitoring dashboards, alerts, thresholds, and performance baselines.
- Identify service degradations and potential incidents before they impact business operations.
- Perform root cause identification through proactive monitoring and correlation of events.
- Analyze operational data, alerts, incidents, tickets, and endpoint health metrics.
- Develop trend analysis, capacity reports, service health dashboards, and predictive insights.
- Design and implement automated remediation workflows for common incidents and endpoint issues.
- Develop scripts and automation runbooks to reduce manual support efforts.
- Support major incident investigations by providing monitoring insights and analytics.
- Identify recurring service issues and contribute to Problem Management initiatives.
- Improve monitoring coverage and alert accuracy.
- Reduce false positives and optimize event management processes.
- Contribute to Experience Management (XMO) and User Experience improvement initiatives.
Krav
- Monitoring & Observability Tools: Grafana, Datadog, Dynatrace, Splunk, Azure Monitor, AppDynamics, Elastic Stack (ELK), Prometheus, SolarWinds, LogicMonitor
- Endpoint Management & Digital Workplace: Microsoft Intune, BigFix, SCCM/MECM, JAMF, Nexthink, Lakeside SysTrack
- Analytics & Reporting: Power BI, Tableau, Excel
- Advanced Analytics: SQL, Azure Data Explorer, Kusto Query Language (KQL)
- Automation & Scripting: PowerShell, Python, Bash Scripting, ServiceNow Workflows, Power Automate, Azure Automation
- ITSM Platforms: ServiceNow, BMC Helix, Jira Service Management
- Experience with AI-driven Operations (AIOps).
- Understanding of ITIL processes.
- Knowledge of Event Management and Correlation Engines.
- Exposure to Cloud Monitoring (Azure, AWS, GCP).
- Experience with Digital Employee Experience (DEX) platforms.
- Strong analytical and problem-solving capabilities.
- Bachelor's Degree in Computer Science, Information Technology, Engineering, or related field.
- ITIL Foundation Certification preferred.
- Relevant vendor certifications in Monitoring, Cloud, Automation, or Analytics are an advantage.
Önskade kvalifikationer
- Proactive Monitoring
- Data Analytics
- Automation & Self-Healing
- Problem Solving
- Incident Analysis
- Stakeholder Management
- Continuous Improvement
- Business Reporting
- Service Reliability Engineering
- Operational Excellence
- Business Outcomes Expected: Reduced Mean Time to Detect (MTTD), Reduced Mean Time to Resolve (MTTR), Increased Service Availability, Improved Digital Employee Experience (DEX), Reduction in Repeat Incidents, Increased Automation and Self-Healing Coverage, Enhanced Operational Visibility and Predictive Insights, Improved Customer Satisfaction and Service Reliability.
Förmåner
- Consolidated revenues as of 12 months ending June 2026 totaled $14.8 billion.
#Monitoring#Event Monitoring#IT Infrastructure#IT Operations#Automation#Analytics#ServiceNow#Grafana#Datadog#Splunk#Azure#AWS#GCP