Senior Administrator - Monitoring Tools, Event Monitoring
Technology, Data & Digital · IT Infrastructure & Security · IT Support · Technical Support · DevOps
In short
Seeking a Senior Administrator for Monitoring Tools and Event Monitoring in Bengaluru. Responsibilities include technical troubleshooting, root cause analysis, batch job optimization, monitoring tool tuning (AppDynamics, Splunk, Big Panda), and incident/problem/change management. This role requires strong Unix/Linux skills, database knowledge, and experience with production support.
Responsibilities
- Handle escalations from L1 and perform in-depth analysis to resolve complex application issues.
- Conduct detailed Root Cause Analysis (RCA) and drive permanent fixes.
- Work with development, infrastructure, and DevOps teams for defect fixes and environment issues.
- Review and optimize batch jobs in Autosys; provide solutioning for recurring failures.
- Tune monitoring thresholds and alerts in tools like AppDynamics, Big Panda, and Splunk.
- Perform log analysis, query analysis, and system health checks.
- Take ownership of high severity incidents (P1/P2) and coordinate with multiple teams for quick resolution.
- Participate in Problem Management to reduce repeat incidents.
- Validate and implement changes during release cycles and maintenance windows.
- Support lower and production environments (DEV/SIT/UAT/PROD).
- Deploy application patches, configuration updates, and troubleshoot integration/API issues.
- Validate releases, perform sanity checks, and support post deployment monitoring.
- Work with cross-functional teams including Development, QA, Infrastructure, and Business SMEs.
- Document troubleshooting steps, SOPs, knowledge articles, and best practices.
- Provide mentorship to L1 analysts and improve support processes.
Requirements
- 7+ years of experience in production support.
- Strong hands-on with Unix/Linux.
- Advanced Autosys experience.
- Ability to write complex queries for Oracle, MSSQL, MongoDB.
- Knowledge of Java/.NET application logs or basic scripting (Python/Shell).
- Ability to manage the shift.
- Strong analytical and debugging skills across distributed systems.
- Ability to manage high severity incidents under pressure.
- Excellent communication and cross-team coordination capabilities.
- Experience working in Agile/ITIL-driven environments.
Desired Qualifications
- Shell scripting is a plus.
- Control-M is optional.
- Kibana is optional.
- AWS basics, CI/CD pipelines, Docker, Kubernetes are good to have.
- Ability to improve processes, automate tasks, and reduce manual work.
Benefits
- Supercharge your potential.
- Find your career.
- Find your spark.
- A place that knows that helping its customers stay on top starts by putting its people first.
#monitoring#event monitoring#application support#root cause analysis#system stability#incident management#problem management#change management#unix#linux#sql#databases#devops#agile#itil