SME - JBoss Application Server, Apache Tomcat
Technology, Data & Digital · IT Infrastructure & Security · DevOps · Software Engineering · Site Reliability Engineering
In short
The SRE L3 Engineer is responsible for ensuring the stability, availability, performance, and reliability of Java-based enterprise applications on WebSphere, JBoss/WildFly, Linux, and Oracle databases. This role involves deep technical troubleshooting, root cause elimination, performance optimization, automation, and technical leadership as the highest level of operational escalation.
Responsibilities
- Own and manage L3 support and technical operations for Java/J2EE applications hosted on WebSphere and JBoss/WildFly platforms.
- Provide expert-level troubleshooting for application, middleware, database, and infrastructure-related issues impacting production services.
- Ensure application availability, resilience, performance, and operational stability against agreed service objectives.
- Act as the L3 technical escalation point for incidents escalated from L1 and L2 support teams.
- Lead technical resolution of critical P1/P2 incidents and ensure rapid service restoration.
- Perform deep technical diagnosis using application logs, thread dumps, heap dumps, JVM metrics, middleware logs, Oracle performance data, and infrastructure telemetry.
- Lead detailed Root Cause Analysis (RCA) for recurring and business-critical incidents.
- Identify and eliminate reliability risks through permanent fixes, automation, configuration improvements, and architecture recommendations.
- Provide advanced administration and troubleshooting for IBM WebSphere Application Server (WAS) and JBoss/WildFly environments.
- Perform and govern application deployments, JVM tuning, thread pool management, datasource configuration, cluster administration, load balancing optimization, and middleware performance tuning.
- Analyze thread dumps, heap dumps, garbage collection logs, and server diagnostics to identify performance bottlenecks.
- Perform advanced troubleshooting of Oracle database-related issues impacting application performance and availability.
- Analyze SQL performance, locking issues, waits, execution plans, and database resource utilization.
- Design, implement, and optimize monitoring, alerting, dashboards, and observability solutions.
- Analyze logs, metrics, traces, JVM behavior, middleware health, and database performance indicators.
- Drive adherence to SLI, SLO, SLA, and Error Budget objectives and continuously improve system reliability.
- Develop and maintain automation solutions using Shell, Python, Ansible, Groovy, or equivalent technologies.
- Automate deployments, health checks, diagnostics, recovery procedures, and operational workflows.
- Mentor and coach L1/L2 engineers on technical troubleshooting, reliability practices, and operational excellence.
- Review operational procedures, SOPs, runbooks, and knowledge articles.
Requirements
- 6–10 years of experience in Java Application Support, Middleware Operations, Application Operations, Production Support, or SRE roles.
- Strong hands-on experience with Java/J2EE enterprise applications.
- Advanced expertise in IBM WebSphere Application Server (WAS), JBoss/WildFly, Oracle Database, and Linux/Unix environments.
- Strong knowledge of JVM internals and tuning, thread dumps and heap dump analysis, application deployment and configuration, middleware clustering and high availability, and performance troubleshooting.
- Strong experience in Production Support, Application Operations, or Site Reliability Engineering (SRE).
- Extensive experience with Incident Management, Problem Management, Change Management, Major Incident Handling, and RCA methodologies.
- Expertise in monitoring and observability tools such as Splunk, AppDynamics, Dynatrace, Grafana, or similar.
- Strong understanding of SLI/SLO/SLA concepts, Error Budgets, Capacity Management, and Reliability Engineering practices.
- Strong analytical and problem-solving skills.
- Ability to lead technical troubleshooting during major incidents.
- Excellent written and verbal communication skills.
- Strong stakeholder management and cross-functional collaboration capabilities.
- Ability to mentor and guide junior engineers.
Desired Qualifications
- Experience supporting large-scale enterprise and distributed environments.
- Experience as a senior support engineer, technical lead, middleware specialist, or SRE is preferred.
- OpenShift, Kubernetes, Docker, or container platform experience.
- CI/CD tools such as Jenkins, GitHub Actions, GitLab, or Azure DevOps.
- Cloud platforms (Azure, AWS, GCP).
- AppDynamics, Dynatrace, Splunk Observability, Prometheus, Grafana.
- Infrastructure as Code and configuration management automation.
- Experience in enterprise transformation, reliability engineering, and platform modernization initiatives.
Benefits
- You'll supercharge your potential.
- You'll find your career.
- You'll find your spark.
- A place that knows that helping its customers stay on top starts by putting its people first.
#SRE#L3 Support#Production Support#Application Operations#Middleware Operations#Site Reliability Engineering#Java/J2EE#WebSphere#JBoss/WildFly#Oracle Database#Linux/Unix#JVM Tuning#Thread Dumps#Heap Dump Analysis#Performance Troubleshooting#Incident Management#Problem Management#Root Cause Analysis (RCA)#Automation#Python#Ansible#Monitoring#Observability#Splunk#AppDynamics#Dynatrace#Grafana#Kubernetes#Docker#CI/CD#Jenkins#Cloud Platforms#Azure#AWS#GCP