Site Reliability, Staff / HPC Infrastructure Engineer
Technology, Data & Digital · IT Infrastructure & Security · DevOps · Systems Engineering · Cloud Engineering
In short
Synopsys is looking for a Staff / HPC Infrastructure Engineer to join their Site Reliability team in Bengaluru. This role focuses on designing, building, and optimizing large-scale HPC compute farm platforms, leading capacity planning, and driving automation. The ideal candidate has extensive Linux/UNIX and HPC administration experience, with strong skills in workload schedulers and scripting.
Responsibilities
- Design, build, and optimize large-scale HPC compute farm platforms that support thousands of engineering workloads across global sites
- Administer and tune IBM Spectrum LSF, Slurm, or equivalent workload schedulers to maximize resource utilization and minimize job queue times
- Lead capacity planning and workload optimization efforts, translating engineering demand into infrastructure requirements and deployment timelines
- Drive automation using Python and Shell scripting to eliminate manual toil, improve reliability, and accelerate incident response
- Lead complex troubleshooting and root cause analysis for platform issues, working across storage, networking, LDAP, NFS, and scheduler layers
- Collaborate with R&D, Cloud, Infrastructure, and Security teams on strategic initiatives including cloud-integrated HPC and AI/ML workload enablement
- Mentor junior engineers and provide technical leadership across global teams, setting standards for operational excellence and engineering rigor
- Participate in 24x5 support operations and lead major infrastructure projects from design through deployment
Requirements
- 8+ years of Linux/UNIX systems administration experience with deep expertise in performance tuning, troubleshooting, and large-scale operations
- 5+ years of hands-on HPC or compute farm administration, including workload scheduling, resource management, and capacity planning
- Strong expertise in IBM Spectrum LSF, Slurm, or equivalent schedulers, including policy configuration, job prioritization, and performance optimization
- Advanced knowledge of LDAP, NFS, DNS, enterprise storage systems, and networking in the context of distributed compute environments
- Proven experience with Python and Shell scripting for automation, monitoring, and infrastructure orchestration
- Solid understanding of monitoring and observability tools such as Grafana, Prometheus, Elastic, or Splunk for proactive incident detection and analysis
Desired Qualifications
- Experience with EDA environments is a strong plus
- Familiarity with cloud-integrated HPC platforms on Azure or AWS
- Familiarity with Kubernetes, Docker, Ansible, or Terraform
Benefits
- Comprehensive medical and healthcare plans
- ETO and FTO Programs
- Maternity and paternity leave, parenting resources, adoption and surrogacy assistance
- Purchase Synopsys common stock at a 15% discount, with a 24 month look-back
- Retirement plans that vary by region and country
- Competitive salaries
#Site Reliability#HPC#Infrastructure Engineer#Linux/UNIX#performance tuning#troubleshooting#large-scale operations#workload scheduling#resource management#capacity planning#IBM Spectrum LSF#Slurm#automation#Python#Shell scripting#monitoring#observability#Grafana#Prometheus#Elastic#Splunk#EDA environments#cloud-integrated HPC#Azure#AWS#Kubernetes#Docker#Ansible#Terraform