Data Center Engineering Operations , DCEO
Technology, Data & Digital · IT Infrastructure & Security · Systems Engineering · Database Administration
In short
AWS Infrastructure Services is seeking a Data Center Engineering Operations Engineer to ensure the optimal operation of data center facilities. This role involves managing critical infrastructure maintenance, ensuring safety and compliance, and coordinating with various teams and vendors.
Responsibilities
- Ownership of all Data Center changes/events/incidents/problems from beginning to end, including post-mortems and root cause analysis.
- Planning and execution of maintenance/repairs for site-critical facility infrastructure or a Data Center.
- Asset and Inventory management.
- Developing and maintaining method statements, standard operating procedures, emergency response procedures, and preventive maintenance programs.
- Ensuring standardization and consistency with best-in-class operating practices.
- Developing a deep knowledge of Data Center systems' design intent, operational alternatives, and contingency plans.
- Managing the engineering aspects of Data Centers related to financial and cost control, code and regulatory compliance, personnel management, and staff training.
- Managing Health & Safety, local statutory requirements, environmental and energy management.
- Developing and delivering regular engineering reports and ensuring adherence to contracted deliverables, SLAs, and KPIs.
- Communicating operating philosophies, technical information, objectives, and expectations to Amazon personnel and vendor critical facilities management teams.
- Providing hands-on facility support, including installation, decommissioning, and replacement of equipment.
- Overseeing technical compliance auditing and the closure of corrective action plans.
- Performing annual operational reviews focusing on compliance with Amazon standards and regulatory requirements.
- Managing the development and delivery of Energy/Environmental Management Programs.
- Reviewing incident reports, documenting trend summaries, and providing recommendations to management.
- Managing information flow during incidents and providing regular updates to management.
- Coordinating with vendors to resolve incidents during emergency situations, potentially requiring on-site dispatch.
- Ensuring all electrical, mechanical, and fire/life safety equipment operates within contract parameters.
- Managing risk and mitigation, corrective and preventative maintenance of critical infrastructure.
- Metric reporting.
- On-site management of contractors, sub-contractors, and vendors.
- Establishing performance benchmarks and preparing reports on data center facility infrastructure operations and maintenance.
- Coordinating projects, managing capacity, and optimizing plant safety, performance, reliability, sustainability, and efficiency.
- Supporting the installation of racks and the provision of power/cooling.
- Assisting in troubleshooting facility and rack-level events within internal Service Level Agreements (SLA).
- Performing rack installs, rack decommissioning, and facility management.
- Providing operational readings and key performance indicators to ensure uptime.
- Performing and overseeing maintenance and operations on all electrical, mechanical, and fire/life safety equipment.
- Working shifts that can be up to 12-hours and may rotate on a predefined schedule; some locations have on-call rotations.
Requirements
- 3+ years of relevant work in a data center or other critical environment experience
- 3+ years of data center engineering or operations experience
- Bachelor's degree in Electrical Engineering, Mechanical Engineering, or a related field
- 4+ years of data center engineering or operations experience
- Experience in electrical or mechanical engineering
Desired Qualifications
- Technical Writing Skills and Automation
- Audits
- The ability to physically be dispatched on to site to investigate and resolve the issue during emergency situations.
Benefits
- Inclusive culture that welcomes bold ideas and empowers you to own them to completion.
- Diverse experiences valued; encouragement to apply even if not meeting all preferred qualifications.
- Pioneered cloud computing and continuously innovating.
- Flexible working culture valuing work-life harmony.
- Inclusive team culture with employee-led affinity groups.
- Endless knowledge-sharing, mentorship, and career-advancing resources.
#Operations#IT#Support Engineering