Storage and Resiliency Operations
Job Description
Job Title: Storage and Resiliency
\nOverview
\nWe are seeking a highly skilled and motivated engineer to serve in our Infrastructure as a Service (IaaS) organization supporting Storage and Resiliency. This role is ideal for a seasoned professional with deep expertise in enterprise storage, backup, replication, disaster recovery, cyber recovery, and recoverability operations across a large scale environment. The ideal candidate will guide junior engineers and drive operational excellence across storage, backup, and resiliency services.
\nKey Responsibilities
\nStorage & Backup Administration
\n- \n
- Install, configure, and maintain enterprise storage, backup, replication, and recovery platforms across private and public cloud environments \n
- Manage lifecycle activities including provisioning, capacity management, upgrades, technology currency, and decommissioning \n
- Monitor storage and backup platform health, performance, availability, recoverability, and capacity through observability capabilities \n
- Perform troubleshooting and root cause analysis for storage, backup, replication, and recovery incidents \n
- Implement platform security hardening, retention standards, access controls, and compliance requirements \n
- Manage storage provisioning, snapshot policies, backup schedules, replication policies, recovery tests, and service reporting \n
Resiliency & Cyber Recovery
\n- \n
- Lead disaster recovery, recoverability validation, backup restoration, replication testing, and operational readiness activities \n
- Partner with Cybersecurity and application teams to support cyber recovery capabilities and recoverability objectives \n
- Drive improvement plans for backup success, restore performance, data protection coverage, and resiliency gaps \n
24x7 Operations & Incident Management
\n- \n
- Support 24x7 storage and resiliency operations including capacity, backup operations, restore support, and vulnerability remediation \n
- Oversee incident response, root cause analysis, problem management, and service restoration for storage, backup, and recovery services \n
Mentorship & Collaboration
\n- \n
- Mentor junior engineers and foster a culture of continuous learning and technical excellence \n
- Collaborate with cross-functional teams including compute, network, security, application, disaster recovery, and service management teams \n
Operational Excellence
\n- \n
- Ensure high availability, scalability, security, recoverability, and compliance of storage and resiliency environments \n
- Develop metrics for capacity, performance, backup success, restore performance, replication health, and recoverability compliance \n
Qualifications
\n- \n
- Strong enterprise storage and backup operations experience in large scale environments \n
- Experience with NetApp, SAN, NAS, object storage, backup platforms, replication, and disaster recovery technologies \n
- Knowledge of backup policies, retention standards, recovery testing, cyber recovery, and recoverability objectives \n
- Scripting and automation experience with PowerShell, Python, Ansible, Terraform, or vendor automation tools \n
- Storage and backup capacity management, performance troubleshooting, vulnerability remediation, and technology currency experience \n
- Networking fundamentals including TCP/IP, DNS, NFS, SMB, Fibre Channel, iSCSI, and firewall concepts \n
- Logging and monitoring tools such as Splunk, vendor management tools, Prometheus, or equivalent \n
Experience with virtualization and/or cloud platforms including VMw