Skip to main content
Posted 09 August, 2026

Sr. Staff Site Reliability Engineer (Linux/Network troubleshooting/Scripting)

Zscaler
Hyderabad, IND Full Time
Reference: 102_705768_5177391007

Role

We are seeking an experienced Senior Staff, Site Reliability Engineer to join our dynamic SRE Cloud Infrastructure & Operations team. In this high-impact role, you will report directly to the Director of Site Reliability Engineering and play a pivotal part in architecting, scaling, and maintaining our next-generation cloud infrastructure. You will bridge the gap between development and operations, ensuring our large-scale distributed systems are highly available, secure, and incredibly resilient.

What You'll Do (Key Responsibilities)

  • Architecting & Automating: Design, implement, and manage various advanced cloud management automations to eliminate toil and accelerate delivery
  • Container Orchestration: Oversee and optimize containerized architectures using EKS and GKE to ensure robust production performance
  • Observability Systems: Lead the creation, deployment, and optimization of highly scalable monitoring and alerting systems
  • Cloud Operations & Incident Management: Own cloud operations, deployments, on-call support, and incident management while continuously designing and tuning Linux and BSD-based systems
  • Cross-Functional Collaboration: Serve as a core member of cross-functional project teams, contributing to technology-based solutions and consulting on concept feasibility for new initiative

Who You Are (Success Profile)

  • You Think at Scale: You connect your day-to-day work to the larger company mission and think globally. You build solutions, processes, and architectures that are not just effective today, but are built to last and support a high-growth organization
  • You are Driven by Innovation: You possess a deep curiosity for how things work and are energized by solving complex technical hurdles. You believe in the power of technology to accelerate transformation and consistently hunt for more secure, scalable methods
  • You are a Problem-Solver: You actively seek out engineering challenges because you are energized by finding solutions, knowing that solving the hardest problems delivers the biggest business and customer impact
  • You are Data-Driven: You lean on data and analytics to uncover engineering truths, measure what matters, and guide informed architectural decisions replacing "I think" with "I know."
  • You are a Learner: You have a true growth mindset and never stop developing your technical or leadership skills. You actively seek and implement feedback to become an exceptional collaborator and teammate

What We're Looking For (Minimum Qualifications)

  • AI & Automation Curiosity: Demonstrated curiosity and active exploration of AI tools, with a proven history of integrating new technologies to enhance daily workflows and augment problem-solving
  • Distributed Systems Experience: Minimum of 7 years of relevant experience in designing, analyzing, and troubleshooting large-scale distributed systems
  • Technical Ecosystem Mastery: Deep hands-on experience with C, Java, GoLang, Python, Terraform, Ansible, Python automation, networking, Kubernetes, and AWS cloud
  • Web Protocols & Security: Comprehensive understanding of web security and core protocols including HTTP, SSL/TLS, DNS, SQL, and networking fundamentals
  • Observability Architecture: Proven experience in observability, building complex dashboards, managing Grafana, and maintaining a sharp understanding of SLIs, SLOs, and error budgets
  • Modern DevOps Expertise: Strong DevOps skills across CI/CD pipelines, Source Control Management (SCM), builds/releases, and Continuous Integration tools/frameworks

What Will Make You Stand Out (Preferred Qualifications)

  • Advanced knowledge of Virtualization, Cloud Architecture, modern Cloud Services, and automated deployment methodologies
  • A proven track record of resolving critical escalations and proactively preventing the reoccurrence of incidents through targeted process, monitoring, and reliability improvements
  • Active contribution to OS/software packaging and distribution, alongside a passion for mentoring others on SRE best practices within the team

#LI-SK3

#LI-HYBRID

Sign up for Job Alerts