Posted 30 June, 2026
Staff Site Reliability Engineer
Okta
Bengaluru, India
Full Time
Reference: 102_700121_7307866
What You'll Be Doing
- Design, build, and operate highly scalable, reliable, and secure infrastructure powering our production systems across AWS and GCP.
- Lead major reliability and modernization initiatives, including container platform migrations (e.g., ECS to EKS/GKE) and microservice enablement across multi-cloud environments.
- Serve as a technical authority in Kubernetes (EKS and GKE), cloud infrastructure (AWS and GCP), and modern CI/CD practices (GitOps, automation pipelines).
- Partner with development teams to architect and enable microservice-based applications, ensuring production readiness, scalability, and observability.
- Implement and manage infrastructure as code (Terraform, Ansible) to automate provisioning, scaling, and configuration management across multiple cloud providers.
- Drive improvements in observability, performance, and cost efficiency through robust monitoring, logging, and alerting systems that span AWS and GCP.
- Champion SRE best practices - defining SLOs/SLIs, conducting blameless postmortems, and continuously improving incident response.
- Lead complex technical projects from conception to completion, managing timelines, and technical dependencies across teams.
- Mentor engineers across teams, fostering a culture of reliability, automation, and continuous learning.
- Collaborate with security and compliance partners to ensure infrastructure adheres to best practices and standards (e.g., IAM Federation, Workload Identity).
- Participate in the on-call rotation, using incidents as learning opportunities to enhance systems and processes.
What You'll Bring to the Role:
- Strong hands-on experience architecting and operating cloud-native distributed systems (AWS and GCP).
- Deep expertise with Kubernetes (EKS and GKE) - design, provisioning, scaling, and advanced troubleshooting in production.
- Proven experience leading ECS to EKS/GKE migrations and driving microservice enablement initiatives at scale.
- Proficiency with Infrastructure as Code tools such as Terraform (multi-provider), Ansible, or CloudFormation.
- Solid coding and scripting ability in Python, Go, or Shell, with a focus on automation, tooling, and operational excellence.
- Advanced understanding of CI/CD pipelines (ArgoCD, GitLab CI, Spinnaker), Linux systems, and networking fundamentals (Direct Connect/Interconnect, DNS, routing, load balancing).
- Experience managing databases and caching systems (e.g., RDS/Cloud SQL, Redis/Memorystore, PostgreSQL, MySQL) in cloud environments.
- Hands-on experience with observability tools (Prometheus, Grafana, ELK, Loki, OpenTelemetry, Google Cloud Operations) for performance and reliability insights.
- Working knowledge of container security, secrets management (HashiCorp Vault, AWS Secrets Manager, Google Secret Manager), and compliance in production environments.
- Strong communication and problem-solving skills, with demonstrated success leading cross-team projects and mentoring peers.
Experience:
- 8+ years in SRE, DevOps, or Infrastructure Engineering roles.
- 3-5 years of experience with Kubernetes (EKS/GKE) and related ecosystem tools (Helm, Karpenter, etc.) in production.
- 3-5 years of experience with AWS and GCP.
- 3-5 years using Terraform to manage multi-cloud infrastructure.
- 5+ years of coding experience in Python, Go, or similar languages.
- Proven track record leading high-impact projects, specifically migration projects (ECS EKS/GKE) and enabling microservice architectures.
- Experience implementing SLOs/SLIs, performing root cause analyses, and improving operational resilience.
- Prior work in SaaS or high-scale, cloud-native environments is a strong plus.
- Strong Linux and security fundamentals.
- Bachelor's degree in Computer Science or equivalent hands-on experience.