Specialist Cloud Site Reliability Engineer
So, what's the role all about?
NICE is looking for a Senior Site Reliability Engineer to join our core Reliability Engineering team, responsible for ensuring the scalability, reliability, and performance of mission-critical systems and observability platforms across multiple environments and regions. This role is ideal for someone who thrives in fast-paced environments, enjoys automation, and has a strong background in cloud-native operations, observability stacks, and incident management. You'll collaborate closely with product, platform, and development teams to drive reliability-first design, proactive observability, and operational excellence.
How will you make an impact?
Reliability & Performance
Design and implement scalable, reliable, and resilient systems across hybrid or multi-cloud environments (primarily AWS/EKS/ECS/Lambda)
Drive improvements in system uptime, latency, and overall service health metrics (SLOs, SLIs, SLAs)
Automation & Infrastructure as Code
Build and manage infrastructure automation using Terraform, Helm, and Kubernetes
Improve CI/CD pipelines using Jenkins, GitHub Actions, ensuring safe and automated rollouts, monitoring, and rollbacks
Observability & Monitoring
Own and enhance the observability stack (Prometheus, Grafana, Loki, Tempo, OpenTelemetry, Mimir, etc.)
Define and implement SLOs and error budgets; enable development teams to monitor and act on reliability metrics
Incident Management
Lead major incident response, root cause analysis (RCA), and blameless postmortems
Partner with product teams to define and enforce operational readiness standards before production releases
Security & Compliance
Ensure platform-level security and compliance with organizational and regulatory standards
Collaborate with InfoSec and compliance teams to maintain a secure and auditable infrastructure
Technical Leadership
Mentor junior SREs and developers on reliability practices, automation, and observability
Contribute to technical roadmaps and reliability-focused design reviews
Have you got what it takes?
-
8+ Years Strong experience with Kubernetes, EKS, ECS and containerized workloads in production
Expertise in AWS services (EC2, Lambda, IAM, RDS, S3, ALB/NLB, VPC, PrivateLink, etc.)
Proficiency with Terraform, Helm, Jenkins, GitHub Actions, and GitOps (ArgoCD or Flux)
Deep understanding of observability frameworks - metrics, logs, traces, and distributed monitoring
Hands-on experience with Prometheus, Grafana, Loki, Tempo, Alloy, OpenTelemetry, or equivalent tools
Strong knowledge of Linux, networking fundamentals, and system performance tuning
Familiarity with Python, Go, or Shell scripting for automation and custom tooling
Practical experience in incident response, RCA, and on-call operations
What's in it for you?
Join an ever-growing, market disrupting, global company where the teams - comprised of the best of the best - work in a fast-paced, collaborative, and creative environment! As the market leader, every day at NICE is a chance to learn and grow, and there are endless internal career opportunities across multiple roles, disciplines, domains, and locations.
Enjoy NICE-FLEX!
At NICE, we work according to the NICE-FLEX hybrid model, which enables maximum flexibility: 2 days working from the office and 3 days of remote work, each week.
Requisition ID - 11465
Reporting into: Tech Manager
Role Type: Individual Contributor