Skip to main content
Posted 08 August, 2026

Senior Software Engineer - Kubernetes & Service Mesh

Roku
Bengaluru, India Full Time
Reference: 102_755645_7957320

What does the team work on?

Our Cloud Compute Platform Engineering team is at the heart of Roku's transformation toward a single, unified platform. We design and scale Kubernetes clusters, service mesh architecture, and supporting systems to ensure all engineering teams speak the same infrastructure language. We partner with internal teams to migrate workloads, enhance CI/CD pipelines, and integrate observability and security into every layer of the stack. Collaboration and impact are core to what we do.

What is the role?

Join us in building Roku's next-generation cloud-agnostic platform that powers Kubernetes and service mesh at scale. If you're passionate about designing resilient infrastructure, automating deployments, and enabling hundreds of workloads across multiple regions, this role is for you. You'll work with cutting-edge technologies like Kubernetes, Istio, Envoy, Terraform, modern observability stacks and collaborate with teams worldwide to deliver a unified hosting experience. If you're passionate about deep platform and infrastructure engineering, solving complex scaling challenges, and enabling teams through elegant, automated solutions-this role is for you.

How will I use AI at Roku?

At Roku, AI agents do most of the keystroke-level coding. Engineers act as the technical leads for those agents - setting context, planning work, verifying outputs, and recovering when the agent goes off course. If you are excited about working this way, we want to talk to you.

At Roku, we're embracing AI as a powerful tool to amplify human creativity, accelerate innovation, and deliver better results for our customers. Across teams, we're exploring how AI can enhance our work-and we're only looking for people who are curious, adaptable, and excited to grow alongside these technologies.

What are the responsibilities of the role?

  • Architect, design, and deploy Roku's next-generation cloud platform and service mesh
  • Build and own solutions to Roku's compute problems using Docker, Kubernetes, Istio/Envoy, Terraform and scripting to evolve our tech stack and deployments
  • Proactively drive the research and implementation of new technologies to enhance scalability, reliability, and developer experience
  • Integrate security best practices into infrastructure design and automation
  • Build tooling to visualize inefficiencies and optimize costs across shared-tenancy clusters, including network traffic insights, cross-cluster communication efficiency, and cost attribution
  • Collaborate with internal teams to migrate workloads to Kubernetes + Istio, leveraging open-source observability tools
  • Work closely with the Observability team to scale monitoring and logging solutions for a holistic view of the platform
  • Leverage SRE principles to maintain high availability and streamline onboarding workflows
  • Mentor team members and help define best practices for infrastructure and automation

What experience would help someone be successful in this role at Roku?

  • Strong hands-on experience with cloud technologies (AWS preferred; GCP or Azure is a plus), specifically in architecting and managing performant, large-scale systems handling significant traffic/data
  • Deep knowledge of Kubernetes (EKS, GKE, AKS, or similar) and service mesh technologies
  • Proficiency in Go or another programming language, Python or another scripting language
  • Experience designing infrastructure and building automation tools, while collaborating with internal team members and external stakeholders
  • Experience building CI/CD pipelines and following modern deployment practices
  • Familiarity with observability tools (Prometheus, Thanos, Loki, Grafana, etc.)
  • Ability to work independently and communicate effectively with technical and non-technical stakeholders
  • Passion for learning and solving complex infrastructure challenges
  • Experience integrating AI tools to improve processes and reduce operational toil (a plus)
  • You have built fluency across the agentic engineering toolchain - coding harnesses like Claude Code or Cursor, MCP servers, custom skills, or agent frameworks, and you can describe projects where you shipped real work with these tools; you know how to drive an agent, verify its output, and ramp on an unfamiliar codebase with an agent helping you
  • Master's degree or equivalent experience (8+ years preferred)
#LI-DN1

Sign up for Job Alerts