Senior MTS - SRE
WHO WE ARE:
Aviatrix is pioneering the Cloud Native Security Fabric - the architecture the Containment Era requires. The Cloud Native Security Fabric governs every workload communication path across every cloud, every VPC, every Kubernetes cluster, and every serverless function, from a single policy plane. One rule. Universal propagation. Enforced at the workload, not at a chokepoint. Trusted by more than 500 of the world's leading enterprises. For more information, visit aviatrix.ai.
About the Role - Senior MTS, Site Reliability Engineering
The Aviatrix SRE team is a small but highly skilled global group of Systems Engineers/SREs dedicated to ensuring the reliability, availability, and performance of Aviatrix's critical systems and services. Our mission is to build and maintain a robust, resilient infrastructure that enables Aviatrix to deliver high-quality services with agility through automation, best practices, and a culture of operational excellence.
As a Senior Member of Technical Staff (Sr MTS) Site Reliability Engineer, you're a proven mid-level engineer who can work independently with some supervision. You'll take on more complex technical challenges while building your leadership and mentoring skills.
Responsibilities
- Kubernetes - Manage application lifecycles, perform troubleshooting, and implement basic monitoring solutions
- Infrastructure as Code: Design and implement laC solutions for infrastructure provisioning and configuration management
- Automation & Development: Build automation tools and enhance existing frameworks in Golang and Python
- Reliability Engineering: Design reliability improvements for individual services; implement basic SLI/SLO frameworks
- Automation Excellence: Build automation tools for routine operational tasks; enhance existing automation frameworks
- Observability: Design and implement monitoring for services; create and maintain alerting rules and basic dashboards
- Incident Management: Lead response for moderate severity incidents; conduct basic post-incident reviews
- Performance Engineering: Analyze performance bottlenecks; implement optimization solutions with measurable impact
- Independent Problem-Solving: Solve technically difficult but well-defined problems with minimal guidance
- Collaboration: Represent SRE perspective in cross-team technical discussions; mentor junior team members
- Mentoring: Provide guidance to junior team members
Requirement
- Experience: 3+ years with BS in designated Engineering field, or 0-3+ years with advanced degree
- Technical Skills: Proficiency in Golang and Python with demonstrated problem-solving ability
- Cloud Expertise: Solid experience with cloud platforms and cloud-native technologies
- Infrastructure as Code: Working knowledge of Terraform for infrastructure management
- Kubernetes: Good understanding of Kubernetes concepts and operations
- Monitoring: Experience with monitoring tools (Prometheus, Grafana) and logging solutions
- System Administration: Solid Linux system administration experience
- Communication: Excellent communication skills for cross-team collaboration
Watch our culture video: glimpse of life at Aviatrix