Principal Engineer - SambaRack
About the team
Join the company that's building the future of AI infrastructure. SambaNova is developing advanced AI systems powered by our RDU (Reconfigurable Dataflow Unit) architecture, combining hardware and software into an integrated platform for large-scale AI deployments.
Our products include SambaRack, a rack-scale AI infrastructure system, and SambaStack, our inference serving platform. Together, they enable customers to deploy and operate production AI environments with performance, reliability, and control.
We are a team of engineers and innovators building next-generation AI infrastructure systems designed for enterprise-scale deployments.
About the role
SambaNova is hiring a Principal Engineer for the SambaRack platform.
You will help build the software that enables customers to monitor, control, update, and diagnose SambaRack systems safely and reliably.
This is a hands-on software development role requiring solid technical expertise in systems and infrastructure, focused on building rack-scale hardware management and infrastructure software that operates in direct contact with hardware across diverse production environments, including customer-managed and air-gapped deployments.
As a Principal Engineer on the SambaRack team, you will:
- Improve the performance, scalability, and reliability of components within rack-scale infrastructure and management software
- Build new features and capabilities for monitoring, control, and diagnostics of SambaRack systems
- Design and build components for hardware-software integration points across rack-scale AI systems
- Support technical execution of infrastructure initiatives, tightly integrated with hardware
This is a high-impact role at the intersection of:
- AI infrastructure
- Distributed systems
- Hardware-software integration
Responsibilities
Some of your responsibilities will include:
- Design and build software components that enable safe, reliable control, monitoring, and diagnostics for rack-scale AI systems
- Contribute to the technical design and implementation of hardware-software integration points and infrastructure services
- Build and improve observability, telemetry, and diagnostics capabilities for the SambaRack platform
- Identify and help resolve performance bottlenecks, reliability gaps, and scaling constraints within owned components
- Collaborate with Hardware Engineering, DevOps, QA, and Product teams to implement infrastructure requirements as sound technical solutions
- Apply and help refine architectural standards, patterns, and best practices within the infrastructure software team
- Partner with senior engineers and architects on technical designs, and share knowledge with peers to support engineering excellence
- Stay current on emerging technologies and industry trends relevant to the platform and assigned work
- Build new systems, components, and capabilities to solve new and interesting problems in the AI inference space
Required qualifications
- 5-8 years of software engineering experience
- Experience building infrastructure, systems, or platform software
- Solid programming experience in Go, Rust, Python, or C/C++
- Good understanding of Linux, networking, concurrency, and distributed systems
- Experience building backend or control-plane services for production systems
- Experience with hardware management systems such as BMCs, Redfish, IPMI, or OpenBMC
- Experience with monitoring, telemetry, or observability systems for infrastructure platforms
- Excellent problem-solving skills and attention to detail
- Ability to collaborate across cross-functional teams
- Knowledge of software development best practices and coding standards
Preferred qualifications
- Experience with rack management, server management, or bare-metal infrastructure platforms
- Experience with Prometheus, Grafana, and metrics exporter design
- Familiarity with fleet management, diagnostics systems, and hardware health monitoring
- Experience supporting enterprise or air-gapped deployments
- Understanding of Kubernetes, Helm Charts, and the Kubernetes ecosystem
- Experience contributing to scalable infrastructure and distributed systems designs
- Familiarity with telemetry systems and infrastructure observability tooling