Build the multicluster operating system
We’re unlocking the future of intelligence by liberating compute.
Open Roles
We're hiring our first Business Operations hire — a generalist role that will work directly with the COO. This is a high-leverage seat for someone who wants broad exposure across a startup's operations: finance, people ops, legal/compliance coordination, and internal systems. You'll help keep the operational engine of the company running smoothly as we scale to a larger team. This role is ideal for someone who is highly organized, resourceful, comfortable with ambiguity, and excited to learn by doing across many different functions.
- Support financial operations: help maintain and update the growth model, track spend, assist with accounts and vendor invoices
- Assist with people operations: coordinate hiring logistics (scheduling, candidate communication, interview loop tracking), help onboard new hires, support benefits administration
- Help manage compliance workflows (e.g., SOC2, GDPR) by tracking action items and coordinating with responsible owners
- Support legal/administrative coordination: organizing documents, tracking contracts and deadlines, liaising with outside counsel on routine matters
- Maintain internal operational systems and tools (Notion, investor databases, shared drives): keeping them organized, current, and useful
- Help Marketing plan and execute internal and external events (team offsites, investor dinners, conference logistics)
- Take on ad hoc research and analysis projects for the COO (market data, competitive info, vendor comparisons, etc.)
- Draft and proofread internal and external communications as needed
- Identify and help fix operational inefficiencies as the company grows
- 2+ years of experience; prior experience in operations or a fast-paced startup/small-team environment is a plus but not required
- Extremely organized with strong attention to detail — things don't fall through the cracks on your watch
- Comfortable juggling multiple workstreams and shifting priorities in a small, fast-moving team
- Strong written and verbal communication skills
- Proficiency with spreadsheets (Excel/Google Sheets); comfort picking up new tools quickly (Notion, Vanta, CRM systems, etc.)
- A builder's mindset — willing to roll up your sleeves on both strategic and administrative work
- High trustworthiness and discretion, given exposure to sensitive information
- Genuine interest in AI infrastructure and startups
- Exposure to basic financial modeling or bookkeeping
- Familiarity with HR/benefits administration or compliance tools (e.g., Vanta, Rippling, Gusto)
- Experience supporting fundraising, investor relations, or board processes
- Ground-floor opportunity to shape operations at a fast-growing AI infrastructure startup
- Direct partnership with the COO across every function of the business
- Broad exposure to fundraising, finance, legal and people ops in a high-growth startup environment
- Mission-driven team building category-defining infrastructure
We're looking for a Platform Engineer who will be instrumental in building and evolving clusterdOS. You'll work directly with our founding team to design systems that abstract away Kubernetes complexity while preserving its power. You'll design GitOps workflows, build Kubernetes operators and controllers, and create the automation that makes cluster management invisible to end users.
- Build and extend clusterdOS core features using Go, including custom Kubernetes operators and controllers
- Design and implement GitOps workflows with ArgoCD that make continuous deployment feel automatic
- Develop infrastructure-as-code patterns using Terraform and Helm that provision and manage clusters seamlessly
- Work on distributed storage solutions using Ceph and WEKA for high-performance, scalable cluster storage
- Create observability and monitoring systems using Prometheus and Grafana to surface cluster health and performance
- Build and optimize container networking with Cilium for network security and observability
- Design and implement federated Kubernetes architectures for multi-cluster management
- Build automation tooling that reduces operational overhead for developers running production workloads
- Contribute to open source components
- Establish platform architecture patterns and practices as an early engineering team member
- Go — Strong proficiency building system-level tooling, controllers, and distributed systems
- ArgoCD — Deep experience with GitOps workflows, declarative infrastructure, and continuous deployment patterns
- Kubernetes — Production experience with Kubernetes internals, custom resources (CRDs), operators, cluster architecture, etcd clusters, including backup/restore procedures and troubleshooting cluster health issues
- Infrastructure as Code — Experience with Terraform, Helm, or similar tools for automating infrastructure provisioning. Experience with Ansible, Kubespray, or similar orchestration frameworks
- Observability — Familiarity with monitoring and logging systems like Prometheus, Grafana, or similar platforms
- Strong understanding of containerization, networking, and cloud-native architectures
- Experience with CI/CD pipelines and automation workflows
- Ability to design systems that are both powerful and simple to use
- Self-directed with good prioritization instincts — you can identify what matters and execute accordingly
- Experience building developer tools or infrastructure products
- Background with service mesh technologies (Istio, Linkerd) or CNI plugins
- Contributions to projects or Kubernetes ecosystem tools
- Familiarity with cloud platforms (AWS, GCP, Azure) and their managed Kubernetes offerings
- Experience with policy-as-code and security tooling for Kubernetes
- Track record shipping infrastructure products that developers love
- Understanding of FinOps and infrastructure cost optimization
- You stay calm under pressure and know when to escalate versus dig deeper yourself
- You communicate clearly with both engineers and non-technical stakeholders, especially when things break
- You're curious about root causes and genuinely interested in preventing problems before they happen
- You work with low ego
We're looking for a Senior Infrastructure/SRE Engineer to join our on-call team for enterprise clients. You'll be the technical expert our clients rely on when things go wrong with extremely valuable clusters—diagnosing complex infrastructure issues, resolving production incidents, and ensuring zero downtime for critical AI workloads. This role requires deep technical knowledge, excellent troubleshooting skills, and the ability to stay calm under pressure.
- Respond to and resolve production incidents across our clients' infrastructure, including Kubernetes clusters, Ceph storage systems, and bare metal servers
- Diagnose complex issues ranging from pod scheduling problems and CNI networking failures to distributed storage performance degradation and hardware issues
- Handle escalations requiring deep expertise in distributed systems, including etcd cluster problems, Ceph RGW authentication issues, and custom networking setups with Cilium and bare metal load balancers
- Work directly with enterprise clients during incidents, providing clear communication about status, timeline, and resolution steps
- Participate in a follow-the-sun on-call rotation with engineers across time zones
- Document incidents thoroughly and improve runbooks based on recurring patterns
- Collaborate with our infrastructure team on long-term reliability improvements and automation to reduce incident frequency
- 5+ years of production experience with Kubernetes in enterprise environments, including deep knowledge of cluster operations, troubleshooting bare metal k8s issues, working with admission controllers, and understanding the control plane architecture
- Strong experience with distributed storage systems—Ceph experience is highly preferred, but deep experience with other systems like Weka, VAST, or similar is highly valuable as well
- Very solid modern Linux systems administration skills
- Experience with bare metal infrastructure management and construction, not just cloud environments. You should be comfortable with IPMI, hardware troubleshooting, networking, etc.
- Experience with at least one CNI plugin in production, preferably Cilium or Calico
- Strong troubleshooting methodology—you know how to systematically narrow down issues across complex distributed systems and can work effectively under pressure during outages
- Production experience with etcd clusters, including backup/restore procedures and troubleshooting cluster health issues
- Experience with GPU infrastructure for AI/ML workloads, including the NVIDIA Kubernetes operator and GPU stack
- Knowledge of infrastructure as code tools like Ansible, Kubespray, or similar orchestration frameworks
- Experience with high-availability load balancers, service mesh technologies, or IPAM systems
- Background working with AI inference or training companies and understanding their unique infrastructure requirements
- You're calm and methodical during incidents, with good judgment about when to escalate versus when to dig deeper yourself
- You communicate clearly with both technical and non-technical audiences, especially during high-stakes situations
- You have intellectual curiosity about root causes and want to prevent problems in advance
About Aranya
We're building Aranya to make production-grade clusters accessible to anyone. Our multicluster OS removes the operational complexity so developers and teams can run serious infrastructure without platform expertise.
At Aranya, different perspectives aren't just welcome, they're essential. The best infrastructure comes from people with diverse experiences and ideas.
.webp)