The missing compute orchestration layer
Today we announced our seed round, and $500M of compute running Aranya - the multicluster OS that orchestrates entire fleets of datacenters into one cohesive, reliable, automated whole. For a company that is 1y old, it might sound like overnight success. But this moment has been building for years.
A Decade of Distributed Compute
I’ve been wrangling masses of GPUs for the last 12y. I was the first GPU eng at Crusoe, now worth somewhere around $10-40B. I was doing this years before the first neocloud was born, though.
Somewhere back in 2014-15, I snuck into the Harvard machine lab and took a bandsaw to the mounting frame of an NVIDIA GTX 750 to get it to fit into a 3D-printed demo chassis. That demo raised my first VC cheque: $25k from Rough Draft Ventures that my cofounders and I turned into a $320k kickstarter. Then Apple caught wind, shut us down and shipped a clone in every store in the US. There went my first mass GPU venture.
Next: I built turingtest.io in 2016 - a sleek p2p ‘AI’ chat app that shortcut its way to near-human communication with the sparsest of compute resources. It worked shockingly well until I pulled the plug because $20/mo felt like an irresponsible amount to spend on compute to fuel an AI chatbot. Quaint, right?
So I hacked together a roughshod, self-healing distributed system to harvest multiple regions worth of idle GPUs (Kepler era!) on the AWS Spot markets at a fraction of list price. That’s when things really clicked.
I started to build Squire: a distributed multiparty GPU supercomputer for AGI.
…a two-sided marketplace for digital commodities, which has the potential to produce a distributed supercomputer technically capable of creating consciousness. This is a worthy goal.
- 2017, from my Digital Commodities paper predicting a distributed future of intelligence
Shortly after Crusoe's seed closed, Chase Lochmiller convinced me to come turn Crusoe's distributed bitcoin infrastructure into GPU datacenters, when there were just 10 of us and the word "neocloud" didn’t exist. A year later, I racked Crusoe's first GPU rig by hand (still have the back injury to show for it)! The seeds of a multibillion-dollar neocloud had been planted, and it was time to get back on the path. Atoka (my next venture) was Crusoe’s first customer, developing datacenter-scale automation in forensic cybersecurity.
Then I joined Hyperbolic as their Founding Platform Engineer, turning a dozen-odd scrappy suppliers into a 200+ cluster Kubernetes federation that could serve production-grade AI on spotty GPUs at 99.9%
So I’ve been building AI supercomputers for a long time. All this history is to ask your credulity when I say:
There is No Chip Shortage
[Compute] is digital oil - the scarce resource in a digital age.
- 2019, from a Squire strategy pdf sent to Chase Lochmiller before joining Crusoe
The trillion-dollar contradiction
For better or worse, civilization now runs on compute. Every prompt, bubble, and chat is rooted deep in the scorching rack of a tightly-packed datacenter, and the demand is exponential.
The hyperscalers gobble up almost the entire production line of new chips years before they even hit the rack, leaving most AI companies to fight for neocloud scraps. They are forced to choose between two bad options - absorb the complexity of heterogeneous GPU hardware themselves by building a platform team from scratch, or pay a steep premium for a generic cluster that poorly fits their architecture.
Even the largest AI companies are now forced to look outside of AWS, Azure, etc for GPUs - at a scale that necessitates multi-datacenter architecture.
So what happens? Precious oil slips through each of these cracks. A GPU sits idle for six days because it needs a reboot and nobody notices it among the millions of data points in their observability system. An overworked SRE gets paged at 1am two times in a row for false alarms, and on the third they sleep through a full network outage. Entire regions of duplicate servers on standby because there is no multicluster orchestration layer.
Now multiply that by every cluster on earth, and you’ll begin to see the scale of leakage. We live in a world where compute is the most scarce resource, and yet enormous quantities sit idle and unused every day. Thousands of new datacenters under construction to serve an asset that already exists.
The compute for AGI is already online - it just needs to be unlocked and connected.
Liberating Compute
To that end, there is no more important effort than improving the accessible supply of compute power and storage. To do so is to fundamentally increase our capacity for intelligence.
- 2017, Digital Commodities
For a full decade, I have dedicated my career to unleashing this bottled-up supply. And because of the incredible team that has come to join this effort at Aranya, it is finally happening. My co-founders Sasi and Arya quickly came to share my fervor, for the last three and nine years respectively. And each of Aranya’s earliest employees - Yoofi, Lily, Weigang, Ayo, Virginia, Tina - carries the torch of that same fire.
We built clusterdOS into a full datacenter-scale operating system for maximum efficiency – and then we open sourced it! The harder problem sits at a layer above.
Aranya is the multicluster OS, leveraging deep-rooted intelligence to orchestrate that efficiency across an entire fleet of clusters at once - smoothly wrangling tens or even hundreds of datacenters into one cohesive whole. From custom drivers to cross-cluster deployments, Aranya automates all the moving parts. Consider the variety of failure modes, heterogeneous hardware, and divergent architecture in each datacenter (if you are familiar with such things) and you will begin to understand the technological breakthrough this team has accomplished.
Federated Kubernetes at hyperscale has been the holy grail of modern distributed computing for the better part of a decade, with an implementation graveyard that stretches back just as long. We are one of a handful of companies that has pulled it off - and then we carefully injected AI swarms and secure multi-tenancy into its core.
A select few recognized what we had created - the kinds of bleeding-edge companies that know what it takes to make truly massive computing architectures work smoothly. Just a year in, Aranya underpins the core architecture of $500M worth of compute across 10 neoclouds and one of the top three inference providers in the world.
What is the multicluster Operating System?
The hope of this dOS concept is to usher in the advent of the Personal Cluster as our next phase of computer use — following the evolutionary pattern of QDOS [Windows] and the personal computer.
- clusterdOS repo (open-sourced 2024, first private commits in 2022)
Back in the ’80s, there was this cool thing called a ‘Computer’ at most companies and it was super powerful, except only a few engineers at work could use it. Then, suddenly, Windows came around and sparked the personal computer explosion. The problem was not the hardware, it was the absence of an OS that could unleash it.
GPU clusters today are stuck squarely in the pre-OS era of computing - there are these powerful things called ‘Clusters,’ but only a few people can use them.
Not at Aranya though, and not for the industry-leading orgs we’ve taken off our waitlist. From SRE to PM to Sales to CFO, Aranya users can securely and reliably work with hundreds of millions of dollars of compute - as seamlessly as a laptop.
The Future of Intelligence
We have created an operating system so powerful, a single person can orchestrate half a billion dollars worth of compute. That has allowed Aranya users to out-scale and out-pace their closest competitors, who still live in the pre-OS era. If that kind of power sounds up your alley, then come check us out.
The install base is exploding - with total fleet value set to cross $1B by the end of 2026. Yet the vast majority of compute is still stuck in the wasteful, gatekept era of pre-OS clusters. That’s why we raised - $11M in total, with First Round Capital leading our seed round.
Thank you to Asylum, Uncommon, Founder Collective, Vermilion Cliffs, Box Group, Parable, and our earliest angels for believing in Sasi, Arya, myself, and the dynamic team at Aranya.
It is time to unlock the future of intelligence by liberating all this compute.
The compute power theorized to be necessary for Artificial Intelligence already exists - it just has to be connected. For this reason, I believe the first digital consciousness will come from a distributed supercomputer.
- 2017, Digital Commodities
Thank you to Arya for co-writing this post, and Sasi for helping edit!
Thank you to Todd and Maggie, Lola and Nick and Mackenzie, Patrick and Chad, David and Jack, Ashley, Adam, Anne, and our earliest angels for betting on us with speed and conviction.