Log in

A flat price, on purpose

One server, split between the people on it, at a price that doesn’t move. It’s the instinct to stop renting cloud and just buy the hardware — without the outlay.

It’s a choice

One price a month. Use the server as much or as little as you like — the number is the same either way, and you know it before you start.

Nothing forced us into that. We chose it, and then built something that can hold it. The rest of this page is why we wanted it and how it holds.

What we’re getting away from

We built this because we were tired of not knowing what the bill would be. Not the size of it — the not knowing.

The same week’s work costs different amounts depending on cache hits, retries, how chatty the agent was, whether something looped overnight. You find out at the end of the month, long after the moment you could have done anything about it.

And the failure mode has no ceiling. A stuck agent doesn’t cost ten per cent more, it costs forty times more, and nothing stops it until something you built yourself notices. So you end up managing the meter instead of the work — budget alarms, capped keys, thinking about token counts in the middle of a refactor. That is effort that produces nothing.

The usual answer to an unpredictable bill is a capped plan, which fixes the number by putting a wall somewhere in your month instead. So there are no rolling five-hour rate limits here, no weekly window and no monthly ceiling. What you get instead is a fair share of one server.

Twelve months of the same work billed two ways. A metered bill varies month to month and spikes badly in one of them. A Share is the same figure every month.
The same work either way. Only one of them is a number you can plan around.

The instinct to just buy a server

If you have been on the wrong end of that, you have probably had the thought: forget it, buy a server. Price one up, work out the monthly cost once, and be done with the whole business.

The instinct is right, and it isn’t nostalgia. A server you can point at has properties a pool doesn’t. You know exactly what it is. You know what it costs in advance and the number doesn’t move. You know how many other people are on it. There is no capacity-management layer between you and the silicon to be surprised by. Most of all, the whole arrangement fits in your head.

The server

GPU
8 × NVIDIA B300
Memory
8 × 288 GB HBM3e (2.304 TB)
Tenancy
dedicated, single-tenant to iqshard

Then you price that, and it is usually where the thought ends. An 8 × B300 is not something you buy on a whim. You would carry the depreciation on hardware that is moving quickly. You would be the ops team — drivers, serving stack, upgrades, monitoring, the 3am reboot. And you would have bought for your peak, so it would sit idle most of the time, which is the thing that made elastic cloud attractive in the first place.

So we bought the server

That is the whole idea. We buy one server, run it, and split it into Shares that people take on a monthly subscription.

You get what made the instinct attractive — a specific server, a price fixed in advance, a small and finite group of people on it, and no capacity layer in the way — without the outlay, the depreciation, the ops or the idle time. Every server is published before it fills, so you can see which one you are joining and how it has been split.

An elastic fleet spreads customers across many servers behind a capacity-management layer, so the overhead moves with the fleet and the bill is how that variance is allocated back. One server has a published split, a fixed group on it, and no capacity layer, so nothing varies from month to month.
Two ways to sell inference. One of them needs a meter to work; the other has nothing for a meter to measure.

That is also why the price cannot drift. There is no elastic pool behind it. We don’t resize anything when the server gets busy — when demand grows we buy another server and split that one too. Nothing about the arrangement changes from month to month, so there is nothing for a meter to measure.

It is a smaller idea than a cloud, and that is the point. An elastic fleet is a sophisticated answer to a genuinely hard problem: spikes absorbed, hardware failures survived, customers far larger than any one server served. All of that is paid for in variable overhead and variable complexity, and the bill is the mechanism that allocates it back to you. We picked the smaller problem, which is exactly why the price can be a single number.

What it costs you

You are not alone on the server. That is the trade for not having bought one outright, and it is worth being plain about.

When it is busy, someone else’s work goes first and yours waits a moment. Next time round it is yours that goes and theirs that waits. Everyone on that server is in the same position on the same terms, nobody is moved aside to make room for anyone else, and nobody can buy their way past the queue — more Shares buy more credits, not more turns.

So what moves under load is ordering, not money. Your bill stays where it is; your place in the queue is the thing that changes. What decides that place, and what keeps any one person from monopolising it, is the subject of Resolving contention.

Take a Share on the next server

One server, one monthly price, and the split published before it fills. Cancel any time.