How it works
You pay one flat monthly price and use the server as much as you like. What a Share buys is precedence — how quickly you get served when the server is busy.
Your monthly price covers all the work you send. When the server is busy, scheduling order stretches response time while your bill stays fixed. The five pages below explain each part in full, with the numbers behind it.
A flat price, on purpose
→Unpredictable bills, the instinct to just buy a server, and how one fixed server makes a fixed monthly price possible.
A fair share
→The whole server when it is free, the three ways to buy inference, and how a Share keeps your work moving under load.
Counting tokens
→A request has three token counts. Credits combine them into one comparable number tied directly to dollars.
Credits that come back
→A maximum and a continuous refill: a balance that stores quiet time for the next burst of work.
Resolving contention
→P1/P2/P3, replenishing buckets, participant interleaving, completion-time charging. The scheduler made legible.
One request, end to end
The whole mechanism in the order you meet it.
- Your request waits in your personal queueRequests are handled in the order they arrive.See how your request reaches the server →
- The scheduling loop looks for work based on who has priority creditsThe loop checks P1 first, then P2, then P3.See how the scheduler chooses the next turn →
- The scheduler picks your request and assigns it a priority class based on what you haveThe class depends on the credits in your buckets when your request is picked.See how priority is assigned →
- It costs you one credit from that classThe credit comes from the bucket for that class.See what one credit represents →
- The request is sent to the backend queue to be executedThe shared backend queue feeds work to the server.See how the backend queue is filled →
- On completion, the exact cost is deducted from your priority class bucketThe charge reflects the work the request actually used.See how completed work is charged →
- Your buckets are always refilling at the rate defined by your SharesMore Shares means faster refill.See how refill restores priority →
At a quiet hour your request can run straight away. At a busy hour the scheduler shares the server by lengthening the gaps between your turns. Your buckets give work priority, their refill brings that priority back, and every request keeps moving toward execution.
Take a Share on the next server
We buy an 8×B300 inference server, split it into Shares, and sell those as monthly subscriptions. Each server's split is published before it fills, and a pledge becomes a subscription once that server is ready to proceed.