Log in

How it works

You pay one flat monthly price and use the server as much as you like. What a Share buys is precedence — how quickly you get served when the server is busy.

Your monthly price covers all the work you send. When the server is busy, scheduling order stretches response time while your bill stays fixed. The five pages below explain each part in full, with the numbers behind it.

One request, end to end

The whole mechanism in the order you meet it.

  1. Your request waits in your personal queueRequests are handled in the order they arrive.See how your request reaches the server →
  2. The scheduling loop looks for work based on who has priority creditsThe loop checks P1 first, then P2, then P3.See how the scheduler chooses the next turn →
  3. The scheduler picks your request and assigns it a priority class based on what you haveThe class depends on the credits in your buckets when your request is picked.See how priority is assigned →
  4. It costs you one credit from that classThe credit comes from the bucket for that class.See what one credit represents →
  5. The request is sent to the backend queue to be executedThe shared backend queue feeds work to the server.See how the backend queue is filled →
  6. On completion, the exact cost is deducted from your priority class bucketThe charge reflects the work the request actually used.See how completed work is charged →
  7. Your buckets are always refilling at the rate defined by your SharesMore Shares means faster refill.See how refill restores priority →
One request from arrival to completion. It waits in your personal queue while the P1 and P2 buckets refill. The scheduling loop looks for work based on who has priority credits, picks your request and assigns it a class based on what you have. One credit is deducted from that class before the request is sent to the backend queue and executed on the DGX B300. On completion, the exact cost is deducted from the same class bucket.
One request waits in your personal queue, is assigned a class, and is charged against that class's bucket. The refill runs throughout.

At a quiet hour your request can run straight away. At a busy hour the scheduler shares the server by lengthening the gaps between your turns. Your buckets give work priority, their refill brings that priority back, and every request keeps moving toward execution.

Take a Share on the next server

We buy an 8×B300 inference server, split it into Shares, and sell those as monthly subscriptions. Each server's split is published before it fills, and a pledge becomes a subscription once that server is ready to proceed.