Log in

Resolving contention

One loop, asking the same question from the top every time. It picks a class for the turn, takes work from every queue that qualifies, and never reconsiders anything already sent.

What the scheduler actually does

Requests arrive and go straight into a queue — one queue per participant, in arrival order. Nothing about them is decided yet. They are not fast requests or slow ones, and they do not belong to a priority class.

The scheduler is a tight loop reading from those queues, and it always wants to be doing P1 work. So it starts by ignoring every queue holding less than one P1 credit. If any queue is left with work waiting, that settles it: this is a P1 pass.

If ignoring them leaves nothing, it asks the same question of P2. If that leaves nothing either, it drops to P3 — which is not a separate pool of work, just the same queues with the credit test switched off. Every queue with something waiting is eligible there, whatever Share it holds.

Now it has a class in mind and a set of queues that qualify. Its cursor interleaves those queues into a turn, deducting a credit for each request it sends to the server. The batch has no fixed size or duration; those move with operating conditions. The loop does not wait for anything to finish. It checks whether the server has room for more; if not it waits a moment and checks again, and if so it goes round and asks the whole question afresh.

One turn of the scheduler loop: it ignores queues without a P1 credit and serves what is left, otherwise repeats the test for P2, otherwise drops to P3 where every queue with work qualifies. It interleaves eligible queues into a batch, charges a credit for each preferential request, dispatches, and checks whether the server has room before going round again.
One turn of the loop. The question is asked again from the top every single time.

Assigning a priority class to a request

When the loop reaches your queue, it assigns your request the highest priority class your current bucket balances support. Your Share decides which buckets you have, and their balance at that moment decides the class.

A request is assigned P1 when your P1 bucket has a credit, P2 when your P2 bucket supports the turn, and P3 when the server has room to take it. Send a burst of heavy work and you will watch your requests move through the classes as your balances change.

And the same thing happens in reverse, because the buckets refill on a clock whatever the server is doing. Work still waiting when a credit lands goes out at the higher class. Dropping down and climbing back up are not two mechanisms; they are one mechanism, seen at two moments.

Why you can’t get stuck

That is what makes waiting bounded rather than open-ended. The thing that ends your wait is your own bucket reaching a credit, and it gets there at the cadence your Share publishes. Nobody else's spending slows your refill down, because eligibility is tested per queue against that queue's own balance.

Pro holds both buckets, so its wait is bounded by its own P1 refill and nothing else. Lite holds P2, so it also needs P1 work across the server to run out before the loop comes down to its class — which it does. That is the difference the two prices buy: not whether your work runs, but how directly you can shorten the wait.

Nothing here refuses anything. There is no state a request can be in from which it never runs.

Taking turns

Within a class the cursor interleaves queues. It resumes where the previous turn left off, so a queue with a hundred requests waiting cannot occupy a class just by having arrived first.

More Shares give you more credits — faster refill, bigger buckets — but they do not give you more turns. Within-class fairness is per participant, not per Share. One Pro Share and ten Pro Shares get the same visits; the ten just have more priority to spend on them.

There is one pleasant consequence of that. The busier the server, the longer between your turns — and the more your bucket has refilled by the time one arrives.

Watch the loop run

Send some work, put some other queues on the server, and run the loop. Watch which queues it ignores, which class it settles on, and what your own requests end up going out under.

Your Share

Your traffic

Simulation

P1 bucket

50%

refill +3.3% / tick

P2 bucket

25%

refill +8.3% / tick

Waiting

1

1 queue · 0 backend

Completed

0

nothing yet

one queue per participantP2P1youq1q2q3q4q5q6Loop startingthe first turn is about to runshared backend queueone row is one turnNVIDIA DGX B300Ready

Turns since reset

P1

N/A

P2

N/A

P3

N/A

You

N/A

This is a scaled-down approximation for illustration, not a production model. Request sizes, bucket sizes, and refill rates use illustrative values to show the mechanism; they are not the exact production settings.

Paying for what you used

A dispatch costs one credit, flat, whatever the request turns out to be. When it finishes, the real resource cost is charged back to the same bucket that admitted it — even if you are on a different class by then. Heavier work costs more, lighter work costs less.

So heavy requests are never blocked in advance. They run, they cost what they cost, and what they slow is your own return to priority. Nothing already dispatched is ever taken back: the loop decides who goes next, never who gets interrupted.

None of that accounting reaches your bill. The bucket regulates burst and keeps things even between everyone on the server, and the price was fixed before you arrived. What the loop decides is the order, not the amount.

What each class holds, and what it costs a month, is on pricing.

Watch the server you’re queueing on

Every server publishes its class split before it fills, and its live metrics once it’s up. Take a Share on the next one.