A fair share
Use the server whenever nobody else is on it, and not a slice of it either. What bounds you is being fair to the people sharing it.
The whole server, whenever it’s free
When nobody else has work waiting, every turn is yours. The scheduler hands your requests to the server until it has no room left, then goes round and does it again. There is no slice of the machine set aside with your name on it. You are using the machine.
When somebody else is working too, you get fewer turns. Not less server, fewer turns. The cursor moves through the queues with work waiting and resumes where it left off, so two people working at once take turns, and the more of you there are at that moment, the longer between yours.
That is the bound, and it is the only one. Not a cap we set or an allowance we issued, but the other people on the server and how much of it they want right now. When they want none of it, you get all of it.
Three ways to buy inference
Pay as you go, at retail. Send what you like, whenever you like, and wait for nothing. This is the industry default and it works. You pay for every token, which for sustained agentic work adds up quickly and unpredictably.
A bundle. A fixed monthly price for a fixed quantity of usage: so many tokens, or so much retail value, per window. Cheaper per unit than retail. Predictable right up until the quantity runs out.
A share. A fixed monthly price for a place on one server, with no quantity attached to it at all. You use it when it is free, and take turns when it is not.
The trade against pay-as-you-go is worth saying plainly. You accept that sometimes you wait. In exchange the number stops moving.
The middle one is the problem
The bundle is the model we won’t sell, and it is the one most of this market has settled on.
Its allowance is denominated in something you cannot hold in your head. OpenCode Go gives you $60 of retail-priced value for $10 a month. Z.AI’s Coding Plan sells 5-hour and weekly credit windows. Retail dollars, tokens, rolling windows: none of them are units you can feel, so there is no way to know how close you are until you are there.
Then it stops you. Not because the server ran out of capacity, but because your allowance did. A developer mid-session does not want to discover a weekly cap. They want the tool to keep working, and a cap turns a resource question into an interruption at the moment interruption costs most.
Groups of private pilots share aircraft on exactly the terms this page describes. Everyone pays a fixed amount each month, any of them can fly whenever nobody else has it booked, and none of them could carry the aeroplane alone. No group has ever radioed a member over the Atlantic to tell them their five hours were up. The idea is absurd there. It is the same idea here.
So nothing already sent is ever taken back, and nothing is ever refused. When the server is busy your request waits its turn. It does not fail, and there is nothing to top up before you carry on.
The wall, and the alternative to it
The same month of work, run twice: once against a published metered allowance, once against a Share. Set how much you send, and how you send it.
The shape matters as much as the total. The same daily volume spread across office hours or concentrated into one long agent run meets completely different limits, and the second one can be over before lunch on the first day.
When you're working
Each strip is one day, midnight to midnight: weekdays on top, weekend below.
How hard your AI is working on it
60 requests an hour
5,958 / 11,220 run
90 working hours with nothing running.
When refused, you can pay extra.
No cap. A fair share, not an allowance.
11,220 / 11,220 run
Nothing refused.
Busy means slower turns. You never pay extra.
Work ranSome of it refusedAll of it refusedYou weren’t working
A working day at 60 an hour hits Z.AI Coding Plan's the five-hour limit on day 1, 12:00. From there, 5,262 of the month’s 11,220 requests never run — not because a server was busy, but because a balance reached zero. A Share runs the same work in full, and never asks for more than the monthly price.
hours bound by: five-hour 28 · weekly 80 · monthly 0
Measured against: Z.AI Coding Plan's published caps (Lite: 2,000 zai-credits/5hr · 10,000/week, no monthly cap). Requests: assumed to be 52,000 cached input, 700 new input and 150 output tokens, based on published OpenCode data. For simplicity the model treats every request as identical; in practice some are lighter than that and some are heavier. Modelled here: a pattern is a schedule of active hours, and the rate applies evenly within each of them. The five-hour and weekly windows roll; the monthly cap is the calendar month. Refused work is simply not done, and one cell is 2 hours. Z.AI Coding Plan charges 50% outside peak, and its peak is 14:00–18:00 UTC+8 on weekdays. On the clock shown here that falls 07:00–11:00. Neither side is a latency claim; this is what each plan does when you ask for more than it wants to give.
What fair takes
Taking turns is easy to say and harder to do. The server has to know whether two requests are comparable before it can be even-handed about them, and a request is not one number.
That is what credits are for. They turn three token counts into one number the server can compare, and one you can price.