The Copilot Credit Cliff: Why Your AI Coding Bill Jumps on 1 September 2026
GitHub Copilot moved to token-metered AI Credits on 1 June 2026, and cushioned the change with promotional credits that run through June, July and August. Those promotional balances expire, and from 1 September Business and Enterprise teams drop back to a base allowance roughly a third the size. Here is how to work out your real burn before the cushion disappears, and why the budget controls that came with the change belong in your compliance programme, not just your finance stack.
On 1 June 2026 GitHub retired premium request units and moved Copilot onto GitHub AI Credits, a token-metered model where every request draws down a balance based on the input tokens, output tokens and cached tokens it consumed, priced at the published API rate for whichever model handled it. One credit equals one US cent, which makes the arithmetic unusually transparent by industry standards. Seat prices did not move: Pro remains 10 dollars a month, Pro+ 39 dollars, Business 19 dollars per user and Enterprise 39 dollars per user, each carrying an included credit allowance equal to the subscription price. What did move is the relationship between what a developer does and what the invoice says, and most teams have not yet seen an unsubsidised month of that relationship.
The reason is a promotion that is about to end. To smooth the transition, GitHub granted enhanced monthly credits for June, July and August: 30 dollars per user on Business against a base of 19, and 70 dollars per user on Enterprise against a base of 39. In credit terms that is 3,000 instead of 1,900, and 7,000 instead of 3,900. From the September billing cycle the enhanced amounts stop and the base allowance applies. For an Enterprise seat that is a 44 percent reduction in included headroom, arriving without any change in seat price, product behaviour or developer habit. Nobody has to do anything wrong for the bill to climb. The cushion simply deflates, and whatever your engineers were already doing starts landing against a smaller allowance.
The second change matters more than the first and has attracted less attention. Under the old premium request model, exhausting your quota fell back to a lower-cost model, so the worst case was degraded output rather than an unbounded invoice. That fallback is gone. When credits run out, what happens next is entirely determined by the policy an administrator has set. Allow additional usage and consumption continues at published per-credit rates with no ceiling except the one you configure. Disallow it and usage stops until the cycle refreshes. There is no longer a safe default in the middle, which means the configuration itself has become the control, and an unset budget is now an open account rather than a soft landing.
GitHub did ship the tooling to manage this. Budgets can be set at the enterprise level, at the cost centre level and at the individual user level, which is a genuinely useful hierarchy if somebody actually populates it. The failure mode we keep seeing is not missing capability, it is unassigned ownership: engineering assumes finance is watching the meter, finance assumes the seat price is the whole cost because that is how developer tooling has always worked, and the first real signal is an overage on a September invoice that nobody forecast. The five weeks before the promotional credits lapse are the cheapest window you will get to find out which of those assumptions is true in your organisation.
The practical exercise is small and worth doing this week. Pull your June and July credit consumption per user and per cost centre, then compare it against the base allowance rather than the promotional one, because the promotional figure is the number currently making everything look comfortable. Any user or team already consuming more than 1,900 credits on Business or 3,900 on Enterprise is, in September terms, an overage that has not happened yet. Sort those users and you will usually find the distribution is nothing like flat: a handful of heavy agent users driving a disproportionate share of token spend, often for legitimate reasons, and a long tail barely touching the allowance. That shape tells you whether the answer is a plan change for a few people, a pooled budget with a cap, or a conversation about which model gets used for which class of work.
This is the same pressure that produced Cursor two-pool seat pricing in early July, and it is not a coincidence. The industry is converging on metered billing because agentic coding consumes tokens in a way that per-seat pricing was never designed to absorb, and every vendor is discovering the same thing at once. Copilot, Cursor, Claude Code and Devin now all expose the real inference cost in some form, which means a decision that used to be procurement, choose a tool and buy seats, has become an ongoing operational one about routing work to the right model at the right price. Teams running more than one of these tools should be comparing effective cost per unit of delivered work, not headline seat price, because the seat price is increasingly the smallest line on the invoice.
There is a compliance dimension here that is easy to miss when the topic is framed as FinOps. A budget control that determines whether an AI agent can keep operating is an availability control, and the credit ledger it draws on is an unusually complete usage record. Auditors working to SOC 2 in 2026 are already asking how AI tooling is authorised, monitored and bounded, and an ISO 42001 management system expects an inventory of AI systems in use with an owner and defined limits for each. Copilot enterprise, cost centre and user budgets map cleanly onto that expectation, and the per-user consumption data is the kind of evidence that platforms like Vanta and Drata can hold against a control rather than leaving it as a finance spreadsheet nobody outside finance ever reads.
The honest summary is that nothing is being taken away on 1 September beyond a promotion that was always described as temporary, and token metering is a fairer model than a flat fee that quietly cross-subsidised the heaviest users. The risk is not the pricing, it is the gap between when the cushion disappears and when anyone notices. Three actions close it: pull the current consumption data and re-baseline it against base allowances, set explicit budgets at enterprise and cost centre level with a deliberate decision on whether overage is permitted, and name one person who owns the number month to month. Do that before the August cycle closes and September is an invoice you predicted. Leave it and it becomes an escalation, in the same week the EU AI Act enforcement powers land and your security team has other things to do.
Editorial note: AES Tech reviews are independent. Some outbound links are affiliate links and are marked sponsored; they never change our rankings. See our disclosure.
Get the next post by email
One short email when something worth knowing ships. No spam, unsubscribe anytime.