Skip to content
Public preview · the console is open

Credits, billing & BYOK

read as .md

Hosted GGUI is one product, paid as you go: every render is charged to a prepaid credit balance at a flat rate per render kind, the same for every account, and a one-time welcome credit on sign-up is the free entry — no card, no plan to pick. The rates, the welcome credit and the top-up presets live in one place, ggui.ai/pricing; this page is the mechanics.

Everything on this page lives at /credits, /keys/providers, and each app’s Billing tab in the GGUI console.

There are no plans to choose between and nothing is metered monthly. Every account renders on the same terms: a render costs its flat per-kind rate, debited from the app owner’s prepaid balance at render time, and a balance at or below zero refuses the next render — on your own provider key too, since the platform lane is still charged — until you top up. There are no subscriptions and no usage bands — every account is on the same terms, and nothing is withheld from a small one. Enterprise is a conversation, not a price: negotiated rates, invoicing and terms through enterprise@ggui.ai.

The free entry. Every new account receives a one-time welcome credit on first sign-in, with no card, and renders on it until it is spent — the same rates, the same product. It does not expire: there is no separate trial state, no render count to run out of and no day clock, and when the welcome credit is gone a top-up continues exactly where you were.

Each render is priced by its kind — a cold generation (the LLM synthesized a component), a cache hit (the render matched an existing blueprint and skipped generation; cheaper, but not free — see the generation pipeline), or a BYOK render (your own provider key paid for the model call, so GGUI charges only its platform lane). The three rates are on ggui.ai/pricing and are the same for every account; premium models are priced by their own rate-card row, which that page does not list (see Premium models), and a pool-funded render that asks for a model with no row is refused with model_not_allowed rather than silently re-priced.

Charges accrue per render and are debited to the app owner’s balance in whole cents, carrying sub-cent remainders forward — a render is never rounded up on its own. An app may also carry an optional hard monthly spending cap, set on request: write to hello@ggui.ai with the app and the monthly amount, and we set, change or remove it for you. Once the app’s render spend for the month reaches the cap, generation is refused (hard_cap_exceeded) until the cap is raised or the month resets.

Pool-funded renders draw on GGUI’s own platform pool, which covers Anthropic (direct or routed through Bedrock), OpenAI, Google, and OpenRouter. A render routed to any provider outside that set is refused with a 402 pointing you at BYOK.

On a pool-funded render the model is resolved from the app’s generation setting, then the request’s infra.model, then the pool default. Any model other than the pool default must have its own rate-card row for your account — one row per model, carrying the three lanes above (cold generation, cache hit, BYOK platform lane). No premium model is enabled by default: every account starts with the pool model only, and a row is seeded for your account on request. A model with no row is refused before anything is read or spent, with model_not_allowed and a fix in the message: send the pool default (or omit infra.model), bring your own key for that model, or ask hello@ggui.ai to enable it. The gate does not apply to BYOK renders — your key, your model — or to guuey-managed apps, which are metered to the platform under their own policy.

The rows seeded for your account are listed in the console: the Premium models card on the app’s Billing tab (console.ggui.ai/apps/<appId>/billing) shows one row per enabled model route, with the cold-generation, cache-hit and BYOK-lane rate per render — the same rows the gate prices from. A render on a listed model bills its row’s rate in place of the standard rate above. When no premium model is enabled the card says so; the pool model is always available, and to run another model on the app you bring your own key for it or write to hello@ggui.ai. The Settings model picker is a different thing: it lists the model registry, not the rate-card rows, so a model picked there is still subject to the gate above.

An app can ask for more effort on its fresh generations: more tries, more evaluation rounds, and a higher quality bar for the generator’s own check. The add-on is charged on top of the fresh-generation rate, at the level the generation applied. An app that sets no level pays no add-on, and neither do generations on your own model key or cache hits.

Level Add-on What it gets
low effort no add-on Up to 5 tries and 1 evaluation round, quality bar 65/100
medium effort no add-on Up to 8 tries and 2 evaluation rounds, quality bar 70/100
high effort +5¢ Up to 8 tries and 3 evaluation rounds, quality bar 75/100
xhigh effort +25¢ Up to 10 tries and 3 evaluation rounds, quality bar 80/100
ultra effort +75¢ Up to 12 tries and 4 evaluation rounds, quality bar 85/100

New accounts receive a one-time welcome credit on first sign-in; it appears as a Free credit row in the transaction log, the amount shows on /credits (and on ggui.ai/pricing), and it does not expire. It is the free entry: renders draw on it at the same rates as on purchased credit, and when it is spent a top-up continues from the same balance.

Before any render runs — pool-funded or on your own provider key, since a BYOK render still charges the platform lane — GGUI checks the app owner’s balance: the owner’s wallet, whichever key made the call. Above zero proceeds; at or below zero is refused with a 402 (insufficient_credit) and a fix hint that points at /credits, rather than a silent failure.

The threshold is deliberately “above zero”, not “enough to cover this render”. A balance of a few cents still admits the next render, which can drive the balance negative and records the true debt; the next check sees a non-positive balance and refuses. You never get a free loop, and you never get a render cancelled halfway for being a cent short.

Every movement is one row: what it was, when, the signed amount, and the balance afterwards. Four kinds appear:

Kind What it is
Free credit The welcome credit on first sign-in
Render The charge for one render, tagged with the session it paid for
Top-up Credit purchased through Stripe
Refund Credit returned to the wallet

Click Top up on /credits, pick one of the preset amounts ($10, $50, $100) or type a custom one, and check out through Stripe.

Credit lands when Stripe confirms the payment, which is usually a few seconds behind the redirect back to the console. The page keeps refreshing the balance for about half a minute after you return, and offers a manual refresh if the confirmation takes longer. A cancelled checkout charges nothing.

A coupon code (cpn_…) is redeemed from the same page. If you belong to any orgs, the redeem form lets you choose where the credit lands — your personal wallet, or any org wallet you are a member of. See Orgs and teams for how shared wallets work.

Store a provider API key and GGUI will call that provider on your account rather than out of its own pool. Four providers are supported, each with the console dialog’s get-a-key path:

  • Anthropic — Claude models. console.anthropic.com → Settings → API Keys.
  • OpenAI — GPT models. platform.openai.com → API Keys.
  • Google — Gemini models. aistudio.google.com → Get API key.
  • OpenRouter — multi-provider routing. openrouter.ai → Settings → Keys.

Submit the plaintext once. It is validated against the provider, encrypted with KMS under an encryption context bound to your account and the provider, and persisted as ciphertext — it is decrypted in memory only, on the request that needs it. Afterwards the console shows the provider, your label, the last four characters, and when the key was last used. The plaintext is never readable again, by you or by GGUI’s own console.

Remove a key at any time from the same card.

/keys/providers holds your account-wide keys. An individual app can override them on its own Keys tab, and the app-level key wins for renders bound to that app — a published app can run on its own provider quota instead of drawing on yours.

The two scopes are encrypted under different contexts, so an app’s stored key cannot be decrypted through the account-scoped path, or the reverse.

When a render actually runs on your provider key, your provider account pays for the model call and GGUI does not charge for the generation — charging you again would be double billing. A BYOK render carries only GGUI’s platform lane rather than the cold-generation or cache-hit rate (rates on ggui.ai/pricing) — and that lane is charged to the same prepaid balance, so the pre-render balance check applies to BYOK renders too: an owner at zero is refused insufficient_credit even on their own key. Bringing a key changes who pays for the model call, not whether the wallet is read.

Resolution is per provider, so mixing is fine: for an app running on its own key, a render routed to a provider you have a key for uses that key, and anything else falls back to the credit pool.