Cost metering + quotas
- Written for
- + Written for
- Deprecated
- + Deprecated
- Applies to
- + Applies to
Cost metering + quotas
How Chatly meters every LLM token, which meter the plan ceiling actually bounds, and where to read your real spend.
Every LLM call in Chatly — copilot suggestions, autonomous agent turns, classifiers, RAG embeddings — is priced against a per-model rate table and written to a spend meter. You can see every cent, and the workspace gets emailed before and when a ceiling is reached.
Warning — Read this before you plan around the plan cap
There are two meters, and the plan ceiling only bounds one of them.
ai:usage:platform:<YYYYMM>is spend on our API key — subsidised spend. That is what the ceiling compares against and the only thing that can refuse a call.
ai:usage:byok:<YYYYMM>is spend on your key. It is recorded, reported and shown to you, and never enforced — refusing a customer for money they paid OpenAI directly would be refusing them for their own spending.Every AI feature described in these docs is BYOK. So in practice the plan ceiling does not stop copilot, autonomous or classifier calls. If you need a hard money bound, set it on your OpenAI or Anthropic account, where the tokens are actually billed.
What gets measured
Each text completion emits a structured log line under the llm_usage key:
{
"workspaceId": "0192...",
"model": "gpt-4o-mini",
"inputTokens": 412,
"outputTokens": 86,
"estimatedUsd": 0.000114,
"latencyMs": 410,
"ts": 1716660000000
}Speech-to-text is duration-priced, not token-priced, so it gets its own record shape under audio_usage — unit: "second", durationSec, provider, and a measured flag recording whether the vendor reported the duration or we trusted the caller's declaration. Inventing a token count so the two could share a type would put a made-up number into a billing log.
Where the numbers live
Parameters
Name | Type | Description |
|---|---|---|
|
| A pino line per call ( |
|
|
|
|
|
|
|
|
|
All three carry a 60-day TTL, so a cross-month read on the 1st still resolves. The period key is YYYYMM in UTC — a calendar month, not a rolling 30-day window, and there is no setting to change that.
Per-plan ceilings
Monthly ceiling on subsidised (platform-key) spend
Name | Type | Description |
|---|---|---|
|
| No enforced ceiling. Free runs on BYOK with zero subsidised spend by construction — there is nothing to cap. |
|
| Warns at 80%. |
|
| Warns at 80%. |
|
| Warns at 85%. |
|
| Contract-sized rather than enforced here. Warns at 90%. |
|
| Not a plan you can buy. It is substituted for a subscription whose status is not |
Info — Why a lapsed subscription is capped rather than dropped to Free
Because
freeis uncapped. "Fall back to Free" reads as the humane answer right up until you notice that cancelling a Business subscription would then upgrade its AI spend ceiling from $250/month to unlimited.lapsedis the lowest paid ceiling, so lapsing can only ever reduce headroom.
past_dueis deliberately still entitled. Stripe dunning runs for weeks and the subscription is live throughout; withdrawing AI on the first failed charge punishes a card that will most likely retry fine.
There is no dashboard control to lower a workspace's own cap below the plan default, no soft-budget threshold, no per-conversation dollar cap, no model allow-list and no per-use model override map. The plan is the ceiling.
How blocking works
Before an LLM call, the caller estimates a worst-case cost — the requested maxTokens priced at the model's output rate, plus a prompt-size estimate at 4 bytes per token — and compares it against the platform meter:
spent = HGET ai:usage:platform:{YYYYMM} {workspace_id}
if spent + estimated > plan_ceiling:
throw QuotaExceededError # HTTP 402The estimate happens after model resolution, not before, so auto never gets priced through the unknown-model fallback and block calls at ~100× their real cost.
The error surfaces as HTTP 402 with code ai_quota_exceeded:
{
"code": "ai_quota_exceeded",
"httpStatus": 402,
"message": "AI budget exceeded (51.12 / 50.00 USD)",
"details": {
"workspaceId": "0192...",
"spent": 51.12,
"budget": 50.00
}
}When the counter cannot be read
This is a distinct outcome from "exceeded", and the two services choose differently on purpose:
Worker (workflow AI actions, autonomous agent) fails open. A Redis blip degrades the cap rather than taking AI down for every customer.
AI service raises
ai_quota_unavailable(HTTP 503) after a 1,500 ms command timeout. Callers surface it as a temporary AI outage and hand off to a human, rather than telling a customer they are out of budget when they may not be.
The short timeout matters: an ioredis client configured with maxRetriesPerRequest: null queues commands while disconnected instead of rejecting, so without the race every LLM call would hang forever and the visitor waiting on a reply would get silence.
Degradation
Copilot: every endpoint returns its empty result — no suggestion, no variants, no articles. The composer never blocks.
Autonomous agent: the run escalates. The visitor gets a reply and a person, the conversation moves to
pending(which is the state bot-routing stands down on, so the next inbound does not start another run), and the workspace owners get an email.Classifiers / summarization: best-effort throughout. The classification is skipped and message ingestion continues.
RAG embeddings: indexing stops. Existing articles stay findable via Postgres full-text search.
Per-model rates
The rate table is PRICING_USD_PER_1K in apps/ai/src/cost/cost-metering.middleware.ts, with a parallel table for the worker's direct-BYOK calls in apps/worker/src/lib/llm-cost.ts. Lookup is longest-prefix, so a dated id (claude-3-haiku-20240307) resolves to its family. There are no pricing-override environment variables — changing a rate means changing the table.
List prices (per 1M tokens)
Name | Type | Description |
|---|---|---|
|
| $0.15 in, $0.60 out |
|
| $0.40 in, $1.60 out |
|
| $2.00 in, $8.00 out |
|
| $0.25 in, $2.00 out |
|
| $1.25 in, $10.00 out |
|
| $1.00 in, $5.00 out |
|
| $3.00 in, $15.00 out |
|
| $15.00 in, $75.00 out |
|
| $0.80 in, $4.00 out |
|
| $0.25 in, $1.25 out |
|
| $0.02 in. Embeddings bill on input only. |
|
| $15 per 1M characters. Character-priced, so it is deliberately never run through the token table. |
|
| $0.006/min. |
Warning — gpt-4o is metered at two different rates
The two tables disagree on
gpt-4o: the AI service prices it at $5.00 in / $15.00 out per 1M, the worker at $2.50 / $10.00. Which number lands on your meter depends on which path made the call. Pin a model from the list above rather thangpt-4oif you want your meter to match your provider invoice.
Info — An unpriced model is charged high, not zero
A model missing from the table is priced at the most expensive row — $15 in / $75 out per 1M — and logged with a one-time warning. Zero would be indistinguishable from "no spend" and would silently disarm the ceiling while real tokens were bought. Erring high costs a customer some headroom; erring low costs them money.
estimatedUsd is computed from the token counts the provider returns in its own response, so for OpenAI and Anthropic it is exact arithmetic over exact counts — but it is still our arithmetic over our snapshot of list prices. The authoritative bill is your provider's invoice.
Reading your spend
GET /v1/ai/quota — Bearer token
Requires workspace:billing:read. This is what Billing → Usage this month renders. It returns all three meters plus the ceiling, so nothing has to guess which number it is showing:
Parameters
Name | Type | Description |
|---|---|---|
|
| Month-to-date subsidised spend — the number |
|
| Spend on your own provider key. Recorded, never capped. |
|
| Every dollar of AI this workspace caused, from the combined legacy counter. |
|
| Always |
|
|
|
|
| True once |
|
| e.g. |
|
| The plan key that drove the budget — |
|
| Linear projection at the current burn rate. |
GET /v1/reports/analytics/ai-deflection — Bearer token
The deflection funnel with token and cost spend, broken down per bot, with a trend. Requires reports:read; takes ?days= up to 365.
Per-run cost is also on every autonomous run — see GET /v1/autonomous-agents/runs/{id}, which reports input tokens, output tokens and cost for that run, and is what the bot Logs tab renders.
Warning — There is no AI-usage CSV export
The CSV export endpoint (
GET /v1/reports/export/{report}, gated onreports:export) serves a closed list:teams,labels,sla-breaches,agents,conversations,channels,response-times. AI usage is not among them, and there is no nightly Postgres rollup table behind it either — the spend counters are Redis-only. For a per-call receipt, ship thellm_usagelog lines to your own store.
Alerts
Two email notices go to the workspace's owner/admin billing recipients:
Approaching the ceiling — at the plan's warn threshold (80%, or 85% on Business, 90% on Enterprise).
Ceiling reached — sent from the refusal path when
QuotaExceededErroris caught.
Each notice is claimed in Redis before it is sent, so a workspace gets one per kind per period rather than one per refused call. The warn notice is suppressed at or past the ceiling, because the exhausted notice is the right one there and warning too would double-mail. An unreadable counter sends nothing rather than guessing.
There is no Slack or Discord destination for budget alerts, and no "Settings → Notifications → AI" screen.
The kill switch
AI_ENABLED is a platform-wide emergency stop, checked before creds are decrypted and before any quota logic runs.
Unset or empty means on — it is a stop switch bolted onto a running platform, not something you have to opt into on every deployment, and a self-hoster who has never heard of it must still get a working product. Once the variable is present, though, only a recognised affirmative (1, true, yes, on, enabled) keeps traffic flowing. Anything unrecognised stops it — so a typo like AI_ENABLED=flase halts calls rather than leaving every provider call running while on-call believes the platform is quiet. The AI service's /health reports whether the value was recognised, so a mistyped stop is distinguishable from an intended one.
When it is thrown, KB retrieval degrades to keyword search rather than failing the run.
Troubleshooting
Warning — My spend bar reads $0 but I know AI is running
Look at
byokSpent, notspent. The progress bar is a ratio of the platform meter against the plan ceiling; rendering a BYOK-heavy workspace as "$412 of $5" would be a lie about what we are billing you for. Your own spending is on the BYOK meter and intotalSpent.
Warning — Spend is higher than expected
Check Bots → (a bot) → Logs for runs with an abnormal hop count and a high cost — a tool that returns a different answer for the same arguments makes the agent chase its tail up to
maxToolCalls. LowermaxToolCallson the bot, or fix the tool's idempotency.
Info — My budget did not reset on the 1st
It does — the period key is UTC
YYYYMMand it rolls at midnight UTC on the 1st. If your local timezone is behind UTC, the reset happens on what is still the last day of the month for you.
Info — An AI call returned 503, not 402
That is
ai_quota_unavailable— the spend counter could not be read within 1,500 ms. You have not exceeded anything; we could not tell. It is a Redis problem, not a billing one.