LLM usage is the primary metered consumption type. Input, output, cache-read, and cache-write tokens are converted to a USD cost and debited from the wallet.
Chat
For each message:
- Reserve - Holds estimated cost for the selected model (not a fixed default model).
- Generate - The selected provider runs the model.
- Finalize - Charges actual tokens using the same pricing rules.
Metering ingest
Developer API metering events with event_type: llm.request are rated using
their model and token counts. Events that do not represent billable usage do
not debit the wallet.
Overage
When grants are exhausted and overage is configured, remainder may be recorded as overage and invoiced through Stripe.