Track LLM costs without changing your provider calls
Choose this path when your backend already calls OpenAI, Claude, Gemini or another provider. Send a separate metrics event after each completed call. No provider credential is stored in HyperProxy for this integration. Create an ingest key on the server and copy actual usage totals from the provider response.
Updated · HyperProxy team
Keep your backend. Connect analytics.
Your backend calls AI providers directly. HyperProxy receives completed request metrics separately; provider keys, prompts and response bodies are not part of this API.
- Create a project and open Overview → Set up analytics.
- Create a key with Send external events only. Keep it on your server.
- Send one event per completed provider call. Use a background worker or queue so telemetry failures do not delay or fail the AI response.
import os, uuid, json
from urllib.request import Request, urlopen
from datetime import datetime, timezone
# Submit from your background worker after a completed call.
event = {
"event_id": str(uuid.uuid4()), # persist and reuse this ID for retries
"occurred_at": datetime.now(timezone.utc).isoformat(),
"provider": "openai",
"model": "gpt-4o-mini",
"status_code": 200,
"duration_ms": 320,
"tokens_in": 1000, # use the actual provider usage totals
"tokens_out": 500,
"client_id": "customer-123",
"session_id": "conversation-456"
}
req = Request(
"https://app.hyperproxyai.com/api/v1/admin/projects/"
+ os.environ["HYPERPROXY_PROJECT_ID"] + "/events",
data=json.dumps(event).encode(),
headers={
"Authorization": "Bearer " + os.environ["HYPERPROXY_INGEST_TOKEN"],
"Content-Type": "application/json",
"User-Agent": "HyperProxy-Telemetry/1.0"
}, method="POST"
)
with urlopen(req, timeout=3) as response:
receipt = json.load(response)
Accepted events return 201 with a request ID; repeated event IDs return 200 without adding usage or consuming another request. A retry must keep the same event ID and payload. Use the returned request ID for annotations with a separate write-scoped token.
Omit both token totals when unavailable. Cost is estimated from the model catalog or your price overrides; missing usage or pricing stays unknown. External events are labelled as server-reported and can be filtered separately in Overview.
Accepted events consume the shared monthly request allowance when received, including failed provider calls. Rate limit: 600 submissions per minute per project. Events can be up to 30 days old or 5 minutes in the future. Use backoff for transient rate limits and 5xx; a monthly quota refusal requires available allowance. Keep pending events in your own durable queue if delivery must survive a process restart.
Gateway features — key protection, request limits, fallbacks, App Attest and runtime prompt injection — apply only to requests routed through HyperProxy. Analytics-only mode cannot enforce them on direct provider calls. Send either gateway telemetry or an external event for a call, not both.
Read recorded usage and cost
- 1Select the traffic
Choose the period and filters for provider, model or client. Gateway requests and external events reflect different connections; narrow the data source when comparing them.
- 2Inspect requests, tokens and coverage
Review totals and model/provider breakdowns. Cost is estimated from recorded usage and catalog rates; incomplete pricing remains marked. Provider spend is separate from your HyperProxy subscription.
- 3Check a contributing request
Open Requests with matching filters and inspect token categories and recorded cost. Chart filters do not change the account's monthly allowance, which is shared across its projects.
Check the estimate's rates
- 1Find the exact model
Search the provider/model catalog for the ID used by your request. A similar display name may have a different price.
- 2Read units and categories
Compare input, output, cache and media rates where available. An absent rate does not mean free usage. A listed price does not grant your provider account access to that model.
- 3Inspect coverage in Requests
Check whether the provider returned the usage details needed for that rate. Estimates can exclude tools, storage or other fees; compare your provider invoice when reconciling actual billing.
Verify it works
The event endpoint returns 201 with a receipt for a new event. Retry the same event ID and payload: it returns 200 without another quota charge. In Overview, filter Source to external events and confirm provider, model, status, latency and cost coverage.
Troubleshooting
401 or 403
Use a project-scoped key with ingest permission on your backend. An app key or a read-only observability token cannot submit events.
422 validation error
Send a timezone-aware occurred_at, a stable event_id and actual numeric usage. Omit both token totals if unavailable; do not replace missing usage with fabricated zeros.
Cost is unknown or duplicated
Use an exact provider/model ID and review price overrides. Send each call through one reporting path: gateway metrics plus an external event would count twice.
Next steps
1,000 requests per month shared across projects. No card required. Provider charges are separate.