hyperproxy
Start free

Documentation / Backend observability

Track LLM costs without changing your provider calls

Choose this path when your backend already calls OpenAI, Claude, Gemini or another provider. Send a separate metrics event after each completed call. No provider credential is stored in HyperProxy for this integration. Create an ingest key on the server and copy actual usage totals from the provider response.

Updated · HyperProxy team

Optional gateway

Keep your backend. Connect analytics.

Your backend calls AI providers directly. HyperProxy receives completed request metrics separately; provider keys, prompts and response bodies are not part of this API.

  1. Create a project and open Overview → Set up analytics.
  2. Create a key with Send external events only. Keep it on your server.
  3. Send one event per completed provider call. Use a background worker or queue so telemetry failures do not delay or fail the AI response.
external-event.py
import os, uuid, json
from urllib.request import Request, urlopen
from datetime import datetime, timezone

# Submit from your background worker after a completed call.
event = {
    "event_id": str(uuid.uuid4()),  # persist and reuse this ID for retries
    "occurred_at": datetime.now(timezone.utc).isoformat(),
    "provider": "openai",
    "model": "gpt-4o-mini",
    "status_code": 200,
    "duration_ms": 320,
    "tokens_in": 1000,  # use the actual provider usage totals
    "tokens_out": 500,
    "client_id": "customer-123",
    "session_id": "conversation-456"
}
req = Request(
    "https://app.hyperproxyai.com/api/v1/admin/projects/"
    + os.environ["HYPERPROXY_PROJECT_ID"] + "/events",
    data=json.dumps(event).encode(),
    headers={
        "Authorization": "Bearer " + os.environ["HYPERPROXY_INGEST_TOKEN"],
        "Content-Type": "application/json",
        "User-Agent": "HyperProxy-Telemetry/1.0"
    }, method="POST"
)
with urlopen(req, timeout=3) as response:
    receipt = json.load(response)

Accepted events return 201 with a request ID; repeated event IDs return 200 without adding usage or consuming another request. A retry must keep the same event ID and payload. Use the returned request ID for annotations with a separate write-scoped token.

Omit both token totals when unavailable. Cost is estimated from the model catalog or your price overrides; missing usage or pricing stays unknown. External events are labelled as server-reported and can be filtered separately in Overview.

Accepted events consume the shared monthly request allowance when received, including failed provider calls. Rate limit: 600 submissions per minute per project. Events can be up to 30 days old or 5 minutes in the future. Use backoff for transient rate limits and 5xx; a monthly quota refusal requires available allowance. Keep pending events in your own durable queue if delivery must survive a process restart.

Gateway features — key protection, request limits, fallbacks, App Attest and runtime prompt injection — apply only to requests routed through HyperProxy. Analytics-only mode cannot enforce them on direct provider calls. Send either gateway telemetry or an external event for a call, not both.

Monitor → Usage

Read recorded usage and cost

  1. 1
    Select the traffic

    Choose the period and filters for provider, model or client. Gateway requests and external events reflect different connections; narrow the data source when comparing them.

  2. 2
    Inspect requests, tokens and coverage

    Review totals and model/provider breakdowns. Cost is estimated from recorded usage and catalog rates; incomplete pricing remains marked. Provider spend is separate from your HyperProxy subscription.

  3. 3
    Check a contributing request

    Open Requests with matching filters and inspect token categories and recorded cost. Chart filters do not change the account's monthly allowance, which is shared across its projects.

Monitor → Model prices

Check the estimate's rates

  1. 1
    Find the exact model

    Search the provider/model catalog for the ID used by your request. A similar display name may have a different price.

  2. 2
    Read units and categories

    Compare input, output, cache and media rates where available. An absent rate does not mean free usage. A listed price does not grant your provider account access to that model.

  3. 3
    Inspect coverage in Requests

    Check whether the provider returned the usage details needed for that rate. Estimates can exclude tools, storage or other fees; compare your provider invoice when reconciling actual billing.

Verify it works

The event endpoint returns 201 with a receipt for a new event. Retry the same event ID and payload: it returns 200 without another quota charge. In Overview, filter Source to external events and confirm provider, model, status, latency and cost coverage.

Troubleshooting

401 or 403

Use a project-scoped key with ingest permission on your backend. An app key or a read-only observability token cannot submit events.

422 validation error

Send a timezone-aware occurred_at, a stable event_id and actual numeric usage. Omit both token totals if unavailable; do not replace missing usage with fabricated zeros.

Cost is unknown or duplicated

Use an exact provider/model ID and review price overrides. Send each call through one reporting path: gateway metrics plus an external event would count twice.

Full error-code reference

Next steps

1,000 requests per month shared across projects. No card required. Provider charges are separate.