hyperproxy
Start free

Documentation / OpenTelemetry · OTLP/HTTP

Export LLM metrics with OpenTelemetry

Use this integration when your backend already emits completed GenAI CLIENT spans. HyperProxy extracts supported LLM metrics from OTLP traces; your instrumented application and exporter still own span creation and delivery. This is an LLM metrics adapter, so use your existing trace backend for general application traces.

Updated · HyperProxy team

Connect OpenTelemetry

Send completed LLM metrics from your backend with an OTLP/HTTP trace exporter. Create a project key with ingest permission in Server API access, then configure:

otel.env
OTEL_EXPORTER_OTLP_TRACES_ENDPOINT=https://app.hyperproxyai.com/api/v1/admin/projects/PROJECT_ID/v1/traces
OTEL_EXPORTER_OTLP_TRACES_PROTOCOL=http/protobuf
OTEL_EXPORTER_OTLP_TRACES_HEADERS="Authorization=Bearer hp_obs_YOUR_SERVER_TOKEN,User-Agent=YourAppTelemetry/1.0"
OTEL_BSP_MAX_EXPORT_BATCH_SIZE=512

Supports HTTP protobuf and JSON, including gzip. Use completed GenAI CLIENT spans with provider, model, input and output token attributes. Trace IDs group sessions; retries with the same span IDs do not duplicate metrics or quota. Limits: 512 spans and 2 MiB per batch, 600 submissions per project per minute.

This adapter collects LLM metrics. It does not store prompt bodies, arbitrary span attributes or general application traces. Unsupported spans receive OTLP partial-success counts. Export each AI call through one path: a gateway request exported again as an external span would count twice.

Keep the ingest key on your backend. Do not embed it in a mobile or browser app. Explore usage and cost reporting

Request details separate ordinary tokens, cache reads, 5-minute and 1-hour cache writes, and reported audio/image/video tokens. Missing prices or ambiguous cache details are marked as incomplete cost. Standard token estimates exclude storage and tool fees. Custom rates can be supplied through the project price override API.

OpenAI background Responses stay pending until you retrieve final usage through the same gateway service. Repeated status checks settle the original generation once, using its original price and billing month. Polls still consume request quota. HyperProxy does not store a complete split key or poll providers autonomously.

Observability

Requests, tokens, dollars

Meter request count, input and output tokens, latency, success rate, and estimated model cost by project and model.

Requests18,420this month
Tokens4.82Minput + output
Cost$36.41catalog priced

Verify it works

Complete a real instrumented AI call and flush the exporter. Inspect the OTLP response, including partialSuccess.rejectedSpans, then confirm the external request in HyperProxy. HTTP 200 alone does not prove that all spans were accepted. Retry with stable span IDs to avoid duplicate usage.

Troubleshooting

Exporter cannot authenticate

Use the project ingest token in Authorization and the project-specific /v1/traces URL. Choose OTLP/HTTP rather than the gRPC exporter.

200 but no LLM metrics

Inspect partial-success counts. The span must be completed, use CLIENT kind and contain supported GenAI provider, model and input/output token attributes. Arbitrary spans are not imported.

413 or 429

Keep batches within 512 spans and 2 MiB, including gzip’s decoded size, and respect the project submission limit. Queue and retry transient failures with backoff.

Full error-code reference

Next steps

1,000 requests per month shared across projects. No card required. Provider charges are separate.