OpenTelemetry
OpenTelemetry Sink (OTLP over HTTP)
The OpenTelemetry sink sends OTLP-formatted telemetry (logs, metrics, traces) to an HTTP endpoint. This guide explains what each parameter does and the operational impact.
Requirements (Very Important)
This sink expects events already conforming to the OpenTelemetry protobuf model (OTLP proto).
Operational meaning:
- If your upstream telemetry is not already OTLP-shaped, the sink will not “magically” convert it.
- You typically need to normalize/shape events upstream so they match OTLP expectations (resources, scopes, attributes, timestamps, etc.).
Buffering and Backpressure
buffer (optional)
Controls how telemetry is queued when the remote OTLP endpoint is slow or unreachable.
buffer.type (optional, default: memory)
- memory: faster; data in the buffer is lost on crash/restart
- disk: more durable across restarts; slower and requires disk capacity/IOPS planning
buffer.max_size (required)
Hard limit for buffer memory/disk usage.
Operational impact:
- Too small → frequent backpressure or drops during outages
- Too large → longer recovery time and larger burst load when the endpoint recovers
buffer.max_events (optional; relevant to memory, default: 500)
Maximum number of queued events (memory buffer only).
Operational impact:
- Adds an event-count safety cap when event sizes vary widely.
buffer.when_full (optional, default: block)
Behavior when the buffer is full:
- block: backpressure upstream (preferred when loss is unacceptable)
- drop_newest: drop incoming events to keep the pipeline moving (preferred when continuity > completeness)
Inputs
inputs (required)
Defines which upstream sources/transforms feed this sink (logs/metrics/traces).
Operational impact:
- Wildcards can unintentionally route too much telemetry (and cause high cost/traffic).
- Ensure you scope inputs carefully and apply any filtering/redaction upstream.
Protocol Configuration (HTTP Transport)
All “how it sends over HTTP” behavior lives under protocol.
protocol (required)
Root block for HTTP transport behavior.
Authentication (HTTP)
protocol.auth (optional)
Controls Authorization behavior for HTTP requests.
Important operational rule:
- Use HTTP-level authentication with TLS/HTTPS, since credentials are transmitted via headers.
protocol.auth.strategy (required when auth is used)
Supported strategies:
- aws: AWS request signing authentication (SigV4-like patterns for AWS-backed endpoints)
- basic: HTTP Basic Authentication
- bearer: Bearer token passed as-is (OAuth2/JWT)
- custom: Custom Authorization header value
Basic Authentication
- protocol.auth.user: Basic Auth username
- protocol.auth.password: Basic Auth password
Bearer Authentication
- protocol.auth.token: bearer token value (passed as-is)
Custom Authorization Header
- protocol.auth.value: direct value inserted as the Authorization header
AWS Authentication (nested)
When strategy is aws, the nested fields control how AWS credentials are sourced and how role assumption works:
- Static creds (access key / secret key)
- Temporary creds (session token)
- Credentials file + profile selection
- STS assume role fields (assume_role, external_id, session_name)
- IMDS behavior (timeouts/retries)
- Region/service name controls (region, service)
Operational impact:
- Misconfigured AWS auth typically manifests as consistent 403/401 responses and infinite retry/backlog.
Batch Behavior (HTTP Payload Grouping)
protocol.batch (optional)
Controls how many events are sent per HTTP request.
protocol.batch.max_bytes (optional, default: 10,000,000 bytes)
Maximum uncompressed size of a batch before flushing.
Operational impact:
- Larger batches improve throughput efficiency but increase memory usage and potential for large retry payloads.
- Oversized payloads may be rejected by gateways or load balancers.
protocol.batch.max_events (optional)
Maximum number of events per batch before flushing.
Operational impact:
- Smaller values reduce per-request payload and latency but increase request rate.
protocol.batch.timeout_secs (optional, default: 1 second)
Maximum age of a batch before it is sent.
Operational impact:
- Lower values reduce latency
- Higher values improve batching efficiency at lower traffic rates
Compression
protocol.compression (optional, default: none)
Controls HTTP payload compression:
- gzip, snappy, zlib, zstd, or none
Operational impact:
- Compression reduces bandwidth but increases CPU usage.
- Useful over constrained links or high-volume telemetry pipelines.
Encoding (What Bytes Are Sent)
protocol.encoding (required)
Controls how events are serialized before sending.
Critical note:
- Even though this is an OpenTelemetry sink, encoding settings still determine how the payload is structured at the byte level.
- For OTLP interoperability, the codec selection must align with what the receiving endpoint expects.
protocol.encoding.codec (required)
Selects the encoding format (for example: otlp, protobuf, json, etc.).
Operational impact:
- Wrong codec = receiver cannot decode payload (typically 400/415 responses and continuous retries).
- Choose a codec that matches the endpoint’s expected OTLP ingestion format.
Common encoding controls (apply depending on codec)
- only_fields / except_fields: reduce payload or exclude sensitive content
- timestamp_format: normalize timestamps (important for traces/metrics correctness)
- Codec-specific sections (Avro/CEF/CSV/GELF/Protobuf/etc.) define schemas or mapping rules and should only be used when required by downstream consumers.
Framing (How Multiple Events Are Delimited)
protocol.framing (optional)
Controls how events are separated when serialized.
protocol.framing.method (required when framing is configured)
Methods include:
- bytes: no delimiter (single blob)
- newline_delimited: newline between events
- character_delimited: custom ASCII delimiter
- length_delimited / varint_length_delimited: prefix each frame with length (often used for protobuf-style streaming)
Operational impact:
- Framing must match what the receiver expects; otherwise it can’t parse event boundaries.
- Length-delimited framing is common when transporting protobuf messages that need explicit boundaries.
protocol.framing.max_frame_length (optional)
Limits the maximum frame size.
Operational impact:
- Protects against runaway payload sizes.
- Too low can cause legitimate payloads to be rejected or truncated upstream.
Headers
protocol.headers (optional)
Adds static custom headers to each request.
Operational impact:
- Useful for routing, tenant identification (if your gateway uses headers), or custom auth schemes.
- Be careful: headers can expose secrets if logged by proxies.
protocol.request.headers (optional)
Adds headers to every request, with support for templating (dynamic values per event context).
Operational impact:
- Powerful for multi-tenant routing and tagging.
- Risk: dynamic headers can increase variability and make troubleshooting harder if overused.
HTTP Method
protocol.method (optional, default: POST)
Controls which HTTP method is used.
Operational impact:
- OTLP over HTTP is typically POST; changing this usually breaks compatibility unless a custom receiver expects otherwise.
URI (Destination)
protocol.uri (required)
Full URI for the HTTP endpoint (scheme, host, optional port/path).
Operational impact:
- Wrong URI is the most common cause of ingestion failure.
- If you template the URI, ensure the template output is strictly controlled; otherwise you can create unpredictable routing and security issues.
Payload Prefix/Suffix (Special Case)
protocol.payload_prefix / protocol.payload_suffix (optional)
Adds a wrapper around the payload, only relevant for certain JSON framing/encoding combinations.
Operational impact:
- Used when the receiver requires a JSON envelope structure.
- Misuse can produce invalid JSON and result in ingestion failures.
Request Behavior (Concurrency, Rate Limit, Retries, Timeouts)
protocol.request (optional)
Controls client-side HTTP behavior.
protocol.request.concurrency (optional, default: adaptive)
- none: fixed concurrency of 1 (predictable, lower throughput)
- adaptive: auto-adjusts concurrency based on observed RTT/performance
Operational impact:
- Adaptive can increase throughput but may oscillate if the endpoint is unstable.
protocol.request.rate_limit_num / rate_limit_duration_secs
Limits outbound request rate.
Operational impact:
- Protects your OTLP receiver from overload.
- If set too low, backlog grows and latency increases.
protocol.request.retry_attempts (optional)
Maximum retries.
Operational impact:
- Very high values can cause long retry tails during outages.
- With at-least-once delivery, retries increase duplicate probability.
protocol.request.retry_initial_backoff_secs (default: 1)protocol.request.retry_max_duration_secs (default: 30)protocol.request.retry_jitter_mode (default: Full)
Controls backoff timing and jitter.
Operational impact:
- Jitter reduces synchronized retry bursts across many agents.
protocol.request.timeout_secs (default: 60)
Request timeout.
Operational impact:
- Too low can create “orphaned requests” and duplicate delivery.
- Too high can slow recovery and consume concurrency slots.
Adaptive Concurrency Tuning (Advanced)
The following are “expert settings” and rarely need changes:
- decrease_ratio, ewma_alpha, initial_concurrency, max_concurrency_limit, rtt_deviation_scale
Operational warning:
- Incorrect tuning can lead to unstable throughput and oscillations.
TLS (HTTP Security)
protocol.tls (optional)
Controls TLS verification and client identity.
Key options:
- ca_file: trust roots
- crt_file / key_file / key_pass: client certificate (mTLS)
- server_name: SNI override
- verify_certificate / verify_hostname: certificate and hostname validation
- alpn_protocols: advanced negotiation options
Operational guidance:
- Disabling certificate or hostname verification weakens security and should be avoided outside controlled test environments.