Prometheus Remote Write
Prometheus Remote Write Sink
The Prometheus Remote Write sink delivers metric data to a Prometheus-compatible remote write endpoint (for example: Prometheus, Cortex, Thanos Receive, Mimir, or managed services that expose a remote write API).
Important Warning: Cardinality
Prometheus strongly discourages high cardinality in:
- Metric names
- Label sets (tags)
High cardinality can cause:
- Increased memory usage and slow ingestion
- Storage amplification
- Query instability and operational incidents
Operationally, you should:
- Limit label dimensions (avoid user/session/device identifiers as labels)
- Prefer aggregation before export where possible
- Apply cardinality controls upstream (worker provides transforms to enforce limits)
Core Connectivity
endpoint (required)
- The full remote write URL (scheme + host + path).
- This is the primary destination for metric batches.
- Misconfiguration typically results in continuous retries and backlog growth (buffer pressure).
Authentication
auth (optional)
Controls how requests authenticate to the remote write endpoint.
auth.strategy (required when auth is used)
Defines which authentication method is used:
- aws: For Amazon Managed Prometheus-style authentication (SigV4 / service-specific patterns)
- basic: HTTP Basic Authentication (username + password)
- bearer: Bearer token sent as-is (OAuth2/JWT/etc.)
Basic Authentication
auth.user (required with strategy = basic)
- Username used for HTTP Basic Authentication.
auth.password (required with strategy = basic)
- Password used for HTTP Basic Authentication.
Bearer Authentication
auth.token (required with strategy = bearer)
- Token placed in the Authorization header as a bearer token.
- Typically used for OAuth2 access tokens or JWT-based gateways.
AWS Authentication (for Managed Prometheus / AWS-compatible endpoints)
AWS auth supports multiple credential sourcing patterns:
auth.access_key_id / auth.secret_access_key (required in some setups)
- Static AWS credentials for signing requests.
auth.session_token (optional)
- Used for temporary credentials (short-lived sessions).
auth.credentials_file (required when using file-based credentials)
- Path to a credentials file containing AWS keys.
auth.profile (optional)
- Selects which profile to use from the credentials file.
auth.assume_role (required when assuming a role)
- IAM Role ARN to assume before sending requests.
auth.external_id (optional)
- Additional identifier used in some cross-account role assumption scenarios.
auth.session_name (optional)
- Name used for the assumed role session.
- Helps auditing and tracing activity in AWS logs.
auth.region (optional)
- Region used to send STS requests (when relevant).
- If omitted, defaults to the service region behavior.
auth.load_timeout_secs (optional)
- Time limit to obtain credentials from the chain (including role assumption).
- Important in environments where credentials may not be immediately available.
auth.imds (optional)
IMDS options apply when credentials are obtained from instance metadata.
- auth.imds.connect_timeout_seconds: Connection timeout for IMDS calls
- auth.imds.read_timeout_seconds: Read timeout for IMDS calls
- auth.imds.max_attempts: Retry attempts for token/metadata retrieval
These settings matter in cloud environments where IMDS can be slow or restricted.
AWS Service Targeting
aws (optional)
Controls the target AWS service region/endpoint behavior.
aws.region (optional)
- Region of the target AWS service.
- Can affect signing scope and endpoint selection.
aws.endpoint (optional)
- Custom AWS-compatible endpoint.
- Used for non-default environments or AWS-compatible services.
Batching Behavior
Remote write is batch-oriented; these settings directly influence throughput vs latency.
batch (optional)
Controls how worker groups metrics into remote write payloads.
batch.aggregate (optional, default: true)
- When enabled, worker aggregates metrics within a batch before sending.
- Typically reduces payload size and improves efficiency.
- If disabled, payloads may be larger and more repetitive (higher ingestion load).
batch.max_events (optional, default: 1000)
- Maximum number of metric events in a batch before flush.
- Smaller values lower latency but increase request rate.
batch.max_bytes (optional)
- Flushes when the uncompressed batch reaches this size limit.
- Useful to prevent overly large payloads (which can cause request failures or longer server processing time).
batch.timeout_secs (optional, default: 1)
- Flushes when the batch reaches a certain age, even if size thresholds are not met.
- Lower values reduce latency; higher values improve batching efficiency.
Distributions: Histograms and Summaries
Prometheus remote write commonly represents distributions as histograms or summaries. These settings define default behavior when worker has to convert distribution metrics.
buckets (optional)
- Default histogram bucket boundaries used when converting distribution metrics into histograms.
- Bucket choice affects:
- Accuracy of latency/size/duration representation
- Query usefulness (e.g., p95/p99 estimations)
- Storage cost (more buckets → more time series)
quantiles (optional)
- Quantiles used when converting distribution metrics into Prometheus summaries.
- More quantiles can increase computational overhead and time series output.
Operational tip: choose either histogram-centric or summary-centric strategy consistently to avoid confusion and unnecessary cardinality.
Buffering and Backpressure
Batch sinks can accumulate backlog quickly when the remote endpoint slows down.
buffer (optional)
Controls internal buffering while awaiting delivery.
buffer.type (optional, default: memory)
- memory: Higher performance, but data loss on crash/restart
- disk: More durable, but slower and requires sufficient disk provisioning
buffer.max_size (required)
- Hard limit for buffer memory/disk usage.
- If too small, you’ll hit backpressure or drops more often.
- If too large (especially on disk), you must ensure adequate I/O and capacity planning.
buffer.max_events (optional; relevant for memory)
- Cap on number of buffered events (memory buffer only).
buffer.when_full (optional, default: block)
- block: Applies backpressure upstream (preferred for lossless pipelines)
- drop_newest: Drops newest events when full (preferred when “keep pipeline flowing” is more important than completeness)
Compression
compression (optional, default: snappy)
Controls request payload compression:
- snappy is common for Prometheus remote write efficiency.
- Stronger compression (zstd/gzip) can reduce bandwidth but increases CPU cost.
- No compression increases bandwidth usage and may affect reliability on constrained links.
Metric Namespace Management
default_namespace (optional)
- Used only when incoming metrics do not already include a namespace.
- When present, namespace becomes a prefix to the metric name, separated by underscore.
- Helps enforce naming consistency and avoid collisions.
Stateful Cache Control
expire_metrics_secs (optional)
- Controls how long incremental metrics remain in the internal cache without updates.
- If not configured and you receive continuously “new” unique incremental series, memory usage can grow indefinitely.
- This setting is critical for long-running agents in dynamic environments (e.g., Kubernetes with changing label sets).
Healthcheck
healthcheck.enabled (optional, default: true)
- Enables startup validation that the remote endpoint is reachable/healthy.
- Helps fail fast during deployment rather than silently buffering and retrying indefinitely.
Inputs
inputs (required)
- Identifies which upstream metric sources/transforms feed this sink.
- Supports wildcards for large topologies.
Proxy Support
proxy (optional)
Routes outbound HTTP(S) requests via proxy infrastructure.
proxy.enabled (optional, default: true)
- Enables proxy usage when proxy settings exist.
proxy.http / proxy.https (optional)
- Proxy endpoints for HTTP and HTTPS traffic.
proxy.no_proxy (optional)
- Host patterns that should bypass proxy.
- Useful for keeping internal traffic direct while proxying external traffic.
Request Behavior (Concurrency, Retries, Timeouts)
This group determines stability under load, failure recovery behavior, and duplicate likelihood (given at-least-once delivery).
request (optional)
request.concurrency (optional, default: adaptive)
- none: One outstanding request at a time (predictable but can bottleneck)
- adaptive: worker automatically adjusts concurrency based on observed performance
Adaptive concurrency is usually best for mixed network conditions, but should be used carefully in very sensitive environments.
request.adaptive_concurrency (optional)
Fine-tunes the adaptive algorithm:
- initial_concurrency: starting point after restart
- max_concurrency_limit: safety ceiling
- decrease_ratio: how aggressively concurrency is reduced when latency increases
- ewma_alpha: how quickly the baseline adapts to new RTT patterns
- rtt_deviation_scale: how tolerant the system is to RTT variance before reacting
Incorrect tuning can lead to unstable oscillations or slow recovery.
request.timeout_secs (optional, default: 60)
- Maximum request duration before aborting.
- If set too low, you can create retry storms and duplicate writes (because the server may still process the original request).
request.retry_attempts (optional)
- Maximum retry count for failed requests.
- Very high values can cause long retry tails and prolonged backlog.
request.retry_initial_backoff_secs (optional, default: 1)
- Initial delay before first retry.
- Subsequent retries follow Fibonacci backoff.
request.retry_max_duration_secs (optional, default: 30)
- Upper bound on delay between retries.
request.retry_jitter_mode (optional, default: Full)
- Randomizes retry delays to reduce synchronized retry bursts across many agents.
request.rate_limit_num / request.rate_limit_duration_secs
- Limits outbound request rate.
- Useful to protect the remote write endpoint or avoid saturating a constrained network link.
Multi-Tenancy
tenant_id (optional)
- Adds the X-Scope-OrgID header to outbound requests.
- Used by multi-tenant remote write backends (e.g., Cortex/Mimir-like systems).
- Supports templating, allowing tenant to be derived from the metric/event context.
TLS
tls (optional)
Secures outbound connections and controls certificate verification behavior.
Key options include:
- tls.ca_file: additional trusted CA bundle
- tls.crt_file / tls.key_file: client certificate authentication (mTLS)
- tls.key_pass: decrypting an encrypted private key
- tls.server_name: SNI override for TLS
- tls.verify_certificate: validates certificate chain trust
- tls.verify_hostname: validates hostname matches certificate identity
- tls.alpn_protocols: protocol negotiation hints where relevant
Operational guidance: disabling certificate/hostname verification is risky and should be avoided except in tightly controlled test environments.