Postgres
PostgreSQL Sink
The PostgreSQL sink delivers observability data (logs, metrics, and traces) into a PostgreSQL table using batched inserts. This guide focuses on what each parameter does, plus the operational implications you should consider in production.
Important Warnings and Limitations
PostgreSQL Column Defaults Are Not Applied
If the destination table has default values defined for columns, PostgreSQL will not automatically apply those defaults when the incoming event is missing that field. Instead, the sink will insert NULL for that column.
Operational impact:
- You may end up with unexpected NULLs even when the table defines defaults.
- If downstream queries assume defaults, they may break or behave incorrectly.
Common mitigations:
- Enforce NOT NULL constraints for fields that must always exist.
- Use database-level logic (constraints/triggers) to set defaults when missing.
Batch Insert Failure Propagates to Entire Batch
Because insertion uses a recordset-style mechanism, a single problematic event (for example, a unique constraint violation) can cause the entire batch to fail.
Operational impact:
- One bad record can block ingestion and increase retry pressure.
- You may observe repeated failures until the offending event(s) are removed or corrected.
Mitigations:
- Use database-side exception-handling logic (triggers) for known constraint patterns.
- Reduce batch size (in extreme cases insert one event per batch).
SQL Injection Risk in table
The table name is interpolated into SQL and is not sanitized.
Operational impact:
- Never construct the table name from untrusted input.
- Treat the table value as static, controlled configuration.
Core Destination Parameters
endpoint (required)
The PostgreSQL connection string used to reach the database.
What it controls:
- Database host/port and target database name
- Potentially user credentials (if embedded in the string)
- Connection behavior depending on supported connection string options
Operational considerations:
- Connection failures cause retries and buffer growth.
- Credential rotation strategy matters if credentials are embedded in the connection string.
table (required)
The name of the destination table where records are inserted.
What it controls:
- Where all incoming events land
- Which schema constraints apply (NOT NULL, UNIQUE, FK, etc.)
Operational considerations:
- Must match an existing table schema that aligns with your event structure.
- Avoid dynamic/templated table names unless you fully control the input source.
Batching
batch (optional)
Controls how events are grouped into insert operations.
Why it matters:
- Batch size affects throughput, latency, and failure blast radius.
batch.max_bytes (optional, default: 10,000,000 bytes)
Maximum uncompressed size of a batch before it is flushed.
Operational impact:
- Larger batches improve write efficiency but increase memory pressure and failure impact.
- Very large batches can increase transaction time and lock contention.
batch.max_events (optional)
Maximum number of events in a batch before flushing.
Operational impact:
- Higher values improve throughput, but increase the number of events lost/retried when a batch fails.
- Lower values reduce failure impact but can increase database request rate.
batch.timeout_secs (optional, default: 1 second)
Maximum age of a batch before it is flushed, even if size thresholds are not met.
Operational impact:
- Lower values reduce latency (faster visibility in DB).
- Higher values improve batching efficiency at lower traffic volumes.
Failure-domain note: Because one event can fail the entire batch, max_events is a key reliability lever.
Buffering
buffer (optional)
Controls how events are temporarily stored when Postgres is slow or unavailable.
buffer.type (optional, default: memory)
- memory: higher performance, but buffered data is lost if worker crashes/restarts
- disk: more durable (survives restarts), but slower and requires disk sizing
buffer.max_size (required)
Hard cap for how much memory/disk the buffer can consume.
Operational impact:
- Too small → frequent backpressure or drops (depending on when_full)
- Too large → can hide downstream issues and extend recovery time after outages
Disk buffer note:
- Must meet minimum sizing requirements and needs sufficient IOPS for sustained backlog.
buffer.max_events (optional; relevant for memory buffers, default: 500)
Maximum number of buffered events (memory mode only).
Operational impact:
- Helps cap memory usage in event-count terms.
- Useful when event size is unpredictable.
buffer.when_full (optional, default: block)
Behavior when the buffer is full:
- block: applies backpressure upstream (preferred when you want to avoid data loss)
- drop_newest: drops incoming events (preferred when system continuity is more important than completeness)
Connection Pooling
pool_size (optional, default: 5)
Number of database connections maintained in the sink’s connection pool.
Operational impact:
- Too low → throughput bottleneck, higher end-to-end latency
- Too high → can overwhelm Postgres (connections, CPU, lock pressure), especially with many agents
Sizing guidance:
- Match to Postgres capacity and expected write concurrency.
- Consider total connections across all worker instances.
Request Behavior (Concurrency, Timeouts, Retries)
request (optional)
Controls how worker schedules and retries outbound database operations. The retry backoff follows a Fibonacci pattern.
request.concurrency (optional, default: adaptive)
How many concurrent outbound requests are allowed:
- none: fixed concurrency of 1 (most predictable, lowest throughput)
- adaptive: worker automatically adjusts concurrency based on observed performance
Operational impact:
- Adaptive can maximize throughput but may oscillate if the DB performance is unstable.
- Fixed (1) is safer for small databases or fragile schemas but may not meet throughput needs.
Adaptive Concurrency Fine-Tuning
These rarely need changes, but they shape stability under load:
- request.adaptive_concurrency.initial_concurrency: starting point after restart
- request.adaptive_concurrency.max_concurrency_limit: safety ceiling to prevent runaway concurrency
- request.adaptive_concurrency.decrease_ratio: how aggressively concurrency is reduced when latency increases
- request.adaptive_concurrency.ewma_alpha: how quickly the baseline adapts to new latency conditions
- request.adaptive_concurrency.rtt_deviation_scale: tolerance for latency variability before scaling back
Operational caution:
- Incorrect values can create unstable behavior (oscillation, slow recovery, or overload).
Retry Controls
- request.retry_attempts: maximum retries for failed operations Very high values can cause prolonged retry storms during outages.
- request.retry_initial_backoff_secs (default: 1 second): delay before first retry Subsequent delays follow Fibonacci growth.
- request.retry_max_duration_secs (default: 30 seconds): maximum wait between retries
- request.retry_jitter_mode (default: Full): randomization strategy for backoff Full jitter helps avoid synchronized retry spikes across many agents.
Timeouts
request.timeout_secs (optional, default: 60 seconds)
Maximum time a request can run before being aborted.
Operational impact:
- If too low, you risk aborting operations that the DB might still complete, creating retries and extra pressure.
- If too high, stalled operations consume concurrency and slow recovery.
Inputs
inputs (required)
Defines which upstream sources/transforms feed this sink.
Operational impact:
- Mis-scoped inputs can unintentionally send high-volume telemetry to Postgres.
- Wildcards are powerful but can increase blast radius of config mistakes.
Supported Data Types
- Logs: supported
- Metrics: counter, gauge, histogram, distribution, summary, set
- Traces: supported
Operational note: ensure the destination schema and downstream queries can handle the variety of payload shapes implied by logs vs metrics vs traces.
Built-in Telemetry You Should Monitor
Worker emits internal metrics that are especially useful for Postgres sink operations:
Buffer Health
- buffer_byte_size (gauge): how much data is queued
- buffer_events (gauge): how many events are queued
- buffer_discarded_events_total (counter): events dropped due to buffer strategy
Reliability and Errors
- component_errors_total (counter): sink errors (connection, insert failures, etc.)
- component_discarded_events_total (counter): events dropped by the component (intentional or error-driven)
Throughput
- component_sent_events_total / component_sent_bytes_total (counter): delivery volume
- component_received_events_total (counter): ingress to the component
Operational interpretation:
- Rising buffer gauges + rising errors usually indicates DB slowdowns/outages.
- Rising discarded counters indicate loss (intentional drops or failure-driven discards) and should trigger incident investigation.