ClickHouse
ClickHouse Sink
Overview
The ClickHouse sink enables worker to deliver log data into a ClickHouse database using HTTP-based ingestion. Events are written in batches, making this sink suitable for high-throughput logging pipelines where analytical queries, long-term retention, and efficient compression are required.
This sink is commonly used in environments where ClickHouse serves as a central log warehouse for observability, security analytics, compliance reporting, or large-scale time-series analysis.
Supported Input Types
This sink supports logs.
Requirements
A ClickHouse server version equal to or newer than 1.1.54378 is required. The target table must already exist in ClickHouse and be compatible with the selected input format and schema produced by Worker.
Core Configuration Parameters
Endpoint (required) Defines the HTTP endpoint of the ClickHouse server that receives insert requests. This typically points to the ClickHouse HTTP interface and must be reachable from the Worker instance.
Table (required) Specifies the target table into which events are inserted. The table schema must align with the fields and data types emitted by worker after encoding and transformation.
Database (optional) Overrides the database name specified in the endpoint and determines which database contains the target table. This is useful when routing data into different databases dynamically.
Authentication and Security
Authentication (optional) Controls how Worker authenticates HTTP requests to ClickHouse. Multiple authentication strategies are supported, including basic authentication, bearer tokens, AWS-style request signing, and custom authorization headers.
Authentication credentials are sent as HTTP headers and should always be used over secure transport. In production environments, HTTPS is strongly recommended to protect credentials in transit.
TLS Configuration (optional) Controls transport security when connecting to ClickHouse over HTTPS. TLS settings allow customization of certificate validation, trust roots, and server name verification to match enterprise security requirements.
Encoding and Data Format
Encoding (optional) Controls how events are transformed before being sent to ClickHouse. Encoding options determine which fields are included or excluded, how timestamps are formatted, and how structured data is represented.
Timestamp formatting is especially important, as it must align with the ClickHouse column types used in the destination table.
Format (optional) Defines the ClickHouse input format used for inserts. Common formats include row-oriented JSON variants that allow ClickHouse to efficiently parse incoming events. The selected format must match the table definition and expected column mapping.
Batching and Throughput
Batching (optional) Controls how events are grouped before being sent to ClickHouse. Batching reduces request overhead and improves ingestion throughput, especially at high event rates.
Batch size can be tuned by event count, total payload size, or maximum time before a batch is flushed. Larger batches improve throughput efficiency, while smaller batches reduce latency and retry impact.
Buffering and Backpressure
Buffer (optional) Defines how events are buffered before being written to ClickHouse. Buffering protects the pipeline from transient slowdowns or brief outages in the database.
Memory buffering provides higher throughput but does not survive restarts. Disk buffering offers greater durability and can absorb longer outages at the cost of additional I/O overhead.
Backpressure behavior can be tuned to either slow upstream components or intentionally drop newer events when buffers are full, depending on reliability and performance requirements.
ClickHouse-Specific Insert Behavior
Distributed Inserts (optional) When inserting into distributed tables, the sink can be configured to allow ClickHouse to route inserts to a random shard. This can simplify ingestion into clustered deployments but may affect determinism and ordering.
Async Insert Settings (optional) The sink supports ClickHouse asynchronous insert modes, allowing data to be queued and flushed by the server in the background. This can significantly increase throughput but may introduce additional latency before data becomes visible for queries.
Deduplication and waiting behavior for asynchronous inserts can be tuned to match the consistency requirements of the workload.
Unknown Field Handling (optional) Controls how ClickHouse handles fields that are not present in the table schema. This allows pipelines to evolve gradually without failing inserts when new fields appear.
Common Usage Patterns
The ClickHouse sink is widely used as a long-term log storage backend for observability platforms, security analytics pipelines, and compliance-driven environments that require fast, SQL-based querying over large volumes of data.
It is particularly well suited for high-cardinality datasets, time-partitioned tables, and workloads that benefit from columnar storage and aggressive compression.
In complex deployments, ClickHouse is often paired with worker transforms to normalize schemas, enforce data quality, and reduce ingestion-side complexity.