Sematext Logs
Sematext Logs Sink
Overview
The Sematext Logs sink publishes log events to the Sematext Logs platform over HTTP. It is designed for reliable, production-grade log delivery with batching, buffering, retries, adaptive concurrency, and optional end-to-end acknowledgements.
This sink is stateless, supports at-least-once delivery semantics, and integrates tightly with worker's buffering and request middleware to handle bursty traffic and transient downstream failures.
Authentication
token (required)
The Sematext ingestion token used to authenticate requests.
This token uniquely identifies the Sematext Logs application where events are written. If the token is invalid or missing, requests will be rejected by Sematext and events will be retried according to the request retry policy.
Batching
batch
Controls how events are grouped before being sent to Sematext.
Batching reduces request overhead and improves throughput, especially for high-volume log streams.
batch.max_bytes
The maximum uncompressed size, in bytes, of a batch before it is flushed.
This limit is applied before serialization and compression, ensuring predictable memory usage.
batch.max_events
The maximum number of events allowed in a batch before it is flushed.
Useful for controlling batch granularity when events are large or uneven in size.
batch.timeout_secs
The maximum amount of time a batch can remain open before being flushed, even if size thresholds are not met.
This prevents low-volume streams from experiencing excessive delivery latency.
Buffering
buffer
Controls how events are buffered locally before being sent.
Buffering protects against temporary network issues, Sematext downtime, or backpressure from downstream services.
buffer.type
Specifies where buffered events are stored.
- memory High performance, low latency. Buffered events are lost if worker crashes or restarts.
- disk Durable storage. Events persisted to disk survive restarts and crashes, at the cost of reduced performance.
buffer.max_size
The maximum total memory or disk space that the buffer is allowed to consume.
For disk buffers, this must be large enough to accommodate on-disk metadata and synchronization overhead.
buffer.max_events
Limits the number of events stored in memory buffers.
Only applicable when using in-memory buffering.
buffer.when_full
Defines behavior when the buffer reaches capacity.
- block Applies backpressure upstream, slowing ingestion without data loss.
- drop_newest Drops incoming events immediately, prioritizing throughput over reliability.
Encoding
encoding
Controls how log events are transformed prior to serialization and transmission.
This layer does not define the wire format, but instead selects which fields are included and how timestamps are represented.
encoding.only_fields
A whitelist of fields to include in the outgoing event.
All other fields are removed. This is useful for schema enforcement or reducing payload size.
encoding.except_fields
A blacklist of fields to exclude from the outgoing event.
All other fields are preserved.
encoding.timestamp_format
Controls how timestamp fields are serialized.
Supported formats include RFC 3339 and multiple Unix timestamp variants (seconds, milliseconds, microseconds, nanoseconds), allowing alignment with downstream expectations.
Endpoint and Region
region
Specifies the Sematext region where data is sent.
This controls the default ingestion endpoint and should match the region of your Sematext Logs application.
endpoint
Overrides the region-based endpoint with a custom ingestion URL.
This is useful for private endpoints, testing environments, or future region expansion.
When set, this value takes precedence over region.
Proxy Support
proxy
Configures HTTP and HTTPS proxy behavior for outbound requests.
This is commonly used in restricted network environments where direct internet access is not permitted.
proxy.enabled
Globally enables or disables proxy usage.
proxy.http / proxy.https
Specifies proxy endpoints for HTTP or HTTPS traffic.
Each must be a valid URI.
proxy.no_proxy
Defines hosts or address ranges that should bypass the proxy.
Supports exact hosts, wildcard domains, IP addresses, CIDR ranges, and catch-all rules.
Request Middleware
request
Controls how outbound HTTP requests are executed, retried, rate-limited, and parallelized.
This layer is critical for maintaining stable performance under load and during partial downstream failures.
Concurrency Control
request.concurrency
Defines how many concurrent requests may be in flight.
- adaptive Uses worker's Adaptive Request Concurrency (ARC) algorithm to dynamically tune concurrency based on observed latency.
- none Forces a fixed concurrency of one request at a time.
request.adaptive_concurrency
Fine-grained tuning parameters for adaptive concurrency behavior.
These settings influence how aggressively concurrency scales up or down in response to latency changes and should generally be left at defaults unless deep tuning is required.
Rate Limiting
request.rate_limit_num
Maximum number of requests allowed per rate limit window.
request.rate_limit_duration_secs
Time window used to enforce rate limits.
Retry Behavior
request.retry_attempts
Maximum number of retry attempts for failed requests.
request.retry_initial_backoff_secs
Initial delay before the first retry attempt.
Subsequent retries follow a Fibonacci backoff pattern.
request.retry_max_duration_secs
Maximum delay between retry attempts.
request.retry_jitter_mode
Controls whether retry delays include randomness.
Full jitter is recommended to prevent synchronized retry storms.
Timeouts
request.timeout_secs
Maximum duration a request is allowed to run before being aborted.
Setting this too low can lead to duplicate ingestion due to retries against slow but successful upstream processing.