AWS S3
AWS S3 sink
The AWS S3 sink writes observability events (logs / metrics / traces) into Amazon S3 objects. It’s a batch sink, so it accumulates events and periodically uploads them as object files.
Destination
- bucket The target S3 bucket name. This is the only “where” that is mandatory.
- region Which AWS region S3 API calls go to. If wrong, you’ll see auth/signing errors or it’ll try the wrong regional endpoint.
- endpoint Overrides the AWS S3 endpoint. Used for:
- S3-compatible storage (MinIO, Ceph RGW, etc.)
- special AWS partitions / private endpoints If set, worker talks to that URL instead of AWS’s default.
- force_path_style Controls URL style:
- true: endpoint/bucket/key (path-style)
- false: bucket.endpoint/key (virtual-hosted style) Many S3-compatible systems prefer path-style; AWS generally supports both (but behavior can vary with TLS/DNS).
Object naming (this defines your “folder structure”)
- key_prefix Prefix prepended to every object key. Think “directory”. Best used to partition by date/env/cluster/tenant to keep listing/querying sane.
- filename_time_format Controls the timestamp portion in the object name (strftime-like). If you want predictable rotation and time-based grouping, this is the knob.
- filename_append_uuid If enabled, appends a UUID to the generated object name. Meaning: avoids collisions when multiple writers create objects in the same second/time-bucket. Tradeoff: harder to visually scan names, but safer at scale.
- filename_extension Forces a specific file extension (e.g., .json, .log, .txt). Useful if downstream tooling relies on extensions. Otherwise, optional.
Batching (how often objects are uploaded + how big they are)
- batch.max_bytes Upper bound on batch size (pre-compression). Larger → fewer, bigger objects (cheaper + more throughput). Smaller → more objects (more PUT/list overhead, lower latency).
- batch.max_events Upper bound on number of events per object. This protects you from huge objects when events are small.
- batch.timeout_secs Maximum time worker waits before it “flushes” the batch even if size/event limits aren’t hit. Smaller → more near-real-time uploads. Larger → better efficiency, more latency.
Buffering (what happens if S3 is slow/unreachable)
- buffer.type (memory or disk)
- memory: fastest, but you lose buffered data on crash/restart
- disk: slower, but survives restarts
- buffer.max_size Hard cap of buffer storage. When you hit this, you’re “out of cushion”.
- buffer.max_events (only relevant for memory buffer) Limit by event count instead of bytes.
- buffer.when_full (block or drop_newest)
- block: apply backpressure upstream (safer, no intentional loss)
- drop_newest: intentionally drop newest events to keep running (only if you accept loss)
Compression + HTTP metadata
- compression (gzip, zstd, snappy, zlib, none) Controls how payload is compressed before upload.
- gzip: common & widely compatible
- zstd: often best compression ratio/speed balance, but some tooling expects gzip
- none: bigger objects, higher cost
- content_encoding Forces the HTTP Content-Encoding metadata on the object. Normally this matches your compression; override only if you have a special consumer that expects a specific value.
- content_type Sets object Content-Type. Useful for browsers/tools or certain pipelines; not critical for most log pipelines.
Storage class + encryption + access control
- storage_class Sets the S3 storage tier of uploaded objects (STANDARD, IA, Glacier variants, etc.). This affects cost + retrieval latency. Many teams leave STANDARD and use lifecycle rules later.
- server_side_encryption (AES256 or aws:kms) Enables server-side encryption:
- AES256 = SSE-S3 (AWS managed keys)
- aws:kms = SSE-KMS (KMS keys/policies/audit)
- ssekms_key_id Which KMS key to use when server_side_encryption = aws:kms. Critical for compliance / cross-account setups.
- acl (canned ACL) Applies an S3 canned ACL to created objects/bucket. In modern AWS setups, bucket policies + IAM are preferred; ACL is for legacy/cross-account corner cases.
- grant_read / grant_full_control / grant_read_acp / grant_write_acp Adds explicit ACL grants to specific grantees. Only use when you must share object-level permissions via ACLs (rare today).
- tags Adds object tags. Useful for lifecycle policies, cost allocation, and governance.
Outbound request behavior (performance + resilience knobs)
- request.concurrency (none, fixed number, or adaptive) How many uploads/requests worker can have in-flight. More concurrency → more throughput, more pressure on S3/network.
- request.rate_limit_num / request.rate_limit_duration_secs Caps request rate. Used to avoid throttling or to be nice to S3-compatible backends.
- request.timeout_secs Per-request timeout. Too low can cause retries and duplicates; too high can stall recovery.
- request.retry_attempts / retry_strategy How aggressively to retry failures. retry_strategy.type controls whether to retry everything, nothing, or a custom set (like only certain HTTP codes).
- request.retry_initial_backoff_secs / request.retry_max_duration_secs / request.retry_jitter_mode Shapes the retry backoff curve (how fast it ramps and whether retries are randomized to avoid retry storms).
- request.adaptive_concurrency.* Controls worker's “auto-tuning” concurrency. Usually leave defaults unless you’re tuning around a flaky endpoint.
Proxy + TLS (network path / trust)
- proxy.enabled / proxy.http / proxy.https / proxy.no_proxy Routes requests via HTTP(S) proxy; no_proxy bypass list.
- tls.ca_file / tls.crt_file / tls.key_file / tls.key_pass TLS trust and client cert configuration.
- tls.verify_certificate Whether to verify the certificate chain (should generally be on).
- tls.verify_hostname Whether to verify hostname matches cert (should generally be on).
- tls.server_name Override SNI / hostname used for TLS verification (useful behind proxies or S3-compatible endpoints).
- tls.alpn_protocols Advanced TLS negotiation detail; rarely needed unless a proxy/backend requires specific ALPN.
Timezone (only affects templated date fields)
- timezone Controls what timezone is used when your templated strings include date/time specifiers. Important if you partition by day and you want day boundaries in UTC vs local time.