Azure Blob Storage
Azure Blob Storage Sink
Overview
The Azure Blob Storage sink stores worker log events in Azure Blob Storage, turning blob containers into a durable, low-cost archive (or a “data lake landing zone”) for observability data.
Worker writes events in batches, producing blob objects (files) under a configurable key structure (prefix + time component). This is a strong fit when you want long-term retention, offline analytics (Spark/Databricks), compliance archives, or cheap cold storage outside your primary log backend.
Supported Input Types
Logs
Prerequisites
Before configuring this sink, you typically need:
An Azure Storage Account with Blob Storage enabled A target container created (or permissions allowing worker to create/write objects) A valid authentication method (connection string with account key or SAS) A clear blob naming / partitioning strategy (prefix and time format), especially for high-volume workloads
Core Configuration Parameters
Connection String (required) The Azure Storage connection string used to authenticate and reach your storage account.
This can represent either:
Account key authentication (common for service integrations), or SAS-based authentication (useful for scoped/temporary access)
Operational note: SAS must include permissions that allow writes (and depending on your environment, reads as well). If you use a non-account SAS, some automatic checks used by clients may not work as expected.
Container Name (required) The target blob container where worker stores the generated objects.
Think of this as your “bucket” for logs. Many teams create separate containers per environment (prod/dev), per domain (security/app), or per tenancy.
Object Naming, Partitioning, and Collision Avoidance
Blob Prefix (optional) A prefix prepended to all blob keys. This is your primary control for “directory-like” partitioning inside the container.
Typical usage: partition by date and/or workload identity so that downstream query engines can scan efficiently.
A trailing slash matters if you want it to behave like a folder path.
Blob Time Format (optional) Controls the timestamp suffix used in blob keys. By default, worker appends a time-based component so each batch becomes a unique object name.
If set to an empty value, worker stops appending a timestamp — which is only safe if your prefix already guarantees uniqueness per write.
Append UUID (optional) When enabled, worker appends a UUID to the end of the blob key to guarantee uniqueness.
This is particularly useful for high-throughput pipelines where multiple writers or fast flush intervals could otherwise generate the same timestamp-based object name and cause collisions/overwrites.
Payload Encoding and File Format Strategy
Encoding (required) Controls how worker serializes events before writing them as blob content.
This is where you decide the shape of the stored data for downstream consumers:
JSON-like formats are convenient for later parsing and indexing Text formats can be cheaper and simpler for raw archives Structured formats can help when the blob container is used as a data lake source
A practical rule: choose the encoding based on the tool you’ll use later to read/process the blobs (SIEM ingest, Spark, custom ETL, etc.).
Compression
Compression (optional) Controls whether worker compresses blob contents before upload.
Compression is typically enabled by default and is a big win for storage cost and network egress, especially for verbose logs. The trade-off is CPU cost during compression and the need for readers to support decompression.
Batching Behavior
Batch (optional) Controls when worker flushes and uploads a blob object.
Key knobs conceptually include:
How many events go into one object How large an object may become How long worker can wait before forcing a flush
Batching is one of the main tuning levers:
Larger batches → fewer blobs, lower API overhead, better compression efficiency Smaller batches → lower latency, more objects, higher request overhead
Buffering and Backpressure
Buffer (optional) Controls how events are buffered before they are uploaded.
This is important for durability and smoothing traffic spikes:
Memory buffering is faster but loses buffered data on crash/restart Disk buffering is more durable and helps survive transient Azure/network issues
Request Behavior and Throughput Control
Request Settings (optional) Controls HTTP behavior such as concurrency, rate limiting, retries, and timeouts.
This matters in blob storage because:
Upload failures will be retried (which can increase duplicates if your naming is not unique) Rate limiting prevents your pipeline from overwhelming storage APIs in large clusters Concurrency tuning affects throughput and memory pressure
Common Usage Patterns
Data lake landing zone Store compressed, partitioned logs in Blob Storage, then run periodic ETL into a warehouse (Databricks/Spark/Synapse) or into a searchable index.
Compliance / long retention Keep an immutable-ish archive in cheap storage, and only index a smaller “hot subset” elsewhere.
Dual-write patterns Send logs to a primary backend for search/alerting, and also store the raw stream in Blob Storage for reprocessing, replay, or forensic backfill.
Collision-safe high volume Enable UUID suffixing and use a date-based prefix so that parallel writers never fight over the same object name while keeping objects easy to locate by time.