GCP Cloud Storage
GCP Cloud Storage Sink
Overview
The GCP Cloud Storage (GCS) sink stores observability events in Google Cloud Storage as objects. It is commonly used for long-term retention, cost-efficient archival, and downstream batch analytics (for example: BigQuery external tables, Dataproc/Spark, or custom pipelines reading from GCS).
This sink is a good fit when:
- you want cheap storage at scale and don’t need “query-now” semantics,
- you want to keep logs in GCP for compliance / audit,
- you want to build a data lake workflow (GCS → ETL → warehouse).
Supported Input Types
- Logs
Authentication
The sink supports multiple authentication mechanisms. One of the following must be available:
- Service account credentials file (via credentials_path)
- API key (via api_key)
- Application Default Credentials (ADC) via environment / instance identity
If no explicit credentials are configured:
- the GOOGLE_APPLICATION_CREDENTIALS environment variable is checked,
- otherwise, instance-level service accounts are used when running on Google Compute Engine.
Required Configuration
Bucket (required)
Defines the target GCS bucket where objects are written.
Key notes:
- must already exist,
- access permissions must allow object creation.
Encoding (required)
Defines how events are serialized into bytes before writing them to objects.
The encoding configuration selects a codec such as:
- JSON
- Text
- Avro
- Protobuf
- CEF
- CSV
- etc.
Encoding also determines what structure the resulting objects contain and how easily downstream systems can parse them.
Object Naming and Partitioning
Key Prefix
Defines a prefix for object keys. This is the primary mechanism for creating directory-like partitioning in the bucket.
Typical uses:
- date-based partitioning (daily/hourly),
- environment or tenant partitioning,
- source-type partitioning.
Important behavior:
- if you intend to use the prefix as a “folder path”, it must end with / (a trailing slash is not automatically added),
- supports per-event templating, so keys can be dynamically derived from event fields.
Filename Timestamp Formatting
Controls the time component used in the generated object keys.
Key behavior:
- by default, objects include a timestamp (commonly rendered as seconds since Unix epoch),
- changing the format affects how objects are grouped and how predictable the object paths are,
- setting it to an empty string disables timestamp appending.
Filename Extension
Controls the file extension suffix used for written objects.
Behavior:
- if unset, the extension is derived from compression scheme (when applicable),
- if set, the extension is forced to the configured value.
Append UUID to Filenames
Controls whether a UUID v4 token is appended to the generated object key.
Why this matters:
- prevents name collisions in high-volume or concurrent workloads,
- improves safety when multiple writers generate identical time-based keys.
Default behavior is typically collision-safe for parallel pipelines.
Storage and Access Control
Storage Class
Controls the GCS storage class used for created objects:
- STANDARD
- NEARLINE
- COLDLINE
- ARCHIVE
Use this to optimize costs based on access patterns:
- frequent reads → STANDARD,
- infrequent reads / archival → NEARLINE/COLDLINE/ARCHIVE.
Object ACL
Applies a predefined ACL to created objects.
Use this to enforce object-level read access patterns (for example: private vs. project-private vs. public-read) depending on your bucket policy model.
Metadata
Adds custom key/value metadata to created objects.
Typical uses:
- dataset identifiers,
- pipeline provenance,
- compliance tags,
- downstream processing hints.
Delivery Shaping
Batch Behavior
Controls how events are grouped into batches before being written.
Batch sizing influences:
- object size (fewer larger objects vs. many small objects),
- upload overhead and API call rate,
- end-to-end latency from ingest to availability in GCS.
Buffering Behavior
Defines local buffering behavior when writes are temporarily constrained.
Supported buffer strategies:
- Memory buffering for higher throughput,
- Disk buffering for better durability across restarts.
When the buffer is full, behavior can be configured to:
- block upstream (backpressure),
- or drop newest events if prioritizing throughput.
Compression
Controls whether objects are compressed before upload:
- none
- gzip
- zstd
- snappy
- zlib
Compression tradeoffs:
- reduces storage and egress costs,
- increases CPU usage,
- may affect downstream tooling depending on decompression support.
Framing
Controls how multiple events are delimited inside an object.
Framing impacts downstream parsing behavior, especially for:
- newline-delimited JSON,
- length-delimited formats,
- raw byte streams.
Supported framing methods include:
- bytes (no delimiters),
- newline-delimited,
- character-delimited,
- length-delimited,
- varint length-delimited.
Network and Transport Controls
Endpoint
Overrides the GCS API endpoint.
Defaults to the standard Google Cloud Storage endpoint. Useful for special environments, testing, or endpoint routing controls.
Proxy Support
Supports HTTP and HTTPS proxies with optional exclusions.
Used in environments with:
- restricted outbound access,
- centralized egress gateways,
- regulated network paths.
Request Behavior
Advanced request controls allow fine-tuning outbound behavior:
- adaptive or fixed concurrency,
- rate limiting aligned with quotas,
- retry strategies with backoff,
- request timeouts.
These are typically tuned for:
- high-volume pipelines,
- quota-constrained environments,
- multi-sink deployments.
TLS Configuration
Supports TLS settings including:
- custom certificate authorities,
- client certificate and key configuration,
- strict certificate and hostname verification.
TLS should be enabled whenever objects are written over untrusted networks.
Timezone Handling
Timezone
Defines the timezone used when rendering date specifiers in template strings (for example, when using date-based key_prefix patterns).
If unset:
- defaults to the globally configured worker's timezone,
- can be set to a TZ database name or local.
This matters when you partition by day/hour and want partitions aligned to a specific operational timezone.