Kubernetes Logs
Kubernetes Logs Source
Collect logs directly from Kubernetes Nodes by reading Pod log files on the host filesystem. This source automatically enriches log events with Pod, Namespace, and Node metadata using the Kubernetes API.
It is designed to run as a DaemonSet, collecting logs only from the Node it is scheduled on.
Collection Model
- Reads container logs from the host’s /var/log/pods directory
- Tails log files written by the container runtime
- Enriches events with Kubernetes metadata in near real-time
- Streams logs downstream with minimal buffering
This source is optimized for low-latency log shipping rather than guaranteed delivery.
Requirements
- Kubernetes version 1.19 or newer
- Read access to /var/log/pods
- When running inside Kubernetes, this access is typically provided via a hostPath volume
- Sufficient RBAC permissions to read Pod, Namespace, and Node metadata
Platform Limitations
- Only tested on Linux-based Kubernetes nodes
- Behavior on Windows-based clusters is not guaranteed
Partial Log Handling
auto_partial_merge (optional)
Controls whether partial log messages are automatically merged.
Some container runtimes split long log lines into multiple fragments. When enabled, these fragments are reassembled into a single event before forwarding.
This is enabled by default and recommended for application logs.
State and Checkpointing
data_dir (optional)
Defines the directory where file read positions are persisted.
Operational notes:
- Defaults to the global data_dir if not explicitly set
- The running user must have write permissions
- Allows worker to resume log collection after restarts without data loss
File Lifecycle and Cleanup
delay_deletion_ms (optional)
Specifies how long metadata entries are retained after a Pod deletion event is observed.
A longer delay allows:
- Continued enrichment of logs emitted shortly after Pod termination
- Reduced risk of unenriched events during Pod churn
If metadata expires before log ingestion completes, logs are forwarded without enrichment and a warning is emitted.
File Selection and Filtering
include_paths_glob_patterns (optional)
Defines which log files are eligible for collection.
By default, all files under the Pod log directory are included.
exclude_paths_glob_patterns (optional)
Defines glob patterns for files that should be ignored.
Common exclusions include:
- Compressed files
- Temporary files created during rotation
ignore_older_secs (optional)
Skips files whose last modification time exceeds the specified age.
Useful for:
- Avoiding ingestion of stale logs
- Reducing startup load on nodes with large log histories
Pod and Namespace Filtering
extra_field_selector (optional)
Applies an additional Kubernetes field selector when watching Pods.
This works alongside the built-in Node filter, which ensures only Pods scheduled on the current Node are observed.
extra_label_selector (optional)
Filters Pods based on Kubernetes labels.
This allows selective log collection without modifying workloads.
extra_namespace_label_selector (optional)
Filters Namespaces based on labels before Pod discovery occurs.
This is useful in large clusters to limit API load and memory usage.
File Fingerprinting and Polling
fingerprint_lines (optional)
Defines how many lines are read to generate a file checksum.
If a file contains fewer lines than this value, it will not be read at all.
This helps distinguish between files with similar paths during rotation.
glob_minimum_cooldown_ms (optional)
Controls how frequently the filesystem is scanned for new files.
Lower values increase responsiveness but may introduce additional filesystem overhead.
Timestamp Handling
ingestion_timestamp_field (optional)
Overrides the field name used to store the ingestion timestamp.
This enables latency analysis between:
- Log write time
- Log collection time
- Downstream processing stages
timezone (optional)
Defines the default timezone for timestamps that do not include explicit timezone information.
Metadata Enrichment Controls
insert_namespace_fields (optional)
Controls whether Namespace-level metadata is fetched and attached to events.
Disabling this:
- Reduces load on the kube-apiserver
- Lowers memory usage
- Removes Namespace label fields from events
Enabled by default.
Annotation Field Mapping
The following configuration groups control how Kubernetes metadata is mapped into event fields. Each field can be disabled by setting its target to an empty value.
Namespace Annotation Fields
Controls enrichment with Namespace metadata, including labels.
Node Annotation Fields
Controls enrichment with Node metadata, including node labels.
Pod Annotation Fields
Controls enrichment with Pod-level metadata, including:
- Pod name and namespace
- Pod UID and owner references
- Container name, image, and IDs
- Pod IPs and labels
- Pod annotations
This enrichment is central to correlating logs with Kubernetes workloads.
Line and Size Limits
max_line_bytes (optional)
Defines the maximum size of a single log line.
Lines exceeding this limit are discarded to protect against malformed or unexpected input.
max_merged_line_bytes (optional)
Defines the maximum size of a merged log line after partial events are combined.
This only applies when partial merge is enabled.
max_read_bytes (optional)
Limits how many bytes are read from a single file before switching to another file.
This helps distribute read capacity across many active log files.
Read Order and Rotation
oldest_first (optional)
Controls whether older files are fully drained before newer files are read.
When enabled, this prioritizes backlog reduction over near-real-time logs.
read_from (optional)
Defines where reading starts when a file is first discovered:
- From the beginning of the file
- Or from the current end
rotate_wait_secs (optional)
Specifies how long file handles are kept open after log rotation.
This prevents premature closure when log writers are slow to release files.
Kubernetes API Access
kube_config_file (optional)
Specifies an explicit kubeconfig file.
If not set, in-cluster configuration is used automatically.
use_apiserver_cache (optional)
Controls whether Kubernetes API requests may be served from a local cache.
Enabling this can reduce API load but may slightly delay metadata updates.
Node Identification
self_node_name (optional)
Defines the name of the Node the source is running on.
By default, this is populated from an environment variable injected by Kubernetes at Pod creation time.
Internal Metrics
internal_metrics (optional)
Controls emission of internal file-related metrics.
An optional file tag can be enabled, but this introduces unbounded cardinality and should be used with caution.
Reliability and Delivery Semantics
- Logs are delivered on a best-effort basis
- No acknowledgements or retries are performed at the source level
- Data loss is possible during crashes or restarts
- Designed for high-throughput, low-latency scenarios
Common Use Cases
- Cluster-wide log collection via DaemonSet
- Shipping application and infrastructure logs
- Enriching logs with Kubernetes context for correlation
- Feeding SIEM, log analytics, or object storage pipelines
- Replacing or complementing Fluentd / Fluent Bit setups