Aws S3
aws s3 source collect logs, metrics, and traces from objects stored in amazon s3 this source ingests data from s3 buckets by consuming s3 event notifications delivered through amazon sqs , enabling reliable and scalable batch based ingestion it is commonly used for ingesting archived logs , cloud service outputs , and object based telemetry pipelines collection model s3 bucket events are delivered to an sqs queue worker polls the queue for object creation notifications referenced objects are fetched from s3 object contents are decoded into events events are forwarded downstream with acknowledgement support this model decouples object creation from ingestion, enabling high durability and retry safety requirements an s3 bucket configured to emit event notifications an sqs queue subscribed to those notifications appropriate iam permissions to read from s3 poll and manage messages in sqs assume roles or retrieve credentials as configured typical architecture patterns s3 based log archival ingestion cloud service log pipelines (cloudtrail, elb, vpc flow logs) batch to stream telemetry conversion disaster recovery or delayed ingestion multi account aws ingestion pipelines acknowledgement handling acknowledgements (deprecated) source level acknowledgement configuration is deprecated important notes acknowledgement behavior is controlled globally or at the sink level enabling or disabling acknowledgements at the source has no effect ensures consistent end to end delivery semantics across pipelines authentication model the aws s3 source supports multiple aws authentication strategies, aligned with standard aws sdk behavior supported authentication methods static credentials (access key / secret key) temporary credentials with session tokens iam role assumption (sts) instance metadata service (imds) shared credentials files and profiles credential loading follows the configured strategy and supports timeout and retry controls iam role assumption when assuming a role a role arn must be provided optional external ids can be used for cross account access session names are auto generated if not specified this is the recommended approach for production grade aws deployments aws region and endpoints region specifies the aws region used for service interactions if not set defaults to the region configured for the service or is inferred from the execution environment endpoint allows specifying a custom s3 compatible endpoint useful for aws compatible object storage private s3 endpoints local or on prem s3 compatible systems object retrieval and addressing force path style controls whether bucket name appears in the hostname (virtual hosted style) or as part of the request path (path style) path style addressing improves compatibility with some s3 compatible services compression handling compression controls how objects are decompressed after retrieval supported modes automatic detection explicit compression formats no decompression compression is inferred from metadata and object key suffixes when auto mode is enabled event decoding the aws s3 source supports a wide range of decoders to transform raw object bytes into structured events supported decoding types json syslog gelf avro protobuf otlp influxdb line protocol native vector formats raw bytes vrl based custom decoding decoding may emit logs metrics traces depending on the selected codec signal type detection for multi signal formats the decoder attempts parsing in priority order signal detection can be restricted for performance optimization framing model framing defines how individual events are separated within object content supported framing methods include newline delimited character delimited length delimited octet counting varint length delimited chunked gelf raw byte frames frame length limits can be enforced to protect against malformed or unbounded inputs multiline aggregation multiline processing allows multiple lines to be combined into a single logical event supported aggregation modes continue through continue past halt before halt with this is commonly used for stack traces multi line application logs structured log blocks timeouts ensure buffered messages are eventually flushed sqs integration queue consumption model messages are polled in batches each message references one or more s3 objects visibility timeouts protect against duplicate processing messages are deleted after successful ingestion concurrency and throughput polling concurrency scales with cpu by default batch size and poll intervals are tunable designed to balance ingestion throughput and sqs visibility constraints deferred processing the source can defer processing of older events events exceeding a defined age threshold forwarded to an alternate sqs queue enables delayed or backfill pipelines reliability characteristics at least once delivery semantics integrated retry handling via sqs visibility timeouts end to end acknowledgements supported stateless ingestion model this makes the aws s3 source suitable for durable, large scale ingestion pipelines tls and proxy support tls can be configured for aws api communication custom ca bundles and certificates are supported http and https proxying is supported with fine grained bypass rules common use cases ingesting cloudtrail, elb, alb, and vpc flow logs processing archived logs stored in s3 batch to stream observability pipelines cross account or cross region ingestion hybrid cloud and on prem object storage ingestion