AWS CloudWatch Metrics
AWS CloudWatch Metrics sink
This sink takes worker metrics (counters, gauges, histograms) and writes them to CloudWatch Metrics using PutMetricData-style batching. Your “main” knobs are:
- default_namespace: where metrics land in CloudWatch (critical for discoverability)
- storage_resolution: 1s (high-res) vs 60s (standard) per metric name
- batch + request: latency vs API cost/throttling behavior
- buffer: what happens when AWS is slow/unreachable (backpressure vs loss)
- gauge warning: gauges are stateful between flushes; “delta gauge” needs care
Parameters
default_namespace (required)
What it does: Sets the CloudWatch Namespace for metrics that don’t already have a namespace.
Why it matters:
- In CloudWatch, metric name alone is not unique; namespace is a primary “container”.
- If you send the same metric names from multiple systems, namespace is how you avoid collisions / confusion.
Typical patterns:
- service / kubernetes / telemetry / kron_tp (something product or domain-scoped)
- Keep it stable; changing namespace looks like a “new metric family” in CloudWatch.
Metric resolution & “1-second metrics”
storage_resolution (optional)
storage_resolution is a map: metric_name -> resolution, where resolution is:
- 60 = standard resolution (default if unset)
- 1 = high resolution (more granular, typically more cost)
Example idea (conceptually):
- Set 1 only for “fast-changing / alert-critical” metrics (queue depth, packet drops, error rates)
- Leave everything else at 60
Important behavior: This mapping is by metric name, so be consistent with naming. If the same name appears with different semantics, you can’t choose different resolutions without renaming.
Batching knobs (latency vs API efficiency)
CloudWatch metrics are batched. Your batch settings decide how often worker flushes and how large each flush is.
batch.max_events (optional, default: 20 events)
Flush when the batch reaches this many metric events.
- Increasing can reduce API calls (cheaper, more efficient)
- Too high can increase latency and amplify retry bursts on failure
batch.timeout_secs (optional, default: 1s)
Flush after this time even if max_events isn’t reached.
- Lower = lower latency, more API calls
- Higher = fewer API calls, more “bursty” delivery
batch.max_bytes (optional)
Upper bound by uncompressed batch size (if supported/used by the sink).
- Think of it as a “don’t accidentally build a massive payload” guardrail.
Rule of thumb:
- If you care about “near-real-time” dashboards/alerts → keep timeout_secs low (1–5s range)
- If you care about cost and don’t need quick visibility → raise timeout_secs and/or max_events
Buffering (what happens when CloudWatch is slow)
buffer.type (optional, default: memory)
- memory: fastest; loses buffered metrics on crash/restart
- disk: durable; keeps buffered data across restarts (safer)
buffer.max_size (required)
Hard cap for the buffer. When this is hit, behavior depends on buffer.when_full.
buffer.when_full (optional, default: block)
- block: applies backpressure upstream (prevents loss but can slow your pipeline)
- drop_newest: keeps pipeline moving but drops newest metrics
Practical guidance:
- For reliability (especially compliance-ish environments) → disk + block
- For “best-effort metrics” in very high-volume situations → memory + consider drop_newest
Reliability semantics (acknowledgements)
acknowledgements.enabled (optional)
Controls end-to-end ack behavior.
- If enabled and your upstream supports it, upstream will only ack once CloudWatch sink acks.
- Useful if you want “don’t say it’s delivered until AWS accepted it” semantics.
- Can reduce throughput if CloudWatch is slow (because it propagates backpressure).
AWS access / identity
auth.* (optional)
How Kron TP gets AWS credentials:
- static access keys (access_key_id / secret_access_key)
- assume_role (recommended in many orgs; adds STS dependency)
- credentials_file / profile
- imds (best for EC2/EKS node roles, when available)
region (optional)
AWS region for CloudWatch Metrics API.
endpoint (optional)
Override AWS endpoint (mostly for AWS-compatible/testing scenarios).
Request behavior (throttling, retries, concurrency)
request.concurrency (optional, default: adaptive)
- adaptive: Worker auto-tunes concurrency based on latency/feedback
- none: forces concurrency=1 (safe but slow)
- Integer: fixed concurrency
Why you tune it: CloudWatch APIs can throttle. Too aggressive concurrency can cause retry storms.
request.rate_limit_num / request.rate_limit_duration_secs (default: 150 per 1s)
Hard cap on request rate.
- Lower it if you see throttling
- Raise carefully if you know your limits and need throughput
request.timeout_secs (default: 60s)
Timeout per request. Too low can cause unnecessary retries; too high delays failure detection.
request.retry_* (attempts/backoff/jitter/max_duration)
Controls how persistent retries are and how “spiky” they are. Jitter helps avoid synchronized retry storms.
Compression
compression (optional, default: none)
Compresses payloads to AWS.
- none is simplest and default here.
- Turn on only if bandwidth is constrained and CPU is available.
Healthcheck
healthcheck.enabled (optional, default: true)
Startup-time connectivity/auth check.
- Disable only if you have known edge cases (e.g., restricted permissions during bootstrap) and it blocks startup.
The warning about Gauges (very important)
CloudWatch is fine with gauges, but Kron TP's internal handling is what the warning is about:
- Gauge values are persisted between flushes.
- On worker startup, each gauge is assumed to start at 0.0 until you set it.
- “Delta gauges” (increment/decrement style) are powerful but risky in distributed setups:
- restarts reset baseline
- multiple instances updating the same logical gauge can create misleading values unless designed carefully
Safe approach:
- Prefer absolute gauges (“this is the current value now”) rather than deltas, unless you’re intentionally doing distributed delta aggregation.