Sampling & Advanced Sampling
Sampling
In Kron TP, sampling refers to the practice of selecting a subset of data points from a larger dataset for analysis or processing. The sampling function enables users to adjust the sampling rate based on specific conditions, events, or patterns. This adaptability enhances the responsiveness of telemetry pipelines. Sampling is employed to manage and optimize the handling of telemetry data, especially when dealing with large volumes. Here are some key purposes and benefits of sampling in a telemetry pipeline:
Data Volume Reduction:
Sampling helps reduce the overall volume of telemetry data by selecting a representative subset. This is particularly valuable when dealing with high-frequency data streams, to avoid overwhelming processing and storage systems.
Performance Optimization:
Sampling improves the overall performance of telemetry processing by focusing on a representative portion of the data. This is especially important for real-time or near-real-time applications where quick response times are essential.
Reducing Noise:
Telemetry data may contain noise or repative that are not essential for analysis. Sampling helps reduce this noise, allowing analysts to focus on meaningful patterns and anomalies.
Difference Between Sampling and Advanced Sampling
The distinction between advanced sampling and normal sampling lies in their capabilities and the types of data they can process. In normal sampling, operations can be performed without a structured data format using regular expressions (regex)and regular string based searches. This allows for more flexible handling of data, especially when the structure is not well-defined.
On the other hand, advanced sampling is applicable only to parsed data, and it uses a different language called VRL instead of regex expressions. VRL provides a more specialized and structured way of expressing conditions on parsed data, making it suitable for scenarios where data has been parsed and follows a specific structure.
Additionally, advanced sampling can be used to express conditions for excluding certain scenarios where sampling should not occur (exclude). This feature allows for more fine-grained control over the sampling process, specifying conditions for excluding specific data points based on parsed attributes.

- rate: The rate at which events are forwarded, expressed as 1/N.
Example: rate = 1500 means 1 out of every 1500 events is forwarded.
- Exclude
Purpose: Specifies a logical condition to exclude events from sampling.
- key_field: The name of the field whose value is used to hash and group events for sampling.
Exclude condition example:
"exclude": ".status == 500"
In summary, while normal sampling with regex or string base is more versatile and can be applied to unstructured data, advanced sampling with VRL is designed for parsed data and allows for the creation of more complex and expressive conditions. The use of advanced sampling extends beyond just selecting a subset of data; it can also be employed to exclude specific scenarios from the sampling process. The choice between them depends on the specific requirements and nature of the data being processed in the telemetry pipeline.