Lossless or Liability: Why You Cannot Sample Industrial Telemetry

,
Lossless or Liability - Why You Cannot Sample Industrial Telemetry

Most observability and historian strategies quietly rest on one assumption: that you can throw some of the data away. Sample every tenth reading. Aggregate to one-minute averages. Keep the last 90 days at full resolution and roll up the rest. It feels reasonable, and in industrial systems it is a liability waiting to surface.

The event you dropped is the one you needed

Industrial failures live in the transients. A pressure spike that lasts 400 milliseconds. A motor current signature that flutters for two seconds before a bearing lets go. A brief comms dropout that precedes a batch going out of spec. Averaging to one-minute windows erases exactly these signals. You are left with a smooth line that tells you everything was fine right up until it was not.

Sampling optimizes for storage cost and sacrifices the rare, short-lived events that root-cause analysis, anomaly detection, and AI models depend on. In a control environment, those events are the whole point.

Lossless is a compliance and forensics requirement

In regulated industries, the standard is not “representative.” It is “complete.” When an auditor, an investigator, or a safety review asks what the plant was doing at 02:47:13 on a specific night, “we had a one-minute average” is not an answer. Pharmaceutical batch records, energy and pipeline reporting, and incident investigations all assume the underlying data actually exists at the resolution it was produced.

Once a signal is downsampled, it is gone. You cannot reconstruct a transient from an average. Lossless capture is the only posture that survives a serious question after the fact.

“But full-resolution data is too expensive to keep”

That has historically been true, and it is the real reason teams sample. It is also the assumption Valak is built to break. Sasquatch Labs’ patent-pending, multi-layer, property-aware compression treats telemetry as the structured, repetitive, slowly-changing signal it actually is, rather than as opaque bytes.

Industrial tags are extraordinarily compressible when you respect their structure: values that move in small deltas, timestamps that advance predictably, tags that repeat the same shape forever. Exploiting that has produced double-digit lossless compression ratios on real telemetry in our own testing. The economics that forced you to sample change when keeping everything costs a fraction of what it used to.

Lossless by default, queryable on demand

Capture is only half of it. Data you cannot query is just a more expensive archive. Valak pairs lossless capture with a queryable surface, so the full-resolution history is not a tape in a drawer. It is something an engineer, or an agent, can actually ask questions of: pull the raw waveform around an event, compare this run to the last hundred, correlate a defect with the exact conditions that produced it.

The standard should be: keep everything

“We sample to save money” is a decision made under old constraints. For industrial telemetry, where the rare event is the valuable one and “complete” is often a legal requirement, the right default is lossless. Capture every signal, keep it affordably, and make it queryable. That is the foundation the rest of industrial AI has to stand on, and it is what Valak is built to deliver.