Skip to content

Telemetry Monitors ​

Beta

Telemetry monitors are in beta, on paid plans. Limits are subject to change.

A telemetry monitor watches your own data: the number of error logs, a service's p95 latency, a metric crossing a line, or a job that stops reporting. It runs every minute (or less often), keeps a state per group, and alerts the same contacts your check monitors do.

Create one under Telemetry → Monitors → New Monitor, or from anywhere you already have a query: Create Monitor in Explore, on a log pattern, or on a derived metric.

The live preview ​

As you edit, the editor evaluates the unsaved monitor against your real data:

  • Right now: what each group would be, and the closest one to its threshold.
  • The chart: the value over time with the down and degraded thresholds (and their recovery values) drawn across.
  • Backtest: how many times it would have gone down or degraded over the chart's range, and when.
  • Cost: what one evaluation reads, so a heavy query shows up before it runs every minute.
  • Sample alert: the SMS, email and webhook it would send, rendered from a real group's value.

What to measure ​

Each monitor has up to eight queries, a to h, over one Stream:

SourceMeasuresExample
Logs, traces, eventscount, count_distinct, sum, avg, min, max, p50 to p99, count_per_min of a field, over a searchcount where level:error
MetricsA metric read as rate, increase, avg, p95 and so on, combined across serieshttp.server.duration as p95

With more than one query, a formula combines them, for example an error rate: a / b * 100 where a counts errors and b counts all requests.

One alert per group ​

For each (group by) up to four fields to get one state, and one alert, per distinct value: service alerts separately for checkout and search. A group that stops reporting is forgotten after a day: its incident closes with a note, and no recovery alert goes out, since nothing recovered.

When it alerts ​

SettingWhat it does
Over the lastThe window each evaluation aggregates, 1 minute to 1 day.
EveryHow often it evaluates (the plan sets a 1 minute floor).
DelayEvaluate this far behind now, so late data has landed first.
Down / DegradedThe thresholds, above or below (optionally "or equal"). Degraded is optional and sits before down.
Recover atStay at a level until the value passes back through this point, so a value hovering at the line doesn't flap.
Alert / recover afterEvaluations in a row that must cross (or clear) before the state changes.

No data ​

What a group with no data does is up to you: keep its last state (the default), treat it as OK, evaluate it as 0, or go down after a silence you choose. Counts and sums are already 0 over an empty window, so a silent group reads 0 for them either way.

Heartbeats ​

A heartbeat alerts when something that should keep arriving stops: a nightly job, a cron, a queue consumer. Pick Heartbeat, a search that matches the job's "done" record, and how often it should appear. The monitor goes down when a window passes without one. With For each host, every host has its own heartbeat.

Composite monitors ​

A composite is a condition over other monitors, down while it holds:

checkout_errors && (latency || !payments_up)

Each reference is true while its monitor is down (or degraded, if you count degraded too). and, or and not work as well as &&, || and !. A composite can combine check and telemetry monitors, and lives with the kind of page it was created on. Use one to page only when a symptom and its likely cause agree. It's checked once a minute, so Down after and Recover after are the minutes the condition must hold, or stay clear, in a row.

Notifications ​

Every monitor, check or telemetry, shares the same alert settings:

  • Priority, P1 (most urgent) to P5. P1 and P2 open major incidents. A contact method can take only P1, or P1 to P3, and so on.
  • Alert on degraded, and when it recovers. A recovery alert goes out when a monitor that alerted is healthy again. Without degraded alerts, a drop from down to a warning stays part of the same incident until it's healthy. With them, that drop is its own alert ("no longer alerting"), and going back to healthy from a warning sends a recovery too.
  • While down, alert once, or every 15 minutes to every day, at most a number of times.
  • Mute a monitor, or one group, for an hour to a week or until unmuted. States and incidents keep updating; only the alerts are held back.

Message templates ​

Write your own alert text with variables in double braces. It replaces the default wording on every channel.

{{#is_recovery}}Resolved: {{/is_recovery}}{{monitor.name}} for {{group.service}} is {{state}}: {{value}} {{unit}} (limit {{threshold}})
Variable
monitor.name, monitor.url, monitor.priorityThe monitor, a link to it, and its priority
state, previous_stateDOWN, DEGRADED or UP, and the state before
causeWhy, in words
value, threshold, unitThe evaluated value and the threshold it crossed
group, group.<label>The group (service=checkout), or one of its labels
targetThe host a check monitor watches
duration, atHow long it was (or has been) down, and when

Sections show text only when a condition holds, and {{^...}} when it doesn't: is_down, is_degraded, is_recovery, is_renotify, is_nodata, is_error.

Webhook payload ​

Webhooks receive JSON with the fields a receiver needs to route and dedupe:

json
{
  "kind": "incident.down",
  "event": "down",
  "message": "Checkout errors is alerting for service=checkout: 67 > 40",
  "monitorId": "mon_...",
  "monitorName": "Checkout errors",
  "monitorKind": "TELEMETRY",
  "priority": 2,
  "state": "DOWN",
  "previousState": "DEGRADED",
  "cause": "67 > 40",
  "group": { "key": "service=checkout", "labels": { "service": "checkout" } },
  "value": 67,
  "threshold": 40,
  "incidentId": "inc_...",
  "url": "https://console.latency.app/...",
  "at": "2026-09-28T23:42:25.000Z"
}

kind is incident.down, incident.recovered, incident.renotify, monitor.warning (degraded) or monitor.error (the monitor couldn't be evaluated).

Monitor detail ​

A monitor's page shows its value against the thresholds (per group, if you pick groups to chart), each group's state and how far it is into confirming or recovering, the run log with what each evaluation read, its incidents, and runs live as they happen.

Limits ​

PlanTelemetry monitors
Hobby5
Pro, Founders25
Team200
Enterprise1,000

Composites count toward the kind of page they were created on.

API ​

Everything the console does is in the API (and the CLI):

GET /v1/telemetry/monitorsList telemetry monitors
POST /v1/telemetry/monitors/create, /updateCreate or change one
POST /v1/telemetry/monitors/previewEvaluate, chart and backtest a definition without saving
GET /v1/monitors/groups, /runsA monitor's groups, and its evaluations per group
POST /v1/monitors/muteMute a monitor or one group
POST /v1/monitors/createComposite, /previewCompositeComposites
POST /v1/monitors/previewAlert, GET /v1/monitors/templateHelpRender a sample alert, and the template variables
© 2026 Latency Labs LLC