Appearance
Telemetry Monitors
Beta
Telemetry monitors are in beta, on paid plans. Limits are subject to change.
A telemetry monitor watches your own data: the number of error logs, a service's p95 latency, a metric crossing a line, or a job that stops reporting. It runs every minute (or less often), keeps a state per group, and alerts the same contacts your check monitors do.
Create one under Telemetry → Monitors → New Monitor, or from anywhere you already have a query: Create Monitor in Explore, on a log pattern, or on a derived metric.
The live preview
As you edit, the editor evaluates the unsaved monitor against your real data:
- Right now: what each group would be, and the closest one to its threshold.
- The chart: the value over time with the down and degraded thresholds (and their recovery values) drawn across.
- Backtest: how many times it would have gone down or degraded over the chart's range, and when.
- Cost: what one evaluation reads, so a heavy query shows up before it runs every minute.
- Sample alert: the SMS, email and webhook it would send, rendered from a real group's value.
What to measure
Each monitor has up to eight queries, a to h, over one Stream:
| Source | Measures | Example |
|---|---|---|
| Logs, traces, events | count, count_distinct, sum, avg, min, max, p50 to p99, count_per_min of a field, over a search | count where level:error |
| Metrics | A metric read as rate, increase, avg, p95 and so on, combined across series | http.server.duration as p95 |
With more than one query, a formula combines them, for example an error rate: a / b * 100 where a counts errors and b counts all requests.
One alert per group
For each (group by) up to four fields to get one state, and one alert, per distinct value: service alerts separately for checkout and search. A group that stops reporting is forgotten after a day: its incident closes with a note, and no recovery alert goes out, since nothing recovered.
When it alerts
| Setting | What it does |
|---|---|
| Over the last | The window each evaluation aggregates, 1 minute to 1 day. |
| Every | How often it evaluates (the plan sets a 1 minute floor). |
| Delay | Evaluate this far behind now, so late data has landed first. |
| Down / Degraded | The thresholds, above or below (optionally "or equal"). Degraded is optional and sits before down. |
| Recover at | Stay at a level until the value passes back through this point, so a value hovering at the line doesn't flap. |
| Alert / recover after | Evaluations in a row that must cross (or clear) before the state changes. |
No data
What a group with no data does is up to you: keep its last state (the default), treat it as OK, evaluate it as 0, or go down after a silence you choose. Counts and sums are already 0 over an empty window, so a silent group reads 0 for them either way.
Heartbeats
A heartbeat alerts when something that should keep arriving stops: a nightly job, a cron, a queue consumer. Pick Heartbeat, a search that matches the job's "done" record, and how often it should appear. The monitor goes down when a window passes without one. With For each host, every host has its own heartbeat.
Composite monitors
A composite is a condition over other monitors, down while it holds:
checkout_errors && (latency || !payments_up)Each reference is true while its monitor is down (or degraded, if you count degraded too). and, or and not work as well as &&, || and !. A composite can combine check and telemetry monitors, and lives with the kind of page it was created on. Use one to page only when a symptom and its likely cause agree. It's checked once a minute, so Down after and Recover after are the minutes the condition must hold, or stay clear, in a row.
Notifications
Every monitor, check or telemetry, shares the same alert settings:
- Priority, P1 (most urgent) to P5. P1 and P2 open major incidents. A contact method can take only P1, or P1 to P3, and so on.
- Alert on degraded, and when it recovers. A recovery alert goes out when a monitor that alerted is healthy again. Without degraded alerts, a drop from down to a warning stays part of the same incident until it's healthy. With them, that drop is its own alert ("no longer alerting"), and going back to healthy from a warning sends a recovery too.
- While down, alert once, or every 15 minutes to every day, at most a number of times.
- Mute a monitor, or one group, for an hour to a week or until unmuted. States and incidents keep updating; only the alerts are held back.
Message templates
Write your own alert text with variables in double braces. It replaces the default wording on every channel.
{{#is_recovery}}Resolved: {{/is_recovery}}{{monitor.name}} for {{group.service}} is {{state}}: {{value}} {{unit}} (limit {{threshold}})| Variable | |
|---|---|
monitor.name, monitor.url, monitor.priority | The monitor, a link to it, and its priority |
state, previous_state | DOWN, DEGRADED or UP, and the state before |
cause | Why, in words |
value, threshold, unit | The evaluated value and the threshold it crossed |
group, group.<label> | The group (service=checkout), or one of its labels |
target | The host a check monitor watches |
duration, at | How long it was (or has been) down, and when |
Sections show text only when a condition holds, and {{^...}} when it doesn't: is_down, is_degraded, is_recovery, is_renotify, is_nodata, is_error.
Webhook payload
Webhooks receive JSON with the fields a receiver needs to route and dedupe:
json
{
"kind": "incident.down",
"event": "down",
"message": "Checkout errors is alerting for service=checkout: 67 > 40",
"monitorId": "mon_...",
"monitorName": "Checkout errors",
"monitorKind": "TELEMETRY",
"priority": 2,
"state": "DOWN",
"previousState": "DEGRADED",
"cause": "67 > 40",
"group": { "key": "service=checkout", "labels": { "service": "checkout" } },
"value": 67,
"threshold": 40,
"incidentId": "inc_...",
"url": "https://console.latency.app/...",
"at": "2026-09-28T23:42:25.000Z"
}kind is incident.down, incident.recovered, incident.renotify, monitor.warning (degraded) or monitor.error (the monitor couldn't be evaluated).
Monitor detail
A monitor's page shows its value against the thresholds (per group, if you pick groups to chart), each group's state and how far it is into confirming or recovering, the run log with what each evaluation read, its incidents, and runs live as they happen.
Limits
| Plan | Telemetry monitors |
|---|---|
| Hobby | 5 |
| Pro, Founders | 25 |
| Team | 200 |
| Enterprise | 1,000 |
Composites count toward the kind of page they were created on.
API
Everything the console does is in the API (and the CLI):
GET /v1/telemetry/monitors | List telemetry monitors |
POST /v1/telemetry/monitors/create, /update | Create or change one |
POST /v1/telemetry/monitors/preview | Evaluate, chart and backtest a definition without saving |
GET /v1/monitors/groups, /runs | A monitor's groups, and its evaluations per group |
POST /v1/monitors/mute | Mute a monitor or one group |
POST /v1/monitors/createComposite, /previewComposite | Composites |
POST /v1/monitors/previewAlert, GET /v1/monitors/templateHelp | Render a sample alert, and the template variables |