Appearance
Heartbeat Monitors
Beta
Heartbeats are telemetry monitors, in beta on paid plans.
A heartbeat turns alerting around. Instead of Latency asking your job whether it's alive, your job tells Latency each time it runs, and you're alerted when those messages stop. It catches the failures that make no noise: a cron job that never started, a backup that quietly stopped, a queue consumer that died, a scheduled deploy that didn't happen.
A beat is an ordinary event (or log line). The monitor counts the beats it finds in a window of time and goes down when there are none.
Set one up
Create a Stream and an ingest key. An Events Stream is the simplest home for beats; follow the Quick Start.
Send a beat from the job when it finishes: examples below.
Create the monitor under Telemetry → Monitors → New Monitor → Heartbeat:
Field Example Meaning Name Nightly backupWhat the alert says Expect name:backup.doneThe Stream, and a search that matches the beat Over the last 1 dayHow long it may stay silent before it goes down Every 15 minutesHow often it checks Then pick who to notify and a priority, and click Create.
The editor describes what you set up in words ("At least one matching event every 1 day. It goes down when one doesn't arrive.") and checks it against your real data as you go. Send a first beat before you save: with nothing in the window a heartbeat is down at once, and the preview says so.
Sending a beat
Send the beat from the job itself, after it succeeds. A separate pinger that only proves a server is up can't tell you the job ran.
Every field you send can be searched, and the body is the same JSON the HTTP API takes. name is the event name; the rest is yours. Add what you'd want to know when it's late, like a job name, a host, a duration_ms or a row count.
bash
# crontab: run the backup at 02:00, and beat only if it succeeded
0 2 * * * /usr/local/bin/backup.sh && /usr/local/bin/beat nightly-backup
# /usr/local/bin/beat
#!/bin/sh
curl -fsS -m 10 --retry 3 -X POST https://ws.latency.app/v1/ingest \
-H "Authorization: Bearer lit_XXXXXXXXXXX" \
-H "Content-Type: application/json" \
-d "{\"name\":\"backup.done\",\"job\":\"$1\"}" > /dev/nullini
# backup.service: ExecStartPost runs only if ExecStart succeeded
[Service]
Type=oneshot
ExecStart=/usr/local/bin/backup.sh
ExecStartPost=/usr/bin/curl -fsS -m 10 --retry 3 -X POST https://ws.latency.app/v1/ingest -H "Authorization: Bearer lit_XXXXXXXXXXX" -H "Content-Type: application/json" -d '{"name":"backup.done","job":"nightly-backup"}'js
// At the end of a successful run. A missed beat is the monitor's job to
// notice, so a failed request shouldn't fail the run.
await fetch('https://ws.latency.app/v1/ingest', {
method: 'POST',
headers: { Authorization: `Bearer ${process.env.LATENCY_KEY}`, 'Content-Type': 'application/json' },
body: JSON.stringify({ name: 'import.done', job: 'nightly-import', rows: 1204 }),
signal: AbortSignal.timeout(10_000),
}).catch(() => {});python
import json, os, urllib.request
req = urllib.request.Request(
"https://ws.latency.app/v1/ingest",
data=json.dumps({"name": "import.done", "job": "nightly-import", "rows": 1204}).encode(),
headers={"Authorization": f"Bearer {os.environ['LATENCY_KEY']}", "Content-Type": "application/json"},
)
urllib.request.urlopen(req, timeout=10)yaml
on:
schedule:
- cron: "0 2 * * *"
jobs:
sync:
runs-on: ubuntu-latest
steps:
- run: ./scripts/sync.sh
# Steps stop at the first failure, so this only runs after a good sync.
- name: Tell Latency it ran
run: |
curl -fsS -m 10 --retry 3 -X POST https://ws.latency.app/v1/ingest \
-H "Authorization: Bearer ${{ secrets.LATENCY_KEY }}" \
-H "Content-Type: application/json" \
-d '{"name":"sync.done","job":"nightly-sync"}'yaml
apiVersion: batch/v1
kind: CronJob
metadata:
name: nightly-report
spec:
schedule: "0 2 * * *"
jobTemplate:
spec:
template:
spec:
restartPolicy: OnFailure
containers:
- name: report
image: registry.example.com/report:latest # needs curl
command: ["/bin/sh", "-c"]
args:
- >
/app/report &&
curl -fsS -m 10 --retry 3 -X POST https://ws.latency.app/v1/ingest
-H "Authorization: Bearer $LATENCY_KEY"
-H "Content-Type: application/json"
-d '{"name":"report.done","job":"nightly-report"}'
env:
- name: LATENCY_KEY
valueFrom:
secretKeyRef: { name: latency-ingest, key: key }-f makes curl fail on an error response, --retry 3 retries a 429 or 503 for you, and -m 10 stops a stuck request from holding the job open. Leave the timestamp out: the beat is stamped when it arrives. See the HTTP API for keys, batching and responses.
Choosing the window
Over the last is how long the job may stay silent before the monitor goes down, so it has to be longer than the gap between beats. Take the next size up from how often the job runs:
| The job runs | Over the last | Every |
|---|---|---|
| every minute | 5 minutes | 1 minute |
| every 5 minutes | 15 minutes | 1 minute |
| every 15 or 30 minutes | 1 hour | 5 minutes |
| hourly | 4 hours | 15 minutes |
| every few hours | 12 hours | 15 minutes |
| daily | 1 day | 15 minutes |
The longest window is 1 day, so a heartbeat suits jobs that run at least daily. For a rarer job, have a daily check beat while the last run is recent enough.
A window as long as the job's own period alerts as soon as a run is a little late. To allow for that, raise Alert after (in a row): each extra check gives a late run one more Every. A daily job checked every 15 minutes with Alert after 4 is given about 45 minutes before it alerts.
Examples
A nightly backup
The job beats {"name":"backup.done","job":"nightly-backup"} once a night.
| Expect | name:backup.done |
| Over the last | 1 day |
| Every | 15 minutes |
| Alert after | 4 |
A job that runs but fails
A beat that only says "I ran" misses a job that runs and fails. Beat either way and put the result in the beat:
bash
if /usr/local/bin/backup.sh; then result=ok; else result=failed; fi
curl -fsS -m 10 --retry 3 -X POST https://ws.latency.app/v1/ingest \
-H "Authorization: Bearer lit_XXXXXXXXXXX" \
-H "Content-Type: application/json" \
-d "{\"name\":\"backup.done\",\"job\":\"nightly-backup\",\"result\":\"$result\"}"Expect name:backup.done AND result:ok. A failed run isn't a beat, so the monitor goes down once the window passes. To hear about a failure straight away, add a threshold monitor that counts name:backup.done AND result:failed and alerts above 0.
A fleet, one heartbeat per host
Each host beats once a minute: {"name":"agent.alive","host":"web-1"}.
| Expect | name:agent.alive |
| For each | host |
| Over the last | 5 minutes |
| Every | 1 minute |
Every host gets its own state and its own alert, and a new host is watched from its first beat.
A job that already logs
On an OpenTelemetry Stream a heartbeat reads logs or traces, so a job that already writes a line when it finishes needs no new code. Pick Logs and search for that line, for example service:billing AND "invoices sent".
When it goes down
When a window passes with no matching beat, the monitor goes down, an incident opens and your contacts are alerted, with the priority, repeats and message you set (see Notifications). The next beat brings it back up and sends the recovery alert.
- A heartbeat with no beat yet is down. A silent window counts as zero, so send a first beat before you create the monitor.
- One monitor per job stays down until its next beat, however long that takes. That's the right shape for a job you can't miss.
- With For each, a group is only known once it has sent a beat, so a host that never reported isn't watched. A group silent for a day is forgotten: its incident closes with a note and no recovery alert goes out, which is what you want for a host you retired.
- Mute the monitor, or one group, for planned downtime. The state keeps updating; only the alerts are held.
What it costs
Each beat is one stored event on your plan's monthly events. A beat every minute is about 43,000 a month, so beat every few minutes (or once per run) unless you need the speed.
Create one with the API
bash
curl -X POST https://ws.latency.app/v1/telemetry/monitors/create \
-H "Authorization: Bearer $LATENCY_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"namespace": "prod",
"name": "Nightly backup",
"intervalSeconds": 900,
"failureLimit": 4,
"target": {
"kind": "query",
"streamId": "strm_...",
"queries": [{ "id": "a", "source": "events", "aggFn": "count", "where": "name:backup.done" }],
"windowSeconds": 86400,
"thresholds": { "down": { "op": "lt", "value": 1 } },
"noData": { "action": "alert", "afterSeconds": 600 }
}
}'A heartbeat is a count that must not drop below 1 in the window. Add groupBy (["host"]) for one per host, and groupExpirySeconds to change how long a silent group is kept (a day by default). Authenticate with a token; the Telemetry Monitors page lists the rest of the monitor endpoints.
Not alerting when it should?
- Is the beat arriving? Open Explore and search
stream:Backups name:backup.done. A202from the endpoint means it was stored;401is the key and429is a limit. - Does the search match? The monitor's preview shows how many beats it found in the window. Match on the field you send:
name:backup.done, not the text of the command. - Is the Stream sampled? A Stream that samples can drop a beat. The editor warns when it does; send heartbeats to a Stream that doesn't.