← All heartbeat examples · What Heartbeat is
Heartbeat for background workers and job queues
The web tier is the part everyone monitors and the part that rarely dies. The worker is the part that
dies quietly: killed for memory, left stopped after a deploy that only restarted the web, cut off from
Redis, or alive and wedged on one poison job. Jobs pile up, emails stop, exports never finish, and
/up stays green the whole time. Users notice
first. A heartbeat from inside the job system notices in two minutes.
1. The pattern: a one-minute job whose only work is the ping
Schedule a job every minute that does nothing but GET
the slot's URL. It runs only when two things are true: the scheduler is alive, and a worker for that queue
picked the job up and finished it. If either stops, the pings stop, and after two silent minutes Fresh Jots
emails you (and posts the JSON webhook, if you set one). Below, the same job in four job systems; the
heartbeat() helper for each language is on the
project code page.
Solid Queue (Rails 8)
The recurring entry is read by the scheduler process; a dead scheduler or a dead worker both mean silence.
# app/jobs/heartbeat_job.rb
class HeartbeatJob < ApplicationJob
queue_as :critical
def perform = Heartbeat.ping
end
# config/recurring.yml
production:
heartbeat:
class: HeartbeatJob
queue: critical
schedule: "* * * * *"
Sidekiq (sidekiq-cron)
sidekiq-cron loads config/schedule.yml on its own and
recognises an Active Job class by itself; active_job: true
just says so explicitly.
# config/schedule.yml
heartbeat:
cron: "* * * * *"
class: "HeartbeatJob"
queue: critical
active_job: true
Celery
The task goes into beat_schedule with the queue under
options; the beat process
(celery -A celery_app beat) has to be running, which is
the point: the slot now watches beat as well as the workers.
# celery_app.py
import os, urllib.request, logging
from celery import Celery
app = Celery("celery_app", broker=os.environ["CELERY_BROKER_URL"])
@app.task
def heartbeat():
url = os.environ.get("HEARTBEAT_URL")
if not url:
return
try:
urllib.request.urlopen(url, timeout=10).close()
except Exception as e: # never let monitoring break the work
logging.warning("heartbeat ping failed: %s", e)
app.conf.beat_schedule = {
"heartbeat": {
"task": "celery_app.heartbeat",
"schedule": 60.0,
"options": {"queue": "critical"},
},
}
BullMQ (Node)
A job scheduler on the queue produces the job every minute; the worker that drains the queue handles it beside the real jobs.
import { Queue, Worker } from "bullmq";
const connection = { host: "127.0.0.1", port: 6379 }; // your Redis
const queue = new Queue("critical", { connection });
await queue.upsertJobScheduler("heartbeat", { every: 60_000 }, { name: "heartbeat" });
new Worker("critical", async (job) => {
if (job.name === "heartbeat") return heartbeat(); // the helper from the project code page
return handle(job);
}, { connection });
2. Silence means dead or drowning
Put the heartbeat job on the queue you care about, not on a quiet one of its own. Now a backlog shows up
too: when the critical queue is more than two minutes deep, the heartbeat job waits its turn behind the jobs
already queued, the ping is late, and you get the same report as for a dead worker. For that to hold, give
the heartbeat job the same priority as the real jobs on that queue; a higher priority lets it jump the line
and hides the backlog. One slot tells you the queue is
either not being served or not keeping up, and the ON
line that follows tells you when it caught up.
One slot per queue you would want a page about: critical, mailers, default. Give each slot the queue's name as its label, and the Heartbeats page becomes a queue status board that costs nothing to run.
3. Hand-written consumers: Kafka, SQS, a poll loop
Without a scheduler the ping goes inside the loop, after each successful pass. Ping after an empty poll too: the loop is doing its job when there is nothing to do, and skipping the ping there makes a quiet queue look dead at three in the morning. A consumer that is connected but stuck inside a batch never reaches the ping, which is the failure this catches.
while True:
messages = receive(wait_seconds=20) # returns [] on an empty queue
for message in messages:
handle(message)
acknowledge(message)
heartbeat() # after every pass, empty or not
If one pass can legitimately take longer than two minutes (a huge batch, a slow upstream), ping between items instead of between passes, or split the batch. The threshold is fixed at two minutes; the loop has to fit it.
4. Several worker processes
Four workers on one queue sharing one slot mean "at least one worker is serving this queue", which is usually the right question for a queue. When each process matters (one per host, one per tenant), give each its own slot and URL, with the host name in the label. Either way, say which guarantee the slot gives in its description, so the person reading the alert at night knows what it means.
5. What the record shows
A deploy that restarts the workers in under two minutes writes nothing. A worker that was down from
03:17 to 03:41 is an OFF /
ON pair in a note nobody can edit, and the
Settings → Uptime page turns those pairs into a 7- or 30-day score per queue. When the question is
"how often is the mailer queue actually stuck", the answer is already written down.
Point the slot's webhook at your on-call channel so a dead queue pages the same way a dead host does; the payload and headers are under what you get back.