Skip to content

← All heartbeat examples · What Heartbeat is

Heartbeat for background workers and job queues

The web tier is the part everyone monitors and the part that rarely dies. The worker is the part that dies quietly: killed for memory, left stopped after a deploy that only restarted the web, cut off from Redis, or alive and wedged on one poison job. Jobs pile up, emails stop, exports never finish, and /up stays green the whole time. Users notice first. A heartbeat from inside the job system notices in two minutes.

1. The pattern: a one-minute job whose only work is the ping

Schedule a job every minute that does nothing but GET the slot's URL. It runs only when two things are true: the scheduler is alive, and a worker for that queue picked the job up and finished it. If either stops, the pings stop, and after two silent minutes Fresh Jots emails you (and posts the JSON webhook, if you set one). Below, the same job in four job systems; the heartbeat() helper for each language is on the project code page.

Solid Queue (Rails 8)

The recurring entry is read by the scheduler process; a dead scheduler or a dead worker both mean silence.

# app/jobs/heartbeat_job.rb
class HeartbeatJob < ApplicationJob
  queue_as :critical
  def perform = Heartbeat.ping
end

# config/recurring.yml
production:
  heartbeat:
    class: HeartbeatJob
    queue: critical
    schedule: "* * * * *"

Sidekiq (sidekiq-cron)

sidekiq-cron loads config/schedule.yml on its own and recognises an Active Job class by itself; active_job: true just says so explicitly.

# config/schedule.yml
heartbeat:
  cron: "* * * * *"
  class: "HeartbeatJob"
  queue: critical
  active_job: true

Celery

The task goes into beat_schedule with the queue under options; the beat process (celery -A celery_app beat) has to be running, which is the point: the slot now watches beat as well as the workers.

# celery_app.py
import os, urllib.request, logging
from celery import Celery

app = Celery("celery_app", broker=os.environ["CELERY_BROKER_URL"])

@app.task
def heartbeat():
    url = os.environ.get("HEARTBEAT_URL")
    if not url:
        return
    try:
        urllib.request.urlopen(url, timeout=10).close()
    except Exception as e:  # never let monitoring break the work
        logging.warning("heartbeat ping failed: %s", e)

app.conf.beat_schedule = {
    "heartbeat": {
        "task": "celery_app.heartbeat",
        "schedule": 60.0,
        "options": {"queue": "critical"},
    },
}

BullMQ (Node)

A job scheduler on the queue produces the job every minute; the worker that drains the queue handles it beside the real jobs.

import { Queue, Worker } from "bullmq";

const connection = { host: "127.0.0.1", port: 6379 };   // your Redis
const queue = new Queue("critical", { connection });
await queue.upsertJobScheduler("heartbeat", { every: 60_000 }, { name: "heartbeat" });

new Worker("critical", async (job) => {
  if (job.name === "heartbeat") return heartbeat();   // the helper from the project code page
  return handle(job);
}, { connection });

2. Silence means dead or drowning

Put the heartbeat job on the queue you care about, not on a quiet one of its own. Now a backlog shows up too: when the critical queue is more than two minutes deep, the heartbeat job waits its turn behind the jobs already queued, the ping is late, and you get the same report as for a dead worker. For that to hold, give the heartbeat job the same priority as the real jobs on that queue; a higher priority lets it jump the line and hides the backlog. One slot tells you the queue is either not being served or not keeping up, and the ON line that follows tells you when it caught up.

One slot per queue you would want a page about: critical, mailers, default. Give each slot the queue's name as its label, and the Heartbeats page becomes a queue status board that costs nothing to run.

3. Hand-written consumers: Kafka, SQS, a poll loop

Without a scheduler the ping goes inside the loop, after each successful pass. Ping after an empty poll too: the loop is doing its job when there is nothing to do, and skipping the ping there makes a quiet queue look dead at three in the morning. A consumer that is connected but stuck inside a batch never reaches the ping, which is the failure this catches.

while True:
    messages = receive(wait_seconds=20)      # returns [] on an empty queue
    for message in messages:
        handle(message)
        acknowledge(message)
    heartbeat()                              # after every pass, empty or not

If one pass can legitimately take longer than two minutes (a huge batch, a slow upstream), ping between items instead of between passes, or split the batch. The threshold is fixed at two minutes; the loop has to fit it.

4. Several worker processes

Four workers on one queue sharing one slot mean "at least one worker is serving this queue", which is usually the right question for a queue. When each process matters (one per host, one per tenant), give each its own slot and URL, with the host name in the label. Either way, say which guarantee the slot gives in its description, so the person reading the alert at night knows what it means.

5. What the record shows

A deploy that restarts the workers in under two minutes writes nothing. A worker that was down from 03:17 to 03:41 is an OFF / ON pair in a note nobody can edit, and the Settings → Uptime page turns those pairs into a 7- or 30-day score per queue. When the question is "how often is the mailer queue actually stuck", the answer is already written down.

Point the slot's webhook at your on-call channel so a dead queue pages the same way a dead host does; the payload and headers are under what you get back.