Green Process, Stuck Work: How to Tell Whether a Local Worker Is Actually Healthy
A worker can exist, its container can be marked healthy, and its queue can still be going nowhere. That is not necessarily a contradiction: process status, readiness to accept work, and evidence of useful progress answer different questions. This guide lays out those distinctions, shows which signals help diagnose a stalled queue worker, and offers a low-risk workflow for local development debugging. The goal is not to add more checks for their own sake, but to make each health signal answer a clear operational question.
In practice, that distinction matters most when a worker looks ordinary from the outside. The process exists. Its container is running. A health status might even be green. Yet the work that matters - jobs completing, state changing, a local request getting the expected result - has stopped moving.
That situation is easy to misread because "healthy" sounds like a single verdict. In a real system, it can refer to several different things: a process responds to a check, a service can accept work, or the work itself is making progress. Those are related signals, but they are not interchangeable.
Separating them gives local debugging a better starting point. Instead of restarting processes until the symptom disappears, first establish what the health signal measures, whether work is available, and whether successful work is still completing.
"Running" is not the same as "working"
A process listing answers a narrow question: is a process present? A container's status answers another narrow question: is the container running, or has a configured health command reported a particular result? Neither statement, on its own, proves that a background worker is consuming jobs successfully.
It helps to separate three conditions:
Liveness: Is the process sufficiently responsive to keep running, or should it be restarted?
Readiness: Is the worker initialized and able to accept the work it is responsible for?
Progress: Is useful work actually advancing and completing?
Docker Compose lets a service define a healthcheck: a command used to determine whether its container is healthy. Docker's documentation describes options such as the check interval, timeout, retries, and start period. Those settings control when and how the configured test runs; the meaning of "healthy" still depends on what the command checks. A command that only confirms a process exists can pass while job execution is stuck.
Kubernetes documents a similar separation through startup, readiness, and liveness probes. A startup probe gives an application time to initialize; readiness indicates whether it should receive traffic; and liveness can cause a container to be restarted when it is considered unhealthy. These probes serve distinct purposes. They are not a general promise that a queue worker is completing business-level tasks.
So when a dashboard is green and a queue is quiet, the useful question is not "Which status is wrong?" It is "Which condition did this status actually measure?"
Choose signals that reveal useful work
For a worker, progress is usually easier to understand when it is tied to work the system expects to happen. A few signals can help build that picture:
Queue activity: How many jobs are waiting, active, completed, or failing? A worker may be idle because there is nothing to process. If jobs are waiting but completions are not advancing, that is a different symptom.
Last successful completion: When did the worker last finish a job successfully? This is often more informative than the time of the last log line, but it must be read in context: a quiet queue can make an old completion time perfectly normal.
Heartbeat or lock renewal: Does the queue system have evidence that the worker still owns an active job? Where available, this helps distinguish an active process from an active job execution.
Errors and retries: Is work being picked up and repeatedly failing, or is it not being picked up at all? Those point toward different parts of the execution path.
Log timestamps: Are relevant events fresh? Logs provide context, but silence does not prove a worker is broken, and frequent output does not prove that jobs finish.
BullMQ's documentation offers a concrete example of why process existence and job progress differ. When a worker starts a job, BullMQ places a lock on it; the worker must periodically renew that lock. If the lock is not maintained, BullMQ can mark the job as stalled and return it for processing. The important diagnostic idea is broader than any one queue library: a process can still exist while the system's evidence of an active job has expired.
Prometheus recommends instrumenting libraries, subsystems, and services so their performance can be understood. Applied to a local worker, that means choosing measurements that explain behavior: work entering, work completing, errors, and relevant worker state. A compact set of signals with clear meanings is more useful than a large dashboard of counters that do not help explain the symptom.
Build health checks around explicit questions
Before adding or editing a check, write down the question it is meant to answer. Is this check about whether the process responds? Whether the worker can accept new jobs? Whether a job that has already started is still progressing? If one check tries to answer all three, it is likely to be confusing in both directions: it may report trouble when the worker is behaving as designed, or report health while useful work has stopped.
Keep liveness focused on the failure it is meant to recover from. If a temporary queue connection problem makes liveness fail immediately, the restart policy may repeatedly recycle a process that would otherwise reconnect. That may be the intended recovery strategy in some systems, but it should be a deliberate choice rather than an accidental side effect of a broad health command.
Readiness has a different job. It should represent whether the worker is initialized and able to receive the work assigned to it, taking account of dependencies it actually needs. A worker that is not ready should not be treated as proof that an already-running job has stalled; nor does readiness establish that completed jobs are arriving.
Progress needs its own evidence. Depending on the queue and application, that could include queue depth, completed-job activity, the time of the last successful completion, failure or retry state, and heartbeat or lock status. Interpret those signals together. An old "last completed" timestamp means little if the queue is empty; it matters more when jobs are waiting and the worker reports no recent completions.
For Docker, inspect the actual health check command and its failure conditions. For Kubernetes, use startup, readiness, and liveness probes according to their documented roles. In both cases, name the gap: if you need to know whether jobs are completing, a process-oriented health check is not enough.
A low-risk local debugging workflow
When a local worker appears stuck, preserve the evidence before changing the environment. A short, repeatable sequence makes it easier to distinguish a real failure from an idle worker or a misleading status.
Record the current state. Note which worker is affected, the process or container status, whether jobs are waiting, and what the queue reports for active, completed, or failed work. Avoid restarting first if doing so will erase useful context.
Check recent worker output. Look for initialization, dependency connections, job pickup, completion, and errors. Pay attention to timestamps, but do not use log volume as the health verdict.
Compare work available with work completed. If no jobs are waiting, a lack of new completions may be expected. If jobs are waiting and completion evidence is absent, inspect whether the worker is receiving jobs and what happens after pickup.
Inspect active-job evidence. If the queue exposes heartbeats, leases, or locks, check whether an active job is still being renewed. For BullMQ, its stalled-job documentation explains how a missing lock renewal affects job state.
Change one thing at a time. After preserving state, try a reversible action on the affected worker rather than resetting the entire local stack. Then check the same signals again: did the worker become ready, did it acquire work, and did successful completions resume?
Use caution with destructive actions such as clearing a queue or deleting local state. They can remove the very evidence needed to understand whether the problem was job acquisition, execution, a dependency, or recovery behavior. A targeted restart may be reasonable after capturing useful information, but it should test a hypothesis - not substitute for one.
Interpret symptoms before choosing a fix
Different combinations of signals suggest different places to look:
The process is missing or unresponsive: Inspect startup output, runtime errors, and the conditions used by its liveness check. Preserve the relevant failure details before restarting.
The process is present but not ready: Check initialization and the dependencies required before it can accept work. This is a readiness question, not yet proof that job execution has stalled.
Jobs are waiting, but completions are not advancing: Follow the path from job acquisition to execution and completion. Inspect errors, retries, and active-job heartbeat or lock renewal where available.
Completions look normal, but logs are quiet: Do not restart just to create activity. Check whether the worker is simply idle and whether the observed completion rate fits the work available.
A healthcheck is green, but the symptom remains: Read the check itself. It may be doing exactly what it was configured to do while measuring a narrower condition than the one you care about.
Prometheus's alerting guidance recommends keeping alerts simple, alerting on symptoms, and giving people a way to pinpoint causes. That principle also helps with local tooling. A useful warning should indicate a condition that needs attention and provide enough context to investigate it. A noisy warning with no next step becomes another status developers learn to ignore.
Make the next diagnosis easier
Good observability is not about collecting everything. It is about making the worker's state legible when something stops behaving as expected. Give workers identifiable names, and keep lifecycle events, job activity, and failures distinguishable in logs or metrics. When possible, make it easy to see queue state alongside the most recent successful work and active-job signals.
Document what each health indicator means. "Healthy" should be traceable to a check: process responsiveness, initialization, dependency availability, job acquisition, or completed work. That small bit of precision prevents a green status from being mistaken for a broader guarantee.
Revisit checks when worker responsibilities change. A command that once verified a meaningful condition can become shallow or misleading after a new dependency, execution mode, or initialization step is introduced. Treat health checks as code with assumptions that need maintenance.
Finally, avoid borrowing universal thresholds without evidence. The documentation cited here explains healthchecks, probes, instrumentation, alerting, and job locks; it does not define one correct number of idle minutes or missed heartbeats for every local worker. Set expectations from the workload and recovery behavior of the system being debugged.
Conclusion
A green process is one useful signal, not a verdict on whether a worker is healthy. Separate liveness, readiness, and progress; pair queue activity with successful completions and heartbeat or lock evidence; and treat log freshness as context rather than proof. With those distinctions in place, local debugging becomes a sequence of low-risk checks instead of a guess-and-restart loop. The best health design makes clear what is working, what is not, and what to inspect next.