A worker offline
This runbook helps you bring back the workers when no worker takes jobs.
Symptom
Section titled “Symptom”- The steering report says that no worker was seen in the last 10 minutes.
/healthz/readyreports the workers check as degraded while jobs wait.
What shows it
Section titled “What shows it”casebox_workers_seenis 0.casebox_jobs_queuedrises.
- Check the worker process. Kubernetes:
kubectl get pods -l app.kubernetes.io/component=workerand the pod log. A host: the service that runscasebox worker. - Read the first lines of the worker’s log. They say what the worker can run: the analysis model, the sandbox provider and the agents. A line that ends with a CBX code names the fix.
- A worker that cannot reach the server logs
CBX012. CheckCASEBOX_SERVERand the network path to port 8080. - A worker whose token was revoked logs
CBX010. Issue a new token (casebox token create --kind worker --name <name>) and updateCASEBOX_WORKER_TOKEN. - When the worker runs but takes no job, compare its kinds with the queued jobs’ kinds (see A stuck job).