RB-SCH-001: Scheduler Daemon Exit Cascades and Job Triage
Metadata
- Severity: P1 / High
- Component:
cmd/zqk/scheduler,pkg/scheduler - Owner: Automation Stewards
Symptoms & Alerts
- Error message:
scheduler daemon is not runningorscheduler exited with code 1 - Warning:
materialized view watermark exceeds toleranceduringwhats-nextqueries. - Periodic jobs (
SCH-retention-tolerance,SCH-val,SCH-objcount-report) not updatinglast_run_at.
Root Cause Analysis
- Zombie or stale PID file at
.zqk/state/scheduler.pidpointing to dead process. - Missing or malformed config in
.zqk/specs/configs/scheduler_maintenance_config.yaml. - Unhandled panic in a background job runner that escaped the recovery wrapper.
- Missing host execution permissions on scheduler job scripts (exit code 126/127).
Step-by-Step Remediation Procedure
Step 1: Check Scheduler Status & PID
Inspect the live status:
./bin/zqk scheduler status
If reported running but unresponsive, inspect the PID file:
cat .zqk/state/scheduler.pid
ps -p $(cat .zqk/state/scheduler.pid 2>/dev/null)
Step 2: Clear Stale PID File
If the process does not exist but the PID file remains:
rm -f .zqk/state/scheduler.pid
Step 3: Check Scheduler Logs
Examine the last 50 lines of daemon output:
tail -n 50 .zqk/state/logs/scheduler.log
Check for panics, syntax errors in scripts, or missing dependencies.
Step 4: Ensure Survival Jobs & Restart
Ensure all mandatory survival jobs are configured:
./bin/zqk system ensure-retention-jobs
Restart the scheduler daemon:
./bin/zqk scheduler start
Verify the daemon is alive:
./bin/zqk scheduler status
./bin/zqk scheduler list
Step 5: Verification Gate
Trigger an immediate test job:
./bin/zqk scheduler trigger SCH-val
./bin/zqk scheduler status
Ensure scheduler daemon status reports Running and job completes with status succeeded.