RB-LCK-001: Concurrency Lock Contention and Deadlock Resolution
Metadata
- Severity: P1 / High
- Component:
pkg/concurrency,pkg/storage/locknames - Owner: Runtime & Concurrency Stewards
Symptoms & Alerts
- Log warning:
lock held longer than threshold: lock=<LOCK_NAME> duration=>100ms - Operations stalling or timing out with
context deadline exceededduring storage transactions. - Goroutine dump showing multiple threads waiting on
sync.(*Mutex).Lockinconcurrency.RunInLockWithLogger.
Root Cause Analysis
- Inverted lock acquisition order between object-level locks and store-wide index locks.
- Long-running I/O or external network calls executed while holding an in-memory mutex.
- Panic occurring inside a critical section where mutex unlock was not deferred properly.
Step-by-Step Remediation Procedure
Step 1: Capture Goroutine Stack Trace
Dump active goroutines to locate blocking locks:
./bin/zqk scheduler dump
Inspect generated diagnostics under .zqk/scheduler/diagnostics/ for mutex contention.
Step 2: Identify Contended Lock Names
Filter the recent daemon logs for slow lock acquisitions:
grep "lock held longer than threshold" .zqk/state/logs/scheduler.log | tail -n 20
Note the specific LockName* reported.
Step 3: Clear Stale Lockfiles (if filesystem-based)
If locks are backed by file locks (flock) under .zqk/locks/:
-
Check for stale lock holder PIDs:
bash lsof +D .zqk/locks/ -
If the owning process has terminated ungracefully:
bash rm -f .zqk/locks/*.lock
Step 4: Verification Gate
Run concurrent read/write validation:
go test -race -v -run TestAdversarialConcurrency ./pkg/storage
Verify zero deadlocks and zero data race warnings.