Solo Operator Kickoff Prompt & A/B Benchmark Protocol
Purpose: Standardized, deterministic prompt template and execution protocol for solo-operator runs and reproducible A/B evaluation benchmarks across independent agent swarms.
| Specification Metadata | Value |
|---|---|
| Framework Version | CEF v0.1.0 |
| Governance Tier | Authoritative Core (Kickoff Template) |
| Target Roles | Solo Operators, Evaluation Evaluators, Benchmark Harnesses |
| Determinism Guarantee | Identical baseline prompt ensures objective comparison across model generations |
1. Standardized Solo Operator Prompt
Copy and paste the following prompt verbatim when executing an evaluation or initializing an agent session. Only alter the declared agent_run_id and output_home parameters:
You are executing the Codebase Evaluation Framework (CEF) v0.1.0 as a SOLO OPERATOR on this repository.
## Identity (FILL-IN — only field that may differ between A/B runs)
- agent_run_id: AGENT_A # or AGENT_B
- output_home: docs/quality/cef-runs/2026-08-13-AGENT_A
# AGENT_B must use a different directory, e.g. docs/quality/cef-runs/2026-08-13-AGENT_B
## Mandatory first actions
1. Read, in order:
- docs/quality/codebase_evaluation/CONSTITUTION.md
- docs/quality/codebase_evaluation/DIAMOND_SCALE.md
- docs/quality/codebase_evaluation/OPERATOR.md
- docs/quality/codebase_evaluation/WAVE_PLAN.md
- docs/quality/codebase_evaluation/DIAGRAM_CONTRACT.md
- docs/quality/codebase_evaluation/HANDOFF_SCHEMA.md
- docs/quality/codebase_evaluation/prompts/_SHARED_PREAMBLE.md
2. Create output_home. Copy
docs/quality/codebase_evaluation/templates/run_scope.example.yaml
→ <output_home>/run_scope.yaml
3. Freeze run_scope.yaml EXACTLY as follows (do not widen/narrow without writing a waive in run_log.md):
cef_version: "0.1.0"
mode: truth_map
repo_root: "."
include_globs:
- "**/*"
exclude_globs:
- ".git/**"
- "vendor/**"
- "node_modules/**"
- "dist/**"
- "build/**"
- "**/*.min.js"
- "**/*.pb.go"
- "docs/quality/cef-runs/**"
- "docs/quality/codebase_evaluation/**"
languages_detected: []
adapters_enabled: ["go"]
top_n_default: 10
output_home: "<your output_home from Identity>"
notes: "A/B CEF compare run; truth map only; no launch triage."
## Mission
Produce a cold truth-map package for this codebase using ALL waves in WAVE_PLAN.md through Wave 4 (Integrator).
## Hard rules
- Do NOT modify application source code.
- Do NOT mint process/backlog/tracker objects.
- Do NOT bias toward product launch / open-core / near-term roadmaps.
- Prefer existing tooling; record missing tools in tooling_gaps.md.
- Tests only if absolutely necessary: narrow package + named test + timeout ≤60s; never full suites.
- Evidence grades E0–E3; accepted handoff findings must be E3 (or E2 with explicit waive in run_log.md).
- Density: D-HIGH exhaustive within scope; D-LOW Top N=10; D-MED per rubric.
- Every analysis lens: specialist pass THEN a separate adversarial pass attacking the specialist artifact (you may play both seats, but keep distinct artifacts).
- Write ALL artifacts under output_home only — never edit CEF framework files under docs/quality/codebase_evaluation/.
## Required package layout under output_home
run_scope.yaml
preflight.json
tooling_gaps.md
run_log.md
diagrams/**
findings/<lens_id>.jsonl
adversarial/<lens_id>.jsonl
adversarial_resolutions.jsonl
scorecard.json
findings.jsonl
handoff_manifest.json
EXECUTIVE_NARRATIVE.md # ≤200 lines, cold
## Wave order (do not skip)
Wave 0 L-PREFLIGHT
Wave 1 L-ARCHITECTURE, L-SECURITY, L-SUPPLY-RELEASE
Wave 2 L-RELIABILITY, L-OBSERVABILITY, L-CONCURRENCY, L-PERFORMANCE
Wave 3 L-CODE-QUALITY, L-TESTING, L-USABILITY, L-DOCS-MODEL
Wave 4 Integrator (prompts/L-INTEGRATOR/specialist.md)
For each lens use:
- rubrics/<lens>.md
- prompts/<lens>/specialist.md
- prompts/<lens>/adversarial.md
Findings must validate against schemas/finding.schema.json.
Scorecard against schemas/scorecard.schema.json.
Diagrams against DIAGRAM_CONTRACT.md + schemas/diagram.schema.json.
## Done criteria
1. scorecard.json has all eight required axes (RDB MNT TST REL OBS RCV SEC ROB) with grade/confidence/drivers.
2. findings.jsonl has no E0; every accepted finding has adversarial resolution.
3. run_log.md lists commands run, timeouts, skips, and any waives.
4. Reply with: path to output_home, axis grade vector, count of findings by severity, and top 10 standing findings (id + title).
Begin with Wave 0 now.
2. A/B Benchmark Evaluation Protocol
When comparing evaluation efficacy between models, agent frameworks, or prompt updates:
- Deterministic Initialization: Initialize Agent A with declared identity
AGENT_Aand target directorydocs/quality/cef-runs/<date>-AGENT_A. - Independent Replication: Initialize Agent B with identical prompt text, changing only
agent_run_id: AGENT_Band its corresponding output directory. - Partitioned Isolation: Prevent Agent B from accessing Agent A's artifacts or findings before its evaluation run is completed.
- Comparative Analysis: Compare results across four primary vectors: - Consistency of 1–5 Diamond Scale axis scores. - Finding overlap and duplication across identical file paths. - Severity calibration and delta distribution. - Evidence-grade rigor and adherence to adversarial challenge.
3. Commit SHA Freezing Invariant
To preserve benchmark validity, both evaluation runs must execute against an identical commit hash:
# Record active commit hash before launching evaluation
git rev-parse HEAD
The resulting commit SHA must be recorded in both run_scope.yaml and run_log.md. Any tree modifications or commits during evaluation contaminate comparison validity.