# Solo Operator Kickoff Prompt & A/B Benchmark Protocol

> **Purpose:** Standardized, deterministic prompt template and execution protocol for solo-operator runs and reproducible A/B evaluation benchmarks across independent agent swarms.

| Specification Metadata | Value |
| :--- | :--- |
| **Framework Version** | CEF v0.1.0 |
| **Governance Tier** | Authoritative Core (Kickoff Template) |
| **Target Roles** | Solo Operators, Evaluation Evaluators, Benchmark Harnesses |
| **Determinism Guarantee** | Identical baseline prompt ensures objective comparison across model generations |

---

## 1. Standardized Solo Operator Prompt

Copy and paste the following prompt verbatim when executing an evaluation or initializing an agent session. Only alter the declared `agent_run_id` and `output_home` parameters:

```text
You are executing the Codebase Evaluation Framework (CEF) v0.1.0 as a SOLO OPERATOR on this repository.

## Identity (FILL-IN — only field that may differ between A/B runs)
- agent_run_id: AGENT_A   # or AGENT_B
- output_home: docs/quality/cef-runs/2026-08-13-AGENT_A
  # AGENT_B must use a different directory, e.g. docs/quality/cef-runs/2026-08-13-AGENT_B

## Mandatory first actions
1. Read, in order:
   - docs/quality/codebase_evaluation/CONSTITUTION.md
   - docs/quality/codebase_evaluation/DIAMOND_SCALE.md
   - docs/quality/codebase_evaluation/OPERATOR.md
   - docs/quality/codebase_evaluation/WAVE_PLAN.md
   - docs/quality/codebase_evaluation/DIAGRAM_CONTRACT.md
   - docs/quality/codebase_evaluation/HANDOFF_SCHEMA.md
   - docs/quality/codebase_evaluation/prompts/_SHARED_PREAMBLE.md
2. Create output_home. Copy
   docs/quality/codebase_evaluation/templates/run_scope.example.yaml
   → <output_home>/run_scope.yaml
3. Freeze run_scope.yaml EXACTLY as follows (do not widen/narrow without writing a waive in run_log.md):

cef_version: "0.1.0"
mode: truth_map
repo_root: "."
include_globs:
  - "**/*"
exclude_globs:
  - ".git/**"
  - "vendor/**"
  - "node_modules/**"
  - "dist/**"
  - "build/**"
  - "**/*.min.js"
  - "**/*.pb.go"
  - "docs/quality/cef-runs/**"
  - "docs/quality/codebase_evaluation/**"
languages_detected: []
adapters_enabled: ["go"]
top_n_default: 10
output_home: "<your output_home from Identity>"
notes: "A/B CEF compare run; truth map only; no launch triage."

## Mission
Produce a cold truth-map package for this codebase using ALL waves in WAVE_PLAN.md through Wave 4 (Integrator).

## Hard rules
- Do NOT modify application source code.
- Do NOT mint process/backlog/tracker objects.
- Do NOT bias toward product launch / open-core / near-term roadmaps.
- Prefer existing tooling; record missing tools in tooling_gaps.md.
- Tests only if absolutely necessary: narrow package + named test + timeout ≤60s; never full suites.
- Evidence grades E0–E3; accepted handoff findings must be E3 (or E2 with explicit waive in run_log.md).
- Density: D-HIGH exhaustive within scope; D-LOW Top N=10; D-MED per rubric.
- Every analysis lens: specialist pass THEN a separate adversarial pass attacking the specialist artifact (you may play both seats, but keep distinct artifacts).
- Write ALL artifacts under output_home only — never edit CEF framework files under docs/quality/codebase_evaluation/.

## Required package layout under output_home
run_scope.yaml
preflight.json
tooling_gaps.md
run_log.md
diagrams/**
findings/<lens_id>.jsonl
adversarial/<lens_id>.jsonl
adversarial_resolutions.jsonl
scorecard.json
findings.jsonl
handoff_manifest.json
EXECUTIVE_NARRATIVE.md   # ≤200 lines, cold

## Wave order (do not skip)
Wave 0 L-PREFLIGHT
Wave 1 L-ARCHITECTURE, L-SECURITY, L-SUPPLY-RELEASE
Wave 2 L-RELIABILITY, L-OBSERVABILITY, L-CONCURRENCY, L-PERFORMANCE
Wave 3 L-CODE-QUALITY, L-TESTING, L-USABILITY, L-DOCS-MODEL
Wave 4 Integrator (prompts/L-INTEGRATOR/specialist.md)

For each lens use:
- rubrics/<lens>.md
- prompts/<lens>/specialist.md
- prompts/<lens>/adversarial.md
Findings must validate against schemas/finding.schema.json.
Scorecard against schemas/scorecard.schema.json.
Diagrams against DIAGRAM_CONTRACT.md + schemas/diagram.schema.json.

## Done criteria
1. scorecard.json has all eight required axes (RDB MNT TST REL OBS RCV SEC ROB) with grade/confidence/drivers.
2. findings.jsonl has no E0; every accepted finding has adversarial resolution.
3. run_log.md lists commands run, timeouts, skips, and any waives.
4. Reply with: path to output_home, axis grade vector, count of findings by severity, and top 10 standing findings (id + title).

Begin with Wave 0 now.
```

---

## 2. A/B Benchmark Evaluation Protocol

When comparing evaluation efficacy between models, agent frameworks, or prompt updates:

1. **Deterministic Initialization:** Initialize Agent A with declared identity `AGENT_A` and target directory `docs/quality/cef-runs/<date>-AGENT_A`.
2. **Independent Replication:** Initialize Agent B with identical prompt text, changing only `agent_run_id: AGENT_B` and its corresponding output directory.
3. **Partitioned Isolation:** Prevent Agent B from accessing Agent A's artifacts or findings before its evaluation run is completed.
4. **Comparative Analysis:** Compare results across four primary vectors:
   - Consistency of 1–5 Diamond Scale axis scores.
   - Finding overlap and duplication across identical file paths.
   - Severity calibration and delta distribution.
   - Evidence-grade rigor and adherence to adversarial challenge.

---

## 3. Commit SHA Freezing Invariant

To preserve benchmark validity, both evaluation runs must execute against an identical commit hash:

```bash
# Record active commit hash before launching evaluation
git rev-parse HEAD
```

The resulting commit SHA must be recorded in both `run_scope.yaml` and `run_log.md`. Any tree modifications or commits during evaluation contaminate comparison validity.
