Idea → deployed feature pipeline. Temporal orchestrates; Pydantic AI agents
think (clarify, architect, plan, QA, quality gate, devops); coding harnesses
do (claude -p, opencode run) inside isolated git worktrees.
Governed by the versioned registry in agents/ (agents/registry.yaml +
one agents/<role>/ folder per role) and validated at worker boot
(src/sdlc/agents/loader.py) — 3 harness + 8 required proposer + 4 optional
proposer roles.
| Role | Kind | Runs as |
|---|---|---|
| clarify | Pydantic AI (TemporalAgent) | activity via TemporalAgent |
| architect | Pydantic AI | activity via TemporalAgent |
| planner | Pydantic AI | activity via TemporalAgent |
| dev / test / devops executor | coding harness | long-running heartbeating activity in a git worktree |
| qa analyst | Pydantic AI + test-suite activity | activities |
| quality gate | DeterministicQualityGate (pure code) + advisory MergeVerdict (Pydantic AI, soft-gate only) |
evaluate_gate activity + TemporalAgent |
| reviewer | coding harness (different model/harness than dev) | activity |
| analyst | Pydantic AI, clean-context | activity via TemporalAgent |
research (optional, research_enabled) |
Pydantic AI, fans out (plan_research → research_subquestion × N → synthesize_brief) |
activities; provider fake (CI) / tavily / exa (ExaSearch + Harness run_code, needs EXA_API_KEY) |
deep_review (optional, deep_review_enabled) |
Pydantic AI, reads the scrubbed harness transcript | activity via TemporalAgent, advisory only |
handoff (optional, handoff_enabled, FR-805) |
Pydantic AI, extracts task→task claims from the scrubbed session | activity via TemporalAgent, best-effort |
adversary (optional, adversarial_review_enabled) |
Pydantic AI, decorrelated second opinion (different model identity than dev+reviewer) on the approving path | activity via TemporalAgent, advisory, fail-open |
Gates: clarify, architecture, plan, merge, deploy — each hard / soft /
off per project (PipelineConfig.gates). Humans interact through signals:
python -m sdlc.cli start --title "Add SSO" --mode brownfield --repo git@...
python -m sdlc.cli status --id feature-add-sso
python -m sdlc.cli answer --id feature-add-sso --q Q1 --text "Use OIDC"
python -m sdlc.cli approve --id feature-add-sso --gate architecture
python -m sdlc.cli benchmark --case cat-cafe # run the eval harness (see docs/BENCHMARK.md)
Scoring stored benchmark runs needs no running Temporal (it reads records on disk):
# score everything on disk (seconds, no Temporal needed)
python -m sdlc.cli benchmark score --all
# one matrix run, re-weighted
python -m sdlc.cli benchmark score --bench <bench_run_id> --weights 0.7,0.2,0.1
# one case across its whole history
python -m sdlc.cli benchmark score --case cat-cafe-monitoring
Local:
temporal server start-devpip install -e .thenpython -m sdlc.workerpython -m sdlc.cli start ...
Docker Compose (docker-compose.yml): brings up Temporal, a real
Hindsight memory backend, and the
worker together, all reading secrets from .env (env_file:) — copy
.env.example to .env first. The worker image installs the logfire extra
so logfire_setup.configure() doesn't crash-loop on a missing module when
LOGFIRE_TOKEN is set. docker compose up.
Agent board API. Optional, read-mostly service over the board the pipeline
writes as it runs ($SDLC_BOARD_DB, default runs/board.sqlite3):
uvicorn interfaces.dashboard.api.main:app --host 127.0.0.1 --port 8500GET /projects/{p} for artifacts + task rollup, /artifacts/{key} for version
lineage, /tasks?status=, /events for the change log, /stats for board
counters. Agents claim work with POST /projects/{p}/tasks/{id}/claim and an
If-Match: <row_version> header. Bind to localhost — there is no auth yet,
and the X-Actor header identifying a writer is self-asserted (ROADMAP OQ-11).
Deploy (stage 13). Off by default. Enable per project with
PipelineConfig.deploy — adapter: compose (reference) or script
(make deploy / make rollback / make version). The stage applies a
frozen DeployPlan, runs its smoke checks, and auto-rolls-back on any check
that is not passed, then opens a deploy_failed gate. A check that could
not be evaluated is errored and never counts as a pass.
pip install -e .[dev]thenpython -m pytest(needsgiton PATH).- Importing the workflow/agents currently requires
ANTHROPIC_API_KEY/OPENAI_API_KEY/EXA_API_KEYset (agents are constructed at import, including the shipped research role'sprovider: exaExaSearch client);tests/conftest.pysets dummy values for import-only, sopytestneeds no real keys. - Added a new module and hit
ModuleNotFoundError? Re-runpip install -e .(setuptools' editable wheel doesn't auto-discover new files). - Prompt changes are gated (E-82):
SDLC_PROMPT_EVAL=1 python -m pytest -m prompt_evalA/B-scores each changedagents/<role>/instructions.mdagainst its committed baseline via promptfoo (pip install -e .[eval]; needs Node ≥ 22.22; spends tokens). Deterministic checks — output validates as the role'soutput_type, per-case rubric vetoes (benchmarks/cases/<case>/vetoes-*.yaml), cost/latency budgets — are absolute and gate; the cross-family judge is staged (rubric → evaluation steps → score) and advisory, failing only on a regression past a noise-aware floor. Sensitivity is proven by the mutation suite:SDLC_PROMPT_EVAL=1 python -m pytest -m prompt_eval -k mutations. Ad hoc:python -m sdlc.cli eval clarify --case add-login-greenfield --gate. Results land inruns/prompt_evals/and join the benchmark record stream byprompt_shaonly — they are never merged into the heatmap, matrices, or SC rollup. - See
docs/foundation.mdfor the contracts, activities, and the deterministic gate, anddocs/architecture-review-2026-07.mdfor the design decisions and implementation status. - See
docs/BENCHMARK.mdfor the benchmark & evaluation design — the four measurement axes (harness / model×role / memory / case), how success criteria SC-1..6 get their numbers, and the E-30…E-37 increments. - Self-contained schema docs (no build step, open directly in a browser),
each checked against actual code:
docs/roadmap.html(every FR/NFR/SC/US/ADR + the 15-stage DAG vs code),docs/architecture-schema.html,docs/agents-schema.html(registry lifecycle, every role, ADR-6/adversary model-inequality checks),docs/research-stage-schema.html(the research fan-out stage, provider seam, ExaSearch wiring),docs/benchmark.html/docs/benchmark-analysis.html.
- Payloads through Temporal stay small (claim-check for specs/diffs/logs).
- Agent names / toolset ids are activity names — never rename in prod.
- Harness sessions are resumed across fix-loop attempts (claude
--resume, opencode-s), so the fixer keeps its context. - Every harness run also emits a canonical
HarnessSession— a normalised, scrubbed, claim-checked transcript (ADR-16) — so how a diff was reached is a first-class signal, not just the diff. The default reviewer never reads it; the opt-indeep_reviewlens does (ADR-6). - The agent board (ADR-21) persists what only Temporal history used to
hold:
requirements/architecture/planversioned per project with lineage, plus task status and an append-only change log. Two status columns — the workflow writesauthoritative_status, agents may only move the livestatus. Stats and scoring read the former, so a confused agent corrupts the live view and nothing else, and replay stays the source of truth. - Cross-harness review: configure
roles["reviewer"]with a different harness/model family thanroles["dev"]. - Harness is a config axis, not a fork:
claude -pandopencode runare registry entries (HARNESSES), so a third adapter (e.g.cursor) drops in once it normalises intoHarnessRunResult. The benchmark sweeps this axis. - Hard gates that time out notify (log/webhook adapters) on reminder, escalation, and expiry timers rather than silently auto-rejecting (FR-303); a green run holds for a human rather than being discarded.
- A decorrelated adversary lens (
agents/adversary/) runs only on the approving path as a non-DAG, fail-open second opinion — decorrelated by model identity (model_id()), not provider prefix, so two prefixes over the same weights don't count as independent. Off by default. - Every terminal run emits a
RunSummary(retro stage) and exportsevents.jsonl/report.html/summary.json;sdlc benchmark scoreaggregates across runs into the SC-rollup + heatmap/task/error/waste matrices, plus anagreement_matrixfor the adversary lens. - Memory (Hindsight) defaults to a fake in-process backend; the real client
(
memory/hindsight_client.py) talks to a live Hindsight container (see Docker Compose above).