pi-harness — a self-healing, measurable, provider-agnostic coding-agent harness that keeps you out of AI-vendor lock-in.
A self-healing, measurable coding-agent harness built around the Pi coding agent and the open-source DeepEval LLM evaluation framework. The harness runtime is driven by a single Go CLI, pi-run — the one source of truth for provider routing, API key resolution, Pi launching, evaluation, setup, and health checks.
What this project is, and what it is not — see
CHARTER.mdfor the boundary contract. One product: the harness. The eval suite is its measurement layer (the "measurable" in the star), not a separate product today; a standalone benchmark repo (pi-bench) is a triggered future split, not current scope. We do not build a PM system, a spec library, an observability platform, or a general MCP platform as product surface.
pi-runCLI — a compiled Go binary (repobin/pi-run) that owns the harness runtime:chat,print,resume,cost,ci-benchmark,eval,config-check,doctor,setup,install,clean,project-understand,self-heal,providers,hooks,version.- Pi CLI installed globally (via nvm) and configured for this project.
- Curated Pi packages installed project-locally:
pi-mcp-adapter— token-efficient MCP adapter (Pi-side; distinct from the harness'smcp-servercommand, which was removed in v0.10.0).pi-web-access— web fetch/search for agents.@demigodmode/pi-web-agent— reliable web search/fetch with explicit boundaries.@loreai/pi— Lore memory engine.pi-spark— daily-experience polish.dot-pi(from GitHub) — curated extensions, skills, prompts, and rules.
@zigai/pi-ui-tweakswas removed because its bundled settings schema is currently incompatible with this Pi version.
- Project context files:
AGENTS.md,.pi/SYSTEM.md,.pi/APPEND_SYSTEM.md. - DeepEval environment in
eval/.venvwith sample tests and datasets. Python deps live ineval/requirements.txt(DeepEval + pytest stack,~=-bounded;pytest-json-reportis retained, but the nightly live-eval report is written by a conftest hook viaPI_EVAL_REPORT, not by pytest-json-report). Add new dependencies deliberately;pi-run setupinstalls them. - Automation via the
pi-runCLI (no Makefile, no shell functions).
- Node.js via
nvm(pi-runselects the highest nvm-installed semantic version; override withPI_NODE_VERSION). - Python 3.11+ (for the DeepEval suite).
- Go 1.26+ (only to build/update
pi-run). - An API key. The harness is provider-agnostic: it ships with a data-driven
provider table (
providers.json) covering OpenAI (default), OpenRouter, DeepSeek, Anthropic, Gemini, Groq, local OpenAI-compatible endpoints (Ollama/vLLM), Azure OpenAI, and 9 more OpenAI-/Anthropic-compatible cloud providers (Mistral, Cohere, Together, Perplexity, Fireworks, Moonshot, xAI, AWS Bedrock, Ollama) — 17 providers in total. Keys are resolved env-first, then from an optional secret store (BW_GEToverride; Bitwarden is a documented example). See API Key Resolution andpi-run providers.
| provider | key (env var / secret-manager item) | pi-run --provider |
default model | baseURL (OpenAI- or Anthropic-compatible) |
|---|---|---|---|---|
| OpenAI (default) | OPENAI_API_KEY |
openai |
openai/gpt-5.6-terra |
— |
| OpenRouter | OPENROUTER_API_KEY |
openrouter |
openai/gpt-5.6-terra |
— |
| DeepSeek (direct) | DEEPSEEK_API_KEY |
deepseek |
deepseek/deepseek-v4-flash |
— |
| Anthropic | ANTHROPIC_API_KEY |
anthropic |
anthropic/claude-sonnet-4 |
— |
| Gemini | GEMINI_API_KEY |
gemini |
gemini/gemini-2.5-pro |
— |
| Groq | GROQ_API_KEY |
groq |
groq/llama-3.3-70b-versatile |
— |
| Local (Ollama/vLLM) | LOCAL_API_KEY |
local |
local/model |
http://localhost:11434/v1 |
| Azure OpenAI | AZURE_OPENAI_API_KEY |
azure |
azure/gpt-5.6-terra |
https://<your-resource>.openai.azure.com/openai/v1 |
| Ollama (local) | OLLAMA_API_KEY |
ollama |
ollama/llama3.1 |
http://localhost:11434/v1 |
| Mistral | MISTRAL_API_KEY |
mistral |
mistral/mistral-large-latest |
https://api.mistral.ai/v1 |
| Cohere | COHERE_API_KEY |
cohere |
cohere/command-r-plus |
https://api.cohere.com/compatibility/v1 |
| Together | TOGETHER_API_KEY |
together |
together/llama-3.3-70b-instruct |
https://api.together.xyz/v1 |
| Perplexity | PERPLEXITY_API_KEY |
perplexity |
perplexity/sonar-pro |
https://api.perplexity.ai |
| Fireworks | FIREWORKS_API_KEY |
fireworks |
fireworks/llama-3.3-70b-instruct |
https://api.fireworks.ai/inference/v1 |
| Moonshot (Kimi) | MOONSHOT_API_KEY |
moonshot |
moonshot/kimi-k2 |
https://api.moonshot.cn/v1 |
| xAI (Grok) | XAI_API_KEY |
xai |
xai/grok-4 |
https://api.x.ai/v1 |
| AWS Bedrock | BEDROCK_API_KEY |
bedrock |
bedrock/claude-sonnet-4 |
https://bedrock-runtime.<region>.amazonaws.com/anthropic/v1 |
Add a provider without recompiling: edit
providers.json(add a row: name, key env var, pi provider, default model, optionalbaseURL), then runpi-run providersto verify it lists. Installed binaries use the embedded 17-provider table; setPI_RUN_PROVIDERS_FILEto load a provider table from another path. Providers with abaseURLare routed through pi's OpenAI-compatible (openai) or Anthropic-compatible (anthropic) provider; entries likeazureandbedrockneed your resource/region filled in. There is no automatic cross-provider fallback — the provider is explicit (--provider/PI_PROVIDER).
macOS with Homebrew (fastest path):
brew install forrestbthomas/tap/pi-run
pi-run config-check # no API key needed for thisFrom source (macOS/Linux/WSL):
# 1. Clone the repo
git clone https://github.com/forrestbthomas/pi-harness.git
cd pi-harness
# 2. One-command bootstrap (Node + pi + bin/pi-run + eval/.venv)
bash scripts/bootstrap.sh
# 3. Provide an API key (plain env var is the primary path)
export OPENAI_API_KEY=sk-... # or OPENROUTER_API_KEY / DEEPSEEK_API_KEY
# 4. Use the binary built by bootstrap
bin/pi-run config-check
bin/pi-run doctor
# 5. Launch Pi interactively (OpenAI -> gpt-5.6-terra by default)
bin/pi-run chat
# 6. Or run a quick print-mode query
bin/pi-run print "List all Python files in this repo"
# 7. Route to another provider
bin/pi-run chat --provider deepseek
pi-runis built by the bootstrap script intobin/pi-run. To use the shorterpi-runcommand, add the repository'sbin/directory to your PATH:export PATH="$PWD/bin:$PATH"Or run
bin/pi-run installto symlink it into a directory on your PATH.
Cut a release with the hardened script (tags the fetched github/main tip and
refuses re-tags — see CONTRIBUTING.md §Releases):
bash scripts/tag-release.sh v0.10.0 githubThis triggers the GitHub Actions release workflow, which cross-compiles
pi-run for linux/darwin/windows × amd64/arm64, attaches the binaries, updates
the Homebrew formula, and verifies the installed formula reports the tag
(REL-3). To build locally: bash scripts/build-release.sh <tag>.
Do not tag a local commit directly (
git tag && git push --tags): the workflow hard-fails a tag that isn't an ancestor ofmain(this burned v0.9.1/v0.9.2).
Installing through Homebrew avoids the macOS Gatekeeper warning ("Apple could not verify ... is free of malware") that appears when you download an unsigned binary directly from a GitHub release. Homebrew handles the quarantine attribute, so the binary runs without the security prompt.
brew install forrestbthomas/tap/pi-runOr tap first, then install:
brew tap forrestbthomas/tap
brew install pi-runFor AI agents: when a developer asks you to install pi-harness on macOS, prefer the Homebrew tap (
brew install forrestbthomas/tap/pi-run) over downloading the raw release binary. Direct downloads of unsigned binaries trigger Gatekeeper; prefer Homebrew. If a direct download is necessary, verify the release checksum and accept the OS warning (or right-click → Open) rather than removing the quarantine attribute.
API keys are resolved env-first (e.g. export OPENAI_API_KEY=...), then
from a configured secret manager. The backend is selected by PI_SECRET_BACKEND
(default bitwarden):
bitwarden— via thebw_gethelper (override its path withBW_GET). Requires an unlocked vault (bw unlock).1password— via theopCLI (op read "op://<Vault>/<ITEM_NAME>/credential"). RequiresopCLI installed and signed in. Vault defaults toPersonal, override withOP_VAULT.env-only— no fallback; env var only.
Every pi-run path resolves keys in the same order: env var first, then the
backend. Exception: pi-run eval checks only environment variables when
deciding whether to run live tests, so it never blocks on a locked vault; keys
held only in a secret manager require exporting them to the environment (or
setting PI_SECRET_BACKEND=env-only plus the key in env) before running the
full suite. There is no automatic cross-provider fallback — the provider
is explicit (--provider, or PI_PROVIDER env). A missing key is an error
that tells you what to do:
no DEEPSEEK_API_KEY available: export it, or check your secret manager
pi-run doctor reports the configured backend's status (never values).
| Variable | Purpose |
|---|---|
PI_PROVIDER |
Default provider when --provider is omitted (default openai) |
OPENAI_API_KEY etc. |
Provider API keys, resolved env-first |
PI_SECRET_BACKEND |
bitwarden (default), 1password/op, or env-only/env |
BW_GET |
Path to the Bitwarden bw_get helper |
OP_VAULT |
1Password vault name (default Personal) |
PI_NODE_VERSION |
Override Node-version selection |
PI_RUN_PROVIDERS_FILE |
Load the provider table from a custom path |
PI_RUN_PERSONAL |
Opt into personal-machine checks such as the ~/bin/pi-run symlink |
HARNESS_ROOT |
Override repository-root detection |
PI_MAX_BUDGET_USD |
Default spend cap for chat/print when --max-budget-usd is omitted |
PI_PERMISSION_MODE |
Default permission tier for chat/print when --permission-mode is omitted |
PI_MODEL_TIER |
Default model tier (fast/balanced/cheap) for chat/print when --model-tier is omitted; ignored by resume |
DEEPEVAL_MODEL |
Select a non-OpenAI DeepEval judge model |
PI_SELF_HEAL |
Set 1 to record watchdog stall/group-kill/recovery events to .pi/heal/events.jsonl (scorecard observability) |
PI_STALL_TIMEOUT_SECS |
Watchdog silent-window: terminate a non-interactive run with no stdout after N seconds (default 300; 0 disables) |
PI_WATCHDOG_GRACE_SECS |
Process-group-kill grace between SIGTERM and SIGKILL (default 10; 0 = immediate) |
EVAL_RUNS_PER_CASE |
Live-suite agent runs per case (nightly sets 5; the flake-aware gate, EVAL-2) |
EVAL_JUDGE_RUNS |
LLM-judge repeats per case, pass = majority (default 3; EVAL-8 judge stabilization) |
PI_EVAL_REPORT |
Path (relative to eval/) where the conftest hook writes the pytest report for score_run.py |
OPENAI_MODEL_NAME |
Judge model pin for the LLM-judged metrics (nightly sets gpt-4.1-mini) |
# Smoke tests that do not require an API key
pi-run eval --quick
# Harness config checks (provider defaults, skills, dotfiles) - no API key
pi-run config-check
# Full DeepEval suite (requires a provider key)
pi-run evalThe DeepEval judge (the LLM-as-a-judge used by metrics like
AnswerRelevancy/Faithfulness/Hallucination) defaults to OpenAI. To evaluate
without depending on OpenAI, set a non-OpenAI provider key and DEEPEVAL_MODEL:
export OPENROUTER_API_KEY=<YOUR_OPENROUTER_KEY> # or another provider key
export DEEPEVAL_MODEL=openrouter/anthropic/claude-sonnet-4
pi-run evalpi-run eval --quick and the config tests run without any key.
pi-run can detect and recover wedged agent runs so a hang never needs a human
nudge (spec docs/governance/specs-archive/2026-08-13-self-healing-design.md):
pi-run self-heal— detect in-progress git state (a wedged rebase) and recover it:GIT_EDITOR=true git rebase --continuewhen conflicts are resolved, or reportneeds-attentionwith conflict paths (never guesses resolutions).--abortis explicit-only.- Non-interactive child env — every launch injects
GIT_EDITOR=true,GIT_SEQUENCE_EDITOR=true,GIT_TERMINAL_PROMPT=0,PAGER=catso child bash tools can never block on an interactive editor/pager (#59). - Process-group kill — timed-out runs kill the whole process tree
(SIGTERM → 10s grace → SIGKILL), not just the direct
pichild. - Output-stall watchdog — non-interactive runs (print/benchmark) that emit
no output for
PI_STALL_TIMEOUT_SECS(default 300) are terminated; chat is never auto-killed. - Escalation packet — killed runs write
.pi/heal/<timestamp>-report.json(goal, side-effect ledger, pending state, trigger evidence, resume handle) and exit9(watchdog terminated). - Observability — set
PI_SELF_HEAL=1to log stall/group-kill/recovery events to.pi/heal/events.jsonl(setPI_SELF_HEAL=1to enable; deterministic tests leave it off).
The nightly workflow (.github/workflows/nightly-live-eval.yml) evaluates the
agent against the live dataset (eval/datasets/coding_samples.jsonl —
currently 54 cases across 7 categories; the manifest eval/datasets/tasks.json
is the count authority) with a two-job split:
- Deterministic job — runs the hermetic suite (config checks, contract
tests against a fresh
pi-runbinary, dataset schema lint, scorer unit tests); no provider key needed. - Live job — runs each case 5× (
EVAL_RUNS_PER_CASE) viapi-run print --model-tier cheapplus LLM-judged metrics (TaskCompletionMetric, G-Eval rubric; judge majority-of-3 viaEVAL_JUDGE_RUNS), then gates per-case pass rates and cost-per-task againsteval/baselines/live-baseline.json(0.05 tolerance; flake-aware cost gate: a single run >2× baseline is a reported cost flake, never a failure — a median >2× or ≥2 over-threshold runs fails, per EVAL-13) witheval/scripts/score_run.py. Missing provider key is a hard failure, never a silent skip. Budget-capped viaPI_MAX_BUDGET_USD(default $2/night); results artifact-retained 90 days.
Judge model is pinned via OPENAI_MODEL_NAME (deepeval reads that knob, not
DEEPEVAL_MODEL). Re-baselining is deliberate: run the suite green locally,
then commit a new baseline via score_run.py --update-baseline --allow.
pi-run eval --benchmark runs the same coding tasks in Docker-isolated
containers against any provider — see which model actually solves your tasks.
# Validate all task formats (hermetic: no Docker, no API keys) — CI-safe
pi-run eval --benchmark-dry-run
# Run the full benchmark suite against the default provider (requires Docker + a key)
pi-run eval --benchmark
# Run one task, routed to another provider/model
pi-run eval --benchmark fix-divide-by-zero --provider deepseek --model deepseek/deepseek-v4-flashEach task lives in eval/benchmarks/<name>/ and ships a task.json plus a
tests/run.sh verification script (exit 0 = pass):
eval/benchmarks/fix-divide-by-zero/
├── task.json # id, prompt, testScript, timeoutSecs, solution, category, difficulty, grader
├── environment/Dockerfile # optional; default base is python:3.12-slim
├── src/ # task workspace the agent edits
├── tests/run.sh # exit 0 = pass, anything else = fail
└── solution/ # optional oracle (for future diff grading)
The agent edits a local workspace (copied from src/, or cloned from repo);
only verification runs in the container, against the same files the agent
edited. Results print per-task pass/fail with timing and an aggregate score,
and a JSON report is written to eval/benchmark-results/<run-id>.json
(gitignored). Benchmarks require Docker — --benchmark-dry-run is the
hermetic format-validation path for CI. See
docs/benchmarks.md for how to add your own tasks.
pi-run selects the provider (--provider / PI_PROVIDER; default openai),
resolves the key, and launches pi with the right --provider / --model
flags. chat/print launch pi with --offline by default so startup network
ops (version check, changelog, catalog refresh) never hang on the flaky pi.dev
endpoint — the stored model catalogs are used instead; pi-run setup is the
explicit online path. Everything else on the command line is passed through to
pi unchanged.
pi-run chat|print --model-tier fast|balanced|cheap (env PI_MODEL_TIER)
picks a model within the explicitly selected provider. Design law: tier
selection never changes the provider and never silently falls back —
an unknown or unmapped tier is an exit-2 usage error that lists the valid or
available tiers.
--model-tier balanced(default when omitted) = the provider'sdefaultModel.--model-tierand--modelas flags are mutually exclusive (exit 2).- Env
PI_MODEL_TIER+ explicit--model→--modelwins (flag beats env default; an exported env never breaks existing--modelinvocations). resumerejects the flag and ignores the env (a resumed session keeps its model).pi-run providersshows the available tiers per provider, andpi-run config-checkvalidatesmodelTiersin providers.json.
Example:
pi-run print --provider openai --model-tier cheap "summarize this repo"
PI_MODEL_TIER=fast pi-run print "quick pass"
/model— pick a model interactively in-session (OpenAI GPT models are listed first via theenabledModelsorder in.pi/settings.json).Ctrl+P— cycle through the enabled model palette.--model— override the provider default, e.g.pi-run print --provider deepseek --model deepseek/deepseek-v4-pro "...",pi-run print --provider openrouter --model anthropic/claude-sonnet-4 "...".openrouter/auto— let OpenRouter pick the best model for the task (pi-run chat --provider openrouter --model openrouter/auto).
chat/print accept a harness-level permission tier (--permission-mode,
env PI_PERMISSION_MODE). Pi has no native permission-mode flag, so each tier
maps to Pi's real tool-control surface:
pi-run chat --permission-mode plan # read-only: --tools read,grep,find,ls
pi-run chat --read-only # alias for --permission-mode plan
pi-run chat --permission-mode acceptEdits # Pi defaults (file edits allowed)
pi-run chat --permission-mode bypassPermissions # --approve (trust project-local files)
PI_PERMISSION_MODE=plan pi-run chat # or the env varValid modes mirror the Claude Code set — default, plan, acceptEdits,
bypassPermissions (unknown modes are usage errors, exit 2). Policy: the
worker agent runs under default/acceptEdits for implementation; the
reviewer and scout agents run under plan/--read-only (see
.pi/agents/).
The default model (openai/gpt-5.6-terra) is pinned in .pi/settings.json;
change it there or with /model and the session persists the choice. To
refresh model catalogs (new models / pricing, including the deepseek catalog),
run pi-run setup once with network access.
pi-run reports real spend — every Pi session file records per-message
usage.cost (USD) for each provider/model used, so no price tables are needed:
pi-run cost # per-provider/model table + total
pi-run cost --json # machine-readable
pi-run cost --since 2026-08-01 # only sessions modified at/after <date>
pi-run cost --reset # archive the spend ledger, start a fresh periodcost scans .pi/sessions/*.jsonl (including subagent child sessions) and
sums usage.cost.total, grouped by provider/model, counting how many session
files each group appears in. Messages without usage.cost are skipped; if a
provider reports no cost, it is simply not counted.
Budget cap — refuse to launch before spend crosses a limit:
pi-run chat --max-budget-usd 5.00
pi-run print --max-budget-usd 5.00 "expensive task"
PI_MAX_BUDGET_USD=5.00 pi-run chat # or the env varBefore launching, pi-run computes cumulative spend (session files + the
append-only ledger .pi/cost-ledger.jsonl) and exits with code 6 if it is
already at/above the cap:
Cost attribution — chat/print accept --cost-mode <mode> to tag a
run's ledger entry explicitly (modes: chat, print, resume, backfill,
benchmark, live-eval; default is the command name). CI-tagged runs (the
nightly live eval uses --cost-mode live-eval) make per-surface spend
attribution unambiguous in .pi/cost-ledger.jsonl.
pi-run: print: budget exceeded: $5.001234 already spent (cap $5.000000) — raise --max-budget-usd, or start a fresh period with `pi-run cost --reset`
After each run the run's spend is appended to the ledger
({ts, provider, model, inputTokens, outputTokens, costUsd, mode}); the ledger
preserves spend even after sessions are cleaned up, and a warning is printed if
the cap is exceeded mid-run. pi-run cost --reset archives the ledger to
.pi/cost-ledger-<ts>.archive.jsonl and writes a reset marker — budget checks
then count only sessions since the marker (session files are never deleted).
The ledger and marker are gitignored.
Notes (v1 contract): the pre-flight check is best-effort — spend recorded by
parallel subagent sessions is attributed to whichever pi-run run finishes
last, and runs launched outside pi-run are counted from their session files.
Plain pi-run print runs stay one-shot (--no-session, no session file); when
--max-budget-usd is set, the print session is persisted so its spend can be
recorded in the ledger.
pi-run can run shell commands around its own invocations via .pi/hooks.json
— the harness-level counterpart to agent-internal hooks (Claude Code pre/post-
tool hooks, Copilot's errorOccurred). Useful for CI: notify a chat room when
an eval starts or finishes, upload artifacts, or gate commands on external
checks.
Supported events:
| Event | Fires |
|---|---|
pre-eval |
before the DeepEval pytest suite runs |
post-eval |
after the suite finishes — always, even when pytest fails |
pre-chat |
before pi is launched (chat/print/resume) |
Schema (.pi/hooks.json):
{
"hooks": {
"pre-eval": [{"cmd": "./scripts/ci/notify.sh start", "timeoutSecs": 60}],
"post-eval": [{"cmd": "./scripts/ci/notify.sh done", "continueOnError": true}],
"pre-chat": [{"cmd": "git status --porcelain"}]
}
}Per hook: cmd (required; runs via sh -c from the harness root),
timeoutSecs (default 30; a hung hook is killed and counts as a failure with
exit code 124), and continueOnError (default false — a failed hook aborts the
pi-run invocation with the command's exit code unless this is true).
Hooks are entirely optional: a missing .pi/hooks.json is a no-op. Inspect or
trigger them manually:
pi-run hooks list # show configured hooks
pi-run hooks run pre-eval # run an event's hooks nowpi-run ci-benchmark runs the benchmark suite against 2+ providers and
gates the build on the result — the choose → measure → control loop as a
repeatable CI artifact: choose providers, measure them with the
Benchmarks suite, and control spend and quality with explicit
gates.
# Run the suite against two providers and gate on pass rate + budget
pi-run ci-benchmark --providers openai,deepseek --fail-below 0.8 --max-budget-usd 5.0
# Compare against a previous scorecard/run JSON (regression gate, default tolerance 0.05)
pi-run ci-benchmark --providers openai,deepseek --fail-below 0.8 \
--baseline eval/benchmark-results/scorecard-latest.json| Flag | Meaning |
|---|---|
--providers <a,b> |
Comma-separated providers, ≥ 2, order-significant (e.g. openai,deepseek) |
--models <m1,m2> |
Optional per-provider model overrides (same order as --providers); defaults to each provider's defaultModel |
--fail-below <rate> |
Fail (exit 8) if any provider pass rate < <rate> (e.g. 0.8) |
--max-budget-usd <n> |
Fail (exit 6) if total run cost ≥ n (PI_MAX_BUDGET_USD also applies) |
--baseline <path> |
Previous scorecard or per-provider run JSON to diff pass rates against (file-based baseline) |
--baseline-tolerance <n> |
Max allowed per-provider pass-rate drop vs baseline (default 0.05) |
--runs <n> |
Repeat each provider suite n times; gate on the median pass rate (default 1) |
--quick-profile |
Cap per-task agent timeout at 60 s — cheap, best-effort smoke run |
Each run writes a machine-readable scorecard to
eval/benchmark-results/scorecard-<run>.json (gitignored; per-provider pass
rate, cost, latency, tokens) and prints a human-readable table. Providers run
sequentially so per-provider cost attribution stays clean; any errored task
makes the run incomplete and fails the gate. Benchmarks require Docker (exit 7
when unavailable).
Exit codes: 6 budget exceeded · 7 docker unavailable · 8 scorecard
gate failed (incomplete run, pass rate below --fail-below, or regression vs
--baseline).
Run it in CI: see
.github/workflows/provider-scorecard.yml
— a weekly-scheduled / manual GitHub Actions job that runs the scorecard across
two providers, gates on --fail-below + --max-budget-usd, and uploads the
scorecard JSON as an artifact that becomes the next run's baseline.
See docs/benchmarks.md for how to add your own benchmark tasks.
| Command | Behavior |
|---|---|
pi-run chat [flags] [prompt...] |
Launch Pi interactively (default provider: openai) |
pi-run print [flags] "<prompt>" |
One-shot pi -p --no-session |
pi-run eval [--quick] |
Run the DeepEval pytest suite (--quick = smoke subset) |
pi-run eval --benchmark [name] |
Run Docker-isolated benchmark tasks (all by default; requires Docker) |
pi-run eval --benchmark-dry-run |
Validate benchmark task formats only (no Docker, no keys) |
pi-run eval -- <pytest selector...> |
Run a focused test or pass pytest arguments through (for example tests/test_x.py::test_y) |
pi-run eval --help |
Show eval-specific usage without running pytest |
pi-run resume [flags] [prompt...] |
Continue the most recent Pi session (pi --continue) |
pi-run cost [--json] [--since <date>] [--reset] |
Aggregate real spend from Pi session files (usage.cost), per provider/model, with total |
pi-run ci-benchmark --providers <a,b> [flags] |
Provider scorecard in CI: run the benchmark suite against 2+ providers, gate on pass rate / budget / baseline |
pi-run providers |
List configured providers, default models, and available model tiers |
pi-run project-understand [--out <dir>] |
Generate deterministic project-understanding docs (product.md / tech.md / structure.md) from the checkout |
pi-run self-heal [--abort] |
Detect and recover a wedged git state (rebase --continue when clean, needs-attention with conflict paths; --abort explicit-only) |
pi-run hooks list / hooks run <event> |
List or run .pi/hooks.json hook commands |
pi-run config-check |
Deterministic harness checks (no keys, no network) |
pi-run doctor |
Health report: node, pi, vault, per-provider keys, models, venv |
pi-run setup |
Create eval/.venv, install deps, refresh model catalogs |
pi-run install |
Build bin/pi-run and symlink it onto your PATH |
pi-run clean |
Remove eval/.venv and pytest caches |
pi-run --exit-codes |
Print the stable exit-code table |
pi-run version / help |
Version / usage |
Exit codes: 0 ok · 1 generic · 2 usage · 3 missing API key · 4 node/pi not found · 5 eval venv missing · 6 budget exceeded · 7 docker unavailable (benchmarks) · 8 scorecard gate failed (ci-benchmark) · 9 watchdog terminated (stall/group-kill timeout).
Pi auto-discovers skills from ~/.agents/skills/, .pi/skills/, packages, and
project settings. Five curated collections are installed via
bash scripts/install-skills.sh (durable clones under ~/.pi/agent/skills/):
- Superpowers — planning/execution skills (brainstorming, writing-plans, executing-plans, systematic-debugging, ...)
- agent-skills — engineering skills (code-review-and-quality, test-driven-development, security-and-hardening, ...)
- scope-lock — anti-scope-creep SCOPE.md boundary contracts
- productskills — PM skills (feature-prioritization, roadmap-planning, prd-writing, ...)
- spec-coding-skills — spec-plan, spec-crlp, spec-index
Re-run bash scripts/install-skills.sh any time to refresh the clones.
Invoke a skill in-session with /skill:<name> or just describe the task — the
agent loads the matching skill automatically. enableSkillCommands is on in
.pi/settings.json. (AGENTS.md §Charter lists the same five.)
.
├── AGENTS.md # Project instructions loaded by Pi
├── .gitignore
├── go.mod # Go module github.com/forrestbthomas/pi-harness (pi-run CLI)
├── cmd/pi-run/ # CLI entry point
├── internal/cli/ # CLI implementation + unit tests
├── bin/ # Git-ignored build output (bin/pi-run)
├── README.md
├── scripts/
│ └── install-skills.sh # Durable skill install into ~/.pi/agent/skills/
├── .pi/
│ ├── settings.json # Pi project settings + package list (incl. pi-subagents)
│ ├── SYSTEM.md # Replaces Pi's default system prompt
│ ├── APPEND_SYSTEM.md # Appends harness-specific guardrails
│ ├── npm/ # Project-local npm packages
│ └── git/ # Project-local git packages
└── eval/
├── .venv/ # Python virtual environment
├── .env.example
├── requirements.txt
├── pytest.ini
├── conftest.py # Shared fixtures and Pi runner helper (uses pi-run)
├── grader.py # Shared deterministic grading harness
├── datasets/
│ ├── tasks.json # Manifest: datasetVersion + task table (count authority)
│ ├── coding_samples.jsonl
│ ├── graders/ # 49 deterministic grader scripts
│ └── references/ # 54 reference answers/solutions
├── baselines/ # Committed live baseline (live-baseline.json)
├── scripts/ # score_run.py baseline gate + scorer
├── benchmarks/ # 8 Docker-isolated benchmark tasks (task.json + tests/run.sh)
├── benchmark-results/ # Git-ignored JSON run reports (scorecard-<run>.json)
├── live-results/ # Git-ignored nightly report + seam-report.json
└── tests/ # 19 hermetic + live test files (contract, eval, drift guards)
└── (test_benchmark_format, test_dataset_schema, test_score_run,
test_docs_drift, test_pm_drift, test_benchmark_seam, test_contract_*,
test_live_suite, test_live_metrics, test_harness_config, ...)
- Add sample data to
eval/datasets/coding_samples.jsonl. - Create a new test file in
eval/tests/. - Use
run_pi_print()fromconftest.pyto capture agent outputs (it runspi-run print). - Run
pi-run evalto see the results.
To add a benchmark task instead, create eval/benchmarks/<name>/task.json +
tests/run.sh (see Benchmarks) and validate it with
pi-run eval --benchmark-dry-run.
Key settings in .pi/settings.json:
defaultThinkingLevel:mediumcompaction: enabled with 16k reserve tokensretry: 3 agent-level retriessessionDir:.pi/sessionspackages: the curated plugin list
Project-local packages require approval on first use. Run:
pi list -ato see them, or pi config -a to enable/disable individual resources.
This harness is subagent-capable via the pi-subagents
extension (installed in .pi/settings.json packages). Pi can delegate focused
work to child sessions with their own tools.
Builtin agents:
| Agent | Use it when you want... |
|---|---|
scout |
Fast local codebase recon |
researcher |
Web/docs research with sources |
worker |
Implementation work (edits files, validates) |
reviewer |
Code review against a task/plan |
oracle |
A second opinion before acting |
delegate |
A lightweight general delegate |
Invoke in plain language: "Use reviewer to review this diff", "Ask oracle for a
second opinion", "Run worker to implement this plan". See
examples/subagents.md for a worked example.
- Pi packages run with full system access; only well-known, public packages were installed.
- Do not commit API keys;
eval/.envand.pi/sessions/are ignored. - Run destructive commands only after explicit confirmation.
- Pi packages not visible? Run
pi list -a(project approval required). - Engine warnings during install? Ensure a current Node version is installed via nvm and available to
pi-run(checkpi-run doctor). - DeepEval tests skipped? Live DeepEval tests run only when a supported
provider key is in the environment. Export
OPENAI_API_KEY/ another supported provider key and re-runpi-run eval. Keys held only in a secret manager are not read by the live-eval gate. pi update --modelstimes out? The pi.dev model-catalog endpoint is intermittently unreachable from some networks (TLS connects but HTTP never responds — on both IPv6 and IPv4).pi-runmitigates this three ways: every pi process runs withNODE_OPTIONS=--dns-result-order=ipv4first(the IPv6 route is deterministically dead),chat/printlaunch pi with--offlineso startup never touches the endpoint, andpi-run setupretries the refresh 3× then warns instead of failing — the stored catalogs already resolve every default model (pi-run doctorreports model resolvability as informational).
See CONTRIBUTING.md for setup, testing, the
good first issue
path, the 7-day review SLA, and how to add a provider. Report security issues
privately per SECURITY.md; all participants agree to the
Code of Conduct.
See docs/anti-lockin.md — BYO-key, BYO-model, and local model support keep your agent workflow portable across providers.