A small, provider-neutral CLI that gives you signals about whether an OpenAI- or Anthropic-compatible endpoint is actually serving the model it claims — or quietly downgrading you.
Read this first — what this tool is and is not. It produces heuristic signals, not cryptographic proof. You cannot prove from the outside which weights a provider runs. What you can do is collect repeatable behavioral evidence that catches the common cheats (model substitution, heavy quantization, context truncation, silent fallback). Treat a clean run as reassuring, not a guarantee; treat a flag as a reason to look closer, not a conviction. The Limitations section is not boilerplate — please read it.
- Zero runtime dependencies. Pure Python standard library. Nothing to
pip install, nothing in a lockfile to audit. Clone it and read it. - Your API key is never printed, logged, or stored. It is read only from an
environment variable (there is deliberately no
--keyflag), and every line of output passes through a redaction layer. See Key safety. - Provider-neutral on purpose. Point it at anyone — a relay you're evaluating, the official API, a local model server, or the gateway the authors work on (see Disclosure). The tool doesn't care who you test.
You pay for gpt-4o / claude-sonnet-4 / deepseek-chat. Are you getting it, or
a quantized, context-truncated, or swapped-out substitute? One command turns a run
into a shareable verdict card — a PASS / SUSPICIOUS snapshot of four
behavioral signals, safe to screenshot and post:
# See both a clean PASS and a downgraded SUSPICIOUS card — no key, no signup:
python3 -m llm_honesty_probe --self-test --card
# Then point it at the endpoint YOU pay for and share the verdict:
python3 -m llm_honesty_probe \
--base-url https://your-endpoint/v1 --claimed-model gpt-4o --card| PASS — a clean run | SUSPICIOUS — a downgraded run |
|---|---|
![]() |
![]() |
- Formats for wherever you post it:
--card-format txt \| md \| svg \| html \| png(defaulttxt;mdpastes straight into a forum,svg/png/htmlare for screenshots and social). Write a file with--card-out card.png, or dump every format at once with--card-format all --card-out ./cards. - It won't out your provider. The endpoint host is masked by default, so a
shared card stays about "an endpoint I pay for" — not a public accusation. Add
--card-show-endpointonly if you want the host on the card. - Same key-safety rules. The key is read only from an env var and never touches
the card. And remember the honest framing below: a card is a smoke test, not a
courtroom —
PASSmeans "no red flags in this run", not a guarantee.
Cheap "GPT / Claude / DeepSeek" API relays have a trust problem nobody advertises: some quietly swap in a smaller, quantized, or context-truncated model. The nasty part is the timing — it often works fine on day one and degrades weeks later, right after you've stopped checking. And the obvious check doesn't work:
Asking the model "what model are you?" is useless. That answer comes from a system prompt or fine-tune and is trivially spoofed. You have to test capability and behavior, not self-report.
So this tool automates a fixed battery of behavioral probes, pins the parameters, measures instead of eyeballing, and (optionally) diffs your endpoint against a reference you trust.
No dependencies, so the most trustworthy path is to just clone and run:
git clone https://github.com/seven7763/llm-honesty-probe
cd llm-honesty-probe
python3 -m llm_honesty_probe --self-test # runs against a built-in mock; no key neededRequires Python 3.8+. After the package is published you'll also be able to run it
without cloning (e.g. pipx run llm-honesty-probe ...), but cloning keeps the
"read the code before you run it" property that an honesty tool should have.
The key is read from an environment variable — never passed on the command line:
export OPENAI_API_KEY="sk-...your key..." # never appears in argv or history
# 1) Single endpoint: does this endpoint behave like the model it claims?
python3 -m llm_honesty_probe \
--base-url https://any-provider.example/v1 \
--claimed-model gpt-4o
# 2) Differential (strongest): diff the relay against the official API.
export OPENAI_API_KEY_OFFICIAL="sk-...official key..."
python3 -m llm_honesty_probe \
--base-url https://cheap-relay.example/v1 --claimed-model gpt-4o \
--compare-base-url https://api.openai.com/v1 --compare-model gpt-4o \
--compare-api-key-env OPENAI_API_KEY_OFFICIAL
# 3) Claude Code / Anthropic-shaped endpoints:
export ANTHROPIC_API_KEY="sk-ant-..."
python3 -m llm_honesty_probe \
--base-url https://any-provider.example --protocol anthropic \
--claimed-model claude-sonnet-4
# JSON for CI / cron; pick specific probes; list what's available:
python3 -m llm_honesty_probe --self-test --json
python3 -m llm_honesty_probe --list-probes
# A shareable verdict card instead of the full report (see the section above):
python3 -m llm_honesty_probe --self-test --card # PASS + SUSPICIOUS samples
python3 -m llm_honesty_probe --base-url https://any/v1 --claimed-model gpt-4o \
--card --card-format png --card-out verdict.pngProduced by --self-test --config <(echo '{"degrade": true}') — the built-in mock
made to behave like a downgraded relay, so you can see what a flagged run looks like:
llm-honesty-probe v0.2.0 — heuristic signals, NOT proof
Endpoint : http://127.0.0.1:0/v1 (openai) claimed model: gpt-4o
Mode : self-test (mock endpoint)
Probes : consistency, identity, needle, reasoning, tokenizer
------------------------------------------------------------------
[!!] identity Self-reported identity (low)
Claimed gpt but self-reports llama. Weak signal (spoofable both ways).
[!!] needle Long-context recall (low)
Recall broke at ~8000 chars. Could be silent truncation or the model's real limit; diff to disambiguate.
[!!] reasoning Capability floor (reasoning) (medium)
Failed 4/4 easy reasoning tasks a full-tier model rarely misses.
------------------------------------------------------------------
Summary: 2 consistent · 4 suspicious · 3 inconclusive
Some signals warrant a closer look (diff against the official endpoint).
Reminder: these are heuristic signals, not cryptographic proof. [...] See LIMITATIONS.
[OK] = consistent · [!!] = suspicious · [--] = inconclusive. Each signal
carries a confidence (low / medium / high) because not all signals are equal.
Every probe is designed to (a) test behavior rather than self-report, (b) degrade to inconclusive when a feature is unsupported or noisy, and (c) never decide a provider is lying from a single weak signal.
Different model families use different tokenizers, and the tokenizer leaks through
the usage.prompt_tokens the server reports. We send a battery of crafted strings
(digit runs, whitespace, CJK, emoji, code, URLs, mixed unicode) and record how
many tokens each costs.
To cancel the fixed per-request chat-template overhead, we measure a delta:
tokens(anchor + probe) − tokens(anchor). The resulting vector is close to the
raw token count of each string and is comparable across requests.
- With
--compare(best, reference-free): diff the vector against a reference endpoint serving the same model id. A different vector means a different tokenizer, which means a different model/family — a high-confidence signal. - Single endpoint: classify the vector against a shipped reference table of
known tokenizers (OpenAI's
cl100k_base/o200k_base, generated fromtiktoken— seescripts/build_reference.py, so the numbers are reproducible, not hand-written). If the claimed family has no public reference (Claude, Gemini, …), the probe says inconclusive, use--comparerather than guessing.
A handful of tasks with objectively checkable answers the tool already knows (multi-step arithmetic, string reversal, character counting) plus a strict-JSON format test and a refusal-behavior check on a clearly benign prompt. These are trivial for any full-tier model; a downgraded or heavily quantized substitute is more likely to trip. We weight failure more than success — passing an easy task proves little, but a "flagship" that fails a floor task is worth surfacing.
We generate deterministic filler text, hide a unique passphrase in the middle, and ask the endpoint to read it back — across several context lengths. Because the tool places the needle, it always knows the right answer (no external key needed). Recall that works at short context but breaks at longer context is the classic signature of silent context truncation (a relay capping your window to save tokens).
Repeats one fixed prompt N times at temperature=0 and looks at output
determinism, the stability of the server-reported model field and OpenAI
system_fingerprint, plus p50/p95 latency and error rate (informational). A
server that reports different model values or wildly different outputs for
identical calls may be routing you across different backends.
Asks the model to name itself and compares to the claim. This is included because its absence would be conspicuous, but it is trivially spoofable in both directions, so its confidence is capped at low and it never drives a verdict on its own.
- temperature=0, fixed
max_tokens, fixed system prompt — remove randomness. - Measure, don't eyeball — one call tells you nothing.
- Diff against a reference you trust whenever you can (
--compare). - Re-run on a schedule — silent degradation is a time-series problem, not a launch-day one. The JSON output is meant to be diffed in cron/CI over time.
Please take these as seriously as the features. Overstating an honesty tool defeats its purpose.
- These are signals, not proof. None of this cryptographically establishes which weights a provider runs. A determined provider that mirrors the official tokenizer and matches capability and holds long context and stays stable is, for practical purposes, giving you the model — but this tool cannot prove intent or rule out sophisticated spoofing.
- The probe set is small and opinionated. A provider that knows these exact probes could special-case them. That's why the strongest mode is a differential diff you control, and why PRs that make the probes harder to fake are welcome.
- Token accounting varies. Some providers don't return
usage, estimate it, or bill with quirks; a junction merge can shift a delta by a token. The tokenizer probe uses a tolerance and falls back to inconclusive rather than accuse. temperature=0is not guaranteed deterministic on real hardware, so the determinism signal is intentionally low confidence.- Legitimate reasons for "suspicious". A smaller model can be the correct
answer if that's what you asked for; a short context limit can be the model's
real limit, not truncation; different infrastructure can change latency. Use
--compareto disambiguate. - Refusal behavior is noisy. Alignment differs across honest providers, so the refusal check is low confidence by design.
- This tests an endpoint's behavior over a moment, not its contract. Re-run over time; a single run is a snapshot.
If you audit one thing, audit llm_honesty_probe/redaction.py
and the Endpoint header code in client.py.
- The key is read only from an environment variable named by
--api-key-env(defaultOPENAI_API_KEY, orANTHROPIC_API_KEYfor--protocol anthropic). There is no--keyflag, so your key cannot land in shell history,ps/argv, or a saved command. - There is no code path that prints
api_keyor anAuthorizationheader. - Everything that could reach your screen, a file, or an error message runs through
redact(), which scrubs both the exact secret and anything matching a common key/authorization shape. - The tool makes only the API calls needed for the selected probes. It writes
nothing to disk unless you pass
--out.
--probes tokenizer,reasoning # subset instead of all
--repeats 10 # samples for the consistency probe
--needle-lengths 2000,8000,32000 # approx context sizes (chars) to test recall
--reference path/to/tokenizers.json
--config path/to/config.json # e.g. {"degrade": true} for --self-test
--json --out report.json # machine-readable, good for cron diffs
--card # render a shareable PASS/SUSPICIOUS verdict card
--card-format png # txt (default) | md | svg | html | png | all
--card-out ./cards # a file, or a directory when --card-format=all
--card-show-endpoint # include the tested host (masked by default)To (re)generate the tokenizer reference yourself (recommended — don't trust numbers you didn't compute):
pip install tiktoken # dev-only; not a runtime dependency
python scripts/build_reference.pyThis tool is not tied to any provider. It ships no provider allow-list, no "preferred" endpoint, and no telemetry. The comparison workflow treats the official API and any relay identically. That neutrality is the point: a verification tool is only useful if it will just as happily flag the people who wrote it.
This project was started by people who work on daoxe, an OpenAI-compatible LLM gateway (it also speaks Anthropic Messages natively for Claude Code). We built it because we'd rather you verify a provider than take a vendor's word for it — including ours. Point it at daoxe and at whatever you use today, and compare: daoxe.com. daoxe actively encourages users to run this against its endpoints; if it ever fails these checks, that's a bug report we want.
The tool's usefulness does not depend on daoxe, and nothing here privileges it.
Two sibling projects, if you're evaluating an endpoint end to end:
- llm-gateway-benchmark — a reproducible speed & availability benchmark (success rate, p50/p95 latency, $/1M tokens). It answers "is this endpoint fast and cheap?"; this tool answers "is it actually the model it claims?" The two are complementary, not competing.
- DaoXE-AI — OpenAI-/Anthropic-compatible gateway setup examples for Cursor, Claude Code & Cline (same authors — see Disclosure). Point this probe at it and compare it against whatever you use today.
The best contributions are harder-to-fake probes and additional tokenizer references. If you can think of a behavioral check a downgraded relay can't cheaply spoof, please open a PR or issue.
MIT — see LICENSE.

