R⁴ is a local, CPU-first AI and geometric routing system aligned with the Universal Object Reference Framework, Prism, and uor-addr.
Transformerless inference is core R⁴ functionality. It is not an external
provider adapter or a separate product: the compiler, runtime, tokenizer,
graded store, witnesses, and model-source adapters live under
uor_r4_core::transformerless.
R⁴ does not require Ollama, llama.cpp, OpenAI, Anthropic, or another inference
service at runtime.
R⁴ currently provides:
- a geometric text router, 96-vertex W(3,3) phase field canvas, and browser dashboard with Developer Mode toggle (
#dev-mode-toggle); - a CPU-only transformerless compiler and table-native zero-multiply inference runtime;
FallbackRouterpipeline cascading from primaryr4g1-graphto secondarytransformerless-tla5on unmapped/pathological states;WINDOW = 8Dyadic-Recency context window with zero-allocation stack/slice sliding truncation;- ChatML prompt formatting for instruction-tuned teacher models (
SmolLM2-135M-Instruct); - strict 32-bit
u32integer token stores (parse_store_strict_u32) with automatic legacy cache purging (purge_legacy_store_cache); - BLAKE3
tokenizer_cidheader verification inuor-r4-graph-format(verify_tokenizer_cid); - UOR attestation envelopes (
uor_address,artifact_cid,store_cid,attestation_cid) andPOST /api/uor/verifyvalidation endpoint; - TLA3/TLA4 compatibility and the current TLA5 artifact format (no f32 centroids in deployed containers);
- direct BF16 Safetensors loading for compatible, unsharded Llama-family Hugging Face models, without Candle;
- byte-level BPE tokenizer export;
- resumable compilation with progress reporting;
- content-addressed model objects and manifests using UOR CIDs;
- witnessed prediction, indexing, and generation APIs;
- one-shot and interactive chat as application examples of the core runtime.
The R⁴ holographic graph compiler program (multiresolution, overlapping
semantic graphs with an allocation-free integer runtime) is underway on top of
this engine — see docs/r4_graph_compiler_implementation_plan.md. Already
landed: the R4G1 packed artifact format with two-stage validation
(crates/uor-r4-graph-format), the TLA/TLS1 → R4G1 migration converter, the
observation pipeline (content-addressed sample IDs, deterministic shard
spill/resume), multiresolution cover induction (spherical k-means with
calibrated overlapping memberships), semantic transitions + reverse indexes,
ScoreQ fixed-point residuals, packed NGRAM context rows (trigram → bigram
backoff) plus the optional FWDA forward-anchor section with
anchor-conditioned infill scoring (score_candidates_infill, issue #399),
and the executable proof model (crates/uor-r4-proof-model). CI runs fmt/clippy/tests/no_std/deterministic-
rebuild/audit/fuzz/wasm gates on every push (.github/workflows/ci.yml).
There is an important boundary in the current workflow: compiling a Hugging
Face model produces continuation evidence and runtime artifacts, but it does
not automatically prove instruction-following quality. ask accepts only an
imported instruction-chat manifest with a CID-addressed passing evaluation
report produced by evaluate-report. This guard prevents a fast but
low-quality continuation artifact from being presented as an accurate
question-answering model.
| Operation | Works from a fresh checkout? | Additional state |
|---|---|---|
| Build and test the workspace | Yes (UOR standards are pinned git deps) | None |
| Geometric router and dashboard | Yes | None |
| Legacy TinyStories certification/benchmark | After setup and corpus generation |
Pinned llama2.c checkpoint |
| Compile a compatible Hugging Face source | After downloading the source | Pinned model revision |
ask and interactive chat |
No | Compiled, evaluated, imported instruction-chat manifest |
R⁴ is a research programme as much as an engine, and the engine's direction is set by measurements rather than by intent. This section records where that programme actually stands, so anyone picking up work can see which paths are closed, which are load-bearing, and which are still open. Every claim below traces to a merged measurement with a pre-declared exit rule; the issue numbers are the durable references.
Every substantive claim in this repo is expected to arrive with a pre-declared
exit rule, a null baseline, and a falsifier. Negative results are recorded and
kept, not discarded — several of the entries below are negatives that redirected
the programme, and they are more valuable than the positives they replaced.
Long runs additionally follow the run-contract discipline in AGENTS.md:
compute the reachability ceiling before spending hours, gate on the cheap
instrument first, and pre-declare what each outcome causes. That discipline
exists because we lost days to runs whose result could not have changed the
next action.
The serving stack consults, in order: packed NGRAM context rows (trigram with bigram backoff), then the graph chain with D4 exact-context precedence, then the root prior. On natural text the induced-cover store with observed continuation evidence is the geometry that carries the result; the legacy teacher-hash store with teacher evidence measured 0.1% off-distribution against it.
Two changes improved results by improving evidence quality per key, and both are shipped: full-width content-bearing storage (#434, PR #465) — the storage path had been discarding fifteen sixteenths of an already-full-width content vector, and de-banding moved retrieval MRR from 0.2348 to 0.8948 and router anchor accuracy from 9.3% to 11.4%; and the two-sided calibration gain (#446), which is causally legitimate, sits at top-1 parity, and grows large at scale in bits (latent-mix 15.4778 vs 22.2078).
A-mode infill serving is validated and shipped: the FWDA forward-anchor
artifact section plus score_candidates_infill, infill_fill, and the
r4 graph infill --skeleton CLI (#399, PRs #416/#419). Anchors are inputs, so
the mode is immune to the drift that killed the standalone variant.
Standalone two-pass generation (#399). Refuted twice. With anchors supplied externally the channel gives +4.2pp on its live slice; with the engine supplying its own anchors from a drafted context it goes negative (40.4% vs 41.3%), and a strict per-step confidence gate does not rescue it (42.0% vs 43.0%). Drift is diffuse rather than concentrated in low-confidence steps, so no gate over the draft can filter it. At 2.11M records the inversion reproduces on a non-degenerate configuration (16.4% vs 26.5%, predicted-anchor accuracy 0.0%), so it is not a capacity artifact.
Code-space subdivision as a capacity lever (#460). Measured negative in its strongest possible form. Raising STAGES from 4 to 5 bought exactly the subdivision the hypothesis asked for — occupied full-code keys 47,403 to 90,824, records per key 36.02 to 18.80, clearing the instrument gate — and Rule 1+2 top-1 came in at 25.6% ± 0.44pp against a 26.5% baseline, below it. The store baseline fell alongside (26.4 to 25.4), which is the signature of thinner per-key evidence rather than sharper context. Exact-context dominance barely responded (98.8% to 97.1%).
Construction-time stratification (#435). Three routing designs plus an identity argument, all against pre-declared rules; v3 mass-linear mixing reduces algebraically to flat. The oracle-stratum edge (36.3 vs 35.2) stands as recorded unrecovered signal, but no routing design reached it.
Hopf sector transport as a router-quality lever (#422/#306). The #306 remediation's occupancy gain (16 to 456 of 512 sectors) does not translate into retrieval value: sector-filtered MRR 0.0045 against the pre-remediation projection's 0.0743. Three content-aligned redesign candidates then mapped a clean spread-versus-retrieval frontier without crossing it.
Cayley–Dickson syntactic morphism (#400), FMM far-field (#290), granularity (#393), E8 group-keying (#395). Each measured dead with a scoped record; the CD term executed 0 times out of 1,998 before removal.
Every lever that added key resolution failed — a better-fitted codebook, more cover regions, a finer code space. The only two changes that helped improved evidence quality per key. That is the clearest signal the programme has, and it is why the open work below concentrates on evidence and estimation rather than on subdivision.
| Issue | Question | State |
|---|---|---|
| #460 | Cover split criterion and codebook fit | Two implemented-but-dark levers: SplitCriterion::{RelativeGain,Mdl} with scaled k0 (default-off, harness cover_scaling.rs unrun) and RVQ_SAMPLE_CAP capping codebook training at 0.59% of the split |
| #424 | Bott-Fock O(1) context fold | Ceiling measured, A/B not reachable. Long-range signal on this corpus is worth +1.02pp of top-1 (two thirds of it order-carried); the shipped decay constant >> 2 retains 16% of that, so the lossless upper bound on the fold as shipped is +0.16pp — one standard error. Retuning the decay to >> 7 would recover the ceiling. docs/context_horizon_424.md |
| #434 | VSA / spectral geometry | Item 1 (zeta-grid at scale) done and shipped; item 2 never measured beyond a synthetic smoke test |
| #469 | Vectorize the assign path | The vectorized kernel already exists (simd::dot_argmax) and is simply not wired into assign_for_bundle; κ-pinned, so bit-identity must be proven |
| #471 | Sampled runs still pay full-corpus table builds | derive_right_codes and the two-sided/latent table builds ignore the sample knob |
| #456–#459 | Reconstructability, block search, IPF reconstruction, estimation ladder | Active track |
| #320 | Teacher upgrade (SmolLM2) | P1/P2 rehearsal recorded; migration decision open |
| #273 | Template rebase / claim register | On-hold; no implementation |
Two pieces of tooling exist so the results above stay cheap to reproduce and
hard to fake. Sampled decision runs (R4_GATE_C_SAMPLE) cut Gate C evaluation
from 597s to 60s and report the sample size and standard error beside every
rate, because at 402,802 positions the standard error is 0.07pp while every
exit rule we write is ±2pp — thirty times finer than any decision needs. The
κ-keyed per-record code sidecar (R4_CODES_PATH) cut the instrument from 625s
to 39s by caching a deterministic computation that every consumer had been
recomputing; it loads only when eight fields and a blake3 digest all agree, and
refuses itself entirely in biased-sampling mode so a partial vector can never
poison a later run. The capacity_scaling instrument prints a saturation
verdict per structure and is meant to be run before trusting any measurement
taken on a given configuration.
- AGENTS.md — contributor/agent operating manual: gates, invariants, κ re-pin procedure
- ELI5 explainer
- Undergraduate explainer
- Transformerless design
- Proof and certificate
- Performance comparison
- Local-only runtime contract
- R⁴ graph compiler implementation plan
- Glossary · R4G1 wire format · Baseline · Threat model
- Minimalist terminal client & local vendor API
- Roadmap
- Current stable Rust and Cargo.
hffrom the Hugging Face CLI formodel downloadorcargo run --release -- compile --model.curlandunziponly for the legacy TinyStories setup workflow.
Inference is CPU-only. GPU features, --device, Metal, CUDA, and Candle are
intentionally not part of this runtime.
For zero-setup testing, model compilation, and interactive Q&A out of the box:
# Launch single-command orchestrator & interactive client
./uor-r4-cliThe uor-r4-cli orchestrator automatically:
- Clears the screen and clears standard terminal settings on launch.
- Auto-detects and restores your last used model (
smollm2-135m-instruct,smollm2-360m-instruct, orsmollm2-1-7b-instruct) and active synthesis engine (r4g1orattention). - Handles 4-stage pipeline execution: downloads pinned teacher weights, compiles zero-multiply observation corpora, and builds scored R4G1 residual graph covers automatically when required.
- Launches the local backend server on port
8000and connects the interactive client REPL.
To run uor-r4-cli from any directory in your terminal:
# 1. Symlink to user local bin
mkdir -p ~/.local/bin
ln -sf $(pwd)/uor-r4-cli ~/.local/bin/uor-r4-cli
# 2. Ensure ~/.local/bin is in your PATH (Zsh)
echo 'export PATH="$HOME/.local/bin:$PATH"' >> ~/.zshrc
source ~/.zshrc
# Now type uor-r4-cli from anywhere!
uor-r4-cliType / at the prompt to launch the interactive slash command menu, or run direct commands:
-
/models— View and switch active teacher models in-session with live download/compilation badges ([DL: ✓ | CP: ✓]). -
/engine— Switch active synthesis engine (r4g1sub-ms zero-multiply graph,attentionteacher oracle fallback,r4-attention,geometric). -
/status— View 4-stage R4G1 pipeline compilation status table and readiness metrics. -
/corpus— Manage extra reading corpus datasets and view indexed server files. -
/compile— Trigger automated 4-stage in-session model graph compilation. -
/audit— Inspect UOR coordinates ($\kappa$ ,$\theta_d$ ,$uor_bias$ ),$\kappa$ -pass reproduction status, token provenance traces, and export session logs (.uor-models/audit_log.json). -
/clear— Clear terminal screen. -
/quit— Exit client session cleanly.
To inspect UOR compliance audit traces from previous sessions directly:
# Launch UOR audit inspector from CLI
uor-r4-cli --audit
# Or using the binary directly
r4 auditVerify the workspace:
cargo check --workspace
cargo test --workspaceThe project exposes one executable, r4. With no subcommand it runs the HTTP
server:
cargo run
# equivalent after a release build: ./target/release/r4Open http://127.0.0.1:8000. The geometric router and dashboard work without a chat model. Use another listener when needed:
cargo run -- --host 0.0.0.0 --port 9000Build the browser package with:
wasm-pack build --target webStatic deployments use geometric synthesis in WebAssembly. Transformerless artifact loading and synthesis run through the native server.
The local lifecycle has two artifact lanes:
pinned source -> resumable teacher observation -> TLA5 + TLS1 bundle
| |
| +-> local ask/evaluate/import
+-> multiresolution cover
-> scored R4G1 graph + Gate C report
Downloading and compiling are explicit offline operations. Neither ask nor
the HTTP server downloads a model or contacts an inference provider.
The top-level compile command builds the deployable transformerless bundle
and retains its observation corpus. The graph compiler then consumes those
outputs through transformerless cover and transformerless score. This
separation makes expensive stages reusable and keeps the deployed runtime
independent of the Hugging Face teacher.
cargo run -- download \
--repository HuggingFaceTB/SmolLM2-135M-Instruct \
--revision 7e27bd9f95328f0f3b08261d1252705110c806f8 \
--name smollm2-135m-instructThe default destination is
.uor-models/sources/smollm2-135m-instruct. Override it with --output:
cargo run -- download \
--repository HuggingFaceTB/SmolLM2-135M-Instruct \
--revision 7e27bd9f95328f0f3b08261d1252705110c806f8 \
--name smollm2-135m-instruct \
--output /path/to/model-sources/smollm2-135m-instructThe downloader prints the repository and destination immediately, streams the
hf process, and emits a heartbeat every two seconds with file count, bytes,
and elapsed time.
Compile an already downloaded directory:
cargo run --release -- compile \
--source .uor-models/sources/smollm2-135m-instruct \
--output .uor-models/compiled/smollm2-135m-instruct \
--seconds 300 \
--target 20000 \
--sequence-length 128Or let the compiler download an immutable revision through hf:
cargo run --release -- compile \
--model HuggingFaceTB/SmolLM2-135M-Instruct \
--revision 7e27bd9f95328f0f3b08261d1252705110c806f8 \
--seconds 300 \
--target 20000 \
--sequence-length 128 \
--r4-attention--revision must be a full 40-character commit hash. --seconds limits the
teacher-generation work performed by one invocation, while --target is the
teacher-token goal. --r4-attention enables the experimental 4D softmax-free
Spin(4) attention geometry during teacher generation (omitting this flag runs
standard scaled dot-product attention). Hugging Face compilation defaults to
20,000 tokens and 128-token teacher stories. The bounded story length keeps
attention cost and KV memory proportional to the eight-token deployed runtime
window; increase --target or --sequence-length explicitly for quality
experiments. Repeat the same command to resume an incomplete corpus.
On macOS, offline Hugging Face teacher execution uses Apple Accelerate's
SIMD-optimized CPU BLAS. Linux and Windows use explicit NEON on AArch64 or
runtime-detected AVX2/FMA on x86-64, with a dependency-free scalar fallback.
These compiler accelerators do not add a runtime dependency or change the
allocation-free table-native inference path. Set TLESS_TEACHER_EXACT=1 to
force the slower, reduction-order-preserving scalar path for diagnostic
comparisons. The pinned legacy proof workflow always uses that exact path.
When compilation completes, the output directory contains:
tless_artifacts.bin # TLA5 teacher projection and class tables
tless_store.bin # TLS1 graded continuation evidence
tokenizer.bin
corpus.meta # observation-corpus metadata
corpus.records # deterministic observation records
hamming_calibration.json
hierarchical_codes.json
space_manifest.json
Interactive terminals show progress bars. Redirected output receives periodic
progress: lines suitable for build logs. Source loading and compilation may
allocate; the allocation-free guarantee applies to the deployed prediction hot
path, which uses fixed and caller-owned buffers.
The graph compiler turns the retained observation corpus and TLA5 artifact into a multiresolution, overlapping semantic graph. First induce and measure the cover:
cargo run --release -- transformerless cover \
--corpus-meta .uor-models/compiled/smollm2-135m-instruct/corpus.meta \
--corpus-recs .uor-models/compiled/smollm2-135m-instruct/corpus.records \
--artifacts .uor-models/compiled/smollm2-135m-instruct/tless_artifacts.bin \
--out .uor-models/compiled/smollm2-135m-instruct/graph-coverThis writes cover.r4g1 and cover_report.json. The cover report now emits a
versioned objective block (objective.config.schema) with separate train and
held-out components for predictive entropy (H(A|R)), future-state entropy
proxies (H(S_future|R)), teacher-loss proxy, runtime/artifact/bytes/structure
costs, and information-bottleneck proxy terms (I(Z;X) - βI(Z;Y_future)),
plus a bounded top-64 between-region distinctiveness term against the global
next-token prior (default weight 0, preserving the default cover), and
auditable split decisions. Objective versions migrate by appending new fields
under objective while keeping Gate C and predictive-sufficiency reports as
separate reproducible artifacts. Then compile semantic
transitions, fixed-point emission residuals, and exact-evidence carryover:
cargo run --release -- transformerless score \
--corpus-meta .uor-models/compiled/smollm2-135m-instruct/corpus.meta \
--corpus-recs .uor-models/compiled/smollm2-135m-instruct/corpus.records \
--artifacts .uor-models/compiled/smollm2-135m-instruct/tless_artifacts.bin \
--cover .uor-models/compiled/smollm2-135m-instruct/graph-cover/cover.r4g1 \
--out .uor-models/compiled/smollm2-135m-instruct/graphThe result is graph/score.r4g1, a stage-validated packed graph containing
regions, refinement/neighbor/forward edges, ScoreQ emission tables, and the
EXCT evidence section. graph/score_report.json records artifact and corpus
kappas plus held-out Gate C top-1 agreement, bits/token, and witness replay.
Passing --cover reuses the measured cover. It may be omitted to re-induce
the default cover deterministically during scoring. For experiments, cover
also accepts --depths, --k0, --regions-budget, and --memory-budget.
The public ask and chat library paths still load the TLA5/TLS1 files from
step 2. The native HTTP server auto-loads graph/score.r4g1 beside
tless_artifacts.bin when present (or accepts --r4g1-artifact), validates it,
and uses it for the transformerless engine before falling back to TLA5/TLS1.
When the dashboard is served by that native process, its Compile / Refresh
R4G1 Graph button runs the same cover → score pipeline against the bundle's
corpus.meta and corpus.records, validates the new graph, and hot-swaps it
into the running server. Static WASM deployments cannot run this compiler.
Static deployments still use the geometric WASM fallback because they have no
native filesystem-backed graph loader yet.
The native dashboard also exposes Download Hugging Face Weights. Its input
defaults to the pinned owner/repository@commit from
models/smollm2-135m-instruct.json, but accepts any repository paired with a
full 40-character commit. Downloads go into .uor-models/sources/; nothing is
downloaded until the button is pressed. Afterward, run the bundle compiler and
then the R4G1 graph compiler. If the downloaded source is present and the
compiled bundle is not, the native Compile / Refresh R4G1 Graph action now
runs the bundle compiler first, then cover and score compilation, as one server
job.
Compilation produces a directly loadable local bundle. On first use, R⁴
content-addresses the artifact, store, and tokenizer in .uor-models/objects:
cargo run --release -- ask "why is the sky blue?"This direct path verifies container integrity but does not claim that the compiled approximation has passed an instruction-quality evaluation. The CLI logs that distinction. Compilation success and answer quality are separate properties.
Run held-out instruction and grounding evaluation against the compiled bundle and retain a machine-readable report:
cargo run --release -- evaluate-report \
--source .uor-models/sources/smollm2-135m-instruct \
--compiled .uor-models/compiled/smollm2-135m-instruct \
--report .uor-models/compiled/smollm2-135m-instruct/instruction-eval.jsonThe report file stores an envelope with the held-out D3 metrics (top-1
accuracy, teacher-argmax agreement, Witten–Bell bits/token vs the teacher
floor), source/artifact/store/tokenizer/corpus CIDs, and
report_cid_of_report_bytes for the inner metrics payload. Do not mark an
artifact as passing merely to bypass the chat quality gate.
cargo run -- import \
--name my-chat-model \
--source-model HuggingFaceTB/SmolLM2-135M-Instruct@7e27bd9f95328f0f3b08261d1252705110c806f8 \
--capability instruction-chat \
--artifacts .uor-models/compiled/smollm2-135m-instruct/tless_artifacts.bin \
--store .uor-models/compiled/smollm2-135m-instruct/tless_store.bin \
--tokenizer .uor-models/compiled/smollm2-135m-instruct/tokenizer.bin \
--evaluation-report /path/to/instruction-eval.json \
--instruction-eval-passed \
--grounded-answer-rate 0.80 \
--repetition-rate 0.01The model store defaults to .uor-models; set UOR_MODEL_STORE to relocate
it. Objects are stored once under objects/blake3/<digest>. Reads verify both
the declared byte length and UOR CID. The import command prints the manifest
CID.
Continuation-only bundles may be imported for certification and benchmarking,
but ask refuses to load them.
One-shot ask calls the R⁴ library directly without a server or network hop:
cargo run --release -- ask \
--model my-chat-model \
"why is the sky blue?"Interactive chat retains turn history:
cargo run --release -- chat --model my-chat-model--model is optional. Selection order is TLESS_MODEL, the newest JSON
descriptor in models/, then smollm2-135m-instruct. A descriptor selects a
name; R⁴ first uses an imported manifest and otherwise falls back to a complete
local bundle under .uor-models/compiled/<name>.
Library consumers can use the chat example directly:
use uor_r4_wasm_router::chat::ChatEngine;
let mut chat = ChatEngine::builder().model("my-chat-model").build()?;
let answer = chat.ask("why is the sky blue?")?;
println!("{}", answer.text);
# Ok::<(), Box<dyn std::error::Error>>(())Chat is an application of transformerless R⁴, not a separate crate or inference layer.
The pinned llama2.c TinyStories path remains available for proof reproduction and same-machine performance comparison.
cargo run --release -- setup
cargo run --release -- gen 300 150000
# repeat gen until it reports done=1
cargo run --release -- certify
cargo run --release -- compare
cargo run --release -- compare-report
cargo run --release -- scenarioscertify performs the compile, store, certificate, and census steps
internally. The bare compile and store subcommands belong to the HF
graph-compiler path — compile requires --model or --source, and store
depends on a prior graph compile — and are not part of the legacy chain.
gen output is not byte-reproducible across machines or eras: story and
held-out counts can differ slightly from the certified stream (e.g. 754
stories / 30,036 held-out against the certified 757 / 30,192), which bounds
how exactly downstream figures reproduce.
Its default files are:
/tmp/tless_artifacts.bin
/tmp/tless_store.bin
/tmp/ref/tokenizer.bin
This is a TinyStories continuation artifact, not an instruction-chat model.
Use compare-report for the recorded certificate without loading the source
checkpoint:
cargo run --release -- compare-reportSee COMPARISON.md for the measured quality and throughput evidence.
Build once and invoke r4 directly, or use the equivalent Cargo commands:
cargo build --release
./target/release/r4 --help
./target/release/r4 ask "why is the sky blue?"
cargo run -- --help
cargo run -- ask --help
cargo run -- compile --help
cargo run -- download --help
cargo run -- import --help
cargo run -- transformerless cover
cargo run -- transformerless scoreAll subcommands support -v, -vv, and -vvv for info, debug, and trace
logging. Tracing uses a dependency-light subscriber.
The server defaults are:
| Option | Environment | Default |
|---|---|---|
--host |
UOR_R4_HOST |
127.0.0.1 |
--port |
UOR_R4_PORT |
8000 |
--manifold-cache |
UOR_R4_MANIFOLD_CACHE |
manifold_cache_rust.json |
--tless-artifacts |
TLESS_ARTIFACTS |
/tmp/tless_artifacts.bin |
--tless-store |
TLESS_STORE |
/tmp/tless_store.bin |
--tless-tokenizer |
TLESS_TOKENIZER |
/tmp/ref/tokenizer.bin |
--r4g1-artifact |
R4G1_ARTIFACT |
<compiled>/graph/score.r4g1 when present |
--tless-corpus-meta |
TLESS_CORPUS_META |
Beside the configured artifact when present |
--tless-corpus-recs |
TLESS_CORPUS_RECS |
Beside the configured artifact when present |
flowchart LR
Source["Pinned local model source"] --> Compiler["R⁴ transformerless compiler"]
Compiler --> Corpus["Deterministic observation corpus"]
Compiler --> Artifact["TLA5 artifact"]
Compiler --> Store["TLS1 graded store"]
Compiler --> Tokenizer["Tokenizer"]
Corpus --> Cover["Multiresolution cover induction"]
Artifact --> Cover
Cover --> Score["Transitions + ScoreQ residuals"]
Store --> Score
Score --> Graph["Validated R4G1 graph"]
Graph --> GraphRuntime["Integer graph scorer (evaluation path)"]
Artifact --> Runtime["Allocation-free CPU prediction kernel"]
Store --> Runtime
Tokenizer --> Runtime
Prompt["Prompt"] --> Router["R⁴ geometric router"]
Router --> Runtime
Runtime --> Witness["UOR CID + Grounded witness"]
Runtime --> Apps["r4 ask / r4 chat / HTTP API"]
The workspace has one public package and ten internal implementation crates:
| Package | Responsibility |
|---|---|
uor-r4-wasm-router |
Public facade, UOR witness integration, HTTP server, WASM surface, and the single r4 executable |
uor-r4-core |
Core R⁴ mathematics and transformerless compiler/runtime/tokenizer/certifier |
uor-r4-router |
Manifold state, indexing, geometric routing, and router witnesses |
uor-r4-graph-format |
Canonical R4G1 serialization, two-stage validation, and borrowed graph views |
uor-r4-graph-compiler |
Offline graph-compiler stages: observation pipeline, cover induction, routing/residual packing |
uor-r4-graph-certify |
Offline certification and measurement: Gate C scoring harness (score), reference scorer (score_runtime), certificates, comparison |
uor-r4-graph-runtime |
no_std allocation-free R4G1 graph runtime (engine, routing programs, packed kernels, patch chains) |
uor-r4-graph-cli |
r4 transformerless … CLI stage dispatch (convert-r4g1, scenarios, corpus tools) |
uor-r4-api |
Typed compile + engine library facade for downstream library consumers |
uor-r4-model-source |
Teacher forward-pass port (llama2.c-exact) and pinned Safetensors adapter |
uor-r4-proof-model |
Executable graph-compiler proof obligations and proof-status matrix |
The public tless_uor module provides TlessAxis,
UorTlessModel, CID addressing, and per-prediction Grounded certificates.
The root chat and model modules are
application-level consumers of the core runtime.
Routes and synthesizes a prompt. The engine parameter selects the generation mechanism:
"transformerless": Run allocation-free table-native codebook retrieval (sub-millisecond latency on CPU)."r4g1": Run the validated R4G1 graph scorer when ascore.r4g1artifact is loaded."attention": Run standard scaled dot-product attention on the loaded teacher model (generates up to 256 tokens)."r4-attention": Run experimental 4D Spin(4) softmax-free attention on the loaded teacher model (generates up to 256 tokens, yielding ~25% computation speedup on CPU by bypassing standard softmax exponents)."geometric": Route purely geometrically and decode directly from the manifold resonance.
Example request payload:
{
"text": "dry season aquifer depth in the Gambia",
"identity": "tenant-alpha",
"engine": "transformerless"
}The browser dashboard (http://127.0.0.1:8000) includes an engine selector dropdown to easily swap between these modes. The dashboard displays a Speed metric (in tokens/sec) under the telemetry card that persists after generation completes, allowing for easy execution speed profiling and audit comparison across the attention and transformerless pathways.
Returns initialization metrics, uptime, and UOR validation state.
Starts and monitors the native server's cover → score compilation job. The
resulting graph is loaded only after validation succeeds; the status response
also includes the generated score_report.json when available.
Starts and monitors an explicit download. The optional JSON body is
{"model":"owner/repository@<40-character-commit>"}; omitting it uses the
pinned source defined by models/smollm2-135m-instruct.json. The server
requires the hf CLI and rejects unpinned revisions.
Indexes text into the geometric manifold:
{
"corpus": "Text to index into the manifold.",
"identity": "tenant-alpha"
}Export or restore router vocabulary, prime products, and sentence manifolds.
Runs one witnessed token prediction:
{ "window": [1, 298, 263, 221, 437, 238, 15, 1979] }Only the eight most recent token IDs are read. The response includes the token,
resolution depth, graded code, evidence count, operation census, artifact and
store kappas, UOR address, and Grounded witness metrics.
Adds tokenized text to the graded store:
{ "text": "Once upon a time, there was a little dog named Rex." }The response includes token count, evidence positions written, and the updated store kappa.
Runs attributable greedy generation:
{ "text": "Once upon a time, there was a little", "max_tokens": 24 }The response includes generated text and tokens plus each step's resolution depth and evidence count.
cargo fmt --all -- --check
cargo check --workspace --all-targets --offline
cargo test --workspace --all-targets --offline
cargo clippy --workspace --all-targets --offline -- -D warningsReproduce the transformerless proof witnesses with:
cargo test -p uor-r4-core
cargo test -p uor-r4-core --release --test kappa_reproduction -- --ignoredThe ignored reproduction and real-SmolLM2 adapter tests require their external model fixtures.
- Downloaded source data is not a compiled chat bundle: run the exact
compile --source ...command printed byask. - Compiled bundle has no quality attestation: local
askcan still run it and content-addresses its files. Evaluate andimportit before presenting its output as instruction-quality validated. - Manifest not found: select an imported manifest name/CID or set
TLESS_MODEL. - Tokenizer or transformerless state unavailable: build the legacy files or
pass the three
TLESS_*paths explicitly. - No
metalfeature or--device: expected; inference is CPU-only.