A lightweight control plane for AI services — service registration, background health checking, network-aware policy-based routing, per-tenant rate limiting, observability, and canary rollout. Applies control-plane / data-plane separation, topology-aware routing, weighted traffic splitting, and quota enforcement from traditional networking to AI infrastructure.
This project grew out of a real need inside the Enterprise AI Business Intelligence Platform: reliably knowing whether an LLM provider or downstream service was actually reachable before routing a request to it. The BI Platform is the first service registered with this control plane, but the control plane itself is generic — it can register and govern any HTTP-based AI service.
- Service Registry: register, list, fetch, and deregister downstream services via REST API.
- Background Health Checking: APScheduler polls every registered service on a configurable interval — the same way a router marks a BGP neighbor up or down based on consecutive missed keepalives.
- Status Model:
UNKNOWN → HEALTHY / DEGRADED → UNHEALTHYbased on a configurable consecutive-failure threshold. - Admin Status Override: manually force a service into any health state — equivalent of
shutdown/no shutdownon a network interface.
- Policy Engine: evaluates active routing policies in priority order — mirroring route-map clause evaluation in traditional network policy-based routing.
- Network-Aware Routing: services carry topology metadata (region, latency zone, network tags); policies can constrain routing to specific regions or latency classes — analogous to BGP community filtering and OSPF link-cost preference.
- Automatic Failover: if the primary target fails health or topology checks, the engine transparently falls back to the configured secondary.
- Policy Fallthrough: when a constrained policy finds no eligible service, evaluation continues to the next policy in priority order — mirroring route-map clause fallthrough.
- Resolution Codes: every routing decision returns a typed outcome (
primary,fallback,no_policy,no_healthy_service).
- Per-Tenant Quota: each tenant gets a configurable request ceiling within a time window (
max_requests / window_seconds). - Redis-Backed Counting: fixed-window counter using
INCR+EXPIRE— lightweight and sub-millisecond. - JWT Tenant Extraction:
tenant_idis read from the Bearer token before each route resolution; unauthenticated requests fall under a sharedanonymousbucket. - Graceful Defaults: tenants without a quota record, or with
is_active=False, pass through without counting. - 429 with Headers: quota-exceeded responses include
Retry-After,X-RateLimit-Limit, andX-RateLimit-Remaining. - Quota Management API: create, inspect (with live Redis counter), update, and reset quotas via REST.
- Request Logging: every call to
/routeis logged asynchronously with tenant, request type, resolved service, resolution code, and latency — without adding to the critical path. - Summary Endpoint: snapshot of service health counts, active policy count, and request volume in the last hour.
- Traffic Metrics: distribution of resolved services and resolution codes over a configurable time window.
- Error Metrics: breakdown of error-class resolutions (
no_policy,no_healthy_service) per service. - Latency Metrics: average routing latency per resolved service, ordered fastest first.
- Weighted Traffic Splitting: each policy has a
weightfield; policies sharing the same priority form a canary group and receive traffic proportional to their weight — mirroring weighted ECMP routing. - Instant Rollback: set
weight=0on any policy to remove it from traffic split without deletion. - Dedicated Weight Endpoint:
PATCH /policies/{id}/weightfor fast canary promotion or rollback without a full policy update. - Observability Integration: traffic endpoint reports
policy_nameandpolicy_weightalongside resolution data.
Request → JWT decode → tenant_id
│
Redis counter
INCR + EXPIRE
│
quota check (DB)
┌──────┴──────┐
pass 429
│
Policy Engine
priority groups
→ weighted selection (canary)
→ region / latency filter
→ health check
→ primary / fallback / fallthrough
│
Downstream AI Service
│
BackgroundTask
→ RequestLog (DB)
┌──────────────────────────────────────────────────────────┐
│ FastAPI Application │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌───────────┐ │
│ │ Registry │ │ Policies │ │ Quotas │ │ Observe │ │
│ │ API │ │ + /route │ │ API │ │ API │ │
│ └──────────┘ └──────────┘ └──────────┘ └───────────┘ │
│ ┌──────────────────────────────────────────────────┐ │
│ │ APScheduler — Health Check Cycle │ │
│ └──────────────────────────────────────────────────┘ │
│ ┌──────────────────────────────────────────────────┐ │
│ │ Policy Engine + Rate Limiter + Observer │ │
│ │ priority → weight → topology → health │ │
│ └──────────────────────────────────────────────────┘ │
└───────────────────┬──────────────────────────────────────-┘
│
┌────────────┴────────────┐
▼ ▼
PostgreSQL + Alembic Redis 7
(services, policies, (rate limit counters)
quotas, request_logs)
| Attribute | Type | Analogy |
|---|---|---|
region |
string (eu-west, on-premise) |
BGP community — restricts routing to certain zones |
latency_zone |
enum low / medium / high |
OSPF link cost — lower cost paths preferred |
network_tags |
list of strings (gpu, air-gapped) |
BGP extended communities — free-form route filtering |
| Layer | Technology |
|---|---|
| API Framework | FastAPI 0.115 |
| ORM | SQLAlchemy 2.0 (async) |
| Database | PostgreSQL 18 |
| Migrations | Alembic 1.18 |
| Cache / Rate Limiting | Redis 7 (async via redis-py) |
| Auth | python-jose (JWT) |
| Scheduling | APScheduler 3.x |
| HTTP Client | httpx (async) |
| Validation | Pydantic v2 |
| Testing | pytest + pytest-asyncio + fakeredis |
| Containerization | Docker / docker-compose |
ai-control-plane/
├── alembic/
│ ├── env.py
│ └── versions/
├── app/
│ ├── api/v1/
│ │ ├── __init__.py
│ │ ├── registry.py # service CRUD + status override
│ │ ├── policies.py # policy CRUD + /route + /weight
│ │ ├── quotas.py # quota CRUD + live counter + reset
│ │ └── observe.py # summary, traffic, errors, latency
│ ├── core/
│ │ ├── config.py
│ │ ├── database.py
│ │ ├── redis.py
│ │ └── security.py # JWT → tenant_id
│ ├── models/
│ │ ├── service.py # Service + ServiceStatus + LatencyZone
│ │ ├── policy.py # Policy with weight + network match conditions
│ │ ├── quota.py # Quota per tenant
│ │ └── request_log.py # RequestLog for observability
│ ├── schemas/
│ │ ├── service.py
│ │ ├── policy.py # includes PolicyWeightUpdate
│ │ ├── quota.py
│ │ └── observe.py
│ ├── services/
│ │ ├── health_checker.py
│ │ ├── policy_engine.py # priority groups → weighted selection → topology → health
│ │ ├── rate_limiter.py
│ │ └── observer.py
│ └── main.py
├── tests/
│ ├── conftest.py
│ ├── test_health_checker.py # 7 tests
│ ├── test_policy_conflicts.py # 4 tests
│ ├── test_policy_engine.py # 13 tests
│ ├── test_rate_limiter.py # 13 tests
│ ├── test_observe.py # 6 tests
│ └── test_canary_rollout.py # 8 tests
├── pytest.ini
├── requirements.txt
├── Dockerfile
├── docker-compose.yml
└── .env.example
- Python 3.12+
- PostgreSQL 18+
- Redis 7+ (or
docker run -d --name redis -p 6379:6379 redis:7-alpine)
cp .env.example .env
# edit DATABASE_URL, REDIS_URL, JWT_SECRET_KEY, RATE_LIMIT_ENABLED in .env
pip install -r requirements.txt
alembic upgrade head
uvicorn app.main:app --reloaddocker-compose up --buildAPI docs: http://localhost:8000/docs
python -m pytest tests/ -v --timeout=15
⚠️ Production note: never deploy with default credentials from.env.example. Always set strong values forDATABASE_URL,REDIS_URL, andJWT_SECRET_KEY.
| Method | Path | Description |
|---|---|---|
POST |
/api/v1/registry |
Register a service (with region, latency_zone, network_tags) |
GET |
/api/v1/registry |
List all services + aggregate health summary |
GET |
/api/v1/registry/{id} |
Fetch a single service |
PATCH |
/api/v1/registry/{id}/status |
Override a service's health status |
DELETE |
/api/v1/registry/{id} |
Deregister a service |
| Method | Path | Description |
|---|---|---|
POST |
/api/v1/policies |
Create a routing policy (with weight) |
GET |
/api/v1/policies |
List all policies ordered by priority |
GET |
/api/v1/policies/{id} |
Fetch a single policy |
PATCH |
/api/v1/policies/{id} |
Update a policy |
PATCH |
/api/v1/policies/{id}/weight |
Fast canary weight adjustment |
DELETE |
/api/v1/policies/{id} |
Delete a policy |
POST |
/api/v1/route |
Resolve which service handles a request |
| Method | Path | Description |
|---|---|---|
POST |
/api/v1/quotas |
Create a quota for a tenant |
GET |
/api/v1/quotas/{tenant_id} |
Fetch quota + live Redis counter |
PATCH |
/api/v1/quotas/{tenant_id} |
Update quota limits or active flag |
DELETE |
/api/v1/quotas/{tenant_id}/counter |
Reset the Redis counter for a tenant |
| Method | Path | Description |
|---|---|---|
GET |
/api/v1/observe/summary |
Snapshot: service health, active policies, request volume |
GET |
/api/v1/observe/traffic?hours=1 |
Traffic distribution by service, resolution, and policy weight |
GET |
/api/v1/observe/errors?hours=1 |
Error breakdown by service |
GET |
/api/v1/observe/latency?hours=1 |
Average latency per service |
| Method | Path | Description |
|---|---|---|
GET |
/health |
Liveness check for the control plane itself |
- ✅ Phase 1 — Service Registry & Health Checking
- ✅ Phase 2 — Policy-Based Routing with Network-Aware Constraints
- ✅ Phase 3 — Rate Limiting & Quota per Tenant
- ✅ Phase 4 — Observability Dashboard
- ✅ Phase 5 — Canary Rollout with Weighted Traffic Splitting
Rate limiting — fixed-window algorithm
The current implementation uses a fixed-window counter. This means up to 2 × max_requests can pass through at the boundary between two windows (burst at the end of window N + start of window N+1). A sliding-window or token-bucket algorithm would eliminate this. Acceptable for the current scale; noted here for transparency.
Single-instance scheduler The APScheduler health-check job runs inside the app process. Running multiple replicas will cause each instance to independently probe downstream services and write status updates concurrently — producing redundant probes and potential write contention. For HA deployments, the scheduler should be extracted to a separate worker or use leader-election.
Decision plane only — no data plane
/route returns the name of the target service; it does not proxy the request. Callers are responsible for forwarding traffic to the resolved service. This is intentional for the current scope but means the control plane is advisory, not intercepting.
In-memory cache absent
Every /route call hits PostgreSQL for policy and service lookups. Under high request rates, the DB becomes the bottleneck. A short-lived in-memory cache (e.g. per-process TTL cache on policies) would significantly improve throughput.
Canary routing is stateless
Traffic splitting uses random.choices() per request with no session affinity. A single user may be routed to both stable and canary versions across consecutive requests. For use cases requiring sticky routing, a hash of tenant_id or session identifier should replace the pure random selection.
'@
Add-Content C:\ai-control-plane\README.md $limitation
MIT