Start here

↑ ↓ to choose · Enter to open · Esc to close

DEVELOPER DOCS / API V1

Change the route.
Keep the request.

Route calls to any model, pay in USDG or with a Stock Token, and verify every receipt.

The router serves this site, so your base URL is this site’s address followed by /api/v1. Create a key in the dashboard, deposit USDG to it and call it from any OpenAI- or OpenRouter-compatible client.

What’s new

Agent rules, Telegram linking and approvals, opt-in public profiles and encrypted chat are switched on at anyroute.tech. Sealed agent hosting is available, but no sealed agent is registered at anyroute.tech yet. Agreements are switched on, with the escrow and dispute contracts deployed on Robinhood Chain; automatic jury rulings are switched on: three models on attested hardware rule by two of three, and the signed ruling is posted on-chain; a hung jury goes to the panel, and anything unruled after 30 days settles 50/50. The network is open for early hosts running the approved build, with automatic admission and probation. Live network statistics and host bond indexing are switched on; payouts, fee buy-and-burn and slashing are not switched on at anyroute.tech yet.

Create capped keys and restrict access.

Create and update keys with include_byok_in_limit; responses echo the stored boolean. usage remains the router’s billed USD total. byok_usage is zero: provider-side BYOK expenditure is not measured here, and router charges appear only in usage. The selection does not change billing or budgets. DELETE adds a top-level deleted: true and keeps the data envelope; it disables the key and retains its records.

List keys with offset and limit (1–1000). An offset without a limit uses pages of 100; no pagination parameters preserves the full list. Disabled keys remain included; past the end the data array is empty. Order is newest first, with key hash breaking ties. Pages reflect current rows rather than a frozen snapshot.

USD floats keep their existing meaning and can round in JSON. usage_pico_usd and usage_period_pico_usd are exact integer strings (one USD = 10¹² pico-USD). usage_micro_usd is an integer string rounded up from the total to micro-USD (one USD = 10⁶ micro-USD); it is not an exact sub-micro amount. BYOK integer fields are zero. Read usage after disabling a key with its management credential. Disabling does not drain accepted calls; usage is a current counter, not a final settlement certificate.

Inference-only key provisioning is switched on at anyroute.tech; self-hosted routers enable it with INFERENCE_KEYS_ENABLED (default false). Where enabled, a management key may create a child with scope: "inference", or PATCH /api/v1/keys/defaults with that scope to make newly minted children default to it. GET that endpoint reads the account default. Existing keys retain access. A management key can explicitly create with scope: "account", including a successor management key; scope is fixed at creation. Inference-only keys cannot have management rights.

An inference-only key may call chat/completions, completions, embeddings, responses and messages, list models, and read its own generations and receipts. Saved account routes, presets and characters cannot be selected through a model name. Other routes return 403, including key administration, balances, teams, agents and account feeds. Stored restrictions stay active even when provisioning is switched off. Calls remain linked to the operator account and the router reads ordinary request text in memory. Public receipts remain accessible by identifier without a key; the scope limits authenticated access and does not make those receipts confidential.

Agent rulebook

Attach a rulebook with PUT /api/v1/agents/:key_hash/policy. It controls allowed models, lanes, declared tools, UTC windows, output tokens and estimated request cost, with rolling hour, day and week caps from charged usage plus open reservations. A session key and its parent must both pass. Only account management keys and team owners or admins can manage rulebooks.

GET /api/v1/agents lists keys; /agents/me describes the calling key and inherited caps. POST /agents/check evaluates an Intent without recording it. POST /agents/:key_hash/kill refuses the next request across router replicas; /resume clears the kill state. GET /agents/:key_hash/events returns a hash chain of metadata decisions. A breach can deny or kill; costs above an approval threshold return agent_approval_required.

AGENT_POLICY_ENABLED defaults to false, with these routes returning 404. Enforcement is deterministic code in AnyRoute’s router for requests through AnyRoute only. The router reads request text in memory on ordinary inference paths. These rules have no on-chain enforcement. Decision events contain model, lane, cost, token and tool identifiers, with 90-day retention when the cleanup worker runs.

Switched on at anyroute.tech.

Read and respect rules from MCP or an SDK

The SDK code is in the repository; npm and PyPI releases are not published yet.

Check before expensive calls and never retry a denied call unchanged. Connect MCP at POST /mcp with the calling key in the Authorization header. anyroute_agent_rules returns that key’s rules, inherited rules, remaining rolling caps and kill state. anyroute_agent_check returns the REST decision and reasons unchanged; token estimates use the current model catalog. Supply est_cost_pico instead to provide an explicit cost estimate.

{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{
  "name":"anyroute_agent_rules","arguments":{}}}
{"jsonrpc":"2.0","id":2,"method":"tools/call","params":{
  "name":"anyroute_agent_check","arguments":{
    "model":"author/model","lane":"public",
    "est_input_tokens":100,"max_output_tokens":32,"tools":[]}}}

Both SDKs return the data object from /agents/me and /agents/check. Costs use decimal strings in pico USD: 10^12 pico equals one USD. A dry run reserves no budget, sends no prompt and creates no policy event; a later request is evaluated again and may be refused. Rulebook reads and dry runs remain available when the key is killed or its tool allowlist excludes these inspection tools.

import { AnyRoute, AgentPolicyDenied, AgentKilled,
  AgentApprovalRequired } from "@anyroute/client";

const client = new AnyRoute({ baseUrl: routerUrl, apiKey });
const rules = await client.agent.rules(); // inherited rules and remaining USD caps
const decision = await client.agent.check({
  kind: "inference", model: modelId, lane: "public",
  est_cost_pico: "1000000000", max_output_tokens: 32, tools: [],
});
if (decision.decision === "allow") {
  try {
    await client.chat.completions.create({
      model: modelId, messages: [{ role: "user", content: "Hello" }], max_tokens: 32,
    });
  } catch (error) {
    if (error instanceof AgentPolicyDenied || error instanceof AgentKilled ||
        error instanceof AgentApprovalRequired) {
      console.error(error.message, error.reasons);
      // Never retry the denied call unchanged.
      // Approval refusals expose approval_id and poll when the router supplies them.
    } else { throw error; }
  }
} else { console.error(decision.reasons); }
from anyroute_client import (AnyRoute, AgentPolicyDenied,
    AgentKilled, AgentApprovalRequired)

with AnyRoute(router_url, api_key) as client:
    rules = client.agent.rules()
    decision = client.agent.check({
        "kind": "inference", "model": model_id, "lane": "public",
        "est_cost_pico": "1000000000", "max_output_tokens": 32, "tools": [],
    })
    if decision["decision"] == "allow":
        try:
            client.chat({"model": model_id, "messages": [
                {"role": "user", "content": "Hello"}], "max_tokens": 32})
        except (AgentPolicyDenied, AgentKilled, AgentApprovalRequired) as error:
            print(str(error), error.reasons)
            # Never retry the denied call unchanged.
    else:
        print(decision["reasons"])

AgentPolicyDenied, AgentKilled and AgentApprovalRequired preserve the router’s message, reasons, status, policy_sha256 and metadata in details. Approval refusals expose approval_id and poll only when supplied; these SDK methods do not create approvals or poll automatically. AGENT_POLICY_ENABLED must be enabled by the operator; it defaults to false. These interfaces do not change who reads inference text: AnyRoute’s router reads ordinary chat text in memory; the encrypted-chat adapter forwards ciphertext.

Switched on at anyroute.tech.

Circuit breakers

An optional rulebook field, breakers, accepts max_spend_usd_per_minute (positive USD up to 1,000,000), max_requests_per_minute, max_denials_per_10min and max_distinct_models_per_hour (positive safe integers). Omitted fields add no restriction. AGENT_POLICY_ENABLED, default false, also controls breakers.

Switched on at anyroute.tech.

Before each policy admission, the router checks recorded rolling counters under the account lock. A value equal to or above a limit trips the kill switch with reason breaker: followed by the field name. That request receives HTTP 403 agent_killed before any reservation or charge. A trip overrides an on_breach action of deny and cannot permit anything another rule denies. These controls apply to requests through AnyRoute only, with no on-chain enforcement.

Spend is charged usage from the last minute plus open usage reservations. Requests and denials count policy admission batches; a batch with several models counts once, and every distinct requested model is observed, including denied models. An MCP tool admission and its routed inference are separate admissions. Approval-required attempts count as requests, not denials. Earlier decision events without a batch marker each count as an observation.

The principal resumes manually. Resume resets breaker observations and the spend window while leaving budget-cap spend intact; in-flight work may finish and later charges can trip the breaker again. Spend breakers also refuse an estimate that would cross the limit, and distinct-model breakers check the complete requested model set before admitting it. Equality is permitted for a new estimate or model set; the next admission trips at the recorded threshold. Final charges can exceed estimates; budget caps remain independent.

Breaker metadata uses the existing rulebook and event tables, with the same 90-day event retention. No prompt or response text is added. Ordinary inference paths still let AnyRoute’s router and the answering provider read request text in memory. The agent page includes breaker fields and a “Tripped by” badge.

Progressive autonomy

An optional autonomy rulebook field defines up to five ascending rungs. Each rung has integer after_days and clean_requests requirements, both nonnegative, and a strictly increasing caps_multiplier above 1 and at most 10. Requirements cannot decrease between rungs. Both requirements apply since reaching the current rung; the request counter starts again at every climb.

{
  "autonomy": {
    "rungs": [
      {
        "after_days": 1,
        "clean_requests": 100,
        "caps_multiplier": 2
      },
      {
        "after_days": 7,
        "clean_requests": 1000,
        "caps_multiplier": 4
      }
    ],
    "demote_on": [
      "deny",
      "kill",
      "breaker"
    ]
  }
}

The agent starts at rung 0 with caps ×1. The selected demote_on events return it to rung 0 and clear progress. A deny means a recorded policy refusal; kill means a recorded kill event; breaker means a recorded agent breaker event. Saving a rulebook restarts progress. Removing autonomy restores the original caps. Existing rulebooks without the field retain their canonical hash and behaviour.

Only per-request and rolling hour, day and week spending caps scale, using integer pico-USD arithmetic with fractional pico values rounded down. Output token limits, models, tools, lanes, UTC windows and approval thresholds never scale. A clean request is counted once per applicable rulebook after all policy checks allow it and its spending reservation succeeds; cached inference and allowed MCP tool checks also count. Dry runs, approval-required decisions, refused requests and failed reservations do not earn progress. This counts routing authorization, not successful provider delivery.

GET /api/v1/agents/me and each applicable rulebook in the agent list report autonomy with the current rung, since timestamp, clean request count, multiplier and next requirements with remaining days and requests; effective_caps reports scaled caps. A session and its parent advance independently and both must pass. The /agents page shows a ladder; configure rungs through the rulebook API. Editing other settings in the page preserves autonomy.

State comes from the metadata event chain, with rung, timestamp and counter checkpoints. Events normally expire after 90 days; active autonomy rulebooks retain their latest checkpoint and following events to preserve progress during inactivity. No prompt text enters these records. AGENT_POLICY_ENABLED defaults to false; it gates autonomy and all rulebook enforcement. These are router rules for requests through AnyRoute, with no on-chain enforcement. AnyRoute’s router reads ordinary chat text in memory on every lane; the encrypted-chat adapter forwards ciphertext.

Switched on at anyroute.tech.

Statements and account export

Monthly statements are switched on at anyroute.tech. Self-hosters enable STATEMENTS_ENABLED (default false) to expose GET /api/v1/statements/:yyyy-mm. Supply a bearer API key; management and owner/admin keys read account totals. Ordinary and session keys read only their own attributed movements, labelled scope: key. This is the same scope rule as Activity.

Example request
GET /api/v1/statements/2026-09?format=json
Authorization: Bearer <your-key>

The response is { data: { payload, alg: "Ed25519", key_id, sig } }. The existing receipt key signs canonical JSON. Verify it in Verify by pasting or dropping JSON. The signature establishes what the router signed; it does not independently prove ledger completeness.

Amounts are exact USDG decimal strings. Opening + deposits + refunds − usage − separate fees + other changes = closing. Each usage grouping sums to charged ledger usage. Month boundaries use UTC, usage follows settlement time, and call counts follow generation time. Missing attribution appears as null. Balances exclude holds and external wallet escrow; embedded usage fees are not added again. Key-only balances are attributed sums, not the account’s shared balance.

Invalid months return 400. Months before account creation or in the future return 404. The current month is marked so_far: true. Disabled statements return 404. Responses use Cache-Control: no-store; later ledger entries can change an earlier statement.

Open Statements under Money for signed JSON or a printable view; use the browser’s Save PDF option. Open Export your data under Account for one browser-built bundle of readable account records with progress and cancel. Existing API scopes and retention apply. Approvals have an existing 100-per-status cap; unavailable sections and per-agent access limits are listed. No key secrets are included. Chat history lives in the browser and is exported from the Harness. Read What we keep for what the router retains beyond this bundle.

Activity & receipts

GET /api/v1/activity is the account-wide list behind Activity in the dashboard: calls (with receipt ids and a link to check each receipt), agent approvals and rule decisions, alerts, deposits and balance changes, and agreement events, newest first in one item shape with exact USDG amounts. Filter with kind, key, model, from and to; page with limit (up to 100) and cursor; format=csv or json exports the filtered page. A key sees its own activity; management, owner and admin keys see the account within their team boundaries; never another account. GET /api/v1/inbox lists what needs attention or is new: pending approvals with their expiry, alerts, deposits, agreement events and status updates for hosts the account operates. The browser keeps the seen time; POST /api/v1/inbox/seen only returns it, and nothing is stored. Both need a key and are switched on at anyroute.tech.

GET /api/v1/agents/:key_hash/ledger lets the principal read a key’s activity. It uses the same account and team permissions as editing that key. GET /api/v1/agents/me/ledger lets an agent or session key read only its own activity. Both require AGENT_POLICY_ENABLED, which defaults to false; off returns 404.

Switched on at anyroute.tech.

Use from and to as ISO 8601 timestamps with an explicit timezone: from is inclusive and to is exclusive. Results are newest first, with up to 100 request rows, an opaque next_cursor and totals_per_day for the full date range in UTC. Send next_cursor as cursor for the next page. format=json returns rows and daily totals; format=csv streams this page, with X-Next-Cursor when another page exists. Downloads on /agents contain the selected page. CSV quotes cells and prefixes spreadsheet formula markers with an apostrophe.

Each correlated request includes its policy decision, model, lane, exact generation token counts, charged USD and pico-USD, signed receipt identifiers and links to the existing /verify receipt view and checker. Denied requests cost zero and have no receipt. Approval IDs are present only when an approval was used. Multiple models or inherited rulebooks share one row, with receipts, event_ids and policy_sha256s arrays; policy_sha256 is the first sorted digest. Costs come from generations, never estimates. An allow decision without a generation is not evidence that inference finished.

Existing records without explicit correlation are returned separately with unlinked=true; their policy, approval, or receipt association is not inferred. Historical generations have an allow label for recorded usage with no recorded rulebook decision. Policy decisions and their correlation metadata are retained for 90 days while the cleanup worker runs; generations keep their existing retention. Requests rejected before policy evaluation or generation creation have no row. The ledger stores no prompt or answer text. AnyRoute’s router reads ordinary chat text in memory, and rulebooks apply to requests through AnyRoute only, with no on-chain enforcement. Verify links do not certify a signature; use the existing receipt checker.

Agent agreements: escrow and model jury

Switched on at anyroute.tech: AgreementEscrow (0xefd8d05f45b8a92aa3b3ef3a7db4c9d3a21f7c96) and DisputeOracle (0xcdeddcea1e039e72868bb8af3206af2647afda5a) are deployed on Robinhood Chain and verified on Sourcify. Automatic jury rulings are switched on: three models on attested hardware rule by two of three, and the signed ruling is posted on-chain; a hung jury goes to the panel, and anything unruled after 30 days settles 50/50. AGENT_AGREEMENTS_ENABLED defaults to false for self-hosters. AGREEMENT_ESCROW_ADDRESS and DISPUTE_ORACLE_ADDRESS must identify distinct nonzero contracts. AGENT_AGREEMENTS_RULINGS_ENABLED defaults to false.

Agreement contracts

The agreement contracts are deployed on Robinhood Chain (AgreementEscrow 0xefd8d05f45b8a92aa3b3ef3a7db4c9d3a21f7c96, DisputeOracle 0xcdeddcea1e039e72868bb8af3206af2647afda5a; Sourcify exact match). They hold USDG for individual milestones: the payee submits a delivery digest, the payer releases payment, and the payee may claim after an unanswered review window. After the deadline, the payer may reclaim undelivered milestones. A dispute locks only its milestone.

A complete signed jury tally can pay the payee, refund the payer or split the milestone. If no verdict reaches the configured majority, a separate panel may decide before dispute expiry. After the immutable dispute timeout, anyone may settle an unresolved dispute with a neutral 50/50 split; the payee receives any rounding remainder. Rulings at or after expiry are refused, even before recovery is called. Only the recorded payer and payee can receive funds. There is no administrator withdrawal, upgrade mechanism or pause.

The optional jury service runs through AnyRoute’s router on the attested lane. The contracts do not verify model execution or hardware attestation. Jury keys and panel decisions remain trusted. Unavailable jury signers or an unavailable panel can delay settlement until expiry. Recovery requires a transaction, and token transfer restrictions can still prevent payouts. The router reads ordinary request text in memory; these contracts add no prompt encryption.

DISPUTE_TIMEOUT_DAYS defaults to 30 in the deployment script, with an allowed range of 7–180 days. The timeout starts when each dispute opens and cannot be extended by the jury or panel. AGENT_AGREEMENTS_ENABLED defaults to false in the configuration loader and deployment script. The service additionally requires configured escrow and oracle addresses. Deployment also requires an explicit operator decision to broadcast. Wallet addresses, amounts, deadlines, digests, jury signatures and rulings are permanently public on chain. Digests can reveal guessable content.

The service uses ABIs generated from the Solidity artifacts and indexes both contracts. GET /api/v1/agreements returns individual milestones, identified as agreement.milestone, for the caller’s wallet, 50 per page with cursor and next_cursor. GET /api/v1/agreements/:id returns canonical state, both parties’ evidence and the signed jury statement. Access requires a wallet-linked authenticated account matching payer or payee.

POST /api/v1/agreements/prepare accepts payee, terms_hash, milestone_amounts_usdg_units (1–64 positive decimal token amounts), and deadline (a future Unix timestamp in seconds). It prepares createAgreement(address,bytes32,uint256[],uint256,address) with the configured oracle. The payer approves USDG and signs the funding transaction. The review window and DISPUTE_TIMEOUT are immutable escrow settings. Rulebooks check the sum of milestone amounts and the counterparty, inherited tools, working hours and kill state. Preparation reserves no funds, aggregates no escrow spending and cannot enforce transactions sent elsewhere. SDK agreements.prepare and MCP anyroute_agreement_prepare use this endpoint; MCP also exposes anyroute_agreement_status and anyroute_agreement_evidence.

POST /api/v1/agreements/:id/evidence accepts JSON, including text as a JSON string. AGREEMENT_EVIDENCE_BYTES defaults to 16384, with at most 32 entries per party. Evidence is encrypted at rest with APP_SECRET. Both parties, the router and the configured jury models can read it. The router reads evidence and ordinary request text in memory. AGREEMENT_EVIDENCE_WINDOW_SECONDS defaults to 86400 after dispute opening. A frozen, deterministic bundle binds the deployment, agreement, milestone, creation, dispute, terms, deliverable and opening-evidence hashes and uploaded material. Its evidenceRoot uses RFC 6962 SHA-256 domains over canonical JSON. Hashes alone cannot establish compliance; the rubric requires abstention when material facts are missing.

AnyRoute runs the jury on attested models. AGREEMENT_JURY_INTERNAL_ENABLED defaults to false; enable it to call providers with AnyRoute’s stored credentials without an API key or customer billing. With AGREEMENT_JURY_API_KEY unset, this internal transport is used; when set, the funded API-key chat transport remains in use. AGREEMENT_JURY_MODELS defaults to auto: the first three distinct currently selectable, freshly attested model IDs in sorted order. An explicit comma-separated list accepts up to 16; the full configured size must be selectable or the dispute stays awaiting jury. AGREEMENT_JURY_THRESHOLD defaults to 2 and requires a strict majority. Votes retain attestation references and router-signed operational receipts in the party-only jury statement. Gateway receipts are checked for signature, request/response digests and upstream attestation; other providers carry a verified attestation reference rather than a provider-signed answer receipt. Internal receipts have no public generation link. Canaries have no separate operational-cost ledger; the jury statement retains provider-list-price usage estimates, not invoices. Failed calls may still incur unmeasured cost. Distinct models may share operators or hardware. Attestation does not establish verdict correctness or resistance to malicious evidence.

Run dry-run first with rulings disabled and no AGREEMENT_JURY_SIGNER_KEYS. An ordinary worker may run agreement-jury alongside indexing and retention; it computes and stores would-be rulings and posts nothing. TLOG_ENABLED is optional for dry-run and required for posting: the statement’s Ed25519 key must be published before a ruling. The router reads evidence in memory and sends it to the selected providers.

When posting is enabled, AGREEMENT_JURY_SIGNER_KEYS is a comma-separated list of one distinct operator-held Ethereum key per ordered model. Production permits these keys only in an isolated WORKER_JOBS=agreement-jury worker, without other transaction signing roles. Startup and every posting pass verify complete oracle jury membership and the configured threshold. Each key signs the oracle’s EIP-712 voteDigest, binding chain, oracle, escrow, agreement, milestone, immutable dispute context, evidenceRoot, jury version and its model’s exact payee bps. Models do not hold or generate these keys. This is a router-run jury: the operator controls all configured signing keys; the contract verifies key authorization, not model execution or hardware attestation.

A complete tally must contain one valid non-abstaining vote per configured signer. Refund is 0 bps, pay is 10000 bps, and split is 1–9999 bps. Exact bps must agree. The oracle records participation and consensus bitmaps in its configured signer order, plus a hash of the signed tally; the statement’s tally_bitmap uses model order. A complete hung tally reaches PanelPending and only the separately configured panel may call postPanelRuling with that evidence root. Failed or abstaining model calls cannot be encoded as refund votes and do not create an on-chain panel path. They remain unresolved until a valid tally or expiry. The service never automatically signs a panel ruling.

After DISPUTE_TIMEOUT, anyone may call resolveStaleDispute for a 50/50 split; an odd base unit goes to the payee. Jury and panel rulings cannot extend this deadline. The index records Settled payout amounts exactly. A successful posting transaction can record PanelPending without moving funds; posted denotes receipt success, while indexed state determines settlement.

AGREEMENT_START_BLOCK defaults to 0 and must include milestone funding and creation. AGREEMENT_FINALITY defaults to finalized (safe is also supported); indexing waits for CHAIN_CONFIRMATIONS, verifies block hashes and replays on reorganization. Run agreement-indexer, agreement-jury and agreement-retention workers as needed. Evidence and model reasons are deleted after canonical settlement plus AGREEMENT_RETENTION_DAYS (default 30) while the index is fresh and retention runs. DELETE /api/v1/agreements/:id/evidence allows each party to delete its own evidence after that interval. Public chain data remains permanent; backups follow the operator’s policy. The agreement schema needs its additive migration before the service can run.

Sign account webhooks

WEBHOOK_SIGNING_ENABLED defaults to false. Signed webhooks are switched on at anyroute.tech. When enabled, Webhooks in your account manages destinations, events, rotation, revocation and the last 100 delivery attempts. An account management key adds destinations. Owner/admin keys outside agent sessions can manage linked destinations scoped to their own key. The disabled flag preserves existing alert behaviour.

New destinations receive a random secret shown once on creation; rotation reveals a replacement once. The router stores it encrypted under APP_SECRET and reads it in memory to sign notices. Existing destinations remain unsigned until rotated. Revocation removes the credential and stops future delivery; a request already in flight can finish. Keep the receiver’s secret outside source code. URLs use guarded HTTPS egress with public DNS answers pinned and redirects refused.

Each enabled delivery carries x-anyroute-event-id. Signed deliveries also carry x-anyroute-signature: t=<unix>,v1=<hex>, where v1 is HMAC-SHA256 using the literal secret string over timestamp, a period and the exact body bytes. Signed bodies include event_id matching the header. Verify before parsing or processing, reject timestamps more than five minutes in either direction, compare in constant time, and atomically reject duplicate event IDs. Retries keep the event ID and receive a fresh timestamp/signature.

Subscribe to spending and agent alerts, approval requests and decisions, credited deposits, funded/disputed/ruled agreements and observed status changes of hosts your account operates. New notices contain only id, type, source reference, time and fixed status; linked spending and agent alerts retain their existing fields with event_id added when signed. Approval references name the approval record. No prompt or answer text is included. Event availability follows the source feature flags.

The minute worker reads 50 destinations per tick, one activity page per destination, and sends at most 100 notices per tick. Creation sets the lower time bound; durable pagination and five minutes of overlap catch ordinary late commits. Source deletion before discovery or commits delayed beyond that overlap can lose notices. Approval requests/decisions and operated host status changes are queued when the router records them, in the same transaction. Changes between worker ticks are retained. Direct database writes outside those router writers do not create notices. Three delivery attempts are separated by five minutes. Delivery metadata stays for 90 days, with no bodies stored. A stopped worker delays delivery. A crash after sending can repeat a notice.

Use GET/POST /api/v1/webhooks, PATCH/DELETE /api/v1/webhooks/{id}, POST /rotate or /revoke, GET /deliveries and POST /test under that destination path. The endpoint check is queued, limited to one per minute and sent even if no event subscription matches. Removing a linked destination requires removing its URL in Spend Watch; revocation can stop it here.

TypeScript verification
import { createHmac, timingSafeEqual } from 'node:crypto';
// rawBody is a Buffer of the original HTTP body, before JSON parsing.
function verify(secret: string, rawBody: Buffer, signature: string,
                eventId: string, now = Math.floor(Date.now() / 1000)) {
  const m = /^t=(\d{1,12}),v1=([a-f0-9]{64})$/.exec(signature);
  if (!m || Math.abs(now - Number(m[1])) > 300) return false;
  const expected = createHmac('sha256', secret)
    .update(m[1] + '.').update(rawBody).digest();
  if (!timingSafeEqual(expected, Buffer.from(m[2], 'hex'))) return false;
  try { return JSON.parse(rawBody.toString('utf8')).event_id === eventId; }
  catch { return false; }
}
// Read x-anyroute-signature and x-anyroute-event-id; verify first.
// Then atomically claim eventId in durable storage before side effects.
// A repeated ID returns success without repeating the side effects.
// Keep processed IDs for at least 90 days; use an outbox for effects.
Python verification
import hashlib, hmac, json, re, time

def verify(secret, raw_body, signature, event_id, now=None):
    # raw_body: exact received bytes, before decoding or JSON parsing.
    now = int(time.time()) if now is None else now
    m = re.fullmatch(r"t=(\d{1,12}),v1=([a-f0-9]{64})", signature)
    if not m or abs(now - int(m[1])) > 300:
        return False
    expected = hmac.new(secret.encode(), m[1].encode() + b"." + raw_body,
                        hashlib.sha256).hexdigest()
    if not hmac.compare_digest(expected, m[2]):
        return False
    try:
        return json.loads(raw_body)["event_id"] == event_id
    except (ValueError, KeyError, TypeError):
        return False

# After verification, atomically claim event_id in durable storage.
# Keep processed IDs for at least 90 days, and use an outbox for effects.
# Return success for repeated IDs without repeating their effects.

Agent alerts

Add alerts: {} to a rulebook to enable owner alerts. Omitted alerts preserve existing behavior and rulebook hashes. Optional at_percent defaults to [80, 100]; denials_in_10min defaults to 5; channels defaults to ["webhook", "telegram"].

Cap percentages use rolling hour, day and week spend, including open reservations. Each threshold has a cooldown of that window’s duration, including across policy edits. Denied request batches trigger once per 10 minutes. Kills, including manual stops and on_breach, and newly created approval requests also add entries. Any breaker that uses the rulebook kill operation produces a kill alert.

Read GET /api/v1/agents/{key_hash}/alerts with the same principal permissions as editing the key, or open the feed on /agents. The feed keeps the newest 100 entries per account for up to 90 days. No prompt, answer, intent or arbitrary kill reason is copied. The router reads ordinary chat text in memory on every lane; the encrypted-chat adapter forwards ciphertext.

Delivery reuses enabled Spend Watch webhook destinations scoped to the agent or its account, and Telegram links authenticated as account management keys or same-team owners/admins. Spend Watch webhooks use guarded HTTPS egress and remain unsigned unless the optional signing feature is enabled and the owner rotates their secret. Read signing instructions. Email alerts are not switched on at anyroute.tech yet. Telegram alert delivery uses AnyRoute’s bot when the owner has linked it. Email has no linked account destination; without a selected linked channel, alerts remain in the feed only. TELEGRAM_LINKING_ENABLED (default false) adds account links through /agents or dashboard settings and /link in the bot; it requires the agent flag and bot token. This linking flow is switched on at anyroute.tech. Unlinking stops delivery through the account link; a separate chat-key connection remains until /forget.

AGENT_POLICY_ENABLED defaults to false; off means no alerts and this endpoint returns 404. Workers attempt delivery every minute, up to 10 notices per account per minute, at most 20 destinations per channel, with three attempts separated by five minutes. Successful destinations are skipped on retry; a crash between delivery and persistence may repeat a notice. Requests through AnyRoute’s router are the enforcement boundary; there is no on-chain enforcement.

Switched on at anyroute.tech.

Your agent asks you first

With AGENT_POLICY_ENABLED=true (default false), costs above a rulebook’s approval.above_usd return HTTP 403 agent_approval_required with metadata.approval_id, expires_at and poll. The router creates one pending approval for an identical intent and requesting key; a repeated request reuses it until expiry.

Switched on at anyroute.tech.

Approve on /agents, or in Telegram after an owner links it (switched on at anyroute.tech).

GET /api/v1/agents/approvals?status=pending lists up to 100 newest approvals visible to a management key or an owner/admin of the same account and team. The Agents page shows “Waiting for you” with Approve and Deny buttons. POST /api/v1/agents/approvals/:id/approve or /deny records that principal’s decision. GET /api/v1/agents/approvals/:id returns only id, status and expires_at to that principal or the requesting key. Session and ordinary agent keys cannot approve themselves.

After approval, retry with X-Agent-Approval: <id>. Approval is single-use, belongs to the requesting key, and binds the model, lane, declared tools and output token limit. The estimated cost may decrease but cannot exceed max_cost_pico. Council and dual requests bind every model and their combined estimate. Other caps, schedules and the kill switch are evaluated again. An approval is consumed with a successful reservation, even if the provider later fails; a cache hit can also consume it. Approval does not bind prompt contents; prompt text is excluded from the intent.

AGENT_APPROVAL_TTL_S defaults to 900 seconds, measured from the request, with bounds of 1–86400. Approval does not extend the deadline. Expiry ends validity; approval rows remain until operator deletion. Approval request, decision and use events enter the existing per-key hash chain, whose retention is 90 days when its worker runs. Approval metadata retains readable model and tool identifiers, estimated cost, token limit, principal key hash and timestamps. AnyRoute’s router reads ordinary chat text in memory on every lane; the encrypted-chat adapter forwards ciphertext.

TELEGRAM_LINKING_ENABLED defaults to false and requires AGENT_POLICY_ENABLED=true, plus TELEGRAM_BOT_TOKEN on the worker that runs the bot and alert jobs (the API only issues and checks link codes). Telegram linking and approvals are switched on at anyroute.tech. When enabled, Link Telegram on /agents or dashboard settings issues a single-use code valid for five minutes, with five issuance attempts per account per minute. Send /link <code> in a private chat with AnyRoute’s bot. No API key is sent to Telegram. Only management keys or account/team owner/admin keys may link; sessions cannot link. There is one Telegram identity per principal key, and one account per Telegram identity. Creating a new code invalidates the account’s previous code. Treat the code as an approval credential.

GET, POST and DELETE /api/v1/telegram/link read this principal’s status, issue a code and unlink or cancel its pending code. Off returns 404. Unlink also works with /unlink in the bot. Linked principals receive pending intent metadata with Approve/Deny buttons through the existing telegram-bot worker. Callbacks recheck the key, role, team scope, link generation and message, then run the dashboard’s approval decision logic. The message updates after a decision, including decisions made on the dashboard. No inference is executed by clicking a button; the agent must retry its bound request before expiry.

Telegram sees link codes, account-link commands, approval metadata and alert messages. AnyRoute stores the code’s hash, link account and principal key hash, Telegram user id, random link generation, timestamps and approval delivery markers, never Telegram message text or API keys for linking. Codes expire after five minutes; delivery markers expire with the approval and are removed on unlink. Expired metadata is purged on enabled polling or code issuance. The key is rechecked for every delivery and callback. A worker interruption can repeat a notification; decisions remain single-use. Chat keys connected through /key are separate and removed through /forget.

Share a signed track record

POST /api/v1/agents/me/record-certificate with your agent’s Bearer key and {"claims":["requests_at_least:100","active_days_at_least:7"]} returns a certificate with a fresh random pseudonym, your requested true claims, issued_at, expires_at, key_id and an Ed25519 signature. Any false or insufficiently supported claim refuses the entire request with HTTP 422 record_claim_unproven. No agent key hash, account id, rulebook hash or underlying activity record is included. Issuance attempts are limited to five per minute per account across its keys.

Signed by AnyRoute’s router, which sees the agent’s activity; not a zero-knowledge proof. This describes records of requests through AnyRoute only. It establishes no on-chain enforcement and no activity outside AnyRoute. The router reads ordinary chat text in memory on every lane; the encrypted-chat adapter forwards ciphertext. This certificate path stores no prompt text, certificate, pseudonym or claims.

requests_at_least:N counts retained generation records for the calling key with a completion reason and no cancellation. active_days_at_least:D counts distinct UTC calendar dates with those records, rather than rounding elapsed time into days. Parent, sibling and child request counts are excluded. Missing or deleted records reduce the count. N and D must be positive integers of at most 15 digits; at most 16 distinct claims are accepted.

no_denials_days:D and no_kills_days:D cover an exact rolling interval of D × 24 hours, up to 90 days. They refer to retained rulebook denials (including denied approvals) and kill events. The key must predate the interval and have a current effective rulebook, including inherited parent rules, unchanged for the whole interval and currently not killed. Recent edits, resumes or rulebook recreation conservatively refuse these claims. Parent events can include sibling activity and can also prevent issuance. Approval-required decisions alone are not denials. These claims describe the router’s retained rulebook history, not all authentication or provider failures. Autonomy rungs are unavailable here; rung_at_least is not accepted.

AGENT_POLICY_ENABLED defaults to false and gates issuance and public verification. Issuance also needs TLOG_ENABLED (default false), and waits for the reused receipt signing key’s receipt_key entry to appear in the key log. Certificates expire exactly seven days after issuance and are snapshots: later denials or kills do not revoke them. Signing key rotation keeps earlier public keys available. Key publication alone does not establish a witnessed checkpoint; check the log’s inclusion proof and witness policy separately.

Switched on at anyroute.tech.

POST /api/v1/agents/certificates/verify with the certificate as its JSON body returns {"data":{"valid":true,"notice":"…"}} only when its shape, signature, signing-key issuance window and expiry pass. No authentication is needed. GET accepts URL-encoded certificate JSON in the certificate query parameter; use POST to avoid placing it in URL history or upstream access logs. Malformed certificates return HTTP 400; expired or tampered certificates return valid:false.

For offline verification, obtain and independently trust or pin the router’s receipt JWKS from /.well-known/anyroute-receipt-keys.json before going offline. Match key_id to kid and to the first 16 hex characters of SHA-256 of the raw public key. For key-log verification, hash that raw key and use /api/v1/tlog/proof?kind=receipt_key&sha256=… with a pinned log verifier and your witness policy. Retain the trusted keys and proof with the certificate.

import { verifyRecordCertificate } from '@anyroute/client';
// trustedKeys is an independently trusted receipt JWKS saved beforehand.
const valid = await verifyRecordCertificate(certificate, { keys: trustedKeys });

The helper verifies Ed25519 over canonical JSON of payload (recursive object keys sorted, no whitespace), requires the fixed certificate type and version, checks issued_at against the key’s valid_from and retired_at, and requires issued_at ≤ current time < expires_at with exactly seven days between them. It makes no network calls. A signature confirms what the router signed; it does not independently establish the underlying activity.

Each certificate has a new 256-bit pseudonym. There is no stable agent identifier across certificates, but the shared issuer signing-key id, selected claims and issuance timing remain visible and may correlate them. The authenticated router knows which key requested issuance; this is not anonymity from the router.

Opt-in public agent profiles

AGENT_PROFILES_ENABLED defaults to false. Switched on at anyroute.tech. With it off, profile APIs return 404 and the directory MCP tool is omitted. Apply the additive agent_profiles migration before enabling it. Rulebook categories require AGENT_POLICY_ENABLED; selected certificates also require TLOG_ENABLED and its existing signing configuration.

An owner or authorised account administrator publishes with PUT /api/v1/agents/:key_hash/profile. GET on the same path reads publication settings; DELETE unpublishes and deletes the row. Use an owner API key; session keys cannot manage profiles. The request accepts name (1–80 characters), description (up to 280), optional HTTP(S) homepage (up to 500, no embedded credentials), capabilities (up to 16 tags, 1–40 characters each), show (selected spending_caps, ask_first, kill_switch categories), and certificate_claims (up to 16 distinct supported record claims). Omitted selections default to empty.

Only selected rulebook booleans are returned, including applicable inherited rules. Exact spending amounts, approval thresholds, model and tool restrictions, policy hashes, kill reasons, key hashes and account identifiers are not automatically published. Owner-written names, tags, descriptions and homepages can identify their owner. Capabilities are supplied by the owner and are not verified abilities.

Certificate claims are checked against the selected key’s retained records and issued with a fresh pseudonym. Updates replace the selected certificate; empty claims remove it. Public reads verify its signature, signing-key issuance window and expiry. Expired or invalid certificates are omitted. The card adds valid:true as read-time metadata; remove valid before passing a certificate to strict certificate verifiers. The profile page’s certificate JSON already removes it. These seven-day snapshots do not revoke on later activity, do not independently prove the underlying observations and are not zero-knowledge proofs. Publishing intentionally links the selected certificate to the public profile. Issuance shares the account’s five-per-minute certificate limit.

GET /api/v1/agents/profiles/:slug returns an A2A-style card with name, description, url, capabilities, provider organization AnyRoute, optional homepage and an anyroute extension containing id, rulebook_summary, certificates and status. This is a discovery card, not an agent invocation endpoint or a complete A2A protocol implementation. The discovery card reports agent hardware attestation as unavailable; inspect registered sealed-agent evidence on /agents separately. Sealed hosting is available, but no sealed agent is registered at anyroute.tech yet.

GET /api/v1/agents/profiles accepts an exact tag, a cursor and limit (1–50, default 20), returning data and next_cursor. The MCP tool anyroute_agent_directory accepts the same fields and needs no API key. Disabled or expired keys are hidden. Public profile slugs are independent random 144-bit identifiers; republishing after deletion creates a new slug. A random slug does not prevent correlation through owner-written fields or timing.

Publish and search profiles; individual pages use /agents/profile/?id=<slug> and fetch current cards. Unpublishing removes router discovery immediately, but cannot remove copies already retained by public readers or database backups. Profiles do not change who can read inference: ordinary chat requests are read in router memory on every lane; the dedicated encrypted-chat path forwards ciphertext.

Sealed agent hosting

Sealed agent hosting is available at anyroute.tech, but no sealed agent is registered there yet. The owner builds and publishes the agent sidecar image. The approved recipe in deploy/agents/sealed runs an owner’s agent image pinned by digest with a credential proxy inside an Intel TDX dstack confidential VM. The proxy derives an application-scoped sealing key from the guest agent using a path containing the measured compose hash, and seals the dedicated AnyRoute API credential with AES-256-GCM. The credential is provisioned through stdin inside the confidential guest, never in the compose manifest. The same measured app can recover it after restarts; manifest changes require fresh provisioning and a fresh dstack app identity. GetKey paths are caller-selected: KMS authorization must restrict release to the approved measurement. Giving replacement code the old application key could let it request the old sealing path.

With the principal’s Bearer key, POST /api/v1/agents/{key_hash}/sealed with attestation_url (HTTPS /attest without query or credentials), agent_image_digest and compose_hash (both sha256: followed by 64 lowercase hex characters). Approve the exact compose manifest and both container image digests independently. The target must be an active key without management permissions. DELETE the same URL removes the registration.

The router obtains a fresh nonce-bound TDX quote over pinned TLS. Its report_data commits canonical JSON of type, agent_image_digest, compose_hash, agent_key_hash (SHA-256 of the API credential), sidecar_version and tls_spki_sha256, followed by the 32-byte nonce. All configured quote verifiers must accept it, and a dstack or phala verifier must recover the approved measured compose hash. Debug-enabled TDX, key mismatch, TLS mismatch and stale nonce evidence are refused. Image identity is established by approving the digest-pinned measured manifest, not by trusting a self-reported image label alone.

AGENT_SEALED_ENABLED defaults to false and requires AGENT_POLICY_ENABLED (default false), plus ATTESTATION_VERIFIERS including a configured dstack or phala verifier. Existing verifier settings apply: DSTACK_VERIFIER_URL or PHALA_VERIFIER_URL and any other selected verifier credentials. The existing attestor worker rechecks at ATTESTATION_INTERVAL_MS (default 600000). A failed check clears the badge; success expires after 30 minutes even without workers. The agent list and GET /api/v1/agents/me include sealed metadata only while the feature is enabled. The badge reads “Sealed · attested · image sha256:…”.

Trust includes Intel TDX, firmware, the guest OS, dstack KMS authorization, quote-verifier services and code inside the measured VM. Measured does not mean audited: guest administrators and approved code can read or exfiltrate secrets. Rollback, denial of service and all side channels are not prevented. Owners may retain a credential copy; this badge does not establish exclusive possession or prove that every later request came from the VM.

The agent container does not receive the API credential. The proxy sends it to AnyRoute as HTTPS Bearer authentication and forwards request and response streams, including receipt, lane and policy headers. Ordinary requests sent to AnyRoute are read by the router in memory on every lane. Only the encrypted-chat transport keeps message plaintext off that router when the agent explicitly uses it; sealed hosting alone does not enable encrypted chat.

The router keeps the registration URL, expected deployment hashes, random revision, fixed result code, verifier names, TLS key hash and check timestamp in its existing key-value store. No quote or API credential is persisted by this registration path. Removing the registration deletes the row; disabled features leave it stored until removal. The guest keeps the encrypted credential file and reads credentials and forwarded bodies in memory. What we keep

Make your first API call.

  1. Get a keyNot done

    Sign in with your wallet to get an API key.

  2. Add fundsNot done

    Open Payments and deposit a listed token.

  3. Make your first callNot done

    Send a POST request with your key and a model.

Sign in from your account to see your progress here.

Set ANYROUTE_API_KEY to your key in your shell before running this command. Choose a model from the model list.

First API call
curl "https://anyroute.tech/api/v1/chat/completions" \
  -H "Authorization: Bearer $ANYROUTE_API_KEY" -H "Content-Type: application/json" \
  -d '{"model":"meta-llama/llama-3.3-70b-instruct","messages":[{"role":"user","content":"Hello"}],"max_tokens":32}'

For ordinary calls, the router reads request text in memory to route it. See what we keep.

Anyroute accepts the familiar chat-completions request. Replace the base URL and key; requests, streaming, tools, provider preferences and usage fields work unchanged. Keys are self-custodial: POST /api/v1/keys (no account) returns a key and the hash to deposit USDG to.

Example request
// Change two lines: the base URL and the key.
const client = new OpenAI({
  baseURL: process.env.ANYROUTE_BASE_URL, // "<your router>/api/v1"
  apiKey: process.env.ANYROUTE_API_KEY,   // "sk-ar-v1-…"
});

const generation = await client.chat.completions.create({
  model: "meta-llama/llama-3.3-70b-instruct",
  messages: [{ role: "user", content: "Hello, Anyroute." }],
});

Make the route explicit.

The provider object supports order, allow_fallbacks, only, ignore, data_collection, zdr, quantizations, sort, max_price, require_parameters and preferred latency or throughput. Model suffixes :nitro, :floor, :free and :private work too. :nitro puts the fastest providers first (by measured throughput) and :floor the cheapest, exactly as provider.sort throughput and price do; a suffix wins over provider.sort. They only order the providers that your other preferences and your lane already allow, so provider.order still goes first and a lane never falls back outside itself. GET /api/v1/models lists them per model in routing_variants. models[] lists fallback models. The private flag is an Anyroute extension: it selects only providers with a fresh TEE attestation. By default providers are weighted by 1/price² × 30-day uptime × canary quality, and a provider with two failures in 30 seconds is skipped.

Routing preferences
{
  "model": "meta-llama/llama-3.3-70b-instruct",
  "models": [
    "qwen/qwen3-32b"
  ],
  "messages": [
    {
      "role": "user",
      "content": "Your prompt"
    }
  ],
  "provider": {
    "allow_fallbacks": true,
    "data_collection": "deny",
    "sort": "latency",
    "private": true
  }
}

Check JSON before using it.

Switched on at anyroute.tech. Self-hosted routers can enable STRUCTURED_OUTPUT_CHECK_ENABLED (default false). Each chat request must also choose anyroute.json_check. Without both switches, the response and billing stay unchanged. Models must still support response_format; this option does not change routing requirements.

Example request
{
  "model": "qwen/qwen3-32b",
  "messages": [
    {
      "role": "user",
      "content": "Return a JSON object with a \"summary\" field."
    }
  ],
  "response_format": {
    "type": "json_object"
  },
  "anyroute": {
    "json_check": "validate"
  }
}

Choose validate to parse the final message without another provider call. With json_object, it must parse as an object. With json_schema, the router also checks response_format.json_schema.schema. The answer, usage and original signed receipt remain intact; the response adds X-Anyroute-Json-Check: valid or invalid and receipt.structured_output with mode, valid and errors (JSON Pointer path and plain reason; the empty path means the root).

Choose repair to allow one correction call when the first answer is invalid. The same routed model and provider receive the original conversation, answer, schema and errors, with a request for corrected JSON only. Both calls are charged at their normal rates and pass the same account and agent rules. Repair requires a bearer API key; a blind token or per-call payment cannot authorize two charges. A single-use agent approval is not reused. Non-streaming repair bypasses the response cache. If authorization, credit or the provider prevents the correction, the first output is returned with valid: false and retry_error; its charge still applies. If the core settles a correction before refusing its attestation, that correction’s signed receipt and charge are included too; the refused answer is not returned.

receipt.structured_output.calls lists each served call with attempt, an exact USDG cost string and its complete signed receipt. total_cost is the sum of those charges; retry_attempted says whether a correction was requested, and initial_errors preserves the first failure after a served correction. Top-level usage and receipt identify the returned call only. A second charge is additional to the first, even if the correction is still invalid. There is never a second correction call.

For streams, the header is pending because it is sent before the answer finishes. The final receipt event carries the completed validation result. A streaming request choosing repair gets pending; repair-unsupported and repair_supported: false: it validates once, without an extra call. Content chunks and their hash-chain comments remain unchanged. An interrupted stream is invalid; a disconnect may prevent delivery of the final receipt.

The schema subset covers objects, arrays, strings, numbers, integers, booleans and null; properties, required fields, additional properties, items, enum, const, anyOf, nullable and type arrays; string length, number bounds, array length and property-count bounds; and local #/ references through $defs or definitions. Unknown assertions, including pattern and format, are refused before a provider call. Schemas are bounded to 4,096 nodes and 64 levels; answer checks to 1 MiB, 100,000 validation steps and 64 levels. Oversized or non-text answers are invalid and cannot be repaired. Only a single chat answer is supported (n: 1); council and dual verification cannot be combined with this option.

Presets can store anyroute: { json_check: "validate" } or repair with their response format. Explicit request settings win. The feature flag still applies.

Validation checks syntax and this schema subset, not factual accuracy. It cannot guarantee a usable answer; callers must inspect valid. The supplemental validation report is unsigned and absent from receipt lookup. It is not written to generation receipts; existing batch-output retention can keep it with the answer. Each call’s payload, signature and v2 receipt remain independently verifiable. The router reads ordinary request and answer text in memory. This option does not hide prompts from AnyRoute.

Presets: config as code, with versions.

A preset is a saved route with a version history. It holds the same fallback models, provider preferences and sampling defaults as @route/<slug>, plus what a route may not hold: a system_prompt (up to 16,000 characters), a response_format and up to 32 tool definitions with a tool_choice. PUT /api/v1/presets/<name> saves the whole document. Each change adds an immutable version, numbered 1, 2, 3 and identified by the SHA-256 of its canonical JSON; saving the same content again adds nothing. Call the latest with model "@preset/<name>", or pin one with "@preset/<name>@3" or a hash prefix of at least 7 characters. The response carries preset: name, version and hash.

The request wins over the preset for every parameter, provider field, models list, tools and response_format it sets. For provider.lane and provider.disclosure the stricter of the two applies, as for a saved route. The system prompt is added as the first message only when the request has no system or developer message; a request with its own keeps it and gets no second one. Legacy /completions has no messages, so a preset with a system prompt, tools or a response_format is refused there (400 preset_unsupported). GET /api/v1/presets/<name>/versions lists the history, /diff?from=1&to=2 returns a JSON diff (JSON Pointer paths with the old and new value), and POST /rollback with { "version": 1 } adds a new version with that content, so nothing is rewritten and every pin keeps resolving until the preset is deleted. A key or agent session allowed "@preset/<name>" may use that preset’s models. Limits: 100 presets per account, 100 versions per preset, 64 KB per version.

PUT /api/v1/presets/support
{
  "description": "Support replies in one line",
  "models": [
    "qwen/qwen3-32b",
    "meta-llama/llama-3.3-70b-instruct"
  ],
  "provider": {
    "data_collection": "deny"
  },
  "params": {
    "temperature": 0.2,
    "max_tokens": 200
  },
  "system_prompt": "You are the support assistant. Answer in one line."
}

Characters: one card, any model.

A character is a Tavern character card kept on the router. POST /api/v1/characters takes the card as JSON (card) or as a PNG that carries it (png, base64): V1 cards with flat fields, V2 (chara_card_v2) and V3 (chara_card_v3). In a PNG the ccv3 text chunk wins over chara. Each character gets an id (ch_ and 24 hex characters) and a card_hash, the SHA-256 of the card as canonical JSON. GET /characters/:id/export?format=json|png&spec=v2|v3 gives the card back in a file SillyTavern and other card tools import.

Visibility is public (listed in GET /api/v1/characters, which needs no key and filters by tag and q), unlisted (readable by anyone with the id, the default) or private. A private card is encrypted on your device with sealCard from @anyroute/client/characters: the router stores only the ciphertext and the card_hash, and never sees the name or the text. To chat with it, your client decrypts it and sends it in card; the router checks it against the card_hash (409 card_hash_mismatch) and uses it for that call only.

Call a character from any OpenAI client with model "@character/<id>" on /chat/completions. The card becomes the prompt; the model that answers is models[0], or the character's default model. POST /api/v1/characters/:id/chat does the same with more options: greeting (0 is first_mes, 1 and up the alternate greetings), regenerate, user_name for {{user}} , session_id, and memory ( { summary, facts } ), which your client decrypted and which is used in memory and never stored. Lanes, budgets and signed receipts apply as on /chat/completions. The lane defaults to attested when the model has an attested provider, else public; x-anyroute-character-lane says which, and x-anyroute-character-note says why when attested was not available. A lane the request asks for wins. POST /characters/group/next picks who speaks next in a group chat (round_robin or named).

The memory ledger keeps a character's long-term memory for you without reading it. Your client seals each entry (summary, fact, lorebook or state) as arm1.<iv>.<ciphertext> and files it under a scope it derives, an HMAC of the character id under your viewing key, so the router cannot tell which character a memory belongs to. POST /api/v1/memory stores an entry, GET lists sizes and kinds without the sealed values, and DELETE ?scope= clears a scope. Search by embedding (POST /memory/search) works only on entries stored with embedding_opt_in: true, and it is off by default: a vector is derived from the memory's text and can leak its topic.

The creator of a public card sees GET /characters/:id/usage: calls and cost per day. Never what was said, and never who said it.

SillyTavern · API Connections
API:                     Chat Completion
Chat Completion Source:  Custom (OpenAI-compatible)
Custom Endpoint:         <your router>/api/v1
Custom API Key:          your Anyroute key (sk-ar-v1-...)
Model ID:                @character/<id>

Tracing: your spans, in your own tools.

Give a key a tracing destination and every public-lane call made with it is sent there as one span that follows the OpenTelemetry GenAI semantic conventions: gen_ai.system, gen_ai.request.model, gen_ai.response.model, gen_ai.usage.input_tokens and output_tokens, gen_ai.request.temperature and max_tokens, gen_ai.response.finish_reasons, plus anyroute.cost_usd, anyroute.receipt_id, anyroute.lane, anyroute.provider and the server latency. Send a W3C traceparent header and the span joins your trace. Set it with PATCH /api/v1/keys/<hash> and a tracing object; tracing: null turns it off.

Prompt and completion text is exported only with include_content: true; it is off by default. Calls on the attested and unlinkable lanes are never exported, whatever the key says: those lanes promise that a call’s model, timing, size and cost do not leave the router tied to a key. The endpoint, headers and keys are encrypted at rest and never returned; GET shows the type, the host, the header names and export counters. Leave a secret out of a later PATCH to keep the stored one. Export runs after the response, from a bounded queue: a full queue drops and counts, failures are retried with backoff, and a destination that keeps failing is paused for a minute. Your calls are never slowed or failed by it. Destinations must be public https:// addresses.

Honeycomb (OTLP)
{
  "tracing": {
    "type": "otlp",
    "endpoint": "https://api.honeycomb.io",
    "headers": {
      "x-honeycomb-team": "YOUR_INGEST_KEY"
    }
  }
}
Grafana Cloud Tempo (OTLP)
{
  "tracing": {
    "type": "otlp",
    "endpoint": "https://otlp-gateway-prod-<region>.grafana.net/otlp",
    "headers": {
      "authorization": "Basic <base64 of instanceId:token>"
    }
  }
}
Langfuse (ingestion API; host defaults to https://cloud.langfuse.com)
{
  "tracing": {
    "type": "langfuse",
    "public_key": "pk-lf-...",
    "secret_key": "sk-lf-...",
    "host": "https://us.cloud.langfuse.com",
    "include_content": true
  }
}
Helicone (custom log API; host defaults to https://api.worker.helicone.ai)
{
  "tracing": {
    "type": "helicone",
    "api_key": "sk-helicone-..."
  }
}

Any OTLP/HTTP collector works: /v1/traces is added to the endpoint unless it is already there, and the body is OTLP JSON. For a self-hosted Tempo or an OpenTelemetry Collector, expose its OTLP/HTTP receiver over https and put its auth in headers.

Network sanctions screening

Optional screening checks provider payout addresses against digital currency addresses published in the OFAC SDN list. A listed address is refused on provider application and skipped during weekly USDG payouts. This is address screening, not identity verification: it collects no KYC or identity data and cannot identify unlisted wallets controlled by a listed person.

SANCTIONS_SCREENING_ENABLED defaults to false. SANCTIONS_LIST_URL selects the official classic SDN XML download and must be explicitly set to an HTTPS URL when screening is enabled in production. SANCTIONS_MAX_AGE_DAYS defaults to 7. An operator schedules sanctions-refresh through the worker job allow-list; registration runs daily and once at worker startup. No production worker configuration is changed by this feature.

The feed is validated before atomically replacing the stored addresses and metadata. Failed downloads preserve the last good list. Screening uses 0x-prefixed 20-byte digital currency addresses across EVM chains and token tickers; other address formats are counted and ignored. Names and other identity fields from the XML are not stored. The source hash is SHA-256 of the downloaded bytes, not an independent signature check.

When the publication date is older than the configured maximum, or the list is missing, new admissions and payouts to addresses without a recorded successful payout pause. Previously paid addresses may continue receiving payouts, but any match in the retained list is always blocked. Skipped settlements remain due for a later run. The logs record skip reasons and the freshness exception. If the database is unavailable, screening pauses both admission and payouts.

Public GET /api/v1/network/sanctions returns the publication date, entry count, source hash, ignored count, refresh time, age and stale status. Screening does not change prompt handling: AnyRoute's router reads request text in memory on every lane.

Teams: anonymous organisations with an audit log.

A team is an organisation with no email and no names. Its owner is the account that made it, a wallet, or a Safe: POST /api/v1/teams/<id>/owner/challenge returns a one-time message, and POST /owner checks the signature by recovery for a wallet or by EIP-1271 for a contract wallet (isValidSignature must return 0x1626ba7e, read on chain). A team made by a wallet-signed-in account is owned by that wallet from the start. The bound address signs in as owner.

Roles: owner, admin (members, invites, keys and budgets), dev (creates keys in the team within its org budget), viewer (reads the team, usage, receipts and the audit log) and agent (API calls only, 403 on every team route). An admin makes a single-use invite (POST /invites, only its SHA-256 is kept). The invitee joins with a passkey (WebAuthn, attestation none, so only the public key is stored) or a wallet signature, and signs in the same way later with POST /api/v1/teams/<id>/sign-in. A sign-in returns a key that lasts 12 hours and has a limit of 0: it manages the team within the member's role, and a dev or admin creates API keys to call models. With budget_usd set, every key in the team needs a limit and the limits together must fit (409 org_budget_exceeded).

Every change to members, invites, the owner, keys, budgets, presets, routes and lane settings is appended to the team's audit log, never prompt text. Each entry is chained to the one before it: hash = sha256(prev_hash bytes, then the canonical JSON of team, seq, at, actor, action, target and detail), starting from 64 zeros, with an RFC 6962 Merkle root per hour. GET /audit pages through it, GET /audit/roots lists the hourly roots and GET /audit/export?format=jsonl|csv downloads the chain. Anyone holding an export can check it with Node alone: an edited, dropped or reordered entry fails.

Verify an audit export offline
node scripts/verify-audit.mjs anyroute-audit-<team>.jsonl --head <hash you saw>
OK  team team_…  42 entries  3 hourly roots  head 9f2c…

Route by what a provider discloses.

Each provider has a disclosure profile its operator documents: retention (attested, policy or logs), jurisdiction, legal-hold status and training use, each with a source and a date. Anything undocumented reads as the conservative default (logs, unknown). GET /api/v1/disclosure/:providerId returns the profile and the class the provider is served under right now. It reports what was documented and, for attested, what the router verified; it is not a guarantee of a provider’s behaviour.

Set provider.disclosure (or the X-Anyroute-Disclosure-Max header) to none, policy or any, the default. none routes only to providers whose retention is declared attested and whose TEE attestation is fresh; policy also accepts a documented no-retention policy with no legal hold. provider.lane (or X-Anyroute-Lane) picks a privacy lane (see below); attested and unlinkable imply none, and if body and header are both set the stricter applies. When no provider meets a disclosure setting the request fails with 409 disclosure_unavailable, or 503 disclosure_provider_unavailable when qualifying providers are down; a lane that no attested endpoint can serve fails with 503 no_attested_endpoint. It is never sent to a provider that does not qualify, and nothing is charged. A request with a disclosure setting or a lane never uses the response cache.

Responses carry X-Anyroute-Disclosure (attested, policy or vendor-forwarded) and X-Anyroute-Lane, and the signed receipt records disclosure and lane. On a stream the header is sent only when every reachable provider shares one class; the receipt always states it. A development attestation is marked attestation_simulated and is refused in production. GET /api/v1/models?lane=attested lists the models that have an attested endpoint now.

Three privacy lanes.

Every request is served on one lane. public, the default, may use any endpoint. attested uses only endpoints whose retention is declared attested and whose hardware attestation the router verified recently. unlinkable adds two things on top of attested: the request arrives through an Oblivious HTTP relay (or over Tor through the router’s onion service, where the router serves the lane that way), and it is paid with a blind token, so the router cannot tie the payer or the address to the prompt. Pick one with provider.lane or the X-Anyroute-Lane header. A key can carry a default (routing.provider.lane on PATCH /api/v1/keys/:hash), which a lane named in the request replaces. A saved route can pin a lane too (config.provider.lane, public or attested); there the stricter of the route and the request applies, as described below. A request that arrives through an independent relay with a blind token and names no lane is served on unlinkable.

On attested and unlinkable there is no fallback. If no endpoint of the model has a fresh, verified attestation, the request fails with 503 no_attested_endpoint and error.metadata.reason none_attested; if attested endpoints exist but are all down, the reason is attested_endpoints_down and Retry-After is set. Nothing is sent to any other endpoint and nothing is charged. provider.order, only and ignore still apply inside the lane, but they cannot bring back an endpoint the lane excludes. On unlinkable, an API key or a wallet is refused with 403 lane_requires_anonymous_auth, because both name the payer; set provider.lane_downgrade (or X-Anyroute-Lane-Downgrade) to attested to be served on the attested lane instead, never on public.

Within a lane the router picks among endpoints by weight: uptime times quality times attested_bonus, divided by the square of the price relative to the cheapest endpoint. attested_bonus is 1.25 on public, so an attested endpoint is preferred at equal price, and 1 on the other two lanes, where every endpoint is attested. Ties break by stake, then provider id. GET /api/v1/models lists lanes for each model and each endpoint, GET /api/v1/models?lane=attested keeps the models that can be served on that lane now, and GET /api/v1/status reports a lanes section with how many models and endpoints each lane has.

Ordinary chat on attested and unlinkable terminates TLS at the router, which reads the prompt in memory before sending it to the enclave over a connection pinned to its attested key. The host outside the enclave cannot read it; the router can. The separate encrypted-chat adapter forwards client-encrypted content through the router to the attested gateway enclave. Direct sidecar HPKE is also available in the SDK; the ordinary chat route does not carry that format.

503 · no attested endpoint
POST /api/v1/chat/completions
{ "model": "<model>", "messages": [...], "provider": { "lane": "attested" } }

HTTP/1.1 503
{ "error": { "code": 503, "type": "no_attested_endpoint",
    "metadata": { "lane": "attested", "reason": "none_attested", "excluded": [...] } } }

A saved route can carry the same settings: provider.lane (public or attested) and provider.disclosure in its provider policy. When a request calls @route/<slug>, the stricter of the route’s and the request’s value applies, so a request can tighten a route but never loosen it, and a route on the attested lane is served by an attested provider or refused. Saving such a route fails with 409 route_lane_unavailable, naming the models, when a model in its list has no provider that meets the setting right now. In the dashboard, Batch Studio can run a whole batch on the attested lane: it sends provider.lane attested with every row, offers only the models GET /api/v1/models?lane=attested lists, shows the lane and receipt id of each row from X-Anyroute-Lane and X-Receipt-Id, and marks a row the router refuses or withholds as failed closed.

The unlinkable lane, where the router enables it.

Lane unlinkable keeps the router from tying together who pays, where a request came from and what it says. It is served only when all three hold: the request arrives through the router’s Oblivious HTTP gateway (RFC 9458) by way of a relay run by an operator other than the router’s own; it is paid with a blind token (Authorization: PrivateToken), never a key or a wallet; and it is routed only to attested providers, the same filter as lane attested. The receipt then says lane unlinkable. GET /api/v1/relays lists the relays by operator. GET /api/v1/ohttp/keys is the gateway’s key configuration and GET /api/v1/ohttp/key-list is the key history, signed with the receipt key and hash-chained, for pinning.

A request that carries an API key or a wallet is refused with 403 lane_requires_anonymous_auth. A direct request for the lane is refused with 403 unlinkable_requires_relay and says what to do; through the gateway without a token it is 401 (unlinkable_requires_token) with the token challenge, and 403 through a relay run by the router’s own operator. With no attested endpoint it is 503 no_attested_endpoint, and the token is not spent. What is hidden: the relay sees your address and an encrypted request; the router sees the request and the relay, never your address; the token’s purchase cannot be tied to its use. What is not: a relay that cooperates with the router can join the two, timing and message sizes can be correlated (responses are padded), and any identifier you put in the request body reaches the provider.

A message/ohttp-req response is encrypted as a whole, so stream: true sent that way is refused with 400 stream_unsupported. Where the router also enables chunked Oblivious HTTP (draft-ietf-ohai-chunked-ohttp), GET /api/v1/relays lists gateway.chunked, and a request sent as message/ohttp-chunked-req is answered with message/ohttp-chunked-res: the response is encrypted and sent chunk by chunk as the router produces it, so a streamed completion arrives token by token while the relay still carries only ciphertext. The relay has to carry it too (RELAY_CHUNKED_ENABLED in relay/); it passes each chunk on as it arrives, under the same size limit and timeout. The last chunk is sealed as the final one, so a client can tell a response that was cut off from a complete one; the TypeScript SDK’s obliviousFetch (@anyroute/client/ohttp) decrypts each chunk as it arrives and fails with ohttp_truncated when the final chunk never came and ohttp_decrypt_failed when a chunk was altered or reordered. Chunk sizes and their timing stay visible to the relay.

TypeScript SDK · a streamed completion on the unlinkable lane, through a relay
import { AnyRoute, TransparencyLog } from "@anyroute/client";
import { obliviousFetch } from "@anyroute/client/ohttp"; // npm install ohttp-ts@0.6.0 hpke@1.1.7

// keyConfig: the base64url-decoded "config" of the current key in GET /api/v1/ohttp/key-list, after checking the
// list's signature against the receipt key you pinned. receiptKeys: the same pinned key set.
const client = new AnyRoute({
  baseUrl: "https://<router>",
  privateToken: "<a blind token>",
  lane: "unlinkable",
  receiptKeys,
  fetch: obliviousFetch({
    relayUrl: "https://<independent relay>/relay",
    keyConfig,
    // Optional: refuse a key configuration the witnessed key log does not include (the log is read directly).
    transparency: new TransparencyLog({ logUrl: "https://<router>", logKey: "<log key>", witnesses: ["<witness key>", "<witness key>"] }),
  }),
});
const stream = await client.chat.completions.stream({ model: "<model>", messages: [{ role: "user", content: "Hello" }] });
for await (const chunk of stream) process.stdout.write(chunk.choices?.[0]?.delta?.content ?? "");
// A response cut off on the way throws AnyRouteError "ohttp_truncated"; an altered chunk, "ohttp_decrypt_failed".
console.log((await stream.meta()).lane); // "unlinkable"

A router can also serve the lane over Tor instead of a relay: see the unlinkable lane over Tor. GET /api/v1/status says which paths it serves in lanes.unlinkable.via ("ohttp", "onion" or both).

Attested providers only
{
  "model": "meta-llama/llama-3.3-70b-instruct",
  "messages": [
    {
      "role": "user",
      "content": "Your prompt"
    }
  ],
  "provider": {
    "disclosure": "none"
  }
}

Reach AnyRoute over Tor.

Where the router runs an onion service, you can call it through Tor and keep your network address from the router and from the network it runs on. GET /api/v1/status publishes the address as onion.address, and it is shown below. Use it as http://<address> from Tor Browser or any client that can use a SOCKS5 proxy and lets the proxy resolve names (curl --socks5-hostname, or torsocks); the API is the same, under /api/v1 on that host. The onion service itself encrypts and authenticates the connection to the router, so plain http:// is correct there. The site’s pages also carry an Onion-Location header, which Tor Browser turns into an “.onion available” prompt.

Reading this router’s status…
Through a local Tor client (SOCKS5 on 127.0.0.1:9050)
curl --socks5-hostname 127.0.0.1:9050 http://<onion address>/api/v1/models

# a call with your key, the same request as on the clearnet
curl --socks5-hostname 127.0.0.1:9050 http://<onion address>/api/v1/chat/completions \
  -H "Authorization: Bearer $ANYROUTE_API_KEY" -H "Content-Type: application/json" \
  -d '{"model":"meta-llama/llama-3.3-70b-instruct","messages":[{"role":"user","content":"Hello"}]}'

Tor hides where you connect from, not what you send: an API key, a wallet signature or a prompt identifies you or your account exactly as it does on the clearnet. For payment that cannot be linked to your requests, use blind tokens, and for the unlinkable lane a relay, or this onion service where the router serves the lane over Tor (below). Requests that arrive over Tor have no address the router can limit, so calls without an API key (unkeyed chat and embeddings, new keys, wallet sign-in challenges) share limits with everyone else using the onion address, and a call with a key is limited per key as usual: use a key or a token for a quota of your own. The first request can take several seconds while Tor builds its circuit. A relay operator can also reach a gateway’s onion address through a SOCKS5 proxy (RELAY_SOCKS5_PROXY in relay/), so the gateway never sees the relay’s address either.

The unlinkable lane over Tor, where the router enables it

Where the router serves it (GET /api/v1/status: lanes.unlinkable.available is true and lanes.unlinkable.via includes "onion"), lane unlinkable also works over Tor, with no relay. Four steps: take the onion address from onion.url in the status; buy blind tokens ahead of time with your key (buyTokens from @anyroute/client/blind, below); send each call to the onion address through Tor with Authorization: PrivateToken token=<one token> and the lane named, in provider.lane or X-Anyroute-Lane; and read X-Anyroute-Lane and the receipt, which say unlinkable. A call that names no lane is served on public, as before, so name it. Streaming works as on the clearnet, because Tor carries ordinary HTTP.

The rules are those of the lane everywhere: attested providers only, with no fallback (503 no_attested_endpoint, and the token is not spent); an API key, a wallet or a per-call payment is refused with 403 lane_requires_anonymous_auth, unless you allow the downgrade to attested; no token is 401 unlinkable_requires_token. The same call to the clearnet address is refused with 403 unlinkable_requires_relay. Only the onion service can mark a call as coming over Tor: it adds a secret header the router checks, and removes any copy a client sends, so sending that header yourself changes nothing.

Buy tokens (TypeScript SDK)
// Buy blind tokens ahead of time with your key (Node or Bun; needs the optional package @cloudflare/blindrsa-ts).
// The purchase names your account; the tokens it returns cannot be tied back to it when you spend them.
import { buyTokens } from "@anyroute/client/blind";

const { tokens } = await buyTokens({ baseUrl: "https://<router>", apiKey: process.env.ANYROUTE_API_KEY, denomination: 10000, count: 10 });
console.log(tokens.join("\n")); // one token pays for one call; keep them private until you spend them
Spend them over Tor (SOCKS5 on 127.0.0.1:9050)
# Is the lane served over Tor here? lanes.unlinkable.via lists "onion", and onion.url is the address.
curl -s --proxy socks5h://127.0.0.1:9050 http://<onion address>/api/v1/status | jq '.data.lanes.unlinkable, .data.onion'

# Models an attested provider can serve on the lane right now
curl -s --proxy socks5h://127.0.0.1:9050 "http://<onion address>/api/v1/models?lane=unlinkable" | jq -r '.data[].id'

# Spend one token on one call, streamed. A SOCKS user name of its own gives the call its own Tor circuit.
curl -N --proxy "socks5h://call-1:x@127.0.0.1:9050" http://<onion address>/api/v1/chat/completions \
  -H "Authorization: PrivateToken token=$TOKEN" -H "X-Anyroute-Lane: unlinkable" -H "Content-Type: application/json" \
  -d '{"model":"<model from the list>","messages":[{"role":"user","content":"Hello"}],"stream":true}'

# The same with torsocks
torsocks curl -N http://<onion address>/api/v1/chat/completions -H "Authorization: PrivateToken token=$TOKEN" \
  -H "X-Anyroute-Lane: unlinkable" -H "Content-Type: application/json" \
  -d '{"model":"<model from the list>","messages":[{"role":"user","content":"Hello"}]}'

What is hidden: Tor takes the place of the relay. Your connection crosses three Tor relays run by volunteers before it meets the onion service, so neither the onion service nor the router ever learns your network address, and on onion calls the router reads no address header and keeps no per-address limit (a blind-token call counts against one bucket shared by everyone using the onion address). The token cannot be tied to its purchase. What is not: the router runs the onion service, sees the request in plaintext and sees its size and timing directly, as it sees the request on the relay path; an observer who can watch both your connection into Tor and the router’s side can match them by timing; calls sent on one circuit can be linked to each other, so give calls you want kept apart their own circuit (a different SOCKS user name, as above, or Tor Browser’s New Identity); buying tokens right before spending them links the two by time; and anything in the body that identifies you reaches the provider. The separate encrypted-chat adapter supports Tor with blind tokens; ordinary chat still exposes text to the router.

Network host policy

When NETWORK_POLICY_ENABLED is on (default false), GET /api/v1/network/policy returns the current policy, and GET /api/v1/network/policy/{version} returns an immutable version. Each response includes the canonical JSON, its SHA-256, an Ed25519 signature over those UTF-8 bytes and the log verifier key. Pin that key independently; a key returned beside a signature does not establish the operator’s identity. The host_policy entry in the transparency log names the version and hash. Check its inclusion under a signed checkpoint and the independent witness or public-log evidence.

Switched on at anyroute.tech.

An admin publishes a complete document through POST /trpc/network.publishPolicy with the operator token. Versions start at 1 and advance consecutively; existing versions cannot be replaced. Publication requires TLOG_ENABLED and the production log signing key and independent checkpoint checks already required by the log. Nothing is seeded automatically. Run scripts/network-policy.ts with --base, --provider and --model-id to draft from the router’s current public attestation record, then supply the reviewed source and engine pins and confirm GPU CC requirements before publication.

The pure check compares verified, quote-bound sidecar image and source hashes, engine images, model IDs and digests, and TEE kind. Development evidence is always refused. GPU models require GPU CC evidence verified separately from CPU attestation. Opt-in SHA-256 sidecar bindings v2 commit the source archive hash, engine image and model ID in addition to image, compose and model digests. Legacy v1 still verifies, but its missing policy-required fields fail admission. The live policy v1 approves deploy/network/approved/tdx-qwen2.5-0.5b: Intel TDX in a supported confidential VM serving Qwen2.5 0.5B. A compose hash is required as an existing binding, but this document does not approve a compose manifest. This policy governs automatic network host admission and renewal; it does not replace other provider admission paths. Digest bindings are software statements inside a verified quote; they do not establish that source produced an image or that running software obeys a data policy. AnyRoute’s router reads ordinary chat text in memory on every lane; the encrypted-chat adapter forwards ciphertext.

Each models entry may include an optional strict offer object: slug (lowercase catalogue namespace/model), name, optional hugging_face_id, context_length, max_completion_tokens, optional quantization, and pricing with positive decimal strings prompt and completion in USD per token. Names and identifiers are bounded to 160 characters, quantization to 32, prices to 32 characters and at most 1000000 USD per token, and token limits to 2147483647. Omitting offer leaves existing canonical bytes and hashes unchanged. With NETWORK_HOSTS_ENABLED, successful admission copies terms only for models both requested and bound to the verified quote. The registry creates shadow offers during probation; models without offer terms create none. Existing lane and attestation freshness checks still apply: the attested lane also requires an operator-declared attested-retention profile; the public lane can select fresh attested network hosts under its existing rules. Probation starts with routing weight 0.1; failed or stale hardware evidence excludes the host. Rejection clears terms and disables retained offers.

End-to-end encrypted chat (Phala gateway)

The dedicated POST /api/v1/e2ee/chat/completions adapter forwards encrypted message content unchanged to the Phala confidential AI gateway. It is switched on at anyroute.tech. It is off by default in the open-source router until its operator sets E2EE_PASSTHROUGH_ENABLED. Ordinary chat routes on every lane still read request text in router memory.

Client encryption ends inside Phala’s attested gateway enclave. The gateway restores content and forwards it to the serving workload over a separate confidential channel. This is not client encryption directly to a GPU. The router sees the model, roles, message count, content sizes, timing, public keys, nonce, timestamp, response framing, finish reasons, usage and API key or blind-token authorization. A correctly encrypting client keeps message content and its private key out of the router.

The JavaScript SDK exports e2eeChat(body, options) and client.e2eeChat(body, options). Supply model, text messages, stream, an optional max_tokens ceiling (default: 512), and a required verifyAttestation callback. Set the options lane to unlinkable when using Tor with blind tokens. The helper fetches GET /api/v1/e2ee/attestation?nonce=… with its own fresh challenge, checks the keyset digest, quote report-data binding, expiry and debug bit, then encrypts each content field using X25519, HKDF-SHA256 and AES-256-GCM with canonical AAD. Your callback must verify the TDX quote signature and Intel collateral, replay and appraise the measured event log, enforce accepted measurements, and assess the deployment and key custody. A router boolean is never evidence.

Both buffered JSON and SSE are supported. The helper verifies every encrypted field tag, then the gateway’s Ed25519 receipt over the exact response wire bytes and restored request. Decrypted streamed chunks remain provisional until iteration finishes and the complete signed receipt verifies; stopping early does not establish a complete answer. A missing sentinel, wrong key, field mutation, reordering or truncated response fails closed. There is no plaintext fallback.

Only exact model identifiers offered by phala-confidential-ai are admitted, on the attested lane or the unlinkable lane over Tor with blind tokens. Gateway-verified routing and zero data retention are required. Tools, files, search, structured content, aliases, presets, caching, content policies and unknown fields are refused. Key balances and blind tokens are supported; wallet and pay-with authorization are refused. The router validates ciphertext framing; it cannot determine the plaintext modality or prove that an arbitrary caller actually encrypted a hex string. The router limits envelopes to 1 MiB and responses to 32 MiB. The SDK retains bounded encrypted wire bytes until it can verify the complete receipt.

Before forwarding, the router places a hold priced from a conservative byte-based input bound plus max_tokens. Ciphertext length is a reservation bound, not observed token usage. Billing trusts the gateway’s reported token counts under the same provider trust model as ordinary chat. On completion or truncation, reported usage settles the charge and releases the unused hold; without reported usage, the hold is charged. Blind-token overages are capped at the hold and unused token value follows the existing forfeiture rules. A provider refusal before a response releases the hold and blind-token claim. A process failure can leave a hold for the existing expiry recovery; expiry recovery cannot reconstruct unseen provider usage.

The router’s signed receipt marks encrypted content and hashes the exact forwarded ciphertext bytes. It keeps counts or billing bounds, cost, timing, byte counts, completion state, gateway keyset reference and response-verification status, with key linkage or blind nullifiers. It keeps no message ciphertext, plaintext, private key or replay nonce. The router cannot reproduce the gateway’s restored plaintext request hash and records that check as false. The SDK can check that hash. Receipts are public by id; hashes of low-entropy plaintext in gateway receipts may permit guessing.

Production startup requires PROVIDERS_FILE to declare the live HTTPS TDX provider, its attestation URL and credential reference, plus the matching configured database provider. Forwarding requires fresh verified gateway attestation and pinned TLS. No existing production independence guard is relaxed. Operators must supply the provider credential, acceptable attestation policy and compatible gateway; repository validation does not establish an encrypted exchange with a live account.

V2 does not pad sizes, hide billing or network metadata, encrypt all clear response metadata, or provide forward secrecy against later compromise of a reused recipient key. Browser-delivered client code must be independently trusted. The gateway’s attested status does not by itself establish GPU attestation. Upstream and GPU claims are gateway assertions, authenticated by its signed receipt and, when cited, a content-addressed session. The client can pin appraised session identifiers.

Make any AI app private in one command.

anyroute-private is a small program that runs on your own computer and looks like the OpenAI API at http://127.0.0.1:8788/v1, so any app that lets you set a base URL can use it. Every call it receives is rebuilt from scratch and sent to AnyRoute’s onion service through your own Tor client, on the unlinkable lane, and paid with a blind token. It is Apache-2.0 licensed, has no dependencies to install, and is one file: node private.mjs.

What this hides, and what it does not. The router still reads every prompt: on this lane it terminates the connection and sees the request text in memory to route it, and an attested provider receives it. What is hidden is who sent the call and who paid for it. Tor keeps your network address from the router, and a blind token cannot be tied to the purchase it came from. Anything in your prompt that identifies you still identifies you. This proxy uses ordinary chat, whose text the router reads. The separate encrypted-chat SDK path forwards ciphertext to the attested gateway enclave.

Get it, and check it.

The program is one file, 172,774 bytes , version 0.1.0 , not minified, so you can read it before you run it. Download it, then compare its SHA-256 with this one:

Download and check (https://anyroute.tech/private.mjs)
curl -fsSLO https://anyroute.tech/private.mjs
shasum -a 256 private.mjs        # or: sha256sum private.mjs

Expected SHA-256: 6059908d1581bfdc8e8646a974562feec728c0d3c42957eacc1591bfabd31dbf

The hash and the file come from the same site, so this catches a damaged or swapped download but not a compromised site. To check the file against the source, build it yourself with Bun (bun packages/private/scripts/build.ts in the repository writes the same bytes); a check in the repository rebuilds it and fails if the served file differs. Downloading from the clearnet shows the site your address, though nothing about what you will do with the file. To avoid even that, fetch it through Tor: curl --proxy socks5h://127.0.0.1:9050 -fsSLO https://anyroute.tech /private.mjs. It needs Node 20 or later. It includes the two libraries it uses, with their licences, at the top of the file.

Four steps.

Set up (macOS or Linux)
# 1. A Tor client on this machine. Skip this if Tor Browser is open: it listens on 127.0.0.1:9150.
brew install tor && brew services start tor       # Debian and Ubuntu: sudo apt install tor

# 2. Buy blind tokens with an API key that has credits. The purchase goes over Tor too.
export ANYROUTE_API_KEY=sk-ar-v1-…
node private.mjs buy --count 20

# 3. Start the proxy. It refuses to start unless Tor answers.
node private.mjs start

# 4. In the app's shell, or its settings: point it at the proxy.
export OPENAI_BASE_URL=http://127.0.0.1:8788/v1
export OPENAI_API_KEY=anyroute-private            # any non-empty value; the proxy discards it

On the first run it asks the router for its onion address at its public name, through Tor (a Tor exit connects, your address never does), checks that the address is a valid version 3 onion address, saves it in ~/.anyroute, and uses it from then on; --onion gives one yourself. Tokens are kept in ~/.anyroute/tokens.json, a file only you can read (mode 0600) in a directory of mode 0700. ANYROUTE_HOME moves them.

What it does to each call.

  • Listens on 127.0.0.1 only, and refuses a request whose Host is not 127.0.0.1 or localhost, or that carries the Origin of a web page, so a page in your browser cannot spend your tokens. On a shared machine, --local-key makes the app present a secret.
  • Sends a fixed set of headers written by the proxy: Host, Accept, Content-Type, Content-Length, Connection, Authorization: PrivateToken and X-Anyroute-Lane: unlinkable. Nothing the app sent is copied: not its API key (an OpenAI key in OPENAI_API_KEY is discarded, never forwarded), its user agent, cookies, referrer, SDK and tracing headers, or forwarded-address headers. The body goes through as it came, except that the OpenAI user field, which names an end user, is removed.
  • Uses one token per OpenAI call, or a budget-covering set for Messages, taken out of the file before the call is sent, so it is never sent twice. A token the router refuses as spent or invalid is dropped and the call is tried with the next; a refusal for another reason (no attested provider for the model, a rate limit, a token too small for the request) keeps the token. A call sent and then lost is marked unconfirmed and its token is not used again.
  • Puts each call on its own Tor circuit, with a fresh SOCKS user name, so two calls are not carried together. --shared-circuit reuses one, which is faster.
  • Has no other way out. The only address it ever connects to is your Tor client’s SOCKS5 port (127.0.0.1:9050, or 9150 for Tor Browser, or --socks), and it asks for the onion service by name, never resolving it. It refuses to start unless a Tor client answers and the onion service reports the unlinkable lane available over Tor. If Tor stops, calls fail with 502; nothing falls back to a direct connection. When the tokens run out, calls fail with 402 and the command to buy more.
One call, and what the router sees
curl http://127.0.0.1:8788/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"<a model from /v1/models>","messages":[{"role":"user","content":"Hello"}],"stream":true}'

# what the router receives, in full:
#   POST /api/v1/chat/completions   Host: <onion address>
#   Accept, Content-Type, Content-Length, Connection: close
#   Authorization: PrivateToken token=<one token>
#   X-Anyroute-Lane: unlinkable

Which apps.

It serves POST /v1/chat/completions (streamed or not), POST /v1/embeddings and GET /v1/models, which lists the models an attested provider can serve on this lane. Point any OpenAI-compatible SDK, command-line tool or editor extension that lets you set a base URL at http://127.0.0.1:8788/v1 with any non-empty API key. Cursor has an Override OpenAI Base URL setting under Settings, Models; Cursor may send requests from its own servers, which cannot reach an address on your computer and would see your prompts, so check that your version calls the API from your computer before relying on it. Claude Code and the Anthropic SDKs can use POST /v1/messages with blind tokens; see Claude Code, unlinkable. The Responses API is not supported by this proxy.

Tokens: cost and expiry.

Size (--denomination) Face value Fits
1000 $0.002 short answers from inexpensive models
10000 (default) $0.02 most chat calls
100000 $0.20 long answers, or expensive models

A token payment pays for one call, whatever the call costs; the rest of its value is not refunded. Messages calls can combine tokens to cover their estimated budget. The router holds the worst case for a call (the prompt, and max_tokens or the model’s maximum, at the model’s price) against the token’s face value. If that is more, it answers 402 token_value_too_low and does not spend the token: lower max_tokens or buy a larger size. Face values are those of the router’s current keys; buy prints what you paid and status what you hold. Tokens expire at the end of the router’s redemption window, one to two weeks after they are bought; buy prints the time and status shows the next expiry, so buy what you will use soon. A purchase names your API key’s account; a token later shown to the router cannot be connected to it, but it hides only among the tokens of the same size bought in the same week, and buying one and spending it at once links the two by time.

Check that everything is ready
node private.mjs status

Tor:              reachable at 127.0.0.1:9050 (Tor daemon)
Onion service:    <56 characters>.onion answers (3.1 s)
Unlinkable lane:  available over onion, <n> models
Blind tokens:     18 usable (18 x 10000 units; $0.36 of face value)
  The soonest expiry is 2026-10-08 00:00 UTC.

Other limits of what this hides: someone who can watch both your connection into Tor and the router’s side can match calls by timing; the size and timing of a call are visible to the router; and a provider that receives a prompt that names you knows who you are. The lane is the one described above as the unlinkable lane over Tor. The tokens are Privacy Pass tokens (RFC 9578) made with RSA blind signatures (RFC 9474), the same ones buyTokens in @anyroute/client/blind buys.

Claude Code, unlinkable.

Claude Code can use the proxy’s Anthropic Messages API over Tor, paid with blind tokens. Start Tor on your computer, download and check private.mjs, and buy tokens with a funded AnyRoute key. The purchase uses the key; inference sends only tokens.

Buy tokens and start the proxy
export ANYROUTE_API_KEY=sk-ar-v1-…
node private.mjs buy --count 20 --denomination 10000
node private.mjs start
# Optional: --shared-circuit reuses a Tor circuit within this proxy session.
# If the operator uses a lower token cap: --max-tokens-per-request <cap>

In another shell, list the currently available attested models through the proxy. Choose models with reliable tool calling and enough context for your code. Replace both model values below with ids returned by that list. The proxy discards the client’s API key, metadata and identifying headers.

Claude Code environment
curl http://127.0.0.1:8788/v1/models
export ANTHROPIC_BASE_URL=http://127.0.0.1:8788
export ANTHROPIC_API_KEY=anyroute-private
unset ANTHROPIC_AUTH_TOKEN
export ANTHROPIC_MODEL='<attested model id from the list>'
export ANTHROPIC_DEFAULT_HAIKU_MODEL='<attested model id from the list>'
export ANTHROPIC_DEFAULT_SONNET_MODEL="$ANTHROPIC_MODEL"
export ANTHROPIC_DEFAULT_OPUS_MODEL="$ANTHROPIC_MODEL"
claude

POST /v1/messages streams the Messages events, tools, refusals and receipt headers from the router. POST /v1/messages/count_tokens is answered on your computer with the router’s estimator: no network call and no token payment. It is an estimate, including tool schemas and images, rather than a model tokenizer.

Before inference, the proxy estimates input tokens plus max_tokens at prices fetched over Tor from GET /api/v1/models?lane=unlinkable, cached in memory for one minute, including the request fee, reasoning price and royalty. It chooses the least total face value covering that estimate, breaking ties by fewer tokens, up to 16 by default. Every selected token is spent for that one request; unused value is forfeited. A changed price or a more expensive routing candidate can cause 402 token_value_too_low without spending the set. Lower max_tokens or buy larger denominations if your available set cannot cover a call.

The header extension is Authorization: PrivateToken token=A, token=B, with each value a base64url Privacy Pass token. Single-token headers and receipts retain their existing format. Sets verify and reserve atomically; receipts add only token_count, token_key_ids and nullifiers as payment identifiers. Those tokens become linked to one request, never to their purchase. A refusal before service releases the whole set; a lost response leaves the sent set unconfirmed on your computer.

Operator configuration: ANYROUTE_FEATURE_BLIND=false, UNLINKABLE_VIA_ONION=false and BLIND_MULTI_TOKEN_ENABLED=false by default. Enable them with the existing onion ingress and attestation configuration; ONION_ADDRESS and ONION_PROXY_SECRET must be configured. BLIND_MAX_TOKENS_PER_REQUEST defaults to 16 (allowed range 1–64). BLIND_UNIT_PRICE_USD defaults to 0.000002 per unit; denominations are 1000, 10000 and 100000. Tor onion access and blind tokens are switched on at anyroute.tech. Downloading the proxy changes no operator configuration.

Tor adds latency, and its first connection can take a minute. Each call uses a fresh SOCKS identity by default; --shared-circuit reuses one for this proxy process, which can reduce setup latency and makes calls share a network circuit. Claude Code’s effectiveness depends on the selected open model and its tool support.

The router reads your code and prompts in memory on this lane, and the attested provider receives them. This path hides who sent and paid for a request, not what it asks. Identifying code, file paths or text can still identify you; timing and request size can correlate calls, especially immediately after a purchase. Tor cannot prevent an observer of both ends from correlating traffic.

Private tokens from your wallet, kept on your device.

The Private tokens page turns a payment from your own wallet into blind tokens that stay in your browser, with no account and no long-lived key. It adds nothing on the router: it uses only the endpoints documented here. Every step runs in the page. It makes a one-time key (POST /api/v1/keys, no credential), keeps it in memory and in that tab’s session storage only, and you pay it from your wallet: USDG to the key’s hash through the Credits contract, or $ANYR to the escrow address. $ANYR is a payment method here, not a claim on anything: escrow credits the wallet that sent it, at the pool’s time-weighted average price (the lower of spot and the average) minus the haircut and up to the per-deposit limit in GET /api/v1/status escrow.anyr, so for $ANYR the page signs you in with one signature to a fresh key on your wallet’s account and turns only the newly credited amount into tokens. When the credit is there it blinds the tokens in the browser (RFC 9474, the suite the router signs with), buys them with POST /api/v1/blind/purchase, unblinds the signatures, keeps the tokens in IndexedDB, and offers them as a token file. Last, it switches the one-time key off at the router (PATCH /api/v1/keys/<hash> with disabled), overwrites the copy in memory and removes the stored copy. A purchase that is interrupted is finished by sending the same request again, which the router does not charge twice; a token is not shown as bought until its signature verifies.

What stays linkable: the payment is a public transaction that names your wallet, the router records that the one-time key (or, for $ANYR, your wallet’s account) spent an amount on tokens, and your network address and the time of purchase are visible unless you use the onion address. What does not: which prompts a token later paid for, because the router signs tokens blind. Tokens hide who pays, not the request: ordinary chat reads request text in router memory; the encrypted-chat adapter forwards ciphertext. A token pays for one call up to its value, the unused part is not returned, and it stops working after its redeem_until. The page works at the router’s onion address because every request it makes is relative to the address it was opened from; a browser wallet is often missing in Tor Browser, so the USDG route also shows the key hash to pay from any wallet app.

The token file (anyroute-tokens.json)
{
  "version": 1,
  "tokens": [
    {
      "token": "AAJ…",
      "key_id": "47aec8f7ed92f6f0c7f516e96db914edb6ceef5faeba45e1c252c38f9b997afe",
      "denomination": 1000,
      "epoch": 2900,
      "value_usd": "0.002",
      "redeem_until": "2026-10-13T01:53:13.999Z",
      "bought_at": "2026-09-30T00:00:00.000Z"
    }
  ],
  "unconfirmed": []
}

The file is one JSON object, the same one the private proxy keeps at ~/.anyroute/tokens.json. version is 1. Each entry in tokens has the finished token (the value after PrivateToken token=: 354 bytes of Privacy Pass token type 0x0002, base64url), the key_id it was signed under, its size and epoch, its value and its redeem_until, and when it was bought (the page writes the day, not the minute). unconfirmed holds tokens a tool sent without ever learning the outcome; they are not used again. The file holds no key, wallet address or account. A reader checks each token’s layout and that its key_id is the one inside the token, and refuses a version it does not know. A tool that spends a token removes it from the file. Keep it readable by you only (mode 0600): anyone who has the file can spend its tokens. @anyroute/client reads and writes it with no optional dependency, and mergeTokenFiles adds a download to a file you already have.

Read the file and spend a token (TypeScript SDK)
import { readFile, writeFile } from "node:fs/promises";
import { AnyRoute, parseTokenFile, serializeTokenFile, withoutTokens } from "@anyroute/client";

const path = `${process.env.HOME}/.anyroute/tokens.json`;
const file = parseTokenFile(await readFile(path, "utf8")); // throws TokenFileError with a reason if it is not a token file
const [next] = file.tokens; // one token pays for one call

const client = new AnyRoute({ baseUrl: "http://<onion address>" }).withPrivateToken(next.token);
// ... send the call, then take the token out of the file so it is not used twice:
await writeFile(path, serializeTokenFile(withoutTokens(file, [next.token])), { mode: 0o600 });
Merge a download and spend one token with curl (SOCKS5 on 127.0.0.1:9050)
# Add a downloaded file to the one you already have (do not copy it over: that would drop the tokens already there)
jq -s '.[0] + {tokens: (map(.tokens) | add | unique_by(.token))}' ~/.anyroute/tokens.json anyroute-tokens.json > merged.json \
  && chmod 600 merged.json && mv merged.json ~/.anyroute/tokens.json

# One token from the file, spent on one call over Tor
TOKEN=$(jq -r '.tokens[0].token' ~/.anyroute/tokens.json)
curl -N --proxy socks5h://127.0.0.1:9050 http://<onion address>/api/v1/chat/completions \
  -H "Authorization: PrivateToken token=$TOKEN" -H "X-Anyroute-Lane: unlinkable" -H "Content-Type: application/json" \
  -d '{"model":"<model from the list>","messages":[{"role":"user","content":"Hello"}]}'

A witnessed log of every key, where the router enables it.

Where the router runs its transparency log, every key and configuration a client encrypts to or verifies against is appended to one append-only Merkle log: receipt signing keys, Oblivious HTTP key configurations, blind-token issuer keys, measurement bundles once their Sigstore entry checks out, and the key bindings of each sidecar whose hardware quote the router’s quote verifiers accepted. The log uses the C2SP formats: tiles at /tlog/tile/, the newest checkpoint at /tlog/checkpoint as a signed note, and witnesses that cosign a checkpoint (cosignature/v1) only after checking a consistency proof from the last one they cosigned. A history shown to one user that differs from the one the witnesses saw cannot collect their cosignatures, unless the quorum of witnesses colludes with the log. GET /api/v1/tlog names the log, its witnesses and the quorum; scripts/tlog-witness.ts in the repository is a minimal witness anyone can run.

The check is opt-in in @anyroute/client. With transparency set, a receipt verifies only if its signing key is in the log under a checkpoint cosigned by the quorum of the witnesses you pinned, and consistent with every checkpoint the client saw before; client.transparency.requireLogged() does the same for any other key and refuses a key that is not logged. Two checkpoints of the same size with different roots, or a newer tree that does not extend an older one, raise SplitViewDetected with both notes as evidence. Mirrors of the checkpoint can be added as a second path. Pin the log key and the witness keys from a source other than the log: the log’s own description of itself proves nothing.

TypeScript · witnessed keys
import { AnyRoute, SplitViewDetected } from "@anyroute/client";

// Pin the log key and the witness keys from somewhere other than the log itself.
const client = new AnyRoute({
  baseUrl: "https://<router>",
  apiKey: process.env.ANYROUTE_API_KEY,
  transparency: { logKey: "<origin>+<key id>+<key>", witnesses: ["<witness key>", "<witness key>"], quorum: 2 },
});

const res = await client.chat.completions.create({ model: "meta-llama/llama-3.3-70b-instruct", messages: [{ role: "user", content: "Hello" }] });
res.anyroute.receiptVerification.checks.find((c) => c.id === "key_logged"); // fails if the receipt key is not logged

// Any other key or configuration, by kind: ohttp_key_config, blind_issuer_key, attestation_binding, ...
try {
  await client.transparency.requireLogged("ohttp_key_config", keyConfigBytes);
} catch (e) {
  if (e instanceof SplitViewDetected) console.error("two views of the log", e.evidence);
  throw e; // not_logged, not_witnessed, bad_proof: do not use the key
}

A log can also be checked through Sigstore’s public Rekor log instead of witnesses (TLOG_REKOR_ENABLED). Each time the checkpoint changes, and at most once every ten minutes, the router records it in Rekor as a hashedrekord entry: the SHA-256 of the checkpoint as the log signed it (the checkpoint text, a blank line and the log’s own signature line) with an ECDSA P-256 signature by a dedicated anchoring key. It verifies the entry’s inclusion proof before it serves it. GET /api/v1/tlog/rekor lists the anchored checkpoints with each entry’s uuid, log index, integrated time, inclusion proof and signed entry timestamp, and links each one to search.sigstore.dev; GET /api/v1/tlog/rekor/key serves the anchoring key. Rekor is run independently of Anyroute, so showing some users a second history needs a second entry under the anchoring key, which anyone can see there, and which the router’s own list does not explain. That detects a split view after the fact; unlike a witness quorum, it does not stop the log from signing a bad checkpoint before clients see it, and it helps only when clients or monitors check Rekor. In production the router starts the log only with the witness quorum or with anchoring on.

With transparency.rekor set to the anchoring key and Rekor’s key, @anyroute/client accepts a checkpoint on a verified Rekor entry in place of cosignatures: the entry must hold the hash of exactly that checkpoint, carry the pinned anchoring key and a valid signature by it, and be included in Rekor under a checkpoint that Rekor’s pinned key signed. The consistency check against the last checkpoint the client accepted, and the refusal of keys that are not logged, stay the same. With witnesses listed as well, both must hold.

TypeScript · Rekor-anchored keys
import { TransparencyLog } from "@anyroute/client";

// Rekor anchoring: a verified Rekor entry for the checkpoint counts in place of witness cosignatures.
// Pin both keys out of band: the log's anchoring key (the one published for this log; the router also serves it at
// GET /api/v1/tlog/rekor/key) and Rekor's own key (https://rekor.sigstore.dev/api/v1/log/publicKey).
const log = new TransparencyLog({
  logUrl: "https://<router>",
  logKey: "<origin>+<key id>+<key>",
  rekor: { anchorKey: "-----BEGIN PUBLIC KEY-----\n…", rekorKey: "-----BEGIN PUBLIC KEY-----\n…" },
});

const logged = await log.requireLogged("receipt_key", receiptKeyBytes); // not_anchored, anchor_mismatch, bad_anchor: refused
console.log(logged.rekor.uuid, logged.rekor.logIndex); // the entry, also on search.sigstore.dev

What we keep, generated from the schema.

The What we keep page lists every table and column the router stores, each with what it holds, whether it is recorded per request, how long it lives where the code says, and, for any column whose name or type suggests request content or a network address, a written review. It also lists the Redis keys with their lifetimes and which ones contain a caller’s address, what the application log can carry, and every place in the source that reads a request body or a caller’s address. It is generated from the schema and the descriptions in src/privacy when the site is built, and the router’s suite fails when a table or column has no description, a description names something that is not there, or a column that looks like request content or an address has no review.

GET /keep/inventory.json is the exact file the page is built from, as canonical JSON. Its SHA-256 is shown on the page with the commit the site was built from, and, where an operator sets TLOG_DATA_INVENTORY next to TLOG_ENABLED, the router appends it to the key log as a data_inventory entry the first time a router with that inventory starts. The page shows the log entry, and the Rekor entry of a checkpoint that includes it, when they exist. The inventory says what is kept, not who can read a request in flight: on every lane ordinary chat reads request text in router memory; encrypted chat forwards ciphertext to the attested gateway.

Shell · check the inventory
# the inventory, byte for byte, and its hash
curl -s <your router>/keep/inventory.json | sha256sum

# the same hash in the router's key log (a data_inventory entry), where the operator has switched that on
curl -s "<your router>/api/v1/tlog/proof?kind=data_inventory&sha256=<the hash>"

Open-weights variants, and paying their creators.

Every model has a variant. mainstream keeps the publisher’s own alignment. native_low_refusal (trained to refuse little) and abliterated (refusal behaviour removed from the weights after training) are restricted variants. GET /api/v1/models reports variant, variant_source, license, base_model, weights (source, revision, digest) and creator_handle, and takes ?variant= (a comma list) next to ?lane=. variant_source says whether an operator declared the variant; a model nobody has classified whose name says its refusals were removed is treated as abliterated.

A restricted variant is served only by a provider that is served under attested retention with a fresh attestation and whose attestation reported the in-enclave hard-block classifier as enabled (classifier_enabled in GET /api/v1/providers). That is a property of the model, not a request option: no lane or disclosure setting, and no provider.only, sends it anywhere else, and the response cache is never used for it. Its listing shows only the endpoints that qualify, and a request with none qualifying fails with no_providers before anything is sent. The router takes the flag only from what a verified attestation commits to (a development report counts only outside production) and treats anything unknown as off, so a provider whose attestation says nothing about a classifier does not qualify. The flag shows what the attestation reports; it does not prove how the classifier behaves.

When an operator enables the day-zero pipeline (DAYZERO_ENABLED, with DAYZERO_BASE_MODELS naming the base models), the router watches Hugging Face for new fine-tunes of those models whose name or tags contain a configured keyword (abliterated, uncensored, decensored and unfiltered by default), and keeps those whose model card carries an allowed license (MIT or Apache-2.0 by default; the base model’s own card is checked too). Each candidate is evaluated on a provider’s endpoint with 16 benign prompts that base models often over-refuse (fiction, security education, medical and legal information; none is harmful or illegal), 12 capability prompts with exact checks and the router’s canary set, and the scores are stored. A model becomes servable only after an operator approves the evaluated candidate and an attested provider that reports the classifier serves it. Until then it is routed to no one, and a candidate that later fails an evaluation is withdrawn. The operator endpoints are under /api/v1/lane/candidates.

The uploader of a model’s weights can claim its royalty. POST /api/v1/creators/claims with the model and a payout address returns a challenge. Commit it, on its own line, to the file named in the response on the main branch of the Hugging Face repository the weights come from, then POST /api/v1/creators/claims/ {id} /verify. The router reads the repository’s owner and the file through the Hugging Face API. On a match the address is recorded as the model’s royalty recipient (5% of the notional price unless the router is configured otherwise, at most 20%), registered in the royalty contract where one is deployed, and shown as creator and royalty_bps. Every later call to the model includes the royalty as its own cost line, and each hourly settlement streams it in USDG to the recipient. A claim needs a Hugging Face weights source recorded by the operator, and it proves control of that repository and nothing more.

Claim a royalty: request, then response
{
  "model": "<author>/<model>",
  "address": "0x…your payout address"
}

{
  "data": {
    "id": "claim-…",
    "model": "<author>/<model>",
    "hugging_face_id": "<owner>/<repository>",
    "handle": "<owner>",
    "status": "pending",
    "file": "anyroute-claim.txt",
    "file_content": "anyroute-claim-…\n",
    "expires_at": "2026-…",
    "royalty_bps": 500
  }
}

One settlement unit. More ways to pay.

How do I get USDG?

USDG is a US dollar stablecoin used for prepaid calls. For AnyRoute deposits, it must be on Robinhood Chain (chain ID 4663). A token on another chain does not fund your account here.

You can buy USDG where it is available, or bridge it from a supported chain using a service that supports USDG on Robinhood Chain. Check the destination chain, supported token, fees and withdrawal details before confirming. You also need ETH on Robinhood Chain for wallet transaction fees.

Sign in, open Add funds, and use the token and destination addresses returned by the router. A USDG Credits deposit needs approval and a deposit call; sending tokens directly to the Credits contract does not credit your key.

$ANYR and listed stock tokens are also accepted through escrow where the router enables them. Send from the wallet you signed in with. Credits wait for confirmation and a current rate, with the haircut and any per-deposit limit shown in the payment instructions. Tokens remain in escrow and credits pay for API calls.

Route Behavior
Prepaid USDG Deposit to your key’s hash on the Credits contract. 0% router fee. Withdraw any time with your key’s signature.
Agent per-call No key: the router answers 402 with a quote. Pay with CallPay (gas can be sponsored) or sign a gasless USDG authorization (see x402 below), then retry with X-Payment. 1% margin, including gas.
Stock Token A wallet-capped session. Calls accrue in USDG; at $1 or 24 hours the router swaps exactly what is owed at the Chainlink fair value and records the token units on each receipt. Every swap carries the wallet’s EIP-712 signature (a bounded allowance or the charge itself) naming the receipts it pays.

Providers are paid their list price minus a 2% settlement fee; creator royalties appear as a separate cost line. Stock Token swaps are slippage-bounded and never exceed your daily cap. Prices are per token and come from each provider.

x402: pay per call with no account.

Per-call payment is not switched on at anyroute.tech yet: calls without a key get 401 there, and GET /api/v1/status shows per_call.configured. Where the router has x402 enabled, any x402 client or agent can pay for a call in USDG on Robinhood Chain (chain id 4663) with no account and no API key. Send the request without credentials: the 402 response is an x402 v1 body (x402Version, error, accepts) with one exact-scheme requirement. Sign the USDG authorization it describes, retry the identical request with an X-PAYMENT header, and the router verifies the signature, amount, recipient, time window and nonce, relays the transfer (it pays the gas) and serves the call. Chat, completions and embeddings all work this way; GET /api/v1/status reports per_call.x402.configured.

  • Requirement. scheme exact, asset USDG, payTo the router’s receiving address, maxAmountRequired in USDG base units (6 decimals) from the same per-call price as the CallPay quote (worst case for this request, including the 1% margin), and extra carrying USDG’s EIP-712 name and version plus the numeric chainId.
  • Network name. x402 v1 names chains with lowercase hyphenated strings and has no Robinhood Chain entry, so the requirement says robinhood-chain (and the router also accepts the CAIP-2 name eip155:4663 in the payment). A client that only knows built-in networks needs that name mapped to chain 4663, or a v2 client with an eip155 handler.
  • Response. The paid response carries X-PAYMENT-RESPONSE (base64 JSON with success, transaction, network and payer) and the signed receipt records the same transaction as payment_tx.
  • Change. The whole payment is credited to the paying wallet’s account and the call is charged from it, so anything unused stays there as credit you can spend with X-Wallet-Auth. Paying more than maxAmountRequired is allowed; nothing is refunded on-chain.
  • Rejections. A payment that fails verification gets a 402 whose error is a reason such as invalid_exact_evm_payload_authorization_value (underpaid), invalid_exact_evm_payload_recipient_mismatch, invalid_exact_evm_payload_authorization_nonce_used (replayed) or invalid_exact_evm_payload_signature, together with fresh requirements. Nothing moves on-chain for a rejected payment.
x402 client (JavaScript, viem)
import { privateKeyToAccount } from "viem/accounts";
import { toHex } from "viem";

const account = privateKeyToAccount(process.env.AGENT_KEY);
const url = "https://<router>/api/v1/chat/completions";
const headers = { "content-type": "application/json" };
const body = JSON.stringify({ model: "meta-llama/llama-3.3-70b-instruct", messages: [{ role: "user", content: "Hello" }] });

// 1. Ask with no key. The router answers 402 with x402 payment requirements.
const { accepts } = await (await fetch(url, { method: "POST", headers, body })).json();
const req = accepts[0];

// 2. Sign a USDG transfer authorization (EIP-3009) to req.payTo. The router relays it and pays the gas.
const authorization = {
  from: account.address,
  to: req.payTo,
  value: BigInt(req.maxAmountRequired),
  validAfter: 0n,
  validBefore: BigInt(Math.floor(Date.now() / 1000) + req.maxTimeoutSeconds),
  nonce: toHex(crypto.getRandomValues(new Uint8Array(32))),
};
const signature = await account.signTypedData({
  domain: { name: req.extra.name, version: req.extra.version, chainId: req.extra.chainId, verifyingContract: req.asset },
  types: { TransferWithAuthorization: [
    { name: "from", type: "address" }, { name: "to", type: "address" }, { name: "value", type: "uint256" },
    { name: "validAfter", type: "uint256" }, { name: "validBefore", type: "uint256" }, { name: "nonce", type: "bytes32" },
  ] },
  primaryType: "TransferWithAuthorization",
  message: authorization,
});
const xPayment = btoa(JSON.stringify({
  x402Version: 1, scheme: "exact", network: req.network,
  payload: { signature, authorization: Object.fromEntries(Object.entries(authorization).map(([k, v]) => [k, String(v)])) },
}));

// 3. Retry the identical request with X-PAYMENT.
const paid = await fetch(url, { method: "POST", headers: { ...headers, "X-PAYMENT": xPayment }, body });
const completion = await paid.json(); // usage + signed receipt; receipt.payload.payment_tx is the settlement
const settlement = JSON.parse(atob(paid.headers.get("X-PAYMENT-RESPONSE"))); // { success, transaction, network, payer }

The response is only the beginning.

Every generation returns normalized usage and a signed receipt with hashes of the request and response (never their content), in two encodings: v1 (JSON, Ed25519) and v2 (a COSE_Sign1 signed EdDSA with the same key, with token counts in buckets such as 512-1024 and no payer). A stream commits to every event as it goes: after each one comes a comment line, : anyroute-chain <i> <hash>, that SSE parsers skip, and the v2 receipt signs the last hash, so a cut or altered stream shows. Receipts are rooted in hourly Merkle batches and GET /api/v1/receipts/{id}/proof returns the path. A root is posted to ReceiptAnchor on Robinhood Chain only where the router runs with a configured chain; otherwise it stays off chain and the proof says anchored: false. Signed is not the same as anchored: the dashboard and /api/v1/receipts/verify report each separately.

Receipts a provider’s sidecar signs with its enclave key can be anchored per host. Where the router runs with host anchoring on, it collects each attested host’s receipt leaves once an interval (an hour by default) over the connection pinned to that host’s attested certificate. It takes the receipt key from the host’s boot quote only when SHA-256 of that quote is the attestation reference it verified and the quote commits to the key, keeps only leaves whose signature verifies under that key and that name that attestation, and roots them per host and interval. Each root is stored with the attestation reference and, where a chain is configured, posted with ReceiptAnchor.anchorAttested under keccak256 of the provider id; otherwise it stays off chain. GET /api/v1/host-anchors/proof/{leaf}, or POST /api/v1/host-anchors/proof with the receipt, returns the root, the path, the attestation reference, the receipt key and the status: anchored: true with the transaction, block and attested anchor index once the root is on chain, anchored: false while it is not.

Response shape
{
  "id": "gen-1790461071-M1D5SJxd7YpD5A",
  "model": "meta-llama/llama-3.3-70b-instruct",
  "provider": "DeepInfra",
  "choices": [
    {
      "message": {
        "role": "assistant",
        "content": "…"
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 120,
    "completion_tokens": 64,
    "total_tokens": 184,
    "cost": 0.00003248,
    "is_byok": false,
    "cost_details": {
      "upstream_inference_cost": 0.00003248,
      "royalty": 0,
      "margin": 0
    },
    "prompt_tokens_details": {
      "cached_tokens": 0,
      "cache_write_tokens": 0
    },
    "completion_tokens_details": {
      "reasoning_tokens": 0
    }
  },
  "receipt": {
    "id": "gen-1790461071-M1D5SJxd7YpD5A",
    "sig": "<base64 Ed25519 signature>",
    "key_id": "ab214b090f5922fe",
    "alg": "Ed25519",
    "payload": {
      "v": 1,
      "model": "…",
      "provider": "…",
      "tokens": {
        "prompt": 120,
        "completion": 64
      },
      "cost": "0.00003248",
      "mode": "prepaid",
      "request_sha256": "…",
      "response_sha256": "…"
    },
    "anchor_hint": "Anchored on chain 4663 within the hour; GET /api/v1/generation?id=… returns the merkle proof.",
    "paid_with": {
      "token": "NVDA",
      "raw_units": "<units>",
      "fair_price": "<18-decimal USD>",
      "swap_tx": "<tx once swapped>"
    }
  }
}

Response headers

Every chat, completion and embeddings response names its receipt in X-Receipt-Id, the same id as the body’s id and receipt.id, and again in Inference-Id, the header Hugging Face inference clients read. X-Anyroute-Lane is the lane the request was served under: public, attested or unlinkable. X-Anyroute-Policy-Hash is sent only when the endpoint that served the call has a fresh, verified attestation that binds the hash of the policy its in-enclave classifier enforces; the router relays that value and never computes one. On a stream these headers arrive before the first chunk, so the policy hash is sent there only when every endpoint the request can reach attested the same one. Browsers can read all of them.

Response headers
X-Receipt-Id: gen-1790461071-M1D5SJxd7YpD5A
Inference-Id: gen-1790461071-M1D5SJxd7YpD5A
X-Anyroute-Lane: attested
X-Anyroute-Policy-Hash: sha256:<64 hex, only when the endpoint attested one>

GET /api/v1/models adds, per model, an attestation object for its strongest live endpoint and a datacenter_region. best is that endpoint’s disclosure class. manifest_ref points at the transparency-log entry and the on-chain registry transaction of its measured image, each only once the router checked it. policy_hash is the attested classifier policy. Anything the router has not verified is null, never filled in; exec_profile_id stays null until an attestation reports one. datacenter_region is set only when every endpoint reports the same single region.

Both GET /api/v1/models and /v1/models add provider_names and capabilities: vision (Reads images), imageOut (Makes images), audio, tools, longContext (at least 128,000 tokens), attested (Proven hardware), network (AnyRoute network) and encrypted (Encrypted chat). Network tags require a live, freshly attested host offer for its admitted model. Encrypted chat requires the separate enabled gateway path with fresh evidence and the exact model ID; ordinary Harness chat remains readable by the router. Older responses derive only the tags their fields establish.

GET /api/v1/models (new fields)
"attestation": {
  "best": "attested",
  "manifest_ref": { "rekor_entry": "<log entry uuid>", "registry_tx": null },
  "exec_profile_id": null,
  "policy_hash": "sha256:<64 hex>"
},
"datacenter_region": null

What we saw: a privacy label for every receipt.

The output facet reads usage.unit_type and usage.units: token (the default for older receipts) covers text and embeddings, image_mp images, video_sec video, audio_sec voice and audio, call tools or search, and gpu_sec computation or fine-tuning. Units describe metering, not every input modality or which operation ran. A token call can include image references. Ordinary chat reads request text in router memory on every lane; the encrypted-chat adapter forwards ciphertext. For media, generation records keep request and output hashes, not the media itself; metering does not establish retention outside that record or by a provider. A call unit alone does not prove a search query was sent, and GPU time does not prove training data or weights were deleted. Unfamiliar units report content and address handling as unrecorded. Council labels include members and the judge; RAG has separate embedding and chat receipts. The response cache keeps the encrypted reply, never the prompt itself; semantic caching keeps a hashed word vector. Batch content is kept sealed outside the database; requests are deleted when the batch finishes and answers when results expire. Failed-provider errors can retain a short sanitized fragment of a rejected request, and unexpected library errors can quote fragments in logs. See What we keep for these limits.

GET /api/v1/receipts/{id}/privacy reads a receipt back in plain English: who could read the prompt, who saw your network address, how the call was paid, what the router kept and what hardware answered. It is public by id, like the receipt. The router works it out from the receipt’s signed fields (lane, disclosure, mode, payer, nullifier, attestation and, for an attested gateway, upstream_attestation) and from what its code does for that combination. A field a receipt does not carry is reported as not recorded and never assumed to be favourable. GET /api/v1/receipts/{id} returns the same label as privacy, beside the signed payload and never inside it, so nothing that is signed changes. The MCP chat tool returns the summary with its result, the Telegram bot puts a one-line form above each answer’s footer, and the verify page shows the label for a receipt id at /verify/?r=<receipt id>.

  • Who could read the prompt. The router reads ordinary chat prompts in memory on every lane. With the encrypted-chat adapter, correctly encrypted content is decrypted in the gateway enclave; the router sees ciphertext and clear metadata. The gateway forwards restored content to the serving workload over a separate confidential channel. The label calls that an attested enclave only when the receipt shows the router had verified the provider’s hardware attestation and, for an attested gateway, that the gateway’s signed receipt for the exchange checked out. Otherwise it is a provider that documents a no-retention policy no hardware backs, or one that may keep the prompt under its own terms. When the router withheld a reply because that check failed, the label says the provider had already read the prompt.
  • Who saw your address. On the unlinkable lane, nobody at AnyRoute: the lane is served only to requests that arrive over Tor (or through an independent relay), and a direct request is refused before it is served. On any other lane AnyRoute’s servers saw the address the request came from. Its software writes that address to no table: the record of a call has no column for it. A call with no API key (a blind token, a wallet or an x402 payment) is rate-limited per address, so for about a minute the address is the key of a counter; a call with a key is limited per key and uses no address. What the network provider that hosts the router logs is outside what a receipt can show, and the label says so.
  • How it was paid. An API key’s balance, Stock Token pay-with, your own provider key, a blind token, a wallet or an x402 payment, and so what the receipt names: a key hash, a wallet address, or only the hash of a spent token. A blind token is signed blind, so the token itself cannot be matched to its purchase; the timing and size of purchases and spends can still hint at a link.
  • What was kept. Token counts, cost, timing, SHA-256 fingerprints of the request and the reply, the payer as above, the providers tried and a ledger line for the charge. There is no column for a prompt, a reply or an address. The response cache is opt-in and holds an encrypted copy until it expires; it is never used on the attested or unlinkable lanes or with a blind token. Anyone who has a receipt’s id can read the receipt.
  • What hardware answered. Attested only where the receipt says so and the check behind it passed; the TEE type comes from the router’s attestation record for the receipt’s attestation hash; GPU only where the gateway’s receipt asserted it. A development report is labelled as one and is never called attested.
GET /api/v1/receipts/{id}/privacy · attested lane, paid from a key
{
  "data": {
    "receipt_id": "gen-1790461071-M1D5SJxd7YpD5A",
    "lane": "attested",
    "label": {
      "prompt_readers": {
        "router": true,
        "provider": {
          "id": "attested-gateway",
          "access": "attested_enclave",
          "reply_withheld": false
        },
        "text": "…"
      },
      "network": {
        "hidden": false,
        "via": null,
        "stored": false,
        "counter": "none",
        "text": "…"
      },
      "payment": {
        "kind": "key_balance",
        "identifies": "api_key",
        "text": "…"
      },
      "stored": {
        "prompt_text": false,
        "reply_text": false,
        "client_address": false,
        "fingerprints": true,
        "linked_to": "api_key",
        "cache": "never",
        "records": [
          "…"
        ],
        "public_by_id": true,
        "text": "…"
      },
      "hardware": {
        "attested": true,
        "tee": "Intel TDX",
        "gpu_attested": true,
        "verified_by": "gateway_receipt",
        "development_report": false,
        "text": "…"
      }
    },
    "summary": [
      "Read by: AnyRoute's router (in memory, to route it) and the provider's attested enclave (attested-gateway).",
      "Your IP address: seen by AnyRoute's servers when you connected; not saved with this answer.",
      "Paid with: an API key's balance. The receipt names the key by its hash.",
      "Kept: token counts, cost, timing and hashes of the request and reply. Not kept: the prompt or the reply text.",
      "Hardware: attested (Intel TDX with GPU attestation), checked from the gateway's signed receipt."
    ],
    "short": "Read by: router + proven enclave · IP: seen, not saved · Paid: API key balance",
    "verify_url": "https://router.example/verify?r=gen-1790461071-M1D5SJxd7YpD5A"
  }
}

The label reads a receipt; it does not check it. Verify the signature first (POST /api/v1/receipts/verify, the verify page or an SDK), then compute the label on your side: privacyLabel(receipt) in @anyroute/client gives the router’s label for any receipt you hold (pass teeKind if you know the attestation’s TEE type), and fetchPrivacyLabel(baseUrl, id) reads the router’s.

Ask several models, or the same one twice.

Two opt-in modes, available when the router enables them (the ANYROUTE_FEATURE_COUNCIL setting, off by default). Neither streams. Every call they make is routed, billed and receipted like a request of its own, and the worst case of all of them is held before anything is sent, so your balance and key budget bound the whole request. Your provider preferences, including a disclosure ceiling or lane, apply to every call, the judge included: a member with no provider that meets them refuses the request (409, or 503 no_attested_endpoint on a lane) instead of being dropped or downgraded. The disclosure header and the top-level receipt show the weakest class among the calls; each member’s receipt shows its own.

Council. Set model to anyroute/council and list 2 to 5 members and a judge. The members run in parallel; the judge either picks one answer, returned unchanged (mode judge), or writes a final one (mode fuse, text only). A member that fails is not billed and the council goes on with the rest, as long as at least two answered. The response carries a council field listing each member (model, receipt id, cost, latency) and the judge; the top-level receipt is the judge’s call and its signed payload lists the member receipt ids. Set council.max_cost_usd to refuse the request unless its worst case fits; lower max_tokens to make it fit. A judge is a model’s opinion, not a proof.

Council request
{
  "model": "anyroute/council",
  "messages": [
    {
      "role": "user",
      "content": "Your prompt"
    }
  ],
  "max_tokens": 400,
  "council": {
    "models": [
      "<model-a>",
      "<model-b>",
      "<model-c>"
    ],
    "judge": "<judge-model>",
    "mode": "judge",
    "max_cost_usd": 0.05
  }
}
Council response (abridged)
{
  "id": "<judge receipt id>",
  "model": "anyroute/council",
  "choices": [
    {
      "message": {
        "role": "assistant",
        "content": "<the chosen member's answer>"
      }
    }
  ],
  "usage": {
    "prompt_tokens": 1180,
    "completion_tokens": 402,
    "cost": 0.0021,
    "calls": 4
  },
  "receipt": {
    "id": "<judge receipt id>",
    "sig": "…",
    "payload": {
      "council": {
        "members": [
          {
            "receipt_id": "<member receipt id>"
          }
        ],
        "judge": {
          "receipt_id": "<judge receipt id>"
        }
      }
    }
  },
  "council": {
    "mode": "judge",
    "members": [
      {
        "label": "A",
        "model": "<model-a>",
        "provider": "…",
        "receipt_id": "<member receipt id>",
        "cost": "0.00041",
        "latency_ms": 812,
        "status": "ok"
      }
    ],
    "judge": {
      "model": "<judge-model>",
      "receipt_id": "<judge receipt id>",
      "cost": "0.00062",
      "latency_ms": 390
    },
    "selected": {
      "label": "A",
      "receipt_id": "<member receipt id>"
    },
    "total_cost": "0.0021"
  }
}

Dual verification. Add verify: "dual" to a request for one model. It goes to two different providers of that model at temperature 0 with a fixed seed (yours, if you send seed), and both outputs are compared. The response has a verification field with the two provider ids, agree (true when the outputs match exactly or differ only in whitespace), the two receipt ids and the seed; each receipt carries the same agreement bit. Both calls are billed, and the body is the first provider’s output. If fewer than two providers of the model support temperature and seed under your routing preferences, the answer is 409. Agreement shows two providers gave the same text; it does not show that either is attested, and providers running different quantizations may legitimately differ.

Dual verification request
{
  "model": "meta-llama/llama-3.3-70b-instruct",
  "messages": [
    {
      "role": "user",
      "content": "Your prompt"
    }
  ],
  "verify": "dual"
}

Attested council and attested dual verification. Set council.attested to true (or send provider.lane or the X-Anyroute-Lane header as attested, which asks for the same thing) and every member and the judge are held to the attested lane: a provider whose retention is declared attested and whose hardware attestation the router holds fresh, the same test as any other attested-lane request. Nothing is downgraded and no member is dropped. A member, or the judge, with no attested provider refuses the whole request with a 503 (no_attested_endpoint, and error.metadata.council_seat says which seat); nothing is sent, held or charged. Each call’s signed receipt then has an attestation_ref: the hash of the attestation report the router verified for the provider that served it, and the TLS key its connection was pinned to (tls_pin is null for a provider that did not attest through a self-signed certificate). The response’s council field adds attested and attestation_refs (members in order, then the judge), and both are signed in the top-level receipt, so changing a reference breaks its signature. attested is true only when every call was served under the attested class. These are the router’s own records, the same ones behind GET /api/v1/attestation/ {providerId} ; they show what was running and pinned, not what it did with your prompt.

Attested council request
{
  "model": "anyroute/council",
  "messages": [
    {
      "role": "user",
      "content": "Your prompt"
    }
  ],
  "council": {
    "models": [
      "<model-a>",
      "<model-b>"
    ],
    "judge": "<judge-model>",
    "attested": true
  }
}
Attested council response (abridged)
{
  "id": "<judge receipt id>",
  "model": "anyroute/council",
  "council": {
    "attested": true,
    "attestation_refs": [
      {
        "role": "member",
        "label": "A",
        "receipt_id": "<member receipt id>",
        "provider": "<provider-a>",
        "tee": "tdx",
        "report_hash": "<attestation report hash>",
        "attested_at": "…",
        "tls_pin": {
          "spki_sha256": "…",
          "attestation_ref": "…"
        }
      },
      {
        "role": "member",
        "label": "B",
        "receipt_id": "<member receipt id>",
        "provider": "<provider-b>",
        "tee": "tdx",
        "report_hash": "…",
        "attested_at": "…",
        "tls_pin": null
      },
      {
        "role": "judge",
        "receipt_id": "<judge receipt id>",
        "provider": "<provider-a>",
        "tee": "tdx",
        "report_hash": "…",
        "attested_at": "…",
        "tls_pin": {
          "spki_sha256": "…",
          "attestation_ref": "…"
        }
      }
    ]
  }
}

Dual verification with provider.lane set to attested sends the two calls to two different attested providers of the model. Both receipts carry the agreement bit, and the verification field and each receipt add attested and the two attestation references. If the model has fewer than two attested providers the answer is 409 verification_unavailable with one, or 503 no_attested_endpoint with none, and nothing is sent or charged. Agreement still only shows that two providers gave the same text.

Attested dual verification request
{
  "model": "<model>",
  "messages": [
    {
      "role": "user",
      "content": "Your prompt"
    }
  ],
  "verify": "dual",
  "provider": {
    "lane": "attested"
  }
}

Availability. Production has one attested provider today, and it serves a small (0.5B) model, so an attested council, which needs an attested provider for every member and the judge, and attested dual verification, which needs two for one model, will mostly answer with the 409 above until more attested providers join. That refusal is the intended behaviour, not a fault: the router does not fall back to a provider that is not attested. On a development router the attestation can be a development report; receipts then say so (attestation_simulated, and simulated inside the reference), and a production router never accepts one. Receipts from calls made any other way do not carry attestation_ref, and older receipts verify as before.

Use every model as a tool.

The router hosts a remote MCP server at /mcp (Streamable HTTP, stateless, JSON replies). Connect it to Claude, Cursor or any MCP client with your Anyroute key. Six tools: list_models (live models, context length and price per 1M tokens), chat (call any model; returns the reply, a receipt id, cost and latency), get_receipt and verify_receipt, and two for private work, list_attested_models and verify_provider (below). Chat goes through /api/v1/chat/completions with your key, so balance, limits and signed receipts are the same. Only chat needs a key.

Claude Code
claude mcp add --transport http anyroute <your router>/mcp --header "Authorization: Bearer $ANYROUTE_API_KEY"
Cursor · ~/.cursor/mcp.json
{
  "mcpServers": {
    "anyroute": {
      "url": "<your router>/mcp",
      "headers": {
        "Authorization": "Bearer sk-ar-v1-…"
      }
    }
  }
}
Claude Desktop · claude_desktop_config.json (through the mcp-remote bridge)
{
  "mcpServers": {
    "anyroute": {
      "command": "npx",
      "args": [
        "-y",
        "mcp-remote",
        "<your router>/mcp",
        "--header",
        "Authorization:${ANYROUTE_AUTH}"
      ],
      "env": {
        "ANYROUTE_AUTH": "Bearer sk-ar-v1-…"
      }
    }
  }
}
Check it with curl
curl -s <your router>/mcp \
  -H "content-type: application/json" \
  -H "accept: application/json, text/event-stream" \
  -H "Authorization: Bearer $ANYROUTE_API_KEY" \
  -d '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-06-18","capabilities":{},"clientInfo":{"name":"curl","version":"1"}}}'

Keep a prompt with proven enclaves

list_attested_models lists the models the attested lane can serve now: those with an endpoint whose TEE attestation the router verified itself and whose provider documents no retention. Each carries gpu_attested, true when the latest verified gateway receipt for the model asserted GPU attestation, false when it did not and null before any receipt. chat takes lane (public or attested) and disclosure (none, policy or any) with the meaning of provider.lane and provider.disclosure: on the attested lane the prompt goes only to such a provider, and when none can answer the call fails (503 no_attested_endpoint) with nothing sent and nothing charged. The result reports the lane, the disclosure class the signed receipt records and, when the provider is an attested gateway, its upstream_attestation (attested, gpu_attested, receipt_verified and a reason when it is not attested). If the gateway’s receipt does not show an attested upstream, the reply is withheld and the error carries the receipt id; the upstream had already produced the answer, so the call is billed.

To make every chat call on a connection attested, add ?lane=attested to the /mcp URL, or send X-Anyroute-Lane: attested; ?disclosure= and X-Anyroute-Disclosure-Max set a ceiling the same way. The strictest setting wins: a call can tighten the connection’s setting and never relax it, and a value the router does not recognise refuses the call instead of falling back to public. The unlinkable lane needs a relay and a blind token, so it is not offered over MCP.

Claude Code · every chat call attested
claude mcp add --transport http anyroute-attested "<your router>/mcp?lane=attested" --header "Authorization: Bearer $ANYROUTE_API_KEY"

# the same restriction as a header instead of a URL query
claude mcp add --transport http anyroute-attested <your router>/mcp --header "Authorization: Bearer $ANYROUTE_API_KEY" --header "X-Anyroute-Lane: attested"
tools/call · chat on the attested lane
{
  "jsonrpc": "2.0",
  "id": 2,
  "method": "tools/call",
  "params": {
    "name": "chat",
    "arguments": {
      "model": "<model from list_attested_models>",
      "prompt": "Your prompt",
      "lane": "attested"
    }
  }
}
Result (abridged)
{
  "text": "…",
  "receipt_id": "gen-…",
  "lane": "attested",
  "disclosure": "attested",
  "upstream_attestation": {
    "attested": true,
    "gpu_attested": true,
    "receipt_verified": true,
    "kind": "aci/1",
    "receipt_id": "<gateway receipt id>"
  },
  "…": "cost_usd, latency_ms, model, provider, usage"
}

verify_provider takes a provider id (the provider field of a receipt) and returns the router’s own attestation record in plain terms: the status (attested or unverified; a development report is marked as such and refused in production), the TEE and the verifiers that accepted its quote, whether the router pins the provider’s TLS key, the state of the transparency-log entry for its measurement, and not_checked, the list of what the router does not verify. Attestation shows what code is running, not what a provider does with a prompt. The same record is at GET /api/v1/attestation/:providerId and on the verify page.

The Telegram bot follows the same rule. /private on sends every chat with lane attested, /models attested lists the models with a proven enclave and /model accepts only those while private mode is on. Each answer’s footer says attested, or attested · GPU when the gateway’s receipt asserts GPU attestation, taken from the signed receipt, with a link to the provider’s verify page; an answer whose receipt does not show an attested provider is not delivered. When no attested provider can answer, the bot says nothing was sent and nothing was charged.

Use AnyRoute from Claude Code and the Anthropic SDKs.

The router speaks the Anthropic Messages API: POST /v1/messages (also under /api/v1) and POST /v1/messages/count_tokens. A client written for that API, such as the Anthropic SDKs or Claude Code, works with an Anyroute key and an Anyroute model. The request is converted to a chat completion and sent through /api/v1/chat/completions inside the router, so the key, its balance and limits, the lane, the signed receipt and the response headers are those of a chat call. Anyroute serves open models, not Anthropic’s, so choose a model from GET /v1/models. This is a compatibility layer over the Messages API, not an Anthropic service: how well an agent such as Claude Code works depends on the model you choose, and it needs reliable tool calling and a long context.

Claude Code

Point ANTHROPIC_BASE_URL at the router and put your Anyroute key in ANTHROPIC_AUTH_TOKEN (sent as Authorization: Bearer) or ANTHROPIC_API_KEY (sent as x-api-key); the router accepts both. Claude Code names Anthropic models for its main, subagent and background calls, so set ANTHROPIC_MODEL and the ANTHROPIC_DEFAULT_SONNET_MODEL, ANTHROPIC_DEFAULT_OPUS_MODEL and ANTHROPIC_DEFAULT_HAIKU_MODEL variables (and CLAUDE_CODE_SUBAGENT_MODEL) to Anyroute model ids. Claude Code assumes a 200K context for a model it does not know; if your model has less, set CLAUDE_CODE_AUTO_COMPACT_WINDOW to its window. Keep the key out of a project’s committed .claude/settings.json.

Claude Code · environment
export ANTHROPIC_BASE_URL=<your router>
export ANTHROPIC_AUTH_TOKEN=$ANYROUTE_API_KEY        # sent as Authorization: Bearer. ANTHROPIC_API_KEY sends x-api-key instead; either works.

# Anthropic model names are not served here: name AnyRoute models.
export ANTHROPIC_MODEL=meta-llama/llama-3.3-70b-instruct
export ANTHROPIC_DEFAULT_SONNET_MODEL=meta-llama/llama-3.3-70b-instruct
export ANTHROPIC_DEFAULT_OPUS_MODEL=meta-llama/llama-3.3-70b-instruct
export ANTHROPIC_DEFAULT_HAIKU_MODEL=qwen/qwen3-32b   # session titles and other background calls
export CLAUDE_CODE_SUBAGENT_MODEL=qwen/qwen3-32b

export CLAUDE_CODE_ATTRIBUTION_HEADER=0              # keep Claude Code's attribution line out of the prompt
claude
Claude Code · ~/.claude/settings.json
{
  "env": {
    "ANTHROPIC_BASE_URL": "<your router>",
    "ANTHROPIC_AUTH_TOKEN": "sk-ar-v1-…",
    "ANTHROPIC_MODEL": "<a model from GET /v1/models>",
    "ANTHROPIC_CUSTOM_HEADERS": "X-Anyroute-Lane: attested"
  }
}

The attested lane

Send X-Anyroute-Lane: attested, or {"provider":{"lane":"attested"}} in the body, and the prompt goes only to a provider whose TEE attestation the router has verified and that documents no retention: the same lane, with the same meaning, as on /api/v1/chat/completions. When no such provider can serve the model the call is refused with 409 (503 while they are down) and nothing is sent or charged; an unrecognised lane is a 400, never public. X-Anyroute-Disclosure-Max and provider.disclosure work the same way. In Claude Code, ANTHROPIC_CUSTOM_HEADERS adds the header to every request (one Name: Value pair per line; in a settings file use \n between pairs). Use a model from GET /v1/models?lane=attested. Attestation shows what code is running, not what a provider does with a prompt; see the verify page for what the router checked and did not.

Claude Code · every request on the attested lane
# Every request from this Claude Code session on the attested lane.
# A request no attested provider can serve is refused (409) with nothing sent and nothing charged.
export ANTHROPIC_CUSTOM_HEADERS="X-Anyroute-Lane: attested"
export ANTHROPIC_MODEL=<a model from GET /v1/models?lane=attested>
Python SDK · a call on the attested lane, reading the receipt headers
import anthropic

client = anthropic.Anthropic(base_url="<your router>", api_key="sk-ar-v1-…")   # or auth_token=... for Authorization: Bearer

raw = client.messages.with_raw_response.create(
    model="<a model from GET /v1/models?lane=attested>",
    max_tokens=512,
    messages=[{"role": "user", "content": "Hello"}],
    extra_body={"provider": {"lane": "attested"}},        # or extra_headers={"x-anyroute-lane": "attested"}
)
message = raw.parse()
print(message.content[0].text)
print(raw.headers["x-receipt-id"], raw.headers["x-anyroute-lane"], raw.headers.get("x-anyroute-policy-hash"))
TypeScript SDK
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic({ baseURL: "<your router>", apiKey: process.env.ANYROUTE_API_KEY });

const { data: message, response } = await client.messages
  .create(
    { model: "<a model from GET /v1/models?lane=attested>", max_tokens: 512, messages: [{ role: "user", content: "Hello" }] },
    { headers: { "x-anyroute-lane": "attested" } }, // or provider: { lane: "attested" } in the body
  )
  .withResponse();
console.log(message.content, response.headers.get("x-receipt-id"), response.headers.get("x-anyroute-lane"));
curl
curl -s <your router>/v1/messages \
  -H "x-api-key: $ANYROUTE_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -H "x-anyroute-lane: attested" \
  -d '{"model":"<a model from GET /v1/models?lane=attested>","max_tokens":256,"messages":[{"role":"user","content":"Hello"}]}'

Model names

Any model id in the catalog works as sent, and so do saved routes and a key’s own model aliases. A claude-* name is not in the catalog: unless the operator has mapped it, the call is a 404 (not_found_error) that says so and names a model to use. An operator maps names with ANTHROPIC_MODEL_MAP, a JSON object from a name sent by a client to a catalog model id. A name ending in * matches every name with that prefix (the longest prefix wins, and an exact name wins over a prefix), and a lone * answers any other name. A map that is not valid JSON stops the router starting.

Router configuration
ANTHROPIC_MODEL_MAP={"claude-sonnet-4-5":"meta-llama/llama-3.3-70b-instruct","claude-haiku-*":"qwen/qwen3-32b"}

What is converted

  • Request. model; system as a string or text blocks; messages with text, image (base64 and URL), tool_use and tool_result blocks (a tool result may hold text and images), and document blocks with a text source; system-role entries inside messages stay where they are. max_tokens, temperature, top_p, top_k, stop_sequences and stream map directly; metadata.user_id is sent to the provider as user. top_k is dropped for a provider that does not list it, and a max_tokens above what the model can produce is lowered to its limit, since Anthropic clients ask for large values. thinking blocks in history are dropped.
  • Tools. Custom tools become function tools with their JSON schema, and tool_choice auto, any, tool and none map to auto, required, a named function and none (disable_parallel_tool_use sets parallel_tool_calls to false). Tools hosted by Anthropic (web search, code execution, bash, text editor, computer use) cannot run behind a model here: they are left out of the request and named in the X-Anyroute-Ignored response header, as is a thinking request, which is accepted and not acted on. PDF documents, mcp_servers and container are refused with a 400 that names the field.
  • Reply. A message object with text and tool_use blocks. stop_reason is end_turn, max_tokens, tool_use or refusal (from a content filter); stop_sequence appears only when the provider says which stop string ended the answer, otherwise an answer that hit one is end_turn. usage splits cached prompt tokens out of input_tokens as cache_read_input_tokens and cache_creation_input_tokens, so the three add up to the prompt. cache_control markers are accepted and do nothing.
  • Receipt. The id of the message is the receipt id, and the reply carries an anyroute object with the lane, the disclosure class, the provider, the cost, the gateway’s upstream_attestation where there is one and the signed receipt. X-Receipt-Id, Inference-Id, X-Anyroute-Lane and X-Anyroute-Policy-Hash come back as headers, as on a chat call.
  • Streaming. With stream: true the events are the API’s own: message_start, ping, then for each block content_block_start, content_block_delta (text_delta, or input_json_delta with the tool call’s arguments as they arrive) and content_block_stop, then message_delta with the billed usage and the anyroute object, and message_stop. message_start states an estimate of the prompt tokens; message_delta has the counted figure. A failure after output has begun ends the stream with an error event and no message_stop. A request the router refuses before any output is a real HTTP error, so an SDK can retry it.
  • Counting. POST /v1/messages/count_tokens takes the same body without max_tokens and returns input_tokens. It is the router’s estimate (about one token for every three characters of text and JSON, and 1,600 per image), not a tokenizer’s count, and it costs nothing. It needs a key.
Status error.type When
400 invalid_request_error The request is malformed or names something not supported; the message names the field. A provider that rejects the request also lands here.
401 authentication_error No key, an unknown key, or a disabled or expired key.
402 · 403 · 404 · 413 · 429 billing_error · permission_error · not_found_error · request_too_large · rate_limit_error Not enough balance; a key that may not do this; an unknown model; a body over 16 MB; a rate limit (with Retry-After).
409 · 501 invalid_request_error · api_error The requested lane cannot be served, or is not run by this router. Sent with X-Should-Retry: false so an SDK does not retry it.
502 · 503 · 504 api_error · api_error · timeout_error Every provider for the request failed (nothing is charged), or an attested reply was withheld because the gateway’s receipt did not show an attested upstream (that call is billed).

Every error has Anthropic’s shape, {"type":"error","error":{"type","message"},"request_id"} , plus an anyroute object with the router’s own error type, its metadata, and the receipt id when a refused call was billed. Every response has a request-id header. A browser can call the endpoint directly: the router allows the x-api-key, anthropic-version, anthropic-beta and anthropic-dangerous-direct-browser-access headers.

Use AnyRoute from any Ollama client.

The router speaks the Ollama API under /ollama: GET /ollama/api/tags, /api/version and /api/ps, and POST /api/show, /api/chat, /api/generate, /api/embed and the older /api/embeddings. Point an Ollama client (Open WebUI, Continue, the ollama Python and JavaScript libraries, LangChain’s ChatOllama, editor and notes plugins) at this router’s address followed by /ollama, and give it your Anyroute key. Each call is converted and sent through /api/v1/chat/completions (or /api/v1/embeddings) inside the router, so the key’s balance and limits, the lane, the signed receipt and the response headers are those of a chat call. The models run on hosted providers: nothing is downloaded and nothing runs on your machine.

The key

Send the key as Authorization: Bearer, which most Ollama clients set from an API key or headers setting. There is no ?key= query parameter, so a key never ends up in a URL or a log line. The ollama command line tool sends no key: with OLLAMA_HOST set it lists (ollama list) and describes (ollama show) the catalog, which are public, and a chat needs a client that sends the header.

ollama CLI · OLLAMA_HOST
export OLLAMA_HOST=<your router>/ollama
ollama list                      # the live catalog, as Ollama names
ollama show meta-llama/llama-3.3-70b-instruct:latest
Python · the ollama library, on the attested lane
from ollama import Client

client = Client(
    host="<your router>/ollama",
    headers={"Authorization": "Bearer sk-ar-v1-…", "X-Anyroute-Lane": "attested"},
)
reply = client.chat(
    model="<a model from GET /ollama/api/tags>",
    messages=[{"role": "user", "content": "Hello"}],
)
print(reply.message.content)
Open WebUI
# Open WebUI: the Ollama connection
OLLAMA_BASE_URL=<your router>/ollama
# then in Admin Settings, Connections, give that connection your key (sk-ar-v1-…) as its Bearer key
Continue
# Continue: ~/.continue/config.yaml
models:
  - name: Anyroute
    provider: ollama
    model: meta-llama/llama-3.3-70b-instruct:latest
    apiBase: <your router>/ollama
    apiKey: sk-ar-v1-…            # sent as Authorization: Bearer
    requestOptions:
      headers:
        X-Anyroute-Lane: attested
LangChain · ChatOllama
from langchain_ollama import ChatOllama

llm = ChatOllama(
    model="<a model from GET /ollama/api/tags>",
    base_url="<your router>/ollama",
    client_kwargs={"headers": {"Authorization": "Bearer sk-ar-v1-…"}},
)
curl
curl -s <your router>/ollama/api/chat \
  -H "Authorization: Bearer $ANYROUTE_API_KEY" \
  -H "X-Anyroute-Lane: attested" \
  -d '{"model":"<a model from GET /ollama/api/tags>","messages":[{"role":"user","content":"Hello"}],"stream":false}'

What is converted

  • Models. GET /ollama/api/tags lists the live catalog. Each name is a catalog id with Ollama’s :latest tag, and a routing suffix works as a tag too (:free, :nitro, :private). size is 0, since there are no local weights, and digest is the SHA-256 of the id; details has the family, the parameter size and the precision, read from the name and the live endpoints. With X-Anyroute-Lane: attested the list holds only the models that lane can serve now. POST /api/show adds capabilities (completion, tools, vision, thinking, or embedding) and the context length in model_info. A pull of a listed model succeeds at once; create, copy, push, delete and blob uploads are 501s; /api/ps is always empty.
  • Requests. /api/chat messages (system, user, assistant, tool) map to chat messages, and images (base64, with the media type read from the bytes) become image parts. /api/generate sends system and prompt as a system and a user message. format "json" is JSON mode and a JSON schema is structured output. options temperature, top_p, top_k, min_p, seed, stop, frequency_penalty and presence_penalty map directly, repeat_penalty becomes repetition_penalty and num_predict becomes max_tokens (-1 and -2 set no limit). Options that tune a local runtime (num_ctx, num_gpu, num_thread and the like), suffix, template, raw and context are accepted and named in the X-Anyroute-Ignored header; keep_alive is ignored. A request with no messages or no prompt loads the model as Ollama does: done_reason load (unload with keep_alive 0), with no call and no charge.
  • Tools. tools are function tools, and the model’s calls come back in message.tool_calls with the arguments as an object. A tool message answers the call it names with tool_call_id, else the earlier call with the same tool_name, else the next call without an answer.
  • Streaming. Ollama streams unless stream is false. The reply is NDJSON (application/x-ndjson): one object per line with done false and a piece of message.content (response for /api/generate), tool calls whole in one line, then a closing line with done true, done_reason (stop or length), total_duration, load_duration (0), prompt_eval_count, prompt_eval_duration, eval_count and eval_duration in nanoseconds, and an anyroute object with the receipt id, the lane, the disclosure class, the provider and the cost. Reasoning the provider returns comes back as thinking unless think is false.
  • Embeddings. /api/embed takes input as a string or a list and returns embeddings, in order; /api/embeddings takes prompt and returns one embedding. dimensions is passed on.
  • Receipt and lane. X-Anyroute-Lane, X-Anyroute-Disclosure-Max and provider in the body work as on a chat call, and X-Receipt-Id, Inference-Id, X-Anyroute-Lane and X-Anyroute-Policy-Hash come back as headers.
A streamed /api/chat reply
{"model":"meta-llama/llama-3.3-70b-instruct:latest","created_at":"2026-09-30T12:00:00.000Z","message":{"role":"assistant","content":"Hel"},"done":false}
{"model":"meta-llama/llama-3.3-70b-instruct:latest","created_at":"2026-09-30T12:00:00.041Z","message":{"role":"assistant","content":"lo!"},"done":false}
{"model":"meta-llama/llama-3.3-70b-instruct:latest","created_at":"2026-09-30T12:00:00.052Z","message":{"role":"assistant","content":""},"done":true,"done_reason":"stop","total_duration":412000000,"load_duration":0,"prompt_eval_count":9,"prompt_eval_duration":361000000,"eval_count":3,"eval_duration":51000000,"anyroute":{"receipt_id":"gen-…","lane":"attested","disclosure":"attested","provider":"…","cost_usd":0.0000021}}

Every error is Ollama’s {"error":"..."} with the router’s status: 400 for a bad field (the message names it), 401 or 402 without a usable key, 404 for an unknown model (the message names one to use), 409 or 503 when the requested lane cannot be served, and 429 with Retry-After. A request every provider fails before any output is an HTTP error; a failure after output has begun ends the stream with an {"error"} line and no closing line.

Use AnyRoute with the OpenAI Agents SDK and Codex.

POST /v1/responses (also /api/v1/responses) is the OpenAI Responses API, so the OpenAI Agents SDK, the Codex CLI and other Responses clients work with a change of base URL and key. The base URL is this router’s address followed by /v1 (or /api/v1: the two are the same), and the key is the one you use for chat completions. The endpoint is an adapter: it sends your request to /api/v1/chat/completions inside the router with your credentials and your routing headers, so balance, limits, disclosure ceilings, lanes and signed receipts are exactly those of a chat call, and the answer comes back as a Response object, or as the Responses event stream when stream is true (response.created, response.in_progress, response.output_item.added, response.content_part.added, response.output_text.delta, response.output_text.done, response.content_part.done, response.function_call_arguments.delta and .done, response.output_item.done, then response.completed, or response.incomplete when the answer was cut off by max_output_tokens; a failure after the stream has started is an error event followed by response.failed).

OpenAI Agents SDK (Python)
import asyncio
import os

from agents import Agent, Runner, set_default_openai_api, set_default_openai_client, set_tracing_disabled
from openai import AsyncOpenAI

client = AsyncOpenAI(
    base_url="<your router>/v1",
    api_key=os.environ["ANYROUTE_API_KEY"],  # sk-ar-v1-…
    default_headers={"X-Anyroute-Lane": "attested"},  # optional: attested providers only
)
set_default_openai_client(client, use_for_tracing=False)
set_default_openai_api("responses")
set_tracing_disabled(True)  # otherwise the SDK sends traces to its own tracing service

agent = Agent(name="Assistant", instructions="Answer briefly.", model="<model from GET /api/v1/models?lane=attested>")


async def main():
    result = await Runner.run(agent, "Hello, AnyRoute.")
    print(result.final_output)


asyncio.run(main())
Codex CLI · ~/.codex/config.toml
# ~/.codex/config.toml  (export ANYROUTE_API_KEY first)
model = "<model id from GET /api/v1/models>"
model_provider = "anyroute"

[model_providers.anyroute]
name = "AnyRoute"
base_url = "<your router>/v1"
env_key = "ANYROUTE_API_KEY"
wire_api = "responses"
# optional: send every prompt only to attested providers
http_headers = { "X-Anyroute-Lane" = "attested" }
Check it with curl
curl -s <your router>/v1/responses \
  -H "Authorization: Bearer $ANYROUTE_API_KEY" -H "Content-Type: application/json" \
  -H "X-Anyroute-Lane: attested" \
  -d '{"model":"<model from GET /api/v1/models?lane=attested>","input":"Your prompt","max_output_tokens":200}'

Keep a prompt with proven enclaves

Set the lane the same way as for chat: the X-Anyroute-Lane: attested header (the Agents SDK example above and the Codex config send it on every call) or provider: {"lane": "attested"} in the request body; when both are given the stricter applies. On the attested lane the prompt goes only to a provider whose TEE attestation the router verified itself and that documents no retention. If none can answer, the call is refused with 503 no_attested_endpoint (error.metadata.reason none_attested, or attested_endpoints_down with Retry-After while they are down), nothing is sent to any provider and nothing is charged; the router does not fall back to a public provider. Attestation shows what code is running, not what a provider does with a prompt; see the verify page for what is and is not checked. The response carries the receipt and lane headers of a chat call (X-Receipt-Id, Inference-Id, X-Anyroute-Lane and, where the serving endpoint attested one, X-Anyroute-Policy-Hash), and its metadata has anyroute_receipt_id, anyroute_lane and anyroute_disclosure taken from the same signed receipt. The response id is resp_ followed by the receipt id. Models that can serve the attested lane are listed at GET /api/v1/models?lane=attested.

Response (abridged)
{
  "id": "resp_gen-…",
  "object": "response",
  "status": "completed",
  "model": "<model>",
  "output": [
    {
      "type": "message",
      "id": "msg_…",
      "role": "assistant",
      "status": "completed",
      "content": [
        {
          "type": "output_text",
          "text": "…",
          "annotations": []
        }
      ]
    }
  ],
  "usage": {
    "input_tokens": 12,
    "output_tokens": 5,
    "total_tokens": 17,
    "cost": 0.00002
  },
  "metadata": {
    "anyroute_receipt_id": "gen-…",
    "anyroute_lane": "attested",
    "anyroute_disclosure": "attested"
  },
  "store": false,
  "…": "instructions, tools, tool_choice, temperature, top_p and the other fields Responses clients read"
}

Stateless by design

AnyRoute keeps no conversation or response on the server, so there is nothing to continue from or fetch back. Send the whole conversation in input on every request, including the function_call and function_call_output items of a tool round trip; the Agents SDK and Codex already do this. store defaults to false and store: true is refused with a 400, as are previous_response_id, conversation, background, stored prompts and references to stored items or files. GET, DELETE and cancel under /v1/responses/:id answer 404 and say why. A call’s signed receipt, which holds no prompt or answer, is at GET /api/v1/receipts/:id with the id from X-Receipt-Id.

A refusal
{
  "error": {
    "code": 400,
    "type": "previous_response_id_not_supported",
    "param": "previous_response_id",
    "message": "`previous_response_id` is not supported. AnyRoute stores no responses, so there is no earlier turn to continue from. Send the whole conversation in `input` on every request…"
  }
}

What is supported

Request: model, instructions, input (text, or message items with input_text and input_image, function_call and function_call_output items, custom_tool_call and custom_tool_call_output items; earlier reasoning items are ignored), function and custom tools with tool_choice and parallel_tool_calls, max_output_tokens, temperature, top_p, text.format (text, json_object or json_schema, which needs a model and provider that support it), reasoning.effort, metadata (echoed back, never sent to a provider), user, and provider for routing. Other options, such as include, truncation and service_tier, are accepted and ignored. Response: a message with output_text and one function_call or custom_tool_call item per tool call, and usage with input_tokens, output_tokens and total_tokens (plus cost in USD). Tools that run on the API provider’s servers (web_search, file_search, code_interpreter, computer_use, image_generation, hosted mcp) are refused with a 400 that names the tool, because AnyRoute hosts none: give the model a function tool and run the work in your own code, and switch off any client feature, such as web search in Codex, that depends on one. File inputs are also refused.

Codex and apply_patch: custom tools

The Codex CLI offers apply_patch as a custom tool: freeform text, not JSON fields, with an optional grammar. AnyRoute passes a custom tool to the model as a function with one string argument, input, and puts the tool’s description, a line saying the tool takes freeform text, and the grammar if there is one into the function’s description. The grammar is guidance for the model only: nothing checks or enforces it, and a model can still produce input that does not follow it, which the client then reports as a failed call. When the model calls the function, the call comes back as a custom_tool_call item with the text as input (in a stream, response.custom_tool_call_input.delta as it is generated and response.custom_tool_call_input.done with the whole text), and the custom_tool_call_output item you send on the next turn goes back to the model as the tool result. Models differ in how well they follow a tool description, so how reliably apply_patch works depends on the model you choose. Codex’s shell and plan tools are ordinary function tools.

Answer from your documents without storing them.

POST /api/v1/rag (also /v1/rag) answers a question from documents you send with it. The router cuts them into overlapping chunks, embeds the chunks and the question through /api/v1/embeddings, ranks the chunks by cosine similarity in memory, and asks a chat model, through /api/v1/chat/completions, to answer from the best top_k of them, citing them by number. Each of those is an ordinary call with your key: billed to it, limited by it, routed by lane and signed as a receipt, and the response lists every receipt. It needs a prepaid API key; per-call payment and blind tokens are not accepted.

Ask a question of two documents, on the attested lane
curl <your router>/api/v1/rag \
  -H "Authorization: Bearer $ANYROUTE_API_KEY" -H "Content-Type: application/json" \
  -d '{"documents":[{"id":"handbook","text":"Refunds are issued within 14 days of the return arriving. …"},{"id":"faq","text":"…"}],"question":"How long do refunds take?","model":"<chat model>","provider":{"lane":"attested"}}'

There is a page for it too: Ask your files reads .txt, .md, .csv, .json, .html, .docx and .pdf files in your browser (a PDF for its text layer, page by page; scanned pages are not read) and sends the text to this endpoint with your question.

Request. documents is a list of strings, or of objects with an id and a text; an id defaults to doc-1, doc-2 and so on, and ids must be unique. By default one request may carry 200 documents, 2 MiB of text, 2,000 chunks and 64 embeddings calls; above any of them it is refused with 413 before anything is sent (the router’s RAG_MAX_DOCUMENTS, RAG_MAX_BYTES, RAG_MAX_CHUNKS and RAG_MAX_EMBEDDING_CALLS settings). chunk.size (100 to 8,000 characters, default 1,000) and chunk.overlap (default 15% of the size, at most half of it) set the chunking, and a chunk ends at a paragraph, line or sentence boundary where it can. top_k (1 to 20, default 4) is how many chunks go into the prompt; a request whose largest possible prompt would not fit the chat model’s context is refused (400 context_too_small) before anything is embedded. embedding_model defaults to qwen/qwen3-embedding-8b when the catalog serves it, else the cheapest embedding model, considering models with an attested endpoint first. max_tokens and temperature go to the chat call. Unknown fields are refused, so a misspelt option is never silently ignored.

Response. The answer, and sources in rank order: each names its document, the chunk’s position, its cosine score and where the chunk lies in the document (start and end). The text of a source is returned only if you send include_excerpts as true. receipts has one entry for every embeddings call and one for the chat call, each with its lane, the disclosure class its signed receipt records and, when the provider is an attested gateway, upstream_attestation: attested, gpu_attested and receipt_verified, what the router checked in the gateway’s receipt for that call. Then the totals. The headers are the chat call’s (X-Receipt-Id and the others), X-Anyroute-Lane is the lane every call used, X-Anyroute-Disclosure is the weakest class among all the calls, and X-Anyroute-Policy-Hash is sent only when every call reported the same one. The answer is a model’s output grounded in the sources you were shown, not a proof; the prompt asks for numbered citations so you can check them.

Response (abridged)
{
  "id": "<chat receipt id>",
  "object": "rag.answer",
  "model": "<chat model>",
  "answer": "Refunds are issued within 14 days of the return arriving [1].",
  "finish_reason": "stop",
  "sources": [
    {
      "ref": 1,
      "document_id": "handbook",
      "chunk_index": 0,
      "score": 0.71,
      "start": 0,
      "end": 212
    }
  ],
  "receipts": [
    {
      "step": "embeddings",
      "receipt_id": "<embeddings receipt id>",
      "model": "qwen/qwen3-embedding-8b",
      "provider": "<provider>",
      "lane": "attested",
      "disclosure": "attested",
      "cost": 0.0000011,
      "tokens": {
        "prompt": 105,
        "completion": 0
      },
      "inputs": 3,
      "upstream_attestation": {
        "attested": true,
        "gpu_attested": true,
        "receipt_verified": true,
        "kind": "aci/1",
        "receipt_id": "<gateway receipt id>"
      }
    },
    {
      "step": "chat",
      "receipt_id": "<chat receipt id>",
      "model": "<chat model>",
      "provider": "<provider>",
      "lane": "attested",
      "disclosure": "attested",
      "…": "cost, tokens, upstream_attestation"
    }
  ],
  "embedding_model": "qwen/qwen3-embedding-8b",
  "lane": "attested",
  "lane_source": "request",
  "disclosure": "attested",
  "retrieval": {
    "documents": 2,
    "chunks": 2,
    "top_k": 2,
    "chunk": {
      "size": 1000,
      "overlap": 150
    },
    "embedding_calls": 1
  },
  "usage": {
    "embedding_tokens": 105,
    "prompt_tokens": 240,
    "completion_tokens": 18,
    "cost": 0.0000101,
    "cost_usd": "0.0000101"
  }
}

Lane, and what a refusal looks like

Send provider.lane (public or attested), provider.disclosure (none, policy or any), X-Anyroute-Lane or X-Anyroute-Disclosure-Max, and every call runs under exactly that. So does a lane pinned on your key (its routing.provider), which applies to the embeddings step as well as the chat step. If a step cannot be served under it, whether there is no attested embedding model, the chat model has no attested endpoint, or the attested endpoints are down, the request is refused with that step’s own error (on the attested lane 503 no_attested_endpoint, with error.metadata.reason none_attested or attested_endpoints_down; with only a disclosure ceiling 409 disclosure_unavailable or 503 disclosure_provider_unavailable), and it is never sent on a weaker lane. error.metadata.step says which step stopped, and error.metadata.receipts lists the calls already made, which are billed. If an attested gateway’s receipt for a call does not show an upstream it verified inside a TEE, that call’s output (the vectors, or the answer) is withheld, the call is billed, and the error is 502 upstream_not_attested with the receipt listed and marked withheld. The unlinkable lane is not available here. A model the router resolves itself (a saved route, an alias of your key, a router model such as anyroute/council) can pin a lane of its own, so a request naming one must state provider.lane (400 lane_required).

A chat model without an attested endpoint, on the attested lane (abridged)
{
  "error": {
    "code": 503,
    "type": "no_attested_endpoint",
    "message": "RAG stopped at the chat step: No endpoint for <chat model> that fits this request has a fresh, verified attestation, so lane \"attested\" cannot be served. Nothing was sent to any provider and nothing was charged. …",
    "metadata": {
      "step": "chat",
      "lane": "attested",
      "reason": "none_attested",
      "receipts": [
        {
          "step": "embeddings",
          "receipt_id": "<embeddings receipt id>",
          "lane": "attested",
          "…": "the call that was made, and billed"
        }
      ]
    }
  }
}

Send none of them and the router chooses: the attested lane when the chat model and the embedding model both have an attested endpoint right now, and otherwise the router’s ordinary public lane. The response says which (lane, and lane_source: request or default) and, when it is public, why (lane_note). Setting provider.lane to attested yourself is how you make the request refuse instead of falling back.

Streaming

With stream set to true the embeddings run first, and a refusal up to the start of the answer is an ordinary JSON error. The stream then sends a chat.completion.chunk event with empty choices and rag.object rag.sources (the sources and the embeddings receipts), the chat stream exactly as /api/v1/chat/completions sends it (on the attested lane an attested gateway’s text is held back until its receipt has been checked), a last chunk with rag.object rag.summary (every receipt, the totals, and error if the answer was refused after the stream began), and data: [DONE].

What is kept, and what is not

Nothing of your documents is stored. The documents, their chunks and vectors, the question and the answer exist in the memory of the request that carries them. The endpoint writes none of them to a database, cache or file, does not log them, and never uses the response cache (it sends no cache option to the calls it makes and ignores X-Anyroute-Cache). When the request ends the router lets go of them and the runtime reclaims the memory. The router does not overwrite freed memory, so this is not a claim about what someone with access to the running process could recover while a request is in flight or soon after.

What is recorded. Each embeddings call and the chat call is a generation of your key, exactly as if you had made it yourself: a record and a signed receipt holding ids, model, provider, lane, token counts, cost, timing, the disclosure class, any gateway attestation check, and the SHA-256 of the request and of the response. The hashes do not reveal text, but whoever holds an exact guess of a request can confirm it against one, and a receipt can be read by its id.

What leaves the router. The text has to reach models to be used. The chunks and the question go to the embedding model’s provider; the question and the best chunks go to the chat model’s provider. What that provider sees and keeps depends on the lane. On the attested lane the router sends them only to a provider whose retention is declared attested and whose hardware attestation it verified itself and holds fresh, and for an attested gateway it checks the gateway’s receipt for each call. That shows what code is running, not what it does with your text; GET /api/v1/attestation/:providerId lists what the router does not check. On the public lane a provider’s documented policy applies (GET /api/v1/disclosure/:providerId).

Untrusted text. Documents are treated as untrusted. The prompt tells the model to use only the numbered sources and to ignore instructions inside them, and a source cannot close its own tag. That lowers, and does not remove, the chance that a document steers the answer. The cost of a request is the sum of its calls; each holds its worst case before it is sent, so your balance bounds every step, and a refusal partway leaves the earlier calls billed.

Run many calls at half the price.

POST /api/v1/batches (also /v1/batches) is the OpenAI Batch API without files. Send many chat completions or embeddings requests in one call and the router runs them in the background, in spare capacity, at 50% of the normal price. A batch can take minutes, and lines not run within the 24-hour completion window expire and are not charged. It needs a prepaid key (Authorization: Bearer). Each line is billed, budgeted, rate-limited and routed exactly like a normal call of that key, so lanes, provider preferences, key budgets and key rate limits apply to every line; a failed line is not charged, and every answered line has its own signed receipt. In the dashboard, Batch Studio can send a batch this way (Run on the server).

Send a batch inline, then read the results
curl <your router>/api/v1/batches \
  -H "Authorization: Bearer $ANYROUTE_API_KEY" -H "Content-Type: application/json" \
  -d '{"completion_window":"24h","requests":[{"custom_id":"row-1","method":"POST","url":"/v1/chat/completions","body":{"model":"<chat model>","messages":[{"role":"user","content":"Summarize: …"}],"max_tokens":200}},{"custom_id":"row-2","method":"POST","url":"/v1/chat/completions","body":{"model":"<chat model>","messages":[{"role":"user","content":"Classify: …"}]}}]}'

# Check on it, then read the results once status is completed
curl <your router>/api/v1/batches/$BATCH_ID -H "Authorization: Bearer $ANYROUTE_API_KEY"
curl <your router>/api/v1/batches/$BATCH_ID/output -H "Authorization: Bearer $ANYROUTE_API_KEY" > output.jsonl
curl <your router>/api/v1/batches/$BATCH_ID/errors -H "Authorization: Bearer $ANYROUTE_API_KEY" > errors.jsonl

Request. Send requests, a list of lines, or input_jsonl, the same lines as JSONL text, with completion_window 24h. Each line has a custom_id (unique, at most 64 characters), method POST, a url of /v1/chat/completions or /v1/embeddings, and the body of that request. endpoint, if you send it, must match every line’s url. There is no files endpoint, so input_file_id is refused with 400: send the lines inline. stream: true, council mode (anyroute/council) and verify are refused in a batch. A batch with an invalid line is not created: the 400 invalid_request lists each problem in error.metadata.errors with its line (counted from 1), code and message.

Progress and results. The response is a Batch object with its id (batch_…). GET /api/v1/batches/:id returns it again: status (validating, in_progress, then completed, failed or expired; cancelling, then cancelled, after a cancel), request_counts (total, completed, failed) and cost (usd, the discounted cost so far, and list_usd, what the same calls cost at the normal price). GET /api/v1/batches?limit=20&after=<batch id> lists your batches, newest first. POST /api/v1/batches/:id/cancel stops a batch: lines already running finish and are billed, lines not started are never run or billed. GET /api/v1/batches/:id/output returns one JSONL line per answered request, with the normal response body (usage.cost and the receipt included), and GET /api/v1/batches/:id/errors one per failed or unrun request.

A line of the output file, and one of the errors file (abridged)
{"id":"batch_req_…","custom_id":"row-1","response":{"status_code":200,"request_id":"<generation id>","body":{"id":"<generation id>","choices":["…"],"usage":{"…":"tokens","cost":0.0000162},"receipt":{"…":"the signed receipt of this line"}}},"error":null}
{"id":"batch_req_…","custom_id":"row-2","response":{"status_code":402,"request_id":null,"body":{"error":{"…":"the error a normal call returns"}}},"error":{"code":"insufficient_credits","message":"…"}}

Limits. By default a batch holds up to 1,000 lines and 8 MiB of input, and a key can have 2 unfinished batches at a time; one more is refused with 429 batch_limit. The operator sets these with BATCH_MAX_LINES, BATCH_MAX_BYTES and BATCH_MAX_ACTIVE. Results are kept for 24 hours after the batch finishes (BATCH_RESULTS_TTL) and then deleted; after that the output and errors routes answer 410 batch_results_expired.

What is kept. The request lines and the answers are never written to the database. Until the results expire they are kept sealed (AES-256-GCM, with a key derived from the router’s secret) in Redis, or in the router’s memory where there is no Redis, and then deleted. The seal protects them from someone who can read Redis without the router’s secret; the router itself can open them, which it does to return your results. The database keeps only each batch’s counts, costs and statuses and the generation id of each answered line, and every answered line has the generation record and signed receipt of a normal call.

Rerank documents against a query.

POST /api/v1/rerank (also /v1/rerank) takes the Cohere and Jina request: a model, a query, documents as strings or objects with a text, and optionally top_n and return_documents. It answers with results best first, each with the index of a document you sent and its relevance_score (and the document’s text when return_documents is true), usage with total_tokens and search_units, the cost, and a signed receipt. The scores are the provider’s own: an answer that is not a valid ranking of the documents sent counts as a failed attempt, the next provider is tried, and if none answers the call fails with 502 and nothing is charged. One request may carry 1,000 documents; a long document is truncated or split by the provider.

Rerank three documents, cheapest provider first
curl <your router>/api/v1/rerank \
  -H "Authorization: Bearer $ANYROUTE_API_KEY" -H "Content-Type: application/json" \
  -d '{"model":"<rerank model>:floor","query":"How long do refunds take?","documents":["Shipping takes 3 to 5 days.","Refunds are issued within 14 days.",{"text":"Returns need a receipt."}],"top_n":2,"return_documents":true}'
Response (abridged)
{
  "id": "<receipt id>",
  "object": "rerank",
  "model": "<rerank model>",
  "provider": "<provider>",
  "results": [
    {
      "index": 1,
      "relevance_score": 0.93,
      "document": {
        "text": "Refunds are issued within 14 days."
      }
    },
    {
      "index": 2,
      "relevance_score": 0.41,
      "document": {
        "text": "Returns need a receipt."
      }
    }
  ],
  "usage": {
    "total_tokens": 48,
    "search_units": 1,
    "cost": 0.002,
    "cost_details": {
      "upstream_inference_cost": 0.002,
      "royalty": 0
    }
  },
  "meta": {
    "billed_units": {
      "search_units": 1
    }
  },
  "cost": 0.002,
  "receipt": {
    "id": "<receipt id>",
    "sig": "…",
    "key_id": "…",
    "alg": "Ed25519",
    "payload": {
      "kind": "rerank",
      "search_units": 1,
      "documents": 3,
      "…": ""
    }
  }
}

Only models whose architecture.output_modalities include rerank serve it; GET /api/v1/models?output_modalities=rerank lists them. Any other model, or a router where no provider lists a rerank model yet, answers 404 model_not_found. A rerank model is priced per token (pricing.prompt) and per search unit (pricing.request: one query over up to 100 documents, each split into chunks of 500 tokens), from what the provider reports; when it reports nothing, from the router’s estimate, and usage.estimated says so. Keys, budgets, per-call payment, blind tokens, lanes, the disclosure ceiling, provider preferences and the :nitro and :floor suffixes work as they do for embeddings. The Oblivious HTTP gateway does not carry rerank calls yet.

Verify before you send.

Two client libraries wrap the OpenAI-shaped call and add the checks a plain HTTP client would skip. The TypeScript package, @anyroute/client, has no runtime dependencies and runs on Bun, Node 20 and later, and in browsers. The Python package, anyroute-client (Python 3.10 and later), depends only on cryptography and httpx. Their source is in packages/client and packages/client-py in the repository. Both report every check as pass, fail or not checked, and a check they did not make is never shown as passed.

Receipts

Each response carries its verification: the Ed25519 signature over the receipt’s canonical JSON against the key in the router’s published key list (fetched once, and read again if a receipt names a key it has not seen, as after a weekly rotation), that the key id is the hash of the key, that the receipt is dated inside its key’s window, that its leaf recomputes from the signed bytes and, when the receipt came with an anchor proof, that the leaf is under the stated root. It does not check that the key is registered on chain or that the root was posted there, and says so. Pass pinned keys to skip the fetch. A receipt a provider’s sidecar signed with its enclave key is checked against the receipt key its attestation binds, and must name the same attestation and model digest. verifyHostAnchor checks such a receipt against its host root in the same order: the signature (under the key verifyProvider bound, when you pass it), the path to the root and, with a reader for ReceiptAnchor (readAttestedAnchor over any RPC endpoint you choose), the root, provider and attestation on chain. A root kept off chain is reported as off chain, never as anchored.

TypeScript · receipts, disclosure and lane
import { AnyRoute } from "@anyroute/client";

const client = new AnyRoute({ baseUrl: "https://<router>", apiKey: process.env.ANYROUTE_API_KEY });
const res = await client.chat.completions.create(
  { model: "meta-llama/llama-3.3-70b-instruct", messages: [{ role: "user", content: "Hello" }] },
  { disclosure: "policy" }, // or lane: "attested"; the router refuses rather than downgrade
);

const check = res.anyroute.receiptVerification; // checked against /.well-known/anyroute-receipt-keys.json
console.log(check.valid, check.anchor);          // anchor: "proof_valid" | "proof_invalid" | "no_proof"
for (const c of check.checks) console.log(c.id, c.status, c.detail); // pass | fail | not_checked

privacyLabel(receipt) turns a receipt you have verified into the plain-English label the router serves at GET /api/v1/receipts/:id/privacy (see What we saw above), computed on your side from the signed fields alone; fetchPrivacyLabel(baseUrl, receiptId) reads the router’s.

Attested providers

Give a request an attested option and the client checks the provider first and sends nothing unless every check passes. It reads the router’s record (GET /api/v1/attestation/:providerId) and the provider’s own /attest document, then checks that the router reports the provider attested with a quote it verified recently; that the quote is an Intel TDX quote whose report_data is SHA-256 of the canonical bindings followed by the nonce, so the TLS key, receipt key and the image, compose and model digests are committed in the quote; that a fresh quote for a random nonce the client chose says the same; that the certificate name is derived from SHA-256 of the quote; that, where the runtime can read the connection’s certificate, it carries that name and the attested TLS key; and that the router’s recorded digests equal the provider’s. If you supply the model digest you expect, it must match. Simulated (development) evidence is refused unless you opt in, and is then labelled simulated. On success the request is pinned to that provider (provider.only, no fallbacks, lane attested), and the receipt is checked for naming it.

What it does not check, and reports as not checked: Intel’s signature and certificate chain over the quote (the router does that; pass a quoteVerifier to run your own), what the provider does with your prompt, whether the running software matches its published source, and the transport when the runtime cannot read the certificate, as in a browser. Passing nodeAttestFetcher on Node or Bun reads the certificate from the same connection that served /attest.

TypeScript · verify before send
import { AnyRoute, AttestationRefused } from "@anyroute/client";
import { nodeAttestFetcher } from "@anyroute/client/node"; // Node and Bun: reads the provider's certificate too

try {
  const res = await client.chat.completions.create(request, {
    attested: {
      providerId: "<provider id>",
      attestUrl: "https://<provider>/attest",
      expected: { modelDigest: "sha256:<the digest you expect>" }, // optional, but this is what makes it your model
      attestFetcher: nodeAttestFetcher(),
    },
  });
  console.log(res.anyroute.provider.bound.modelDigest, res.anyroute.servedByVerifiedProvider);
} catch (e) {
  if (e instanceof AttestationRefused) console.error(e.verification.checks.filter((c) => c.status === "fail"));
  else throw e;
}
Python
from anyroute_client import AnyRoute, AttestedOptions, AttestationRefused, ExpectedDigests

client = AnyRoute("https://<router>", "sk-ar-v1-…")
try:
    res = client.chat(
        {"model": "meta-llama/llama-3.3-70b-instruct", "messages": [{"role": "user", "content": "Hello"}]},
        attested=AttestedOptions(
            provider_id="<provider id>",
            attest_url="https://<provider>/attest",
            expected=ExpectedDigests(model_digest="sha256:<the digest you expect>"),
        ),
    )
    print(res["anyroute"]["receipt_verification"].valid)
except AttestationRefused as e:
    print([c for c in e.verification.checks if c.status == "fail"])

Blind tokens and end-to-end encryption

Where the router has blind tokens enabled, buyTokens() from @anyroute/client/blind blinds, buys and unblinds tokens with your key, and client.withPrivateToken(token) makes a client that spends one; the blind-token module needs the optional package @cloudflare/blindrsa-ts, and the main entry point never loads it. For end-to-end encryption to an enclave, sealedPost() encrypts a request with an HPKE implementation you supply, sends it as application/anyroute-hpke and opens the reply. It seals only to an HPKE key the provider’s verified quote commits to, and refuses if the provider did not verify or the quote commits to no such key. The SDK ships no HPKE cipher: the implementation and wire format must match the provider’s sidecar. The Python package covers receipts, provider verification, and disclosure and lane options, and can spend a blind token you already hold; it cannot buy tokens and has no streaming or HPKE support.

The verify page shows what the router has recorded for a provider (/verify/?p=<provider id>), the privacy label of a receipt id (/verify/?r=<receipt id>) and checks a pasted receipt in your browser. It reads the router’s record only; use an SDK to check the provider itself.

Public host records.

Hosts lists providers with a hardware verification record. GET /api/v1/hosts returns public identity, hardware status, probation and models. GET /api/v1/hosts/{providerId} adds the existing attestation status and paged history, current and superseded build measurements with Rekor links, per-host receipt-root counts and the latest root, observed uptime, proof-time and coarse invoice bands.

The API is available when HOST_DASHBOARD_ENABLED=true (default false). The static page reads the host id from /hosts/?id={providerId} or /hosts/#{providerId} and uses the same origin, including over onion. A stale host remains inspectable, with no current hardware verification claim.

Public earnings are lifetime USDG settlement invoice bands after provider fees, including paid invoices and excluding the open hour. They are not a real-time payout balance. Exact lifetime and unpaid invoice units, payout mode and payout address appear only in an operator response. USDG has six decimal places. An unpaid invoice has no recorded payment transaction; a pending payout may already be processing.

To request the operator view, sign anyroute:{unixSeconds}:{sha256(scope)} with personal_sign, where scope is exactly GET /api/v1/hosts/{encodeURIComponent(providerId)}. Send X-Wallet-Auth: {address}:{unixSeconds}:{signature} with the GET. The router uses walletAuth’s five-minute window and replay protection, then compares the recovered wallet with the stored host operator. Signatures for other resources are refused. Detail responses use Cache-Control: no-store; the list is public.

A confirmed root with a recorded transaction is reported as anchored on chain. A root kept off chain says so; pending roots are awaiting confirmation. Uptime is the measured success rate of requests and probes across served models, not continuous availability. Proof-time counts periods holding fresh router-verified hardware evidence, with earlier unrecorded time unknown. Build measurements establish recorded digests, not build admission approval, source equivalence or prompt handling. AnyRoute’s router reads request text in memory on every lane today. These views add no table, money logic, log field or key family; signed access reuses the existing wallet-auth replay marker described in What we keep.

Self-serve host signup

NETWORK_HOSTS_ENABLED defaults to false; signup and status return 404 while disabled. Enabling it requires NETWORK_POLICY_ENABLED and TLOG_ENABLED. Production also requires Redis, hardware quote verifiers over HTTPS, a persistent log signing key and the existing independent checkpoint checks. Publish a signed host policy before admission.

Switched on at anyroute.tech.

Send POST /api/v1/network/hosts with {name, endpoint, payout_address, models, contact?}. Name is at most 60 characters, models contains one to eight distinct IDs, and optional contact is at most 120 characters. Endpoint is the sidecar root URL over HTTPS, without credentials, query or fragment; the router calls its public /attest endpoint. HTTP submissions are recorded as rejected without contacting the endpoint. Private and reserved network destinations are refused by the existing provider transport in production.

Use the existing X-Wallet-Auth: address:unixSeconds:signature header. Sign anyroute:unixSeconds:sha256(canonical JSON body) with personal_sign: recursively sort object keys, retain array order, encode UTF-8 without whitespace. Signatures expire after five minutes and are single-use. Signup is limited to three attempts per wallet and ten per network address per minute; onion requests use the existing shared scaled bucket. The temporary address counter lasts 61 seconds in Redis.

The response contains provider_id, status, plain-English reasons and a dashboard link. A new wallet and endpoint pair returns 201; re-applying to the same pair updates it and returns 200. The wallet becomes the lowercase operator, with USDG payout mode. Fresh hardware evidence must pass the existing attestor, the published policy and sanctions screening when enabled. Development evidence is refused. Success records probation for NETWORK_PROBATION_DAYS (default 7); refusal records rejected. Both operator and payout wallets are screened.

GET /api/v1/network/hosts/ {provider_id} /status returns status, reasons, fresh attestation state, probation end and weight. Weight is the current network selector multiplier, rounded to two decimals: the highest across eligible model offers, not a traffic percentage. It is zero without eligible offers or fresh healthy evidence. With offer terms in the signed policy, the registry creates shadow offers for requested quote-bound models; fresh healthy probation hosts receive a reduced routing multiplier of up to 0.1 through the existing selector. Payouts to network hosts are not switched on yet. A policy refusal can coexist with a successful hardware check; inspect the reasons. The approved recipe, deploy/network/approved/tdx-qwen2.5-0.5b, uses Intel TDX in a supported confidential VM and SHA-256 sidecar bindings v2. These commit the source archive hash, engine image and served model ID alongside existing digests. Admission is automatic after the fresh quote, signed host policy v1 and operator/payout sanctions checks pass. Probation has a public record on Hosts. Legacy v1 evidence still verifies but lacks the fields this admission policy requires.

Signup needs no inference API key from the router. The host operator generates its own high-entropy key and configures its SHA-256 through the sidecar’s auth.keys flow. To supply it securely for later inference, send wallet-authenticated PUT /api/v1/network/hosts/ {provider_id} /credential with {provider_id, api_key} over HTTPS. The signed provider ID must match the path and recovered operator wallet. The router stores it only in the existing AES-GCM encrypted provider-key column, never returns it, and retains it on re-application. Rotate it by repeating this endpoint and updating the sidecar’s hash. Protect the operator’s key and router APP_SECRET.

Download and inspect join.mjs; compare its digest on the registration page. Add --api-key-file /path/to/sidecar.key to signup to supply the credential after a successful response, or use --credential-only PROVIDER_ID --key-file /path/to/operator.key --api-key-file /path/to/sidecar.key later. --api-key-env NAME reads an environment variable instead. The UTF-8 key is trimmed to 16–500 characters; POSIX files must not be group or world readable. --dry-run reads neither key, prints canonical bodies with api_key as <redacted> and uses <provider_id> until signup assigns the ID. A credential failure leaves signup in place; retry with credential-only mode using the same operator wallet.

The router retains the submitted host settings, optional contact, wallet and payout address, admission reasons and existing attestation records. Status is public by ID; contact and credentials are not returned there. Ordinary chat reads request text in router memory on every lane; the encrypted-chat adapter forwards ciphertext. Admission does not change that. Read the data inventory.

With NETWORK_HOSTS_ENABLED=true, scheduled attestation renews probation hosts against the current signed host policy without promoting them to live. A failed renewal blocks selection and appears on Hosts; revoked bindings reject the host and disable its offers. Admission records retention as attested, sourced to the checked host policy version, while jurisdiction, legal hold and training use remain undeclared. Fresh evidence and eligible offers are required for the attested lane. Rejection removes this disclosure profile. The router reads ordinary chat text in memory on every lane; the encrypted-chat adapter forwards ciphertext.

Host bonds

HostBond holds USDG work deposits. The router reads canonical bond and slash events on Robinhood Chain and exposes GET /api/v1/network/bonds. When enabled, host records include total and active base units, unbond requests, index freshness and slash history with explorer links. USDG uses six decimals.

Use keccak256(UTF-8 provider.id) as the contract’s bytes32 host id. Its on-chain operator must equal the provider’s registered wallet. A bond for one host cannot boost another host sharing that wallet. Unbond requests reduce the active amount immediately; a delisted or below-minimum host receives no bond boost.

NETWORK_BONDS_ENABLED=false and NETWORK_SLASHING_ENABLED=false by default. Set HOST_BOND_ADDRESS explicitly; it has no default. HOST_BOND_START_BLOCK defaults to 76855987. NETWORK_BOND_FULL_USDG=5000 sets the amount for the maximum 1.5× multiplier. Probation remains below the host’s unboosted graduated weight. Price, health, attestation and lane rules continue to apply.

Bond indexing: Switched on at anyroute.tech. HostBond on Robinhood Chain is 0x2921d34fd86d3323a5369a270a82814a74250518, with a minimum of 5,000 USDG. Slashing is not switched on at anyroute.tech yet. Payouts are not switched on at anyroute.tech yet.

NETWORK_BOND_FINALITY=finalized defaults to the finalized RPC tag; safe is also supported. The indexer also waits for CHAIN_CONFIRMATIONS. It resumes from its stored cursor and rewinds canonical state after a reorganization. A scan older than 120 seconds or incomplete backfill supplies no boost and stops slash work.

Register host-bond-indexer in the indexing worker’s WORKER_JOBS. Use an isolated host-slasher worker for proposals. With slashing off, that job logs a would-be proposal once per evidence commitment and sends nothing. Live slashing requires SLASHER_PRIVATE_KEY, whose address is checked against HostBond.slasher() at startup and before signing. The public API must never receive this signing key.

Verified quote-bound rejection against a published host policy produces a MeasurementDrift proposal for the current bond, without delisting. Invalid receipts from the pinned host feed produce review-only evidence because HostBond has no receipt-integrity reason. Evidence bundles store source digests and identifiers, not raw quotes, receipt envelopes or inference text; the commitment alone cannot prove fault. Review requires the source evidence.

The independent contract owner must approve each exact proposal. The worker executes only after 72 hours, while pending, undisputed and currently approved. It never approves a proposal. Signed transaction intents are saved before broadcast; retries reuse the same signed bytes. An intent that reverts or remains unresolved requires operator review rather than an automatic replacement.

Network statistics

GET /api/v1/network/stats is public and read-only when NETWORK_STATS_ENABLED is true (default false). Switched on at anyroute.tech. Token activity currently shows “No data yet”. No API key is required. The network page fetches statistics in the browser; its static export renders without the API.

The response has a data envelope: as_of, cache_seconds, hosts (total, probation, live and rejected), attested_hosts, capacity (distinct model count and IDs), tokens, bonds, interest and policy_version. Host status counts include network hosts only. Other statuses contribute to total but not the three named counts. Attested hosts require admission evidence, fresh successful hardware attestation and a fresh successful latest attempt. Models must have eligible offers and be among that host’s admitted model IDs; model count is not GPU memory or throughput.

tokens.public_lane.days_7 and days_30 contain decimal-string lower and upper_exclusive bounds in 100,000-token buckets. They count input plus output tokens from retained public-lane generation records on network hosts, excluding cache hits and future records. These coarse ranges are not differential privacy and cannot establish complete historical coverage after records expire. Null means no records in that window.

tokens.private_lanes is null. The existing DP releases at GET /api/v1/stats cover router-wide private-lane traffic, without a host dimension or 7/30-day history; they cannot be assigned to network hosts. Private-lane billing rows are excluded from this endpoint.

Bonds reuse the bond indexer’s total and active amounts as decimal strings in USDG base units (6 decimals), with freshness and indexed block. Null means indexing is not enabled. Stale index data remains identified as stale and is not shown as a current amount on the site. Interest reuses waitlist counts by role, region and readiness mentions; these are self-reports, not eligibility. Policy version comes from the published host policy, or is null if unavailable or disabled.

Snapshots are cached in router memory for at most 30 seconds. HTTP freshness uses the remaining snapshot lifetime. A failed refresh returns 503 with no stale snapshot. Disabled deployments return 404. No new database rows, Redis keys, request-body readers or address readers are added. The router still reads ordinary inference request text in memory on every lane; these statistics make no additional prompt-privacy claim.

Why this route?

Route explanations are switched on at anyroute.tech. ROUTE_EXPLAIN_ENABLED defaults to false. When an operator enables it, ordinary chat, text completions and embeddings include a compact JSON x-anyroute-route header once the serving provider is known, and an optional route object in the signed receipt. Streams carry it in the header and the final receipt event. With explanations enabled, stream headers wait for the existing routing attempt to find the first meaningful output; no second selection or provider call is made. The Harness shows “Why this route?” only when that receipt evidence is present.

The extension has its own v: 1; the enclosing canonical-JSON receipt remains v1 and the COSE receipt remains v2 where issued (embeddings use v1 only). Both signatures cover the entire optional claim. Older receipts omit it. With the flag off, no field or header is added and the existing signed bytes stay unchanged. Verifiers check the extension’s version, fixed vocabulary, counts and serving-provider identity; a valid signature establishes the router’s statement, not an independent replay of selection.

Fields are provider (the serving provider id), reason, eligible (candidate offers in this model’s final plan, after provider and parameter rules and fallback restrictions, before the attempt limit), skipped (counts grouped by a fixed reason), lane, parameters (required names from a fixed list), and network_host. A fallback adds error-class counts in fallback, including failures on earlier models. It names no failed or skipped provider. Unknown skip or error reasons become other.

Reasons distinguish price, latency or output-speed order, weighted choice, preferred performance, provider order, your own provider key, a single eligible provider, required parameter support and fallback. Weighted choice uses price, health, quality and hardware checks; eligible network hosts can also receive network weights. It does not always choose the cheapest or healthiest provider. Lane rules and disclosure checks still apply. Parameter names describe routing rules; where the existing selector accepts undeclared support, this summary does not establish support. Dual-verification receipts describe each partition; council receipts describe each member or judge’s provider call. A top-level explanation does not describe the whole panel.

Cache replies and encrypted-chat receipts have no new selection explanation. Provider ids outside the safe header vocabulary omit the extension. No prompts, answers, other callers’ records, provider URLs, keys, raw error messages or numeric weights are added. The summary is retained in the existing generation receipt fields under their current retention rules. Ordinary request text remains readable by the router in memory; this feature makes no new prompt-privacy claim.

Run a provider.

A model host runs the sidecar in front of its model server, inside a confidential VM. The sidecar hashes the weights at boot and refuses to start unless the digest is on its allow-list, binds its TLS key, receipt key and the image, compose and model digests into an Intel TDX quote, and signs a receipt for every response. The onboarding command writes all of that for you:

Terminal
git clone https://github.com/AnyRouteRH/AnyRoute.git && cd AnyRoute
bun sidecar/src/cli.ts init

It asks where the model runs (a Phala Cloud CPU or GPU confidential VM, or your own TDX host), measures the weights, and writes sidecar.yaml and a compose file in which every image, the weights and the sidecar source are pinned by hash. It makes the key the router will use: 32 random bytes in a file only you can read, with just their SHA-256 in the configuration. Then deploy, and check the running endpoint before you apply:

Terminal · without questions
bun sidecar/src/cli.ts init --yes --target phala-gpu \
  --weights ./my-model --hf-repo <owner/name> --hf-revision <40-hex commit> \
  --model-image <vllm image>@sha256:<digest> --id my-model

bun sidecar/src/cli.ts doctor --dir anyroute-provider --url https://<your endpoint>
bun sidecar/src/cli.ts apply  --dir anyroute-provider --url https://<your endpoint> --router https://<router> --submit

The doctor command reads /healthz and /attest the way a client would, and checks that the served model digest is your weights, that the quote commits to the keys and digests, that the certificate carries the attestation name and the attested key, and that a response carries a receipt signed by that key. It sends your router key only over a certificate the evidence proves belongs to the attested instance. It does not repeat Intel’s signature check on the quote; the router does. The apply command prints the exact body for POST /api/v1/providers/apply and, with --submit, files it. An operator reviews the application before anything is routed, and the router key goes to them separately unless you pass --include-key.

The sidecar attests the TDX virtual machine and does not collect GPU confidential-computing evidence. The data policy in your application is your own declaration and is shown as declared. Providers the router lists, and what it has verified about each, are on the providers page.

Show your attestation with a badge.

Any site can show an endpoint’s live status with one line. The script has no dependencies and sets no cookies. It reads the router’s public record from each visitor’s browser (the proof-time summary, the attestation record and the disclosure class, and for a model its attestation object and endpoints) and shows Attested only when every check passes: the record was read within five minutes of the visitor’s clock, the last verified attestation is inside the router’s freshness window, the record and the summary agree and name the same measurement, and the policy hash is well formed and, for a model, the same in the model’s attestation object, its endpoint and the record. Any failed check reads Unverified.

HTML · script badge
<script src="https://<router>/badge.js" data-endpoint="<provider id or model id>"
        data-theme="light" async></script>

data-endpoint takes a provider id or a model id. data-theme is light, dark or auto (follows the visitor’s colour scheme). The badge shows the status (Attested, Policy, Vendor-forwarded or Unverified), the first eight characters of the measurement and of the policy hash while attested, and the share of the last 7 days with a fresh attestation, never rounded up. It links to the endpoint’s registry entry. It does not verify the hardware quote itself: the router does that with its configured verifiers, and the badge says so.

For places that do not run scripts, the router serves the same status as an image. The image says what the router’s record says; nothing about it is checked in the viewer’s browser.

HTML and Markdown · image badge
<img src="https://<router>/api/v1/badge/<provider id>.svg" alt="Anyroute attestation status" height="48">

<!-- Markdown, for a README or a model card -->
[![Anyroute attestation status](https://<router>/api/v1/badge/<author>/<model>.svg)](https://<router>/registry/)

The registry lists every attested endpoint, and each entry (/registry/<provider id>/) shows its measurement versions, every check the router ran and the badge snippets for it.

Endpoint map

Endpoint Purpose
POST /api/v1/chat/completions Chat, tools and streaming (OpenAI/OpenRouter shape); X-Pay-With, X-Payment, X-Wallet-Auth headers
POST /api/v1/completions · /embeddings Legacy completions; embeddings (prepaid keys)
POST /api/v1/rerank · /v1/rerank Rerank documents against a query (Cohere and Jina shape: query, documents, top_n, return_documents), billed per token or per search unit, with a signed receipt; 404 model_not_found until a provider lists a rerank model
GET /api/v1/models · /models/:author/:slug/endpoints Catalog, prices, policies, quantization, attestation (best class, manifest reference, policy hash) and datacenter region; per-provider health and attested policy hash
GET /api/v1/generation?id=… · /generations Full generation record with receipt and anchor proof; your recent generations
POST · GET · PATCH · DELETE /api/v1/keys Create a self-custodial key (no auth), or budgeted sub-keys with rpm/tpm, model allowlists, guardrails and a tracing destination
GET /api/v1/key · /credits Current key; balance, held and total usage
POST /api/v1/credits/deposit-tx · /withdraw-request Unsigned wallet transactions to deposit, or a key-signed withdrawal request
GET /api/v1/credits/withdrawal-proof Merkle proof and calldata to finalize a withdrawal
POST /api/v1/paywith/open · /close · /revoke · GET /session · /statement Stock Token sessions and monthly statements
POST /api/v1/paywith/allowance/typed-data · /allowance · GET /charges · POST /charges/{id}/signature Wallet authorizations: a bounded allowance, or a signature per charge
POST /api/v1/byok · /teams Bring your own provider key; team roles
GET · POST · PATCH · DELETE /api/v1/routes Saved Routes: named routing policies you call as model "@route/<slug>", optionally pinned to the attested lane
GET · PUT · DELETE /api/v1/presets/:name · /versions · /diff · POST /rollback Presets: versioned saved routes with a system prompt, tools and response_format, called as model "@preset/<name>" or pinned as "@preset/<name>@<version>"
POST · GET /api/v1/characters · GET · PUT · DELETE /characters/:id · GET /characters/:id/export · /greetings · /usage Characters: Tavern cards (V1, V2 or V3, as JSON or PNG) kept public, unlisted or private (sealed on your device), called as model "@character/<id>"; public discovery by tag and text; export as JSON or PNG; per-day calls and cost for the creator
POST /api/v1/characters/:id/chat · /characters/group/next Chat with a character (chat completions plus greeting, regenerate, user_name and memory), on the attested lane when the model has one; pick the next speaker in a group chat
POST · GET · DELETE /api/v1/memory · GET · PUT · DELETE /memory/:id · POST /memory/search Memory ledger for characters: entries encrypted on your device, stored as ciphertext under a scope the router cannot tie to a character; embeddings only with embedding_opt_in
POST · GET · DELETE /api/v1/sessions · GET /sessions/current Agent Sessions: short-lived, budget-capped keys for agent runs
GET /api/v1/spend · /spend/alerts Spend Watch: totals, projection, breakdowns, key budgets and alert rules
GET /api/v1/disclosure/:providerId A provider’s documented retention, jurisdiction, legal hold and training use, each with a source and date, and the class it is served under now
GET /api/v1/models?variant=… Open-weights variants (mainstream, native_low_refusal, abliterated) with license, base model and weights source; restricted variants list only attested endpoints
POST /api/v1/creators/claims · /claims/{id}/verify Claim a model’s creator royalty by publishing a challenge in its Hugging Face repository
GET /api/v1/blind/keys Blind tokens, where enabled: the issuer keys per epoch and denomination, the challenge every token carries, and prices
POST /api/v1/blind/purchase Blind tokens, where enabled: buy tokens with credits by sending blinded messages; spend one with Authorization: PrivateToken on chat or embeddings
GET /api/v1/ohttp/keys · POST /api/v1/ohttp/gateway Oblivious HTTP, where enabled: the gateway key configuration (application/ohttp-keys), and the gateway that unwraps message/ohttp-req sent by a relay and returns message/ohttp-res (and, where chunked Oblivious HTTP is on, message/ohttp-chunked-req and message/ohttp-chunked-res, sent as it is produced)
GET /api/v1/ohttp/key-list · GET /api/v1/relays Oblivious HTTP, where enabled: the gateway key history signed with the receipt key, and the relays clients may use, by operator
GET /tlog/checkpoint · /tlog/tile/… Transparency log, where enabled: the newest checkpoint (a signed note with the witnesses’ cosignatures) and the C2SP tlog-tiles hash tiles and entry bundles
GET /api/v1/tlog · /tlog/proof · /tlog/consistency · POST /tlog/cosignatures Transparency log, where enabled: origin, log key, witnesses and quorum; an entry with its inclusion and consistency proofs; and where witnesses hand in cosignatures
GET /api/v1/tlog/rekor · /tlog/rekor/key · /tlog/rekor/{size} Transparency log with Rekor anchoring, where enabled: the anchoring key (ECDSA P-256), the anchored checkpoints, and one checkpoint’s Rekor entry with its log index, integrated time, inclusion proof and signed entry timestamp
GET /api/v1/holder $ANYR holders: balance, live tier (higher rate limits, lower fees), the tier ladder and free credits received
POST /api/v1/receipts/verify · GET /receipts/keys Verify a receipt (v1, or v2 with its chain head and Merkle path); signing keys (JWKS)
GET /api/v1/receipts/:id · /receipts/:id/proof A receipt by id, v2 beside v1 (?format=cose for the COSE bytes), with its privacy label beside the signed payload; the Merkle path to its hourly root, with anchored true only once that root is on chain
GET /api/v1/receipts/:id/privacy What we saw: who could read the prompt, who saw the address, how it was paid, what was kept and what hardware answered, in plain English and as fields, computed from the signed receipt. Public by id
GET /api/v1/host-anchors/proof/:leaf · POST /host-anchors/proof Where enabled: the Merkle path from a receipt a provider’s sidecar signed (by its leaf, or the receipt itself) to that host’s root, with the attestation reference and receipt key every leaf in the root was checked against; anchored true only once the root is on chain
GET /.well-known/anyroute-receipt-keys.json The same signing keys at a fixed path, for clients that verify receipts themselves
GET /api/v1/attestation/:providerId What the router has verified about a provider’s hardware attestation: status, verifiers, measurements, transparency-log and on-chain state, and what was not checked
GET /api/v1/badge/:id.svg Attestation badge image for a provider id or a model id (attested, policy, vendor-forwarded or unverified), with the measurement and policy hash while attested and the share of 7 days with a fresh attestation; ?theme=dark. An unknown id is Unverified with a 404
GET /api/v1/attestation/summary · /attestation/:providerId/history Proof-time: per attesting provider, the share of the last 24 hours and 7 days with a fresh attestation the router verified itself, measurement changes and the last failed check; and a provider’s recorded attestor, canary and probe events, newest first, paged by cursor. Failures are codes with fixed messages, never the provider’s own text. Kept for ATTESTATION_HISTORY_DAYS (30); 501 when it is 0
GET /api/v1/measurements/key · /measurements/bundles/:providerId Where enabled: the key that signs measurement bundles (compose hash, source commit and tarball hash, model and image digests, MRTD allow-list), and a provider’s bundles with the transparency-log entry the router verified for each
POST /mcp AnyRoute MCP: list_models, list_attested_models, chat (optionally on the attested lane), verify_provider, get_receipt and verify_receipt as tools for Claude, Cursor or any MCP client
POST /v1/messages · /messages/count_tokens Anthropic Messages API (also under /api/v1) for the Anthropic SDKs and Claude Code: x-api-key or Authorization: Bearer; tools, images and streaming; the lane in X-Anyroute-Lane or provider.lane; the receipt in the reply and in X-Receipt-Id
GET /ollama/api/tags · POST /ollama/api/chat · /generate · /embed Ollama API for Ollama clients (Open WebUI, Continue, the ollama libraries, LangChain): set the host to <router>/ollama and send the key as Authorization: Bearer; NDJSON streaming, tools, images, format and options; also /api/show, /api/version, /api/ps and /api/embeddings; the lane in X-Anyroute-Lane; the receipt in X-Receipt-Id and the closing line
POST /v1/responses · /api/v1/responses OpenAI Responses API for the OpenAI Agents SDK, the Codex CLI and other Responses clients: the chat route’s billing, lanes and signed receipts behind the Responses shape and event stream. Stateless: store must be false, there is no previous_response_id and no GET; function and custom tools only
POST · GET /api/v1/batches · GET /batches/:id · POST /batches/:id/cancel · GET /batches/:id/output · /batches/:id/errors Batch API (also under /v1): many chat or embeddings requests in one call, run in the background at 50% of the normal price within 24 hours; results as JSONL for 24 hours after the batch finishes (prepaid key)
POST /api/v1/rag · /v1/rag Answers from documents you send with the question, ranked in memory and stored nowhere: it embeds, ranks and answers through the embeddings and chat routes, and returns the sources and every call’s receipt (prepaid key; the lane and disclosure options of chat)
GET /api/v1/rankings · /providers · /status Usage rankings and creator payouts; the provider registry with each provider’s attestation status (attestation.status, tee, verifiers, last_verified_at); router configuration, including its onion address where there is one
GET /api/v1/status/slo · /status/incidents · /status/incidents.atom · .rss The status page (/status): per lane and API surface, availability over 1 hour, 24 hours, 7 days and 30 days, p50 and p95 latency, error counts, the SLO target (STATUS_SLO_PUBLIC 99.5%, STATUS_SLO_ATTESTED and STATUS_SLO_UNLINKABLE 99%), the 30-day error budget and 90 daily figures, cached for 30 seconds. The public lane is summed from public-lane request outcomes; the attested and unlinkable lanes use only the noisy hourly releases of /api/v1/stats (source: dp-noised). Incidents as JSON, Atom and RSS. The operator (ADMIN_TOKEN) opens incidents with POST /status/incidents, posts updates with POST /status/incidents/:id/updates and confirms or dismisses the suggestions the router records when a lane stays below target for five minutes (an hour on the private lanes)
GET /api/v1/skills · /skills/:id · /skills/:id/download · POST /skills/import · /skills/:id/install Secured Skills Hub for agent skills (a folder with SKILL.md, scripts and resources). Import from a git repository (https, at a ref, one folder) or an uploaded .tar.gz, .tar or .zip (SKILLS_MAX_BYTES, SKILLS_MAX_FILES). Each skill is normalised into one canonical tar and identified by its sha256, then scanned for data exfiltration (calls to hosts outside the allowlist, secrets or env values sent out, SSH keys, keychains, browser profiles, env dumps), prompt injection (instruction overrides, hidden or zero-width characters, instructions to send secrets), obfuscation, dangerous shell commands, untrusted package indexes and binaries. The report has a score, a level (trusted, caution or dangerous) and each finding with its rule, file, line and excerpt. Only trusted and caution skills download or install (SKILLS_DOWNLOAD_LEVELS); a dangerous or revoked one returns 403 with the report. A paid install debits the key's balance and credits the author 90% and the network fee 10% (SKILLS_FEE_BPS), once per account, with an Ed25519-signed receipt. The operator revokes with POST /skills/:id/revoke; SKILLS_SOURCES lists repositories and registry indexes the mirror job pulls. SkillRegistry records hashes and levels on chain. Scanned, not guaranteed: a clean scan is not proof a skill is safe
POST /api/v1/providers/apply · /creators/claim · /paymaster Provider onboarding; royalty claims; ERC-7677 gas sponsorship

Limits and guarantees.

Each call holds its worst-case cost before routing and settles the metered usage after, so a key never goes past its balance or budget. Keys have a per-minute request limit, and optional token-per-minute limits; n and best_of are capped at 16. If every provider fails, nothing is charged. A cancelled stream is billed only for what was generated.

The router stores key hashes, balances and receipt metadata. It never stores prompts or responses; the optional response cache is encrypted, per-workspace and expires, and a Batch API batch keeps its lines and answers sealed, outside the database, only until its results expire. Provider data policies are listed per provider.

Primary references

OpenRouter’s official API documentation describes the request shape Anyroute is compatible with. Robinhood Chain’s official documentation describes the network. Neither reference implies a partnership.

Network host payments

Payouts and fee buy-and-burn are not switched on at anyroute.tech yet. No payouts are being made.

When enabled, hosts are paid per token served only from receipts included in their confirmed per-host roots. The planned weekly USDG payments deduct the network fee: the planned 5% network fee buys and burns $ANYR. Both accrual and the fee keeper are switched off by default; runtime availability is reported by network.payouts_open in GET /api/v1/status.

NETWORK_PAYOUTS_ENABLED=false requires sanctions screening, host anchoring and a settlement worker signer when enabled. NETWORK_FEE_BURN_ENABLED=false requires a keeper signer, the reviewed buyback oracle, staking configuration and NETWORK_FEE_BURN_ADDRESS. The dedicated executor must be deployed, authorized by the existing adapter and funded with fee USDG; nothing funds it automatically. NETWORK_FEE_BPS=500 is bounded to 0–2000. The configured percentage is shown at runtime.

The keeper uses the same oracle minimum and optional independent TWAP as staking buybacks, plus a daily limit. It transfers acquired $ANYR to the dead address because AnyrToken has no holder burn function; total supply is unchanged. Aggregate fees and recent swap and burn transactions are public through GET /api/v1/network/burns. Sub-unit conversion dust remains unswapped. Periods above the per-run or remaining daily limit wait; unresolved or reverted payouts require operator reconciliation. Ordinary chat reads request text in router memory on every lane; the encrypted-chat adapter forwards ciphertext.