Rumrunner

The three lanes

One endpoint, three places a request can go. The router is a rule list with a paper trail: every decision is written to a log on your machine with the inputs that made it.

Local (“ashore”) — a model running on this Mac, via the supervised engine. Free, and the default. With no key stored this is the only lane there is.
Nearshore — hosted open-weight models on your Together, Fireworks or Groq key. Fractions of a cent a request.
Frontier (“offshore”) — Anthropic, OpenAI or DeepSeek, on your own key, billed by them at their price. Real money, used only when it earns it.

How a lane is chosen

In this order; the first rule that matches wins.

  1. Fences. If the request's context touches a path you have fenced in rules.toml, it goes local. Absolute — nothing below can override it, and a fenced request goes local even when the local engine is down, and fails there.
  2. Per-app rule. Each tool can be set to auto, local_only or local_near in rules.toml, keyed by the X-RR-App header when the tool sends one and its User-Agent otherwise. A lane pinned by the request itself is honoured too, and a local pin means local: it is never escalated.
  3. Capability. More context than the local model's window, or a tool-calling request the local model is not marked as able to serve, escalates.
  4. Shape. Shipped defaults, all adjustable in config.toml: a prompt over 12k tokens goes nearshore; agentic markers (tool calls, or more than two system or tool messages) go nearshore; both together over 60k context go frontier. Everything else stays local.

Escalation is stepwise — local, then nearshore, then frontier — and skipping a lane needs a named reason. A request with no nearshore key stored cannot take the nearshore lane, whatever the rules say about it.

Sessions and escalation

Send an X-RR-Session header (Claude Code's x-claude-code-session-id is accepted too) and the router remembers a latched lane and the hosted targets that refused the session's context for its size. After you truncate, compact or reset the context, send a new id: the old one's memory describes a context that no longer exists.

An escalation lane — a local model judging whether the local answer is stuck, and handing the session up when it is — exists and is off by default. [escalation] enabled = true in config.toml turns it on; it costs a few seconds a turn and will spend on your nearshore key when it escalates.

Reading a decision

rr why

prints the last decisions as lane, model and reasons in plain English — for instance that a prompt exceeded the local threshold, or that a fence applied. The same rows are the log at ~/Library/Application Support/Rumrunner/decisions.jsonl, one line per request, which you can open and read.