English documentation · runtime 7.24.4 · SDK 2.6.7. Content is maintained with runtime development; see each guide's scope and review date.
Execution Modes (Local, Cloud, Hybrid, Council, Arena)
Status: Active
Scope: current-state
Last reviewed: 2026-09-12
Owner: ax-code runtime
AX Code can place work on local inference, hosted/CLI providers, or both (hybrid), and can fan out high-stakes work across multiple connected providers (council review and arena best-of-N). This page documents shipped behavior for those modes.
Source of Truth
When behavior changes, verify against:
packages/ax-code/src/mode/— pure policy, hybrid, council aggregation, arena ranking, debate, budget, memory, worktree policy, implement-arena scoringpackages/ax-code/src/tool/council.ts— multi-provider council toolpackages/ax-code/src/tool/arena.tsandarena-implement.ts— plan and implement arenapackages/ax-code/src/session/prompt/prompt-routing.ts— hybrid placement whenmodes.defaultishybridpackages/ax-code/src/config/schema-impl.ts—modesconfig schemapackages/ax-code/src/command/template/{council,arena}.txt—/counciland/arenaare in the default slash menu
Work mode selector (Agent | Council | Arena)
TUI and Desktop expose a work mode control for multi-model routing. Default is Agent.
| UI selection | Free-text send becomes |
|---|---|
| Agent (default) | Normal single-agent prompt |
| Council | /council {your message} multi-provider review |
| Arena | /arena {your message} multi-model best-of-N |
- TUI default chrome: the footer does not show an Agent chip. Run mode and Sandbox stay.
/work-mode(palette Choose work mode) opens an explicit picker: Agent, Council, and Arena, with cost/semantics on each row. Unavailable ensemble rows are disabled with the reason. - Armed ensemble chrome: after you pick Council or Arena, a chip appears (
Council · 2, or hollowArena (off)if it later becomes unavailable). Click the chip to return to Agent. New chats reset to Agent. - Availability: a mode is available when it is enabled in config, at least two connected providers have a selectable model, and the configured member cap is not 1. Chips and picker rows update live as providers connect or disconnect.
- Pre-submit hint (TUI): a one-line hint above the prompt appears when council/arena is blocked or still checking, and on first use of an available mode (e.g.
Council mode · up to 2 reviewers · advisory · approval on first use). After a successful submit in that mode the chip remains the status and the hint is hidden. Submitting while the selected mode is unavailable is blocked with the reason — the draft is kept and the prompt is never silently downgraded to a single-model run. - Desktop: composer toolbar chip (next to Manual/Autonomous).
- Explicit
/counciland/arenaare never rewritten and remain the one-shot entry points. - Specialist agents (architect, security, …) stay on the separate agent picker.
Placement modes at a glance
| Mode | What it does | Mutates workspace? | Default |
|---|---|---|---|
| local | Prefer AX Engine (or configured local provider) | Yes (single agent) | When you pin local / hybrid places local |
| cloud | Prefer hosted or CLI frontier providers | Yes (single agent) | When local unavailable |
| hybrid | Policy chooses local vs cloud from availability + complexity + privacy | Yes (single path) | Set modes.default: "hybrid" |
| council | Fan out structured review/design; classify consensus / majority / minority / singleton | No (advisory) | Tool + /council or Work mode = Council |
| arena | Multi-model plan comparison or worktree implement best-of-N | Plan: no. Implement: only in worktrees | Opt-in (modes.arena.enabled) + Work mode = Arena |
Keyword specialist routing and complexity tiering (see Auto-Route) are orthogonal to hybrid placement and ensemble modes.
Model effort / thinking level (Fast, Balanced, Deep, Max) is also orthogonal — it is a per-model reasoning budget, not a work mode. See Model Effort.
Configuration
In ax-code.json:
{
"modes": {
"default": "hybrid",
"hybrid": {
"preferLocalWhenAvailable": true,
"escalateOnHighComplexity": true,
"localProviderID": "ax-engine"
},
"council": {
"enabled": true,
"maxMembers": 3,
"timeoutMs": 180000,
"debateRounds": 0
},
"arena": {
"enabled": true,
"maxContestants": 3,
"strategy": "verify_first"
},
"budget": {
"maxEstimatedUsd": 0.5,
"estimatedUsdPerMember": 0.05
}
}
}
| Field | Meaning |
|---|---|
modes.default |
local | cloud | hybrid | arena | council. Unset: hybrid when local fits the policy signals, else cloud for single-path defaults. |
modes.hybrid.* |
Local preference, high-complexity escalate to cloud, local provider id |
modes.council.* |
Enable, member cap, timeout, reasoning-model timeout scale, per-member timeout overrides, debate rounds, opt-in chairman / adaptive fan-out (both default off) |
modes.arena.enabled |
Must be true for the arena tool (default off). Mid-session edits are picked up on the next tool call (Config.getFresh). Or pass enableIfDisabled: true on the arena tool. |
modes.arena.strategy |
verify_first (recommended for implement), diversity, or hybrid_score |
modes.arena.reasoningTimeoutScale |
Timeout multiplier for contestants whose model declares reasoning capability (falls back to modes.council.reasoningTimeoutScale, then 3) |
modes.arena.memberTimeoutMs |
Absolute per-contestant timeout overrides keyed by "providerID" or "providerID/modelID" (falls back to modes.council.memberTimeoutMs) |
modes.arena.judge |
Blinded rubric judge for plan mode (default: true) |
modes.ensembleLedger |
Local JSONL call ledger for ensemble generations (default: true; SHA-256 prompt hashes only, no bodies, no egress) |
modes.budget.* |
Fail-closed cap on estimated USD for ensemble fan-out |
Hybrid placement
When modes.default is hybrid and the user/agent did not pin a model:
- If local provider (default
ax-engine) has a selectable model → prefer local for low/medium complexity. - If complexity is high and
escalateOnHighComplexityis true → cloud. - If privacy requires local and local is available → local.
- If local is unavailable → cloud.
Complexity still uses the existing small/fast model path for low messages when auto-route complexity routing is enabled (Auto-Route). Hybrid does not replace keyword specialist routing.
Local models and memory guidance: AX Engine Model Selection. Providers list: Supported Providers.
Council (consensus mode)
Tool: council
Slash: /council <question>
- Selects diverse connected providers (family diversity — under an unrecognized multi-model gateway the family falls back to the model id; soft bias from outcome memory).
- Fans out a structured review or design prompt in parallel.
- Aggregates issues into consensus (unanimous among successful members at quorum — at least
max(2, ⌈2/3 × attempted⌉)successes), strict majority (more than half of attempted members), minority (at least two), and singleton tiers. Findings disclose support against attempted members (2/6), and low-coverage reports state that consensus labels require quorum. - Optional debate rounds: anonymous (Chatham House) synthesis shared between rounds; no brand attribution. Debate is capped at three rounds and stops early on convergence.
- Returns an advisory markdown report. Does not edit files.
Needs at least two resolved members to run at all — fewer short-circuits with an “insufficient members” preflight before any approval prompt or model call (explicit same-gateway model pairs count as two). Meaningful consensus tiers still need at least two successful members; otherwise the report is marked incomplete.
Evidence admission. Members receive only the supplied question and context. They do not inherit the calling session or read files from paths in the brief. Include the requirements, relevant diff, required original snippets, and verification evidence needed for the stated review scope.
The optional context is accepted verbatim up to 24,000 UTF-16 code units. Larger context returns
context_rejected before member inference; AX Code never silently shortens it. Split the review into explicitly
scoped requests or remove optional background while retaining required evidence.
Before each round, AX Code checks the whole prompt against a 128,000-byte local cap and every resolved member’s known input/context limits, reserving the requested output and fallback instruction plus 2,048 tokens for schema and framing. Input size uses a deliberately conservative UTF-8-byte estimate. It can reject prompts that would fit; it is neither an exact tokenizer count nor a guarantee about provider serialization. Unknown model limits are disclosed and remain subject to the local caps. If a debate round cannot fit, the result is incomplete and retains the last completed round’s report.
contextAdmission records the local context-length gate, supplied size, and a content digest; promptBudget checks
the complete request separately. Both must pass before inference. These fields are independent of successfulMembers
and do not establish semantic completeness, source freshness, or guaranteed review quality.
Timeouts. Each member runs under modes.council.timeoutMs (default 180000 ms); models that declare
reasoning capability get modes.council.reasoningTimeoutScale times that budget (default 3, so 540000 ms).
To give one known-slow member more time without inflating everyone else’s wait, set an absolute
modes.council.memberTimeoutMs override keyed by "providerID" or "providerID/modelID" — the exact
model key wins over the provider-wide key, and either wins over the base/scale computation:
{
"modes": {
"council": {
"memberTimeoutMs": { "deepseek/deepseek-v4-pro": 900000 }
}
}
}
ax-code.json is a protected config file — agents must ask the user to change it.
Optional lanes (default off). modes.council.chairman: true appends one blinded chairman synthesis call after aggregation (and any debate rounds): the chairman receives anonymized findings only (tiers and support counts, never member identities) and returns a verdict, recommended actions, and dissent notes. Deterministic tiering remains the primary output; chairman failure is disclosed and non-fatal. modes.council.adaptive: true starts the fan-out with two members and expands one at a time up to maxMembers while round-1 coverage is below quorum or dissent is material; expansion triggers are harness-tunable constants.
When to use
- Architecture / security / design trade-offs
- High-stakes code review where multi-model agreement raises confidence
- User asks for a multi-model or “second opinion” review
Agent workflow (important)
Call council early once the relevant evidence is available with an explicitly scoped context brief.
Avoid broad multi-explore digs unrelated to that review; gather the required original evidence before asking members for code findings.
If the user asked for council/arena, task_parallel is rejected until the ensemble tool has been the intended primary action.
When not to
- Trivial questions (latency/cost)
- Privacy-sensitive code that must not leave local inference
- Only one provider connected
Arena (best-of-N)
Tool: arena
Slash: /arena <task>
Requires: modes.arena.enabled: true and ≥2 distinct selectable models on connected providers (including a shared gateway)
Evidence admission (shared with council). The optional context is accepted verbatim up to 24,000 UTF-16 code units. Larger context returns context_rejected before any approval prompt, worktree creation, or model call — AX Code never silently shortens it. Split the task into explicitly scoped requests or reduce optional background while retaining required evidence. The approval prompt itself fires only after every no-op preflight has passed (disabled, context admission, implement git preflight, budget, member resolution).
mode: "plan" (default)
- Each contestant proposes an approach, steps, risks, and a calibrated self-assessed risk score (no workspace writes).
- With ≥2 successful proposals, one blinded rubric judge call (the first resolved member; identities stripped, order randomized) scores each proposal on requirement coverage, feasibility, verification plan, and risk evidence (0–10 each, ties allowed). The rubric total (0–40) is the primary ranking signal; self-assessed risk stays display-only. Judge failure or
modes.arena.judge: falsefalls back to self-assessed scoring with a disclosure note. - Ranked with verification tier first, then judge/risk score, then patch-fingerprint diversity (never pure popularity). Plan rankings are advisory and are not execution verification.
- Advisory only.
mode: "implement"
- Requires a primary git worktree with at least one commit and no uncommitted changes, records its exact base commit, and creates a git worktree per contestant from that commit.
- Runs an implement agent in each worktree.
- Snapshots every contestant’s tracked and untracked changes into a durable branch commit, including commits created by the agent itself.
- Runs detected project verification commands (typecheck / test / lint) only after a non-empty patch is captured.
- Ranks with verify-first by default: only completed, non-empty patches that pass verification can win; among passers prefer lower risk and diverse patches.
- Does not auto-merge. Report includes worktree paths, branches, and commit ranges for you to inspect, merge, or cherry-pick.
Implement arena requires a git project.
Ranking rule (research-aligned)
For code candidates: verification first, diversity second, popularity never alone.
Naive majority vote on similar wrong patches is an anti-pattern (popularity trap).
Slash commands
| Command | Purpose |
|---|---|
/council … |
Drive multi-provider advisory review |
/arena … |
Drive plan or implement best-of-N |
Safety and cost
- Sandbox / autonomous still apply to single-agent work (Sandbox, Autonomous).
- Council and plan-arena do not write files.
- Implement arena writers are isolated in worktrees; a dirty primary worktree is rejected so uncommitted input cannot be silently omitted.
- Ensemble fan-out multiplies provider egress and cost; use
modes.budgetand keepmaxMembers/maxContestantssmall. Budget estimates price the worst case: council2 × (debateRounds + 1)calls per member (schema fallback + retry), plan arena 2 per contestant plus one flat judge call, implement arena a documented 12-call per-trajectory estimate. - A local-only ensemble call ledger (
ensemble-calls.jsonlin the global state directory, 2 MB cap) records per-generation outcomes with SHA-256 prompt hashes — never prompt bodies, never credentials, no egress. Disable withmodes.ensembleLedger: false. - Multi-model agreement is evidence, not proof — run tests before shipping.
Related
- Auto-Route — specialist keywords + complexity tier
- Supported Providers — cloud, CLI, AX Engine
- AX Engine Model Selection — local model choice