English documentation · runtime 7.24.4 · SDK 2.6.7. Content is maintained with runtime development; see each guide's scope and review date.
Supported Providers and Models
Status: Active Scope: current-state Last reviewed: 2026-09-19 Owner: ax-code runtime
This page lists the provider presets AX Code exposes in the default setup flows. The source of truth is the runtime provider allowlist in
packages/ax-code/src/provider/default-setup-providers.ts, the bundled model snapshot in
packages/ax-code/src/provider/models-snapshot.json, local runtime presets in
packages/ax-code/src/provider/local-runtime.ts, and the AX Engine definitions in
packages/ax-code/src/provider/ax-engine/constants.ts.
Use /connect in the terminal UI or ax-code providers login <provider-id> for interactive setup. Headless and CI environments can also provide the listed environment variables.
Managing connected providers
The session view keeps a top bar with the session ID (click to copy), title, and live status, followed by a Providers entry showing the connected count; click it or manage to open the provider manager, which is also available as the /providers command. From there you can select a model, disable, or disconnect each provider, or jump to the full /connect flow.
- Replace a key: select the provider in
/connectand choose Replace key, or re-runax-code providers login <provider-id>. - Disable temporarily: choose Disable in
/providersor/connect, or runax-code providers disable <provider-id>. This adds the provider todisabled_providersin the global config and keeps the saved credentials. Disabled providers stay visible in/providersand under Disabled in/connect, and can be turned back on with Enable orax-code providers enable <provider-id>. - Disconnect permanently: choose Disconnect in
/providersor/connect, or runax-code providers logout <provider-id>— this deletes the stored credential fromauth.json.
Hosted model catalogs are bundled with each AX Code release and filtered for usable coding-agent capabilities. Run
ax-code models <provider-id> for the authoritative model IDs in your installed release; the raw registry and copied
web lists can contain models AX Code hides because they lack text output or tool calling. A provider preset does not
imply that every model is free or available on every account.
Runtime choices
Selecting a model without a version
Both the TUI and ax-code run accept these case-insensitive --model family names:
| Family | Model |
|---|---|
deepseek |
deepseek-flash |
glm |
glm-5.3-flash |
qwen |
qwen3.8-flash |
For example, ax-code --model glm opens the TUI with GLM Flash, and
ax-code run --model qwen -- "Review this change" requests Qwen Flash.
Put the prompt after -- (or pass --prompt / --prompt-file); --file attaches
files and is not a prompt file. ax-code models lists usable provider/model IDs.
Family names resolve through the native provider or a connected gateway serving
the same model. You can also name the provider: --model my-gateway/glm.
Use a full ID such as my-gateway/glm-5.3 to request a specific version or tier.
Omitting --model preserves configured, agent, session, and recent model choices.
When no usable choice exists, the shared fallback order starts with DeepSeek Flash,
GLM Flash, then Qwen Flash across connected providers. Provider setup has its own
defaults: Alibaba Coding Plan uses qwen3-coder-plus; Alibaba Token Plan uses
qwen3.8-flash.
Connecting a runtime
The choices in /connect are ordered API Cloud Provider, CLI Provider,
AX-Engine runtime, Local LLM runtime, Private GPU cloud, and AX Trust.
AX-Engine runtime opens local model selection and runtime management directly.
Local LLM runtime lists these choices in order:
| Menu option | Provider id | Default endpoint | Host environment variable |
|---|---|---|---|
| Ollama | ollama |
http://localhost:11434/v1 |
OLLAMA_HOST |
| LMStudio | lmstudio |
http://localhost:1234/v1 |
LMSTUDIO_HOST |
| MTPLX | mtplx |
http://localhost:8000/v1 |
MTPLX_HOST |
| oMLX | omlx |
http://localhost:8000/v1 |
OMLX_HOST |
| AX-Studio | ax-studio |
http://localhost:18080/v1 |
AX_STUDIO_HOST |
| Others | local-llm |
Enter your OpenAI-compatible endpoint | LOCAL_LLM_HOST |
MTPLX and oMLX presets are available in the source checkout; v7.19.3 packaged binaries do not include them. See MTPLX and oMLX setup for tool-capability configuration, optional authentication, and the shared default-port caveat.
Start your local server, select its menu entry, and confirm or edit the endpoint.
AX Code adds /v1 if omitted. Others saves one configurable endpoint under
local-llm; select it again to change the endpoint. LMStudio and Others use the
existing local setup flow without an API-token prompt. LM Studio documents its
model-list endpoint.
Setup choices respect enabled_providers and disabled_providers. Merely opening
the picker does not activate or probe these local servers. Discovered external
local models retain conservative capability defaults; models must support tool
calling to be selectable for coding-agent workflows.
Cloud API Providers
These providers call hosted APIs or hosted account-plan endpoints.
| Provider id | Display name | Credential environment variables |
|---|---|---|
google |
GOOGLE_API_KEY, GOOGLE_GENERATIVE_AI_API_KEY, GEMINI_API_KEY |
|
deepseek |
DeepSeek | DEEPSEEK_API_KEY |
meta |
Meta (Muse Spark) | MODEL_API_KEY, META_MODEL_API_KEY |
groq |
GroqCloud | GROQ_API_KEY |
openrouter |
OpenRouter | OPENROUTER_API_KEY |
huggingface |
Hugging Face | HF_TOKEN |
unorouter |
UnoRouter | UNOROUTER_API_KEY |
alibaba-coding-plan |
Alibaba Coding Plan | ALIBABA_CODING_PLAN_INTL_API_KEY, ALIBABA_CODING_PLAN_API_KEY |
alibaba-coding-plan-cn |
Alibaba Coding Plan (China) | ALIBABA_CODING_PLAN_CN_API_KEY, ALIBABA_CODING_PLAN_API_KEY |
alibaba-token-plan |
Alibaba Token Plan | ALIBABA_TOKEN_PLAN_INTL_API_KEY, ALIBABA_TOKEN_PLAN_API_KEY |
alibaba-token-plan-cn |
Alibaba Token Plan (China) | ALIBABA_TOKEN_PLAN_CN_API_KEY, ALIBABA_TOKEN_PLAN_API_KEY |
github-copilot |
GitHub Copilot | GITHUB_TOKEN |
zai |
Z.AI | ZHIPU_API_KEY |
zai-coding-plan |
Z.AI Coding Plan | ZHIPU_API_KEY |
minimax-coding-plan |
MiniMax Token Plan | MINIMAX_TOKEN_PLAN_API_KEY, MINIMAX_API_KEY |
minimax-cn-coding-plan |
MiniMax Token Plan (China) | MINIMAX_TOKEN_PLAN_CN_API_KEY, MINIMAX_API_KEY |
MiniMax renamed Coding Plan to Token Plan. The provider IDs stay minimax-coding-plan /
minimax-cn-coding-plan to match models.dev and OpenCode. Use a Token Plan key
(sk-cp-…) from platform.minimax.io
(international) or platform.minimaxi.com (China).
huggingfaceis the hosted Serverless Inference Providers router (https://router.huggingface.co/v1). It is unrelated to the local Hugging Face snapshot cache that AX Engine uses to store downloaded local models — connectinghuggingfacenever starts or requires the local engine.
For a no-cost first run, see Free-Tier API Quickstart. It distinguishes a provider’s free offering from AX Code’s compatibility and explains why an external free-API catalog is not itself a support matrix.
Grok is supported only through the grok-build-cli provider listed below. AX Code no longer includes a direct xai cloud API provider; existing Grok API credentials do not create a connection.
github-copilot is the GitHub Copilot account bridge. It is not the separate github-models API listed by some
free-API directories.
OpenAI-compatible and Anthropic-compatible gateways are also supported through custom provider configuration. See Custom and Gateway Providers.
Private GPU cloud
These providers appear in /connect under Private GPU cloud, after Local LLM runtime.
Choose Custom provider to connect your own OpenAI-compatible GPU endpoint.
Enter its URL and API key; AX Code discovers models from /v1/models and stores
the key in encrypted auth storage. The saved connection stays in Private GPU
cloud. Select it again to choose a model, replace the endpoint, or disconnect.
This entry stores one configurable endpoint under custom-private-gpu.
Catalog (API key)
OpenCode-style hosted GPU catalogs. Models come from the bundled models.dev snapshot. Connect with an API key.
| Provider id | Display name | Credential environment variables |
|---|---|---|
nebius |
Nebius Token Factory | NEBIUS_API_KEY |
fireworks-ai |
Fireworks AI | FIREWORKS_API_KEY |
togetherai |
Together AI | TOGETHER_API_KEY |
baseten |
Baseten | BASETEN_API_KEY |
nvidia |
NVIDIA NIM | NVIDIA_API_KEY |
deepinfra |
Deep Infra | DEEPINFRA_API_KEY |
The hosted Hugging Face router (huggingface / HF_TOKEN) stays in Cloud API Providers. Dedicated Hugging Face Inference Endpoints are listed below.
Dedicated (URL + token)
PAI-style dedicated GPU endpoints. Paste the OpenAI-compatible URL and token; AX Code calls GET …/models and uses the deployed model IDs. These are not Alibaba Coding Plan / Token Plan (DashScope) providers.
| Provider id | Display name | Credential environment variables |
|---|---|---|
alibaba-pai |
Alibaba PAI-EAS | ALIBABA_PAI_API_KEY, ALIBABA_PAI_BASE_URL |
runpod |
RunPod | RUNPOD_API_KEY, RUNPOD_BASE_URL |
huggingface-endpoints |
Hugging Face Endpoints | HF_ENDPOINTS_TOKEN, HF_ENDPOINTS_BASE_URL |
sagemaker |
Amazon SageMaker | SAGEMAKER_API_KEY, SAGEMAKER_BASE_URL |
volcengine-ark |
Volcengine Ark | ARK_API_KEY, ARK_BASE_URL |
modelarts |
Huawei ModelArts | MODELARTS_API_KEY, MODELARTS_BASE_URL |
tencent-ti |
Tencent TI | TENCENT_TI_API_KEY, TENCENT_TI_BASE_URL |
custom-private-gpu |
Custom provider | CUSTOM_PRIVATE_GPU_API_KEY, CUSTOM_PRIVATE_GPU_BASE_URL |
sagemaker is for an OpenAI-compatible URL in front of SageMaker (vLLM / TGI / API Gateway). It does not sign AWS SigV4.
AX Trust
AX Trust is the last category in /connect, after Private GPU cloud. Choose
Connect AX Trust, enter the gateway base URL including /v1, and supply
the client API key. AX Code discovers the gateway’s models and stores the key
in encrypted auth storage. Saved connections offer model selection, endpoint
updates, model refresh, and deletion.
Reconnecting an existing URL preserves its provider ID, models, and saved key
when the token is left blank. The provider remains in AX Trust even with a
custom name. Legacy ax-trust and ax-trust-* IDs also appear in this category.
AX Trust continues to own gateway policy, approvals, and audit.
AX Code sends a session affinity header on these connections by default.
CLI Providers
CLI providers reuse a local vendor CLI and its login/session instead of storing a hosted API key in AX Code.
| Provider id | Display name | Required local command | Supported model id |
|---|---|---|---|
claude-code |
Anthropic (Claude Code) | claude |
claude-code |
codex-cli |
OpenAI (Codex CLI) | codex |
codex-cli |
grok-build-cli |
Grok Build CLI | grok |
grok-build-cli |
muse-cli |
Muse Code CLI | muse |
muse-cli |
Run the vendor CLI login first when required, then run ax-code providers login <provider-id>. AX Code probes the CLI command and stores a local marker credential after the probe succeeds.
For Muse Code, install the local muse binary, run muse login (or set META_API_KEY), then ax-code providers login muse-cli. AX Code reuses the Muse CLI session (~/.config/muse) rather than storing a hosted Meta API key. The hosted meta API provider remains a separate connection.
Hosted MiniMax Token Plan remains available as minimax-coding-plan / minimax-cn-coding-plan. There is no bundled MiniMax Code CLI or Kimi Code CLI provider.
AX Engine Local Provider
ax-engine is the built-in local inference provider. It is available only on eligible Apple Silicon Macs and offers only AutomatosX Tiel Coder and Cyber-Tiel Coder 35B A3B MXFP4 MTP development packs. Native capabilities are checked when the model starts.
| Provider id | Model id | Selection | Context | Output |
|---|---|---|---|---|
ax-engine |
tiel-coder-35b-axq-mxfp4 |
Tiel Coder 35B A3B MXFP4 MTP (default) | 65,536 | 8,192 |
ax-engine |
cyber-tiel-coder-35b-axq-mxfp4 |
Cyber-Tiel Coder 35B A3B MXFP4 MTP | 65,536 | 8,192 |
Qwen3.8 27B, Ornith, Qwen3-Coder-Next and all other repositories are excluded from managed selection. Historical records remain readable for status and cleanup.
The default local model is tiel-coder-35b-axq-mxfp4. See AX Engine Model Selection for the exact selected repositories, memory, and disk guidance.
For the existing separately configured loopback attach interface, /v1/models is authoritative: AX Code discovers the live model IDs, context/output limits, modalities, and structured tool-call support. A model that does not advertise structured tool calling is not used for coding-agent requests.
AX Engine uses the compact core tool profile by default (bash, file discovery/read/edit/write, and skills). Set provider.ax-engine.options.toolProfile to full only for a custom deployment with enough context for the complete tool registry.
Installing the engine
Local inference needs AX Engine 7.5.7 or later. Apple Silicon Mac installers already include a self-contained AX Engine sidecar (engine/<version>/ next to the CLI runtime), so a clean Mac does not need Homebrew.
AX Code then resolves the binary in this order:
provider.ax-engine.options.binaryPathinax-code.json(requires a verified version of at least 7.5.7)- the
AX_ENGINE_BINenvironment variable (requires a verified version of at least 7.5.7) ax-engineon yourPATH(requires a verified version of at least 7.5.7)- an AX Code-managed overlay install (
ax-code providers ax-engine install; also requires at least 7.5.7) - the sidecar bundled in the current Mac runtime
It first checks --version and falls back to install.version from ax-engine doctor --json. Doctor’s exit code reports host readiness (Metal toolchain, MLX files), so a non-zero exit still counts when that JSON contains a version. AX Code owns server startup and normally launches ax-engine serve on 127.0.0.1:31418. The optional Homebrew formula remains an alternative for users who want a brew-owned engine; it is not required.
Managed Tiel launches use the catalog window (65,536 context tokens and 8,192 output tokens), which is larger than ax-engine serve’s built-in 16,384-token default so the agent prompt fits. Shrink either budget with provider.ax-engine.options.contextTokens and provider.ax-engine.options.maxOutputTokens (outputTokens is the same output knob). The environment names are AX_ENGINE_CONTEXT_TOKENS and AX_ENGINE_MAX_OUTPUT_TOKENS (AX_ENGINE_OUTPUT_TOKENS is an alias). An override can only narrow the catalog ceiling: a larger or invalid value is reported and ignored, and a Tiel or Qwen3.8 window that is not a multiple of 1,024 is snapped down so prefix-cache blocks stay 1,024 tokens. The same numbers are what the running server and the model limit use. A live /v1/models card can shrink that window further, and it cannot raise it. Configured options win, and AX Code warns when an environment value is ignored. The same warning is emitted when binaryPath ignores AX_ENGINE_BIN, or when connectionMode / baseURL ignores AX_ENGINE_HOST. A smaller window can be too small for the full agent prompt; AX Code then refuses the model instead of sending a prompt the server cannot hold.
Tiel and Cyber-Tiel advertise structured tool calling. A live /v1/models card still overrides that advertisement once the server is running. mtpPolicy: disabled turns MTP off even when the binary version cannot be read; auto and required still need AX Engine 7.4.0 or newer.
A PATH or overlay install that is missing libmlx.dylib / mlx.metallib is rejected. The bundled and overlay archives are the self-contained 7.5.7 macOS payload.
Installing the engine does not download a model. Pick and download a model afterward from the Desktop Models page or with ax-code providers ax-engine prepare. A complete compatible base snapshot already in the Hugging Face cache is accepted for direct decode; prepare --download uses the catalog’s preferred MTP package when available.
Use /connect -> AX-Engine runtime -> Select a model. AX Code configures
and starts the local runtime automatically when the selected model is needed;
no URL or API key is required. View status, Stop local runtime, and
Disable manage the local runtime. Selecting a model also replaces legacy
manual attachment settings with managed local setup.
The engine ships for Apple Silicon macOS only. On other hosts, use a hosted provider or an OpenAI-compatible provider gateway; AX Code servers are local-only.