Model Selection
Hosted providers (Gemini, Codex, Claude, Antigravity) auto-select a sensible default model with automatic fallback on provider-specific conditions — quota/rate-limit exhaustion, and for Antigravity also a rejected shipped model slug (see the table below). Ollama runs locally and uses exactly the model you pull; it never substitutes another. Most users should never override the model parameter; the defaults are tuned for quality.
Defaults & Fallbacks
| Provider | Default | Fallback | Trigger |
|---|---|---|---|
| Gemini | gemini-3.1-pro-preview | gemini-3.6-flash | RESOURCE_EXHAUSTED quota error or "exhausted your capacity" pattern |
| Codex | gpt-5.6-sol | gpt-5.6-terra | Quota errors (rate_limit_exceeded, 429, insufficient_quota) |
| Claude | opus | sonnet | Claude Code native fallback when Opus is overloaded or unavailable |
| Antigravity | gemini-3.1-pro (--effort high) | gemini-3.5-flash; one model-less retry when agy rejects a model whose value equals gemini-3.1-pro or gemini-3.5-flash (reported as agy default). Both retain the effective effort (high, or the ASK_ANTIGRAVITY_EFFORT override) | Subscription rate limit; model-unavailable (shipped slug values only — other rejected models fail actionably) |
| Ollama | qwen3.6:27b | none | Local, no fallback; a missing model returns a clear ollama pull error |
For Gemini, Codex, Claude, and Antigravity, fallback is automatic. Gemini, Codex, and Claude expose the actual model and usage.fellBack in structured output; Antigravity reports the model through the top-level model field only (it returns no usage statistics) — gemini-3.5-flash after a rate-limit fallback, or the literal agy default after a model-less recovery (agy does not reveal which model it picked). Ollama never falls back, so its fellBack is always false.
Codex uses medium reasoning effort for ordinary calls to preserve the previous default behavior. The quality-first /codex-review and /brainstorm skills use high. Direct ask-codex calls can override this with reasoningEffort (low, medium, high, xhigh, or max).
Choosing a Provider
Different providers excel at different things. Pick by what you're doing, not by which is "best":
| Task | Suggested provider | Why |
|---|---|---|
| Targeted code reasoning, refactor critique | Codex | GPT-5.6 Sol is the flagship agentic coding model; Terra keeps the fallback balanced |
| Claude second opinion while working in Codex | Claude | Opus review through Claude Code CLI, with native session continuation and read-only file access |
| Private / air-gapped analysis | Ollama | Runs locally, nothing leaves your machine |
| Subscription-backed second opinion, larger context | Antigravity | agy via your Google AI Pro/Ultra plan, the Gemini CLI successor |
| Whole-codebase review (enterprise seats) | Gemini | 1M+ token context fits what others can't (enterprise-gated from 2026-06-18) |
| "What do they all think?" comparison | Multi-LLM (multi-llm tool or /compare skill) | Parallel dispatch, see all responses side-by-side |
| Code review with verified findings | /multi-review skill | Antigravity + Codex in parallel, then verifies each finding against source |
Overriding the Model
Pass model explicitly when you have a reason to:
Use ask-llm with provider gemini and model gemini-3.6-flash to quickly check this CSS fileOr programmatically:
{ "name": "ask-llm", "arguments": { "provider": "gemini", "model": "gemini-3.6-flash", "prompt": "..." } }For Codex, common overrides:
Use ask-codex with model gpt-5.6-terra to summarize this commitFor Ollama, you can request any model you've pulled:
ollama pull deepseek-coder:6.7bUse ask-ollama with model deepseek-coder:6.7b to review this implementationFor Antigravity, ask-antigravity requires agy ≥1.1.5 and has no per-call model argument — configure it via env vars:
ASK_ANTIGRAVITY_MODEL— an agy model slug fromagy models(e.g.claude-sonnet-4-6). Legacy effort-carrying display strings likeGemini 3.1 Pro (High)still resolve for backward compatibility.ASK_ANTIGRAVITY_EFFORT— reasoning effort,low|medium|high(defaulthigh). Invalid values log a warning and fall back to the default behavior. Note agy limits tiers per model (gemini-3.1-prohas nomedium), and effort-carrying model names reject--effortentirely — so the default effort is only sent with the shipped base slugs and model-less recovery runs, while an explicitly set value is always sent.
Token Limits & Cost
| Provider | Context window | Cost model |
|---|---|---|
| Gemini Pro | ~1M tokens (~250k LOC) | Gemini Code Assist Standard/Enterprise seat (from 2026-06-18) |
| Gemini Flash | ~1M tokens | Cheaper than Pro; fallback target for quota relief |
| Codex GPT-5.6 Sol | Per OpenAI's published context window | Per OpenAI billing |
| Codex GPT-5.6 Terra | Per OpenAI's published context window | Balanced fallback target |
| Ollama | Per model (e.g., 256k for qwen3.6) | Free, runs locally |
Track What You're Spending
Token usage is exposed live via:
- Per-call:
result.structuredContent.usageonask-*tool responses (provider, model, inputTokens, outputTokens, cachedTokens, thinkingTokens, durationMs, fellBack) — exceptask-antigravity, which reports only the top-levelmodelfield (agy exposes no token counts) - Per-session aggregate: call the
get-usage-statsMCP tool, or read theusage://current-sessionMCP Resource for a JSON snapshot - In the REPL: type
/usagefor a markdown-formatted breakdown
This is in-memory only; no persistence to disk, resets when the MCP server restarts.
Recommendations by Use Case
- General code review → defaults are correct; let the fallback chain handle quota
- Whole-codebase analysis →
ask-gemini(Pro) if you have an enterprise seat, otherwiseask-antigravityfor large-context reads without per-token billing - Quick fixes, fast iteration → request Flash or
gpt-5.6-terraexplicitly to skip the flagship→fallback round-trip - Privacy-sensitive code →
ask-ollama, never leaves your machine - Multi-perspective debate →
multi-llmor/brainstormskill; Claude weighs verified vs inferred