Since Copilot CLI version 1.0.94-0 (October 7, 2026), the /model picker lists models from a running local Ollama instance. If Ollama runs on your machine, inference stays on your machine. But a local model does not make the CLI offline: GitHub telemetry (usage metadata, not your prompts or code) keeps flowing, and web tools stay on. Setting COPILOT_OFFLINE=true is what stops that. Microsoft also says Copilot will start routing work between local and cloud models automatically by the end of October. Fair warning: we build a local-first AI coding tool, so we are not neutral. Everything below comes from GitHub's and Microsoft's own pages, linked at the end.
GitHub shipped three related things on October 7: local model discovery in Copilot CLI, general availability of local sandboxing, and an announcement that Copilot will soon decide on its own when to run a task locally. Together they make local models a first-class option in Copilot, not a workaround. They also blur a line that matters if you are running a local model for privacy: where the model runs is one question, and what else talks to the network is another.
What changed in Copilot CLI on October 7?
The /model command now discovers models from a local Ollama instance and lists them next to your configured models and GitHub's cloud models. The details from GitHub's changelog:
- Version. It starts in CLI version 1.0.94-0. Check yours with
copilot --version. If you are on an older build, update first. - Nothing is added silently. You pick a discovered model, review its provider and endpoint, then choose “Add and use for this session” or “Add without switching.” No restart needed.
- No installs. Ollama and the model must already be on your machine. The flow does not install a runtime or download weights.
- Requirements. The model must support tool calling and streaming.
- Errors are visible. If the provider connection fails, the picker says so and explains why.
The Copilot desktop app has the same idea under Settings > Model providers.
How do I use a local Ollama model with Copilot CLI?
Either pick it from /model, or point the CLI at Ollama with two environment variables before you start it. The environment route comes from GitHub's bring-your-own-key (BYOK) setup, and it is the one to script. Ollama speaks the OpenAI Chat Completions format, which is the CLI's default provider type, and a local Ollama needs no API key:
# bash / zsh
export COPILOT_PROVIDER_BASE_URL=http://localhost:11434
export COPILOT_MODEL=your-model-name
copilot
# PowerShell
$env:COPILOT_PROVIDER_BASE_URL="http://localhost:11434"
$env:COPILOT_MODEL="your-model-name"
copilotUse the base URL exactly as GitHub's docs show it, and replace your-model-name with the name you pulled in Ollama. Three things trip people up:
- Context length. GitHub recommends a model with a context window of at least 128k tokens. Don't assume Ollama uses the model's full context by default. Set it explicitly (for example with
num_ctx). Our local LLM setup guide covers how much context to aim for. You can also cap requests withCOPILOT_PROVIDER_MAX_PROMPT_TOKENS. - Tool calling. A model without tool calling or streaming returns an error rather than limping along.
- No silent fallback. If your provider config is wrong, the CLI exits with an error (connection refused, model not found, timeout, auth failure). It does not quietly switch to a GitHub-hosted model. That is the right call, and GitHub deserves credit for documenting it.
One more thing worth knowing: with BYOK, GitHub's docs say you do not need to sign in to GitHub at all. The CLI talks straight to your provider. You lose the features that live on GitHub's servers: /delegate (the cloud agent), the GitHub MCP server, and GitHub Code Search.
Which local model should I use?
Any model you already run in Ollama that supports tool calling and streaming, with enough memory left over for a long context. Some honest notes on the names in the news this week:
- Qwen3.8-27B is the practical pick for a single 24 GB card. Its weights are about 17 GB at Q4, but that is weights only. A long context needs memory for the KV cache on top, so check your budget with the VRAM calculator before you set 128k. Confirm tool calling works with your quant before you rely on it.
- Mellum2.1, released by JetBrains on October 8, is a 12B mixture-of-experts model with 2.5B active parameters under Apache 2.0, built as a fast model for coding agents. It is on Hugging Face now. JetBrains says GGUF builds for Ollama, llama.cpp, and LM Studio are coming soon, so it is not a one-command Ollama pull yet.
- MAI Code 1.1 Flash is Microsoft's local coding model for Copilot: 137B total and 6.8B active parameters. It is announced, not yet available. Microsoft tested it on an NVIDIA RTX Spark laptop (up to 128 GB of unified memory), where the quantized build is about 53 GB and peaks at 75.5 GB of memory at 256k context. Its benchmark numbers are Microsoft's own, measured on its own setup.
For a wider view of what runs on what hardware, see our local LLM rankings.
Does a local model make Copilot CLI offline?
No. Choosing a local model changes where inference runs, not what else the CLI sends to GitHub. GitHub says this plainly in the changelog. Here is how the three setups differ, according to GitHub's docs:
| Setup | Prompts and code | Telemetry to GitHub | GitHub sign-in | web_fetch | Code Search, GitHub MCP, /delegate |
|---|---|---|---|---|---|
| Signed in, local model picked | To your Ollama | Yes (usage metadata) | Yes | On | Available |
| BYOK, not signed in | To your Ollama | Yes (usage metadata) | No | On | Unavailable |
COPILOT_OFFLINE=true + local Ollama | To your Ollama | No, fully disabled | Not attempted | Off | Unavailable |
GitHub's docs say the telemetry does not include your prompts or code, but it does include usage metadata. In offline mode, the CLI only makes network requests to the provider you configured. To turn it on:
# bash / zsh
export COPILOT_OFFLINE=true
# PowerShell
$env:COPILOT_OFFLINE="true"The catch: offline mode is only fully isolated if the provider is local too. If COPILOT_PROVIDER_BASE_URL points at a remote endpoint, your prompts and code context still go over the network to that endpoint, offline mode or not.
What changes when Copilot's Auto routing arrives?
Copilot will decide per task whether to run locally or in the cloud, which means you stop choosing where a given prompt goes. Microsoft's announcement says that, by the end of October, Copilot's Auto mode will route work between local and cloud models across a session, weighing task context and cache state. It did not say which plans get it. The alternative is explicit selection: you pick a specific local model and endpoint yourself.
Microsoft's own framing is the right one: “local inference does not make the session offline.” If a repo must never leave your machine, do not leave that decision to a router. Pick a local model explicitly, set COPILOT_OFFLINE=true (per GitHub's docs, the CLI itself then only talks to your configured provider), and use /sandbox to block internet access for the commands the agent runs. If you are fine with some cloud inference, Auto is a reasonable default. Just know which one you are using.
Where does local sandboxing fit?
Sandboxing controls what the agent's commands can touch, whichever model asked for them. Local sandboxing went generally available the same day in Copilot CLI, the Copilot app, and VS Code sessions that use Agent Host, at no extra cost. It runs on Microsoft Execution Containers (MXC), which is open source and uses Windows' native process containers, Seatbelt on macOS, and bubblewrap on Linux. You configure it with the /sandbox command. With it, you can:
- Limit which files and folders agent-run commands can read or change.
- Control access to the internet, your local network, Git credentials, and GitHub CLI credentials.
- Let an organization require sandboxing and lock policies that developers cannot weaken.
Read the boundary carefully. Shell commands, and by default local MCP servers and language servers, run inside the OS-enforced sandbox. Copilot's built-in file tools run inside Copilot itself: they are checked against the policy, but not isolated by the OS. Remote MCP servers sit outside the local sandbox. That is still a real improvement, and it pairs well with a local model: one controls where thinking happens, the other controls what the agent can do.
How do I check whether any AI coding setup is really local?
Ask four questions, then watch the network instead of trusting a settings screen. This works for Copilot CLI and for any other tool, ours included:
- Where does inference run? A local model behind
localhost, or a remote endpoint? - What else talks to the network? Telemetry, sign-in, web tools, update checks, and remote tool servers are all separate channels.
- Is there a router or fallback? Anything that can move a request to the cloud on its own changes the answer to question 1.
- What can tool execution reach? A local model running an unsandboxed shell can still run
curl.
Then verify. Start a session and list the open connections: netstat -ano or Get-NetTCPConnection on Windows, lsof -nP -i on macOS and Linux. Filter to the tool's process ID (the PID column in netstat -ano, Get-NetTCPConnection -OwningProcess <PID>, or lsof -nP -i -a -p <PID>), and check more than once during a session, since short calls are easy to miss. In a truly local setup, you should only see connections to your model server. For the broader breakdown of local, local-first, and offline, see our offline AI IDE guide.
Where does Bodega One Code fit?
If you already live in Copilot, Copilot CLI with Ollama and COPILOT_OFFLINE=true is now a legitimate local setup. If you want local to be the default rather than a configuration, that is what we build. Bodega One Code is a local-first AI coding IDE: bring your own model, with 10+ provider presets including fully local ones, and air-gap mode blocks network egress across nine layers of enforcement.
Honest caveats:
- Bodega One Code is not open source. So treat our air-gap claim as testable, not trust-me: run the network check above in air-gap mode and confirm nothing leaves.
- If you choose a cloud model provider, your prompts go to that provider. A local model plus air-gap mode is what keeps everything on your machine.
- It is free for everyone during the open beta, and beta software has bugs. Report them on Discord or GitHub.
Sources
- GitHub Changelog - “Discover local models in GitHub Copilot CLI” (Oct 7, 2026)
- GitHub Docs - Adding LLM models to GitHub Copilot CLI (providers, environment variables, offline mode)
- GitHub Docs - Authenticating GitHub Copilot CLI (unauthenticated BYOK use, offline mode)
- GitHub Docs - Application card: GitHub Copilot Agents (telemetry contents, no fallback to GitHub-hosted models)
- GitHub Changelog - “Local sandboxing for GitHub Copilot now generally available” (Oct 7, 2026)
- Microsoft Command Line blog - “Bringing local models and sandboxed tools to Windows and GitHub Copilot” (Oct 7, 2026)
- JetBrains AI Blog - “Mellum2.1 Gets to Work: A Fast Open Model for Coding Agents” (Oct 8, 2026)
Common questions
- Does GitHub Copilot CLI work with Ollama?
- Yes. Since CLI version 1.0.94-0 (October 7, 2026), the /model picker discovers models from a running local Ollama instance. You can also set COPILOT_PROVIDER_BASE_URL=http://localhost:11434 and COPILOT_MODEL before starting the CLI. The model must support tool calling and streaming, and GitHub recommends at least 128k tokens of context.
- Do I need a GitHub account to use a local model in Copilot CLI?
- No. GitHub's docs say that when you bring your own model provider, GitHub authentication is not required and the CLI connects directly to your provider. Without signing in, you lose /delegate, the GitHub MCP server, and GitHub Code Search.
- Does Copilot CLI send telemetry when I use a local model?
- Yes, unless you turn on offline mode. GitHub says telemetry continues when you use your own model provider, and it includes usage metadata but not your prompts or code. Setting COPILOT_OFFLINE=true disables telemetry fully.
- What does COPILOT_OFFLINE=true do?
- It stops Copilot CLI from contacting GitHub. No sign-in is attempted, telemetry is disabled, web tools like web_fetch and Code Search are off, and the CLI only talks to the model provider you configured. It is only fully isolated if that provider is local too.
- Will Copilot Auto mode send my code to the cloud?
- It can. Microsoft says that by the end of October 2026, Copilot's Auto mode will route work between local and cloud models based on task context and cache state. If code must stay on your machine, pick a local model explicitly, set COPILOT_OFFLINE=true, and sandbox the agent's commands so they cannot reach the internet.
Written by the Bodega One team. We build Bodega One Code, the local-first AI IDE, and we write here about local models, AI costs, and what we learn shipping it. More about the team and why we build local-first on the about page.
Related posts
Keep learning
Free, vendor-neutral courses and guides in the Bodega One AI Academy, from what a model is to shipping an app.
Stay in the loop
Build-in-public updates, model picks, and Copilot/Cursor news as it breaks.
Follow @BodegaOneAI on X →