A small local proxy that lets any OpenAI-compatible chat client (Open WebUI, editors, scripts) use the models included in your Mistral Vibe subscription. It is a single Python file with no dependencies beyond the standard library.
It uses the same key the official Mistral Vibe CLI creates when you log in, sends requests to Mistral's public API and returns the answers in the exact shape OpenAI clients expect.
Unofficial project. Not affiliated with, endorsed by or supported by Mistral AI. Read the "Terms of service" section below before using it.
- Exposes
/v1/chat/completionsand/v1/modelson127.0.0.1:8788, so a client only needs a base URL. - Fixes wire-format differences that break OpenAI clients:
- reasoning models return
contentas a list of typed parts; the gateway flattens it to plain text and moves the thinking intoreasoning_content, in both streaming and non-streaming replies - a final
finish_reason: "stop"after tool-call deltas is rewritten totool_calls stream_options.include_usageis added to streaming requests so clients get token countsmax_tokensdefaults to 32768 when the client sends none (Mistral's default is 4096, which reasoning models can use up before writing any visible text)reasoning_effort: "high"is added to GLM turns that answer a tool result, because GLM sometimes ends those turns with empty output otherwisemax_completion_tokensandseedare passed on as Mistral'smax_tokensandrandom_seed, anduseris left out, because Mistral rejects all three- a
reasoning_effortthe model rejects is retried with the closest level it accepts, or without it for models that have no reasoning (see Models)
- reasoning models return
- Handles rate limits politely: on HTTP 429 it waits and retries the same key every 10 s for up to 60 s.
- Optional fallback to a Mistral AI Studio API key when the Vibe key is rate limited or its budget is spent.
- Records daily token usage per model in a local SQLite file (
usage.db), readable at/usage.
The chat models below were tested through the gateway on 2026-10-06, including tool calls and every reasoning_effort level. Mistral changes the line-up from time to time; /v1/models always lists what your key can use.
| Model | Also available as | Context | Images | Reasoning effort |
|---|---|---|---|---|
mistral-large-4 |
mistral-large-4-0 |
1M | yes | none, high |
mistral-medium-latest |
mistral-medium-3-5, mistral-medium-2604, mistral-vibe-cli-latest |
256K | yes | none, high |
mistral-small-latest |
mistral-small-2603, mistral-vibe-cli-fast |
256K | yes | none, high |
zai-glm-5-3 |
zai-glm-latest, zai-glm-5 |
1M | no | low, high, max |
mistral-large-latest |
mistral-large-2512 |
256K | yes | no reasoning |
codestral-latest |
codestral-2508 |
128K | no | no reasoning |
ministral-14b-latest |
ministral-14b-2512 |
256K | yes | no reasoning |
ministral-8b-latest |
ministral-8b-2512 |
256K | yes | no reasoning |
ministral-3b-latest |
ministral-3b-2512 |
256K | yes | no reasoning |
voxtral-small-latest |
voxtral-small-2507 |
32K | no | no reasoning |
All of them handle tool calls. Clients can send any standard reasoning_effort level to any model. When a model does not offer the requested level, the gateway uses the closest one it does offer, and on a tie minimal and low go down while everything else goes up: on Mistral Large 4, Medium and Small, minimal and low become none and the rest become high. Models without reasoning get the request without it. Labs models such as labs-leanstral-1-5 only answer once an admin enables Labs models in the organization settings. Embedding, OCR and transcription models are listed too, but the gateway only serves chat.
Requirements: Python 3.10+ and a Mistral Vibe plan with coding included (Pro, Team or Enterprise).
-
Install the official Vibe CLI (see the Mistral Vibe repository for other install methods):
curl -LsSf https://mistral.ai/vibe/install.sh | bash -
Log in with your Mistral account. Pick the browser login in the setup wizard:
vibe --setup
The CLI stores your plan key in
~/.vibe/.env. The gateway reads it from there, so you never copy the key by hand.Without the CLI: open Code › Extensions in Vibe, expand Advanced on the Vibe CLI card and copy the Vibe API Key. Put it in
gateway.envasVIBE_KEY=.... The same panel shows your monthly usage and lets you rotate the key. -
Get the gateway and start it:
git clone https://github.com/Classic298/mistral-gateway.git cd mistral-gateway python3 gateway.pyThe log shows your plan name when the key works:
INFO vibe plan: ... (...) INFO mistral gateway on 127.0.0.1:8788 (keys: vibe) -
Point your client at it:
- Base URL:
http://127.0.0.1:8788/v1 - API key: anything (or the value of
GATEWAY_KEYif you set one) - Model: any ID from
curl http://127.0.0.1:8788/v1/models
Requests must be sent as
Content-Type: application/json. Browsers cannot call the gateway directly; it is meant for server-side and desktop clients.Quick test:
curl http://127.0.0.1:8788/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{"model": "mistral-small-latest", "messages": [{"role": "user", "content": "Hello"}]}'
- Base URL:
mkdir -p ~/.config/systemd/user
cp mistral-gateway.service.example ~/.config/systemd/user/mistral-gateway.service
# the example assumes the clone is at ~/mistral-gateway; otherwise edit WorkingDirectory and ExecStart
systemctl --user daemon-reload
systemctl --user enable --now mistral-gateway
journalctl --user -u mistral-gateway -fSettings come from environment variables or from a gateway.env file next to gateway.py (see gateway.env.example; gateway.env wins over the environment).
| Variable | Default | Meaning |
|---|---|---|
VIBE_KEY |
MISTRAL_API_KEY from ~/.vibe/.env |
Your Vibe plan key. Set it if you copied the key from Code › Extensions, or to override the CLI's file. |
VIBE_ENV_FILE |
~/.vibe/.env |
Where to look for the Vibe CLI's key. |
STUDIO_API_KEY |
unset | Optional pay-as-you-go fallback key from Mistral AI Studio. |
GATEWAY_KEY |
unset | If set, clients must send it as Authorization: Bearer <key>. Without it the gateway only answers requests addressed to localhost, 127.0.0.1 or [::1], so set it if a client reaches the gateway under another hostname (e.g. from a Docker container). |
BIND |
127.0.0.1 |
Listen address. Keep it on loopback (see below). |
PORT |
8788 |
Listen port. |
RETRY_429_SECONDS / RETRY_429_INTERVAL |
60 / 10 |
How long and how often to retry a rate-limited key. |
COOLDOWN_429 / COOLDOWN_DEAD |
60 / 21600 |
Seconds a key is skipped after rate limiting / after 401 or 402. |
LOG_LEVEL |
INFO |
Python logging level. |
Endpoints: /v1/chat/completions, /v1/models, /health, /pool/status (which key is serving, cooldowns) and /usage?days=N.
Checked against the Mistral AI Terms of Service for EU consumers (effective August 7, 2026) and the Mistral AI Commercial Terms of Service (effective August 5, 2026). Consumers outside the EU have separate terms. This section is the author's reading of those documents and is not legal advice.
Why the gateway is designed to stay within those terms:
- Official key, official API. The key comes from Mistral's own login flow in the official Vibe CLI. The gateway calls the same public API endpoint with the same standard
Authorization: Bearerheader the CLI uses. It does not scrape, impersonate the CLI, reverse engineer anything or touch any security mechanism. Mistral's API key documentation lists Vibe keys as "Keys used by Vibe Code" and states that "Vibe-only users usually do not need API keys unless they also use Studio, the API, Vibe Code, or another developer tool." - Your plan's limits stay in force. Every request counts against your plan exactly like a Vibe CLI request. A rate limit makes the gateway wait and retry the same key; it never switches to another Vibe account to get around a limit. The only fallback is your own optional Studio key.
- One person, one account. The EU consumer terms state: "The creation or use of multiple Mistral AI accounts by a single individual is strictly prohibited, including to bypass rate limits or any other restrictions." The gateway therefore supports exactly one Vibe key.
- Personal use only. The consumer terms also state: "Your account is intended for your individual use only, and you may not share your account with any other person. [...] You may not make your account credentials available to third parties, [...] or resell or lease access to your account." The commercial terms (section 2.2) forbid customers to "buy, sell, or transfer API keys" and to "grant any third party access to the Mistral AI Products without our prior written authorization". The gateway therefore listens on
127.0.0.1only. Do not expose it to the internet or share it with other people; doing so would hand them access to your account.
Mistral has confirmed third-party use. In a stickied comment on the official r/MistralAI subreddit, a moderator with the Mistral team badge wrote: "you are currently able to freely use the Vibe key with any software, including Hermes and any other packages or third-party services. This has always been the case, meaning if you used the Vibe API key via third parties, the Vibe plan budget would be used first before any PAYG." In the same thread, the user who asked reports that Mistral support confirmed this and refunded the API charges they had been billed.
Your responsibility. Mistral can change its terms at any time. Before you use this gateway, and again whenever Mistral announces a change, read the current terms yourself and stop using the gateway if they no longer allow it. Mistral decides how to enforce its own terms, and nobody can promise that an account will not be restricted or suspended.
This software is provided "as is" under the PolyForm Noncommercial License 1.0.0, without warranty of any kind. The author accepts no liability for any damage or loss arising from its use, including suspension, restriction or termination of your Mistral account, lost subscription fees or lost data.
PolyForm Noncommercial 1.0.0: free for personal, hobby, research and other noncommercial use. Commercial use is not permitted.
It is noncommercial because a Mistral Vibe subscription is a consumer plan: Mistral's consumer terms only cover "your personal use as a consumer", and business use falls under their separate Commercial Terms of Service.