What this is
If a service speaks the OpenAI chat-completions format, you can use it in Cortex without waiting for us to add support.
Register the endpoint once in your account, list the models on it, and they appear in your model picker.
Two things stay true throughout:
What works
Anything serving that request shape in OpenAI format. In practice that covers most of the market.
Auth style No key — there is nothing to paste.
Auth style Bearer — paste your key in Cortex.
Works if it exposes an OpenAI-compatible route.
Providers with their own wire format rather than OpenAI's:
- Anthropic's native API
- Google Gemini's native API
- Bedrock and Vertex
Those need a real client in the app, which is why they ship built in. If you add one anyway, Cortex ignores it rather than breaking.
Add a provider
Open My Providers and choose Add provider. Four fields, in this order.
- 01Slug
- A short id like
my-ollama. It becomes part of every model id (dp/my-ollama/<model>) and names the key on your machine, so it cannot be changed later. - 02Base URL
-
Where requests go, up to and including
/v1. Cortex appends/chat/completionsitself.http://localhost:11434/v1the Ollama default - 03Auth style
Bearerfor a normal API key,Custom headerif the service wants something likex-api-key, orNo keyfor a server on your own machine.- 04Key prefix hintoptional
- Shown in Cortex so you recognise the right key, for example
gsk_.
Save, and you land on the models page for that provider.
Add models
A provider on its own gives Cortex nowhere to send a request. Add each model you want on My Models.
The model id is exactly what your endpoint calls it — llama3.1:8b for Ollama, llama-3.3-70b-versatile for Groq. Cortex sends that string straight through; you do not add any prefix yourself.
The rest are hints for the agent, and getting them roughly right matters:
- 01Context window
-
The budget for everything sent up: your files, history and prompt. It is the one field where both directions go wrong.
Too lowWastes capacity About right Too highRefuses mid-taskWhen unsure, set it low and raise it.
- 02Max output tokens
- Caps the reply length.
- 03Supports tools
- Leave this on unless the model rejects tool calls. The agent needs tools for almost everything beyond plain chat.
Use it in Cortex
Two pages in, and the work moves to the IDE. Four steps to a working model.
- Open Settings → Models & Providers. Your provider appears in Provider API Keys alongside the built-in ones.
- Paste your key and let the field lose focus to save it. A
No keyprovider skips this. - Switch the toggle on. That is what puts its models in the chat dropdown.
- Pick a model from the dropdown and start working.
Changes on this website reach Cortex the next time it starts.
CSV import
Adding twenty models by hand is nobody's idea of a good time. Both pages take a CSV, and both work the same way.
slug for providers and model_id for models, so you can export, edit in a spreadsheet, and upload again.
Save the file as CSV UTF-8. Other encodings are rejected with a message rather than importing mojibake.
Column reference
Every column both files accept. Rows carrying the green rail are the ones you cannot leave out.
Providers CSV
Required: slug and base_url. Everything else is optional.
| Column | Accepted values |
|---|---|
| slug | 2–40 chars, lowercase letters, digits, hyphens |
| base_url | No query string, no credentials in the URL |
| display_name | Shown in Cortex. Defaults to the slug |
| description | One line, up to 160 chars |
| auth_style | bearer, header, or none |
| auth_header_name | Only when auth_style is header |
| models_path | Defaults to /models |
| key_prefix_hint | Cosmetic. Shown beside the key field in Cortex |
| accent_color | Cosmetic |
| signup_url | Cosmetic |
| docs_url | Cosmetic |
| is_active | true or false |
Models CSV
Required: model_id. Everything else falls back to a default.
| Column | Accepted values |
|---|---|
| model_id | Exactly what the endpoint calls it |
| display_name | What you see in the picker |
| description | What you see in the picker |
| context_window | Whole number. default 128000 |
| max_output_tokens | Whole number. default 8192 |
| supports_tools | Default true |
| supports_vision | Default false |
| supports_thinking | Default false |
| is_free | Shows a FREE badge. Default false |
| is_active | Default true |
| sort_order | Lower sorts first. default 100 |
Limits & rules
Checked when you save. A save that breaks one of these is refused with the reason, not silently dropped.
- 20 providers per account, 500 models per provider.
- You cannot reuse the slug of a published provider. Model ids would be ambiguous.
- A private provider may use plain
http, but only forlocalhostor a LAN address. That is what makes a local Ollama work. - Anything public must use
https. - No credentials, query strings or fragments in a base URL. The key belongs in Cortex, not in a URL that syncs to our server.
- Reserved headers cannot be set through
extra_headers—Authorization,Cookie,Hostand similar. Cortex builds those itself. extra_bodycannot overridemodel,messages,stream,toolsortool_choice. It is for extras liketop_p.
Troubleshooting
Symptoms in the order people hit them. Find yours and read the one line under it.
The provider is not in Cortex
Restart Cortex — the list is fetched at startup. If it is still missing, check the provider is switched on here, and that you are signed in to the same account in the IDE.
Its models are not in the dropdown
Three things must all be true: the provider is on, at least one of its models is on, and the toggle in Settings → Models & Providers is switched on. The toggle is the one people miss.
404“Not found for account”
The endpoint answered, so your key and URL are fine, but that model id is not available to your account. Check the exact spelling against the provider's own model list.
401 · 403Rejected
The key is wrong, expired, or lacks access. Re-paste it in Cortex.
refusedConnection refused on localhost
The local server is not running, or it is on a different port. Confirm with curl http://localhost:11434/v1/models.
Replies cut off early
max_output_tokens is too low. Raise it on the model.
Errors about context length
context_window is set higher than the model accepts. The error usually names the real limit; put that number in the field.