Skip to main content

LLM Providers

💡TL;DR

Notemd has 36 provider presets across five protocol families: OpenAI-compatible, Anthropic, Google, Azure OpenAI and Ollama. Choose a preset matching your endpoint's protocol, enter the required credentials, and test an accessible model on a small note. Preset defaults are configurable starting values, not promises of current upstream model availability.

Provider Categories​

The following preset values come from the provider registry. Keep the full URL scheme and path. A placeholder model or endpoint must be replaced with your deployment's actual value. Hosted provider access, pricing and retirement schedules belong to the provider.

Cloud Providers​

ProviderPreset modelPreset base URL
DeepSeekdeepseek-v4-prohttps://api.deepseek.com
Qwenqwen3-235b-a22bhttps://dashscope.aliyuncs.com/compatible-mode/v1
Qwen Codeqwen3-coder-plushttps://dashscope.aliyuncs.com/compatible-mode/v1
Doubaoep-xxxxxxxxxxxxxxxxhttps://ark.cn-beijing.volces.com/api/v3
Moonshotkimi-k2-0905-previewhttps://api.moonshot.cn/v1
Xiaomi MiMomimo-v2.5-prohttps://api.xiaomimimo.com/v1
GLMglm-5https://open.bigmodel.cn/api/paas/v4
Z AIglm-5https://api.z.ai/api/paas/v4
MiniMaxMiniMax-M2.7https://api.minimaxi.com/v1
Baidu Qianfanernie-4.5-turbo-32khttps://qianfan.baidubce.com/v2
SiliconFlowQwen/QwQ-32Bhttps://api.siliconflow.cn/v1
Huawei Cloud MaaSDeepSeek-V3https://api.modelarts-maas.com/v1
OpenAIgpt-4ohttps://api.openai.com/v1
Anthropicclaude-3-5-sonnet-20240620https://api.anthropic.com
Googlegemini-2.0-flash-exphttps://generativelanguage.googleapis.com/v1
Mistralmistral-large-latesthttps://api.mistral.ai/v1
Azure OpenAIgpt-4oConfigure your Azure resource endpoint
xAIgrok-4https://api.x.ai/v1
Groqmoonshotai/kimi-k2-instruct-0905https://api.groq.com/openai/v1
Togethermeta-llama/Meta-Llama-3.1-70B-Instruct-Turbohttps://api.together.xyz/v1
Fireworksaccounts/fireworks/models/kimi-k2p5https://api.fireworks.ai/inference/v1
Nebiusopenai/gpt-oss-120bhttps://api.studio.nebius.com/v1
Cerebrasgpt-oss-120bhttps://api.cerebras.ai/v1

Gateway / Proxy Providers​

ProviderPreset modelPreset base URL
OpenRouteranthropic/claude-3.7-sonnethttps://openrouter.ai/api/v1
AIHubMixgpt-4o-minihttps://aihubmix.com/v1
GitHub Modelsgpt-4o-minihttps://models.github.ai/inference
PPIOqwen/qwen3-32bhttps://api.ppinfra.com/v3/openai
New APIgpt-4.1http://localhost:3000/v1
LiteLLMyour-proxy-modelhttp://localhost:4000/v1
Hugging Faceopenai/gpt-oss-120bhttps://router.huggingface.co/v1
Vercel AI Gatewayanthropic/claude-sonnet-4.5https://ai-gateway.vercel.sh/v1
Requestyanthropic/claude-3-7-sonnet-latesthttps://router.requesty.ai/v1
OpenAI Compatibleyour-model-idhttps://your-openai-compatible-endpoint/v1

Local Providers​

ProviderPreset modelPreset base URL
OVMSopenvino-modelhttp://localhost:8000/v3
LMStudiolocal-modelhttp://localhost:1234/v1
Ollamallama3http://localhost:11434/api

China Providers​

See regional provider setup for DeepSeek, Qwen, Doubao, Moonshot, GLM and related presets. The category describes configuration, not a guarantee of geographic availability or Chinese-language quality. Match account region, endpoint and model entitlement.

Per-Task Model Selection​

Enable useMultiModelSettings to select separate providers and model overrides. Without it, tasks use activeProvider. An empty task-model override uses that provider's configured model; invalid saved provider choices are reconciled with the active provider when settings load.

TaskProvider settingModel override
Add linksaddLinksProvideraddLinksModel
Title generationgenerateTitleProvidergenerateTitleModel
Research summaryresearchProviderresearchModel
TranslationtranslateProvidertranslateModel
Concept extractionextractConceptsProviderextractConceptsModel
Mermaid summarysummarizeToMermaidProvidersummarizeToMermaidModel
Original-text extractionextractOriginalTextProviderextractOriginalTextModel

The default active provider and saved task providers are DeepSeek. Verify each configured task before processing private material; changing only the active provider does not override enabled task-specific choices.

API Call Architecture​

Transport Layers​

The plugin uses Obsidian request facilities and protocol-aware fallbacks. Desktop fallbacks can use Node HTTP streams; other environments can use fetch. The provider protocol and runtime transport are different concepts. A Claude model served by an OpenAI-compatible gateway uses that gateway preset, rather than the native Anthropic profile.

Retry Logic​

Stable request mode is off by default, with configured retry defaults of three retries and a five-second interval when applicable. Transient transport failures can enter the shared retry/fallback path. Authentication or invalid model configuration must be corrected instead of relying on retries. Long requests can still fail because of remote timeouts, model limits or quota exhaustion.

Response Caching​

Successful responses can be reused by a bounded in-memory cache. Endpoint, model, generation settings, prompt and content contribute to request identity. Caching is not persistence, a cost meter or a guarantee that a repeat operation will avoid a new paid request.

Reasoning Model Handling​

Reasoning and thinking controls are exposed where supported by the preset and model rules. Do not copy protocol-specific parameters between providers. Validate the exact task with the chosen model; a model that can chat may still reject particular reasoning, token or sampling settings.

Token Estimation​

Text-length estimates and configured output limits are not exact provider token counts. Inspect the provider's actual usage for billing. The plugin does not provide a universal cost estimate.

Model Discovery​

Fetch models discovers models when supported by the preset. Some deployments require a manually entered model or deployment name. Listing permission and generation permission are independent.

Test connection follows the preset's test mode. A models-then-chat preset can report success from the model-list endpoint alone; a chat-only preset performs a small generation request. Always run a small real task to verify the exact model and output shape needed by your workflow.

  • Anthropic and Google have native model-list protocols.
  • Ollama lists locally pulled model tags; LM Studio lists models exposed by its local server.
  • Azure OpenAI requires deployment configuration, including the API version; a generic model list is not a substitute.
  • Gateways may limit model discovery even when a specific model accepts chat requests.

Quick Start​

  1. Select the matching preset in Settings → Notemd.
  2. Enter its endpoint, credentials and an actual model/deployment available to you.
  3. Run Test connection, then process one disposable note.
  4. Inspect the output and progress report before enabling batch work or task-specific routing.

For local operation, download/load the model and start its server first. Ollama has no API-key requirement; LM Studio accepts an optional key and the preset uses EMPTY. Set a real token if your server requires authentication. Local LLM use does not make web search or other network-dependent tasks offline.

Next Steps​