LLM Providers
Notemd has 36 provider presets across five protocol families: OpenAI-compatible, Anthropic, Google, Azure OpenAI and Ollama. Choose a preset matching your endpoint's protocol, enter the required credentials, and test an accessible model on a small note. Preset defaults are configurable starting values, not promises of current upstream model availability.
Provider Categories
The following preset values come from the provider registry. Keep the full URL scheme and path. A placeholder model or endpoint must be replaced with your deployment's actual value. Hosted provider access, pricing and retirement schedules belong to the provider.
Cloud Providers
| Provider | Preset model | Preset base URL |
|---|---|---|
| DeepSeek | deepseek-v4-pro | https://api.deepseek.com |
| Qwen | qwen3-235b-a22b | https://dashscope.aliyuncs.com/compatible-mode/v1 |
| Qwen Code | qwen3-coder-plus | https://dashscope.aliyuncs.com/compatible-mode/v1 |
| Doubao | ep-xxxxxxxxxxxxxxxx | https://ark.cn-beijing.volces.com/api/v3 |
| Moonshot | kimi-k2-0905-preview | https://api.moonshot.cn/v1 |
| Xiaomi MiMo | mimo-v2.5-pro | https://api.xiaomimimo.com/v1 |
| GLM | glm-5 | https://open.bigmodel.cn/api/paas/v4 |
| Z AI | glm-5 | https://api.z.ai/api/paas/v4 |
| MiniMax | MiniMax-M2.7 | https://api.minimaxi.com/v1 |
| Baidu Qianfan | ernie-4.5-turbo-32k | https://qianfan.baidubce.com/v2 |
| SiliconFlow | Qwen/QwQ-32B | https://api.siliconflow.cn/v1 |
| Huawei Cloud MaaS | DeepSeek-V3 | https://api.modelarts-maas.com/v1 |
| OpenAI | gpt-4o | https://api.openai.com/v1 |
| Anthropic | claude-3-5-sonnet-20240620 | https://api.anthropic.com |
gemini-2.0-flash-exp | https://generativelanguage.googleapis.com/v1 | |
| Mistral | mistral-large-latest | https://api.mistral.ai/v1 |
| Azure OpenAI | gpt-4o | Configure your Azure resource endpoint |
| xAI | grok-4 | https://api.x.ai/v1 |
| Groq | moonshotai/kimi-k2-instruct-0905 | https://api.groq.com/openai/v1 |
| Together | meta-llama/Meta-Llama-3.1-70B-Instruct-Turbo | https://api.together.xyz/v1 |
| Fireworks | accounts/fireworks/models/kimi-k2p5 | https://api.fireworks.ai/inference/v1 |
| Nebius | openai/gpt-oss-120b | https://api.studio.nebius.com/v1 |
| Cerebras | gpt-oss-120b | https://api.cerebras.ai/v1 |
Gateway / Proxy Providers
| Provider | Preset model | Preset base URL |
|---|---|---|
| OpenRouter | anthropic/claude-3.7-sonnet | https://openrouter.ai/api/v1 |
| AIHubMix | gpt-4o-mini | https://aihubmix.com/v1 |
| GitHub Models | gpt-4o-mini | https://models.github.ai/inference |
| PPIO | qwen/qwen3-32b | https://api.ppinfra.com/v3/openai |
| New API | gpt-4.1 | http://localhost:3000/v1 |
| LiteLLM | your-proxy-model | http://localhost:4000/v1 |
| Hugging Face | openai/gpt-oss-120b | https://router.huggingface.co/v1 |
| Vercel AI Gateway | anthropic/claude-sonnet-4.5 | https://ai-gateway.vercel.sh/v1 |
| Requesty | anthropic/claude-3-7-sonnet-latest | https://router.requesty.ai/v1 |
| OpenAI Compatible | your-model-id | https://your-openai-compatible-endpoint/v1 |
Local Providers
| Provider | Preset model | Preset base URL |
|---|---|---|
| OVMS | openvino-model | http://localhost:8000/v3 |
| LMStudio | local-model | http://localhost:1234/v1 |
| Ollama | llama3 | http://localhost:11434/api |
China Providers
See regional provider setup for DeepSeek, Qwen, Doubao, Moonshot, GLM and related presets. The category describes configuration, not a guarantee of geographic availability or Chinese-language quality. Match account region, endpoint and model entitlement.
Per-Task Model Selection
Enable useMultiModelSettings to select separate providers and model overrides. Without it, tasks use activeProvider. An empty task-model override uses that provider's configured model; invalid saved provider choices are reconciled with the active provider when settings load.
| Task | Provider setting | Model override |
|---|---|---|
| Add links | addLinksProvider | addLinksModel |
| Title generation | generateTitleProvider | generateTitleModel |
| Research summary | researchProvider | researchModel |
| Translation | translateProvider | translateModel |
| Concept extraction | extractConceptsProvider | extractConceptsModel |
| Mermaid summary | summarizeToMermaidProvider | summarizeToMermaidModel |
| Original-text extraction | extractOriginalTextProvider | extractOriginalTextModel |
The default active provider and saved task providers are DeepSeek. Verify each configured task before processing private material; changing only the active provider does not override enabled task-specific choices.
API Call Architecture
Transport Layers
The plugin uses Obsidian request facilities and protocol-aware fallbacks. Desktop fallbacks can use Node HTTP streams; other environments can use fetch. The provider protocol and runtime transport are different concepts. A Claude model served by an OpenAI-compatible gateway uses that gateway preset, rather than the native Anthropic profile.
Retry Logic
Stable request mode is off by default, with configured retry defaults of three retries and a five-second interval when applicable. Transient transport failures can enter the shared retry/fallback path. Authentication or invalid model configuration must be corrected instead of relying on retries. Long requests can still fail because of remote timeouts, model limits or quota exhaustion.
Response Caching
Successful responses can be reused by a bounded in-memory cache. Endpoint, model, generation settings, prompt and content contribute to request identity. Caching is not persistence, a cost meter or a guarantee that a repeat operation will avoid a new paid request.
Reasoning Model Handling
Reasoning and thinking controls are exposed where supported by the preset and model rules. Do not copy protocol-specific parameters between providers. Validate the exact task with the chosen model; a model that can chat may still reject particular reasoning, token or sampling settings.
Token Estimation
Text-length estimates and configured output limits are not exact provider token counts. Inspect the provider's actual usage for billing. The plugin does not provide a universal cost estimate.
Model Discovery
Fetch models discovers models when supported by the preset. Some deployments require a manually entered model or deployment name. Listing permission and generation permission are independent.
Test connection follows the preset's test mode. A models-then-chat preset can report success from the model-list endpoint alone; a chat-only preset performs a small generation request. Always run a small real task to verify the exact model and output shape needed by your workflow.
- Anthropic and Google have native model-list protocols.
- Ollama lists locally pulled model tags; LM Studio lists models exposed by its local server.
- Azure OpenAI requires deployment configuration, including the API version; a generic model list is not a substitute.
- Gateways may limit model discovery even when a specific model accepts chat requests.
Quick Start
- Select the matching preset in Settings → Notemd.
- Enter its endpoint, credentials and an actual model/deployment available to you.
- Run Test connection, then process one disposable note.
- Inspect the output and progress report before enabling batch work or task-specific routing.
For local operation, download/load the model and start its server first. Ollama has no API-key requirement; LM Studio accepts an optional key and the preset uses EMPTY. Set a real token if your server requires authentication. Local LLM use does not make web search or other network-dependent tasks offline.
Next Steps
- OpenAI, Anthropic, Google, China providers and local models: protocol-specific setup.
- Configuration: task settings and output destinations.
- Troubleshooting: failures, diagnostic limits and safe issue reports.