Skip to main content

Local Models

💡TL;DR

A local provider lets a running model server on your device process task content. Install or load the model first, then configure the matching protocol. Local LLM selection does not make web search, remote endpoints, dependency installation or cloud-routed tasks offline.

Setup​

For Ollama, install the server and pull a model. For LM Studio, load a model and start its local API server. For OVMS, configure the model server deployment before adding its profile to Notemd.

Select the corresponding preset, enter the exact model name exposed by the server, run Test connection, and process a disposable note. Notemd does not install the model or start these servers for you.

Endpoint And Authentication​

PresetDefault base URLModel placeholderProtocol / authentication
Ollamahttp://localhost:11434/apillama3Native Ollama; no API key required
LMStudiohttp://localhost:1234/v1local-modelOpenAI-compatible; optional key, preset value EMPTY
OVMShttp://localhost:8000/v3openvino-modelOpenAI-compatible; configure the deployment's model and access requirements

localhost is the machine running Obsidian. A phone cannot reach a model on your computer through the phone's own loopback address. A LAN endpoint is a network service: use its actual address and appropriate authentication, and do not expose an unauthenticated server to untrusted networks.

Model Discovery​

Ollama lists pulled tags through /api/tags. LM Studio can expose /v1/models. OVMS has its own deployment-aware discovery rules. Listing a model is not a guarantee that it fits memory or can perform a particular structured task; verify an actual request.

Troubleshooting​

  • Connection refused: check server state, listening address and port.
  • Model unavailable: pull/load the model and use the exact name or alias served by the endpoint.
  • Authentication rejected: enter the token required by your server instead of assuming EMPTY is accepted.
  • Slow or incomplete output: inspect memory, context and output limits; start with one request at a time.
  • Malformed output: confirm the protocol, prompt and model capability before changing batch settings.

When To Use​

Use local inference when its privacy boundary, availability and measured output meet your needs. Prepare model files and any optional export dependencies before disconnecting the network. Review every enabled task-specific provider and local-retrieval setting: local excerpts are sent to the selected LLM, which may still be cloud-hosted.

Web research through Tavily or DuckDuckGo remains online. Native exporters may need separately installed tools. A local provider is one part of an offline workflow, not a blanket guarantee for every feature.

Next Steps​