Yesterday’s tip covered Ollama on Apple silicon. LM Studio is the other common Mac path: a desktop chat UI, a Discover browser for weights, and a local server you can point other apps at.
Official system requirements put the floor at Apple silicon (M1 through M4 listed) and macOS 14.0 or newer. Intel Macs are not supported. Docs recommend 16 GB of RAM or more. On 8 GB machines they say stick to smaller models and modest context sizes.
Install LM Studio from the Mac download
Open lmstudio.ai/download and grab the macOS build. The site currently lists LM Studio for macOS 0.4.25. Open the DMG, drag the app into Applications, then launch it. Approve the first-run Gatekeeper prompt if macOS asks.
That is the GUI app. LM Studio also documents a headless daemon called llmster for servers and CI, with a separate install script on the same download page. Most readers only need the DMG.
Keep disk free before you pull models. Weights land under My Models. The same storage squeeze shows up in the Offload Unused Apps tip on iPhone. On Mac, check About This Mac, Storage, first.

Discover, download, then load in Chat
After install, open the Discover tab. Official docs say you can jump there with ⌘2. Search by keyword such as llama or gemma, by a Hugging Face user/model string, or paste a full Hugging Face URL.
Many listings show several files named like Q3_K_S or Q8. Those are the same model at different quantization levels. Docs say choose a 4-bit option or higher if the Mac can hold it. Download finishes into the local models folder.
Switch to Chat. Open the model loader, pick a downloaded model, optionally tweak load settings, then start chatting. Loading means the weights move into unified memory. If the Mac feels tight, unload before you try a bigger file.

MLX runtime on Apple silicon
LM Studio runs GGUF models through llama.cpp on Mac, Windows, and Linux. On Apple silicon it also supports Apple’s MLX runtime. Docs say press ⌘ShiftR to install or manage LM Runtimes.
That is the practical split versus cloud chat. Local weights stay on disk. Replies can stay offline once the model is present. Apple’s own assistant stack is a different pipe, closer to the Siri AI English beta path than a self-hosted GGUF.
Local server for other Mac apps
Developer docs ship OpenAI-compatible and Anthropic-compatible endpoints plus an LM Studio REST API. From the CLI, the documented quick start starts the local server on port 1234 with the lms server start command and a --port 1234 flag.
Then tools on the same Mac can talk to localhost. Keep the server bound to your machine unless you mean to expose it on the network. Cloud ChatGPT or Claude accounts still leave the Mac when you use those apps’ normal sign-in flows.
After a big OS train, confirm the app still opens. Today’s firmware note on Xcode 27.1 beta 27A9269 is a reminder that toolchains move. LM Studio’s own macOS floor remains 14.0.
Start with one small local model in Chat. Confirm replies work with Wi‑Fi off. Add the local server only after that works.
Photo: Apple
Source: LM Studio docs (Get started, Download an LLM, System Requirements, Developer), lmstudio.ai/download, Apple Newsroom (macOS 27)