TOOL LAB Developer Tools Beginner Tested Last tested: October 2026

The Complete Local LLM Setup: From Download to API Endpoint in 10 Minutes

🎯 Why This Matters

Every tutorial tells you to download LM Studio and chat with a model — which is fine until you want to plug it into your own tooling. The real unlock is the OpenAI-compatible endpoint: once it's running, every library, script, and app that already speaks the OpenAI API can talk to your local model by changing one URL. No refactoring, no new SDK, no cloud bill.

📖 Overview

A product manager at a fintech startup got tired of paying per-token for every internal document summary, so she downloaded LM Studio, pointed it at a 7B model, and had her team's internal tools calling a local API endpoint in under ten minutes. The trick isn't the download — it's knowing exactly which three settings to flip so that any tool built for the OpenAI API can talk to a model running on your own laptop.

You need a computer with at least 8GB of RAM (16GB is comfortable), a free LM Studio install, and about 4GB of free disk space for a small model. No account, no credit card, no cloud subscription. If you already have an OpenAI API key somewhere in your code, you can swap it for a local endpoint by changing one line.

By the end, you will have a running local model, a working API endpoint at http://localhost:1234/v1, and a tested curl command that proves it works. You will also know how to swap models without touching your application code.

Expect this to take 10 minutes if the download cooperates and 20 if it doesn't. The finished output is a terminal response showing a JSON object with the model's reply — the same shape you'd get from OpenAI, just without the bill.

🛠️ Step-by-Step Guide

1

Step 1

Download LM Studio from lmstudio.ai and install it. Open the app. You'll land on a search screen. [Screenshot: the model search bar at the top of the main window with a list of model cards below].

2

Step 2

In the search bar, type a small model name. For a first run, search for 'Llama 3.2 3B Instruct' or 'Qwen2.5 7B Instruct'. Look for the quantized versions labeled Q4_K_M — those are the sweet spot between quality and size. Click the download button on the model card and wait for it to finish. A 3B model at Q4 is roughly 2GB; a 7B is roughly 4.5GB.

3

Step 3

Once downloaded, click the chat icon in the left sidebar to load the model. At the top of the chat window, you'll see a model selector — confirm your downloaded model is selected. In the right-hand panel, set Temperature to 0.7 for general use, or 0.1 if you want deterministic outputs for testing. Set Context Length to at least 4096. [Screenshot: the right-hand configuration panel with Temperature and Context Length sliders].

4

Step 4

Click the Developer tab in the left sidebar (it may be labeled with a terminal-style icon). You'll see a server toggle. Turn it on. Confirm the port reads 1234 and the server status shows 'Running'. Enable the CORS toggle — this is the setting that 99% of people miss, and without it, anything running in a browser can't reach your endpoint. [Screenshot: the Developer tab with the server toggle on, port 1234 visible, and CORS enabled].

5

Step 5

Open your terminal. Run this exact command, replacing [your-model-name] with the model identifier shown in the Developer tab (it usually looks like 'llama-3.2-3b-instruct' or similar): curl http://localhost:1234/v1/chat/completions -H "Content-Type: application/json" -d '{"model": "[your-model-name]", "messages": [{"role": "user", "content": "Say hello in exactly five words."}], "temperature": 0.7}'. You should get a JSON response back within a few seconds.

6

Step 6

If you got JSON back, you're done. To use this from Python, install the OpenAI SDK (pip install openai) and point it at your local server: client = OpenAI(base_url="http://localhost:1234/v1", api_key="not-needed"). The api_key can be any non-empty string — LM Studio ignores it. Save this snippet somewhere you can reuse it. To stop the server, flip the toggle in the Developer tab back off.

🚀 Level up your AI toolkit

Understand AI in 5 minutes a day. No jargon. No hype. Unsubscribe anytime.