Ollama
Runs open-weight language models locally with a one-line install and an OpenAI-compatible API.
why this verdict
Keep paying — OpenAI has not replaced it
OpenAI's version overlaps, but does not finish the dev tools job, so this one still earns its $20/mo.
- A named launch, not a vibe named, unlinked
- Gpt-oss open weights plus cheap hosted APIs reducing the reason to run models locally. OpenAI, August 5, 2025 — no announcement link recorded yet.
- How much of the job it covers not the job editorial call
- Parts of it. The job still needs the tool to get finished.
- Is there a free way to do it? yes
- 3 of 3 listed replacements have a usable free tier: LM Studio, llama.cpp and vLLM.
- What the call is worth $240/yr
- $20/mo at the entry paid tier — $240 a year per seat.
- Threatened by
- OpenAI
- Since
- August 5, 2025
- List price
- $20/mo
- Per year
- $240
The backstory
Ollama made running an open model on your own machine as simple as one pull command, with a local API that existing clients can point at. OpenAI releasing gpt-oss weights and cutting hosted prices cuts both ways: hosted inference is cheap, but there are now excellent weights to run. Ollama is the runner people reach for either way, and it added a cloud tier for larger models. Being the default local runtime is a real position.
Escape hatches
Desktop app for downloading and chatting with local models
lmstudio.ai open_in_newThe underlying C++ inference engine, free and hackable
github.com open_in_newHigh-throughput serving engine for production GPU inference
vllm.ai open_in_new3 of 3 replacements have a usable free tier.