llama.cpp
C++ inference engine that runs quantised language models on ordinary hardware.
why this verdict
Keep it — Meta has not replaced it
Meta's version overlaps, but does not finish the dev tools job, so this one is still worth keeping open.
- A named launch, not a vibe named, unlinked
- Cheap hosted APIs making local inference a hobby rather than a requirement. Meta, September 25, 2025 — no announcement link recorded yet.
- How much of the job it covers not the job editorial call
- Parts of it. The job still needs the tool to get finished.
- Is there a free way to do it? yes
- 3 of 3 listed replacements have a usable free tier: Ollama, LM Studio and MLX.
- What the call is worth nothing to cancel
- No paid entry tier tracked, so there is no subscription to cancel.
- Threatened by
- Meta
- Since
- September 25, 2025
- List price
- free
- Per year
- —
The backstory
llama.cpp proved that a capable model could run on a laptop with no GPU, and its GGUF quantisation format became the standard for distributing models people actually run at home. Hosted APIs are cheaper than most people's time, which caps local inference as a mass-market proposition. It endures because privacy, offline capability and zero marginal cost are not features an API can offer, and because almost every consumer-facing local AI tool, Ollama included, is built on top of it.
Escape hatches
Friendly wrapper around this engine
ollama.com open_in_newDesktop app with a model browser and chat UI
lmstudio.ai open_in_newApple silicon framework with better Mac performance
ml-explore.github.io open_in_new3 of 3 replacements have a usable free tier.