Last updated: Sep 23, 2026
#2 in AI Utilities

llama.cpp

Trusted

Efficient local LLM inference in C/C++

llama.cpp, developed by ggml.ai, is an efficient local LLM inference in C/C++. It helps you with GGUF, quantisation and server, and is available on Windows, macOS and Linux.

4.7/ 5
Shujaz editor score
Visit Official Website
  • Free & open source
GGUF
Quantisation
Server
Cross-platform

What is llama.cpp?

llama.cpp, developed by ggml.ai, is an efficient local LLM inference in C/C++. It helps you with GGUF, quantisation and server, and is available on Windows, macOS and Linux. AI utilities are small, focused tools that solve one job well, from summarising any web page in your browser to running AI models locally on your laptop, cleaning up files or generating QR art.

Read the full review, features, use cases, pricing and more.

Key Features

GGUF

GGUF sits at the core of the experience.

Quantisation

With quantisation, the tool takes over a large share of repetitive manual work.

Server

The server capability is one of the reasons people pick this tool over a generic AI…

Cross-platform

Cross-platform is especially useful when you are working against a deadline.

Use Cases

For Everyone

  • Web page summaries
  • Quick answers
  • File cleanup

For Power Users

  • Local AI models
  • Shortcuts
  • Automation

For Students

  • Reading help
  • Translation
  • Notes

For Professionals

  • Email help
  • Document tools
  • Research

llama.cpp Pricing

View Official Pricing

Open Source

$0 / forever

  • Full source code
  • Self-hosting
  • Community support
Get Started

Prices are indicative and may change or vary by region. Always confirm on the official website.

Have you used llama.cpp?

Rate it to help others choose the right AI tool.