We wrote this book —
and we’ll write yours.

Your first book, $9.99.

Commission your book →  ·  ready in 24–48 hours, yours to publish anywhere

Technology · Machine Learning Systems

llama.cpp in the Real World: On-Device LLM Inference, Profiling, and Deployment

by Trex Team

Large language models no longer live only in hyperscale data centers. This rigorous engineering guide treats on-device inference as a discipline shaped by real hardware limits: it explains how llama.cpp is structured as an inference stack, why the GGUF format defines the deployment boundary, and how backend selection, build strategy, quantization, context sizing, and KV-cache planning interact to decide whether a model is feasible -- and how fast it runs -- on the machine in front of you.

Read the full book free (PDF)

Download the EPUB

Free digital edition shared by the publisher. Enjoy the read.

← Back to the bookshelf