We wrote this book —
and we’ll write yours.
Your first book, $9.99.
Commission your book → · ready in 24–48 hours, yours to publish anywhere
Technology · Machine Learning Systems
by Trex Team
Large language models no longer live only in hyperscale data centers. This rigorous engineering guide treats on-device inference as a discipline shaped by real hardware limits: it explains how llama.cpp is structured as an inference stack, why the GGUF format defines the deployment boundary, and how backend selection, build strategy, quantization, context sizing, and KV-cache planning interact to decide whether a model is feasible -- and how fast it runs -- on the machine in front of you.
Free digital edition shared by the publisher. Enjoy the read.