Rearchitecting LLMs

Structural techniques for efficient models

Manning Early Access Program (MEAP), since January 2026. Publication estimated for Spring 2027.

Read it at Manning

Affiliate link: I may earn a commission on purchases.

Cover of Rearchitecting LLMs

About the book

A practical guide to turning large pre-trained language models into small, efficient models specialized for a domain. Instead of treating a model as a black box, the book works on its architecture: removing the layers and neurons that do not contribute to the goal (depth and width pruning), recovering capability through knowledge distillation, and specializing the result with LoRA-based fine-tuning. It also introduces methods of my own, such as fair pruning, which reduces bias at the neuron level, and internal activation interpretation using pair contrastive prompts.

Why I wrote it

It grew out of my own optimization projects, where the techniques were scattered across papers, often without reproducible code. The book unifies them into a coherent, practical pipeline. Read the preface →

Built on published research

The core techniques are similar in spirit to those behind NVIDIA’s Minitron and Mistral’s Ministral model families. The book adapts them to smaller models, limited data and modest GPUs. The main papers behind each chapter are listed below.

Who it’s for

AI, ML and data engineers who know Python and want to go beyond fine-tuning, with some knowledge of PyTorch and curiosity about what happens inside a transformer.

Hands-on

Every chapter comes with notebooks that run on the free tier of Google Colab, using open models such as Llama, Gemma and Qwen.

Chapters published so far

Part 1 · Foundations

Part 2 · Hands-on optimization

Part 3 · Beyond the black box

More chapters are being written and will be added here as they are published.

Papers behind each chapter

1 · Why rearchitecting LLMs matters

2 · An end-to-end rearchitecting project

3 · A blueprint to modern transformers

4 · Building smaller and faster LLMs with depth pruning

5 · Shaping model architectures via width pruning

6 · Knowledge recovery through distillation

7 · Model specialization

8 · Attention optimization

9 · Dynamic routing with Mixture of Experts

10 · Exploring the transformer black box

Companion resources

Cite the book

Martra, P. (2026). Rearchitecting LLMs: Structural techniques for efficient models. Manning Publications. ISBN 9781633434332. (Manning Early Access Program)