Rearchitecting LLMs
Structural techniques for efficient models
Affiliate link: I may earn a commission on purchases.
About the book
A practical guide to turning large pre-trained language models into small, efficient models specialized for a domain. Instead of treating a model as a black box, the book works on its architecture: removing the layers and neurons that do not contribute to the goal (depth and width pruning), recovering capability through knowledge distillation, and specializing the result with LoRA-based fine-tuning. It also introduces methods of my own, such as fair pruning, which reduces bias at the neuron level, and internal activation interpretation using pair contrastive prompts.
Why I wrote it
It grew out of my own optimization projects, where the techniques were scattered across papers, often without reproducible code. The book unifies them into a coherent, practical pipeline. Read the preface →
Built on published research
The core techniques are similar in spirit to those behind NVIDIA’s Minitron and Mistral’s Ministral model families. The book adapts them to smaller models, limited data and modest GPUs. The main papers behind each chapter are listed below.
Who it’s for
AI, ML and data engineers who know Python and want to go beyond fine-tuning, with some knowledge of PyTorch and curiosity about what happens inside a transformer.
Hands-on
Every chapter comes with notebooks that run on the free tier of Google Colab, using open models such as Llama, Gemma and Qwen.
Chapters published so far
Part 1 · Foundations
- 1 · Why rearchitecting LLMs matters — The case for specialized models over generic LLMs
- 2 · An end-to-end rearchitecting project — Full pipeline: prune and recover
- 3 · A blueprint to modern transformers — GLU architectures, attention, and model internals
Part 2 · Hands-on optimization
- 4 · Building smaller and faster LLMs with depth pruning — Block removal, capturing block importance with Python hooks, evaluating pruning
- 5 · Shaping model architectures via width pruning — GLU neuron selection, data-driven pruning strategies
- 6 · Knowledge recovery through distillation — Recovering capability after structural compression
- 7 · Model specialization — LoRA / DoRA fine-tuning and quantization for domain tasks
- 8 · Attention optimization — KV cache, attention bypass, inference acceleration
- 9 · Dynamic routing with Mixture of Experts — Mixture of Experts (MoE) adapted to SLMs
Part 3 · Beyond the black box
- 10 · Exploring the transformer black box — Activation analysis and behavioral interpretability
More chapters are being written and will be added here as they are published.
Papers behind each chapter
1 · Why rearchitecting LLMs matters
- FineScope: Precision Pruning for Domain-Specialized Large Language Models Using SAE-Guided Self-Data Cultivation
- LLM Pruning and Distillation in Practice: The Minitron Approach
2 · An end-to-end rearchitecting project
3 · A blueprint to modern transformers
- Attention Is All You Need
- Fast Transformer Decoding: One Write-Head is All You Need
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
- GLU Variants Improve Transformer
- Exploring GLU Expansion Ratios: Structured Pruning in Llama-3.2 Models
4 · Building smaller and faster LLMs with depth pruning
- ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
- Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods
5 · Shaping model architectures via width pruning
- Dependency-Aware Semi-Structured Sparsity of GLU Variants in Large Language Models
- CFSP: An Efficient Structured Pruning Framework for LLMs with Coarse-to-Fine Activation Information
- Fragile Knowledge, Robust Instruction-Following: The Width Pruning Dichotomy in Llama-3.2
6 · Knowledge recovery through distillation
- Distillation Dynamics: Towards Understanding Feature-Based Distillation in Vision Transformers
- DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs
- LLM Pruning and Distillation in Practice: The Minitron Approach
- Ministral 3
7 · Model specialization
- LoRA: Low-Rank Adaptation of Large Language Models
- QLoRA: Efficient Finetuning of Quantized LLMs
- DoRA: Weight-Decomposed Low-Rank Adaptation
- Fine-Tuning LLMs on Small Medical Datasets: Text Classification and Normalization Effectiveness on Cardiology reports and Discharge records
8 · Attention optimization
9 · Dynamic routing with Mixture of Experts
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- Mixtral of Experts
- Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints
10 · Exploring the transformer black box
Companion resources
- Code and notebooks
- Hands-on labs (GitHub Discussions)
- Companion models (Hugging Face collection)
- OptiPFair library
- Related research: Fairness Pruning, arXiv:2607.28319
Cite the book
Martra, P. (2026). Rearchitecting LLMs: Structural techniques for efficient models. Manning Publications. ISBN 9781633434332. (Manning Early Access Program)