← All contributors
author

Dinesh

@dinesh

Systems engineer learning ML systems from the ground up and writing it down along the way. Interested in how the fundamentals assemble into the models and infrastructure we actually run.

4articles
Architecture · Training Systemsfocus

4 articles

How a Transformer Really Works: Attention, the KV Cache, and Why Inference Eats Memory

Architecture

A from-scratch tour of what's actually inside an LLM: how a transformer turns tokens into predictions, what Query, Key, and Value really mean, and how generating text one token at a time builds the KV cache — the growing pool of memory that makes inference so expensive.

Jul 21, 202610 min

Every Mask in a Transformer, Untangled

Training Systems

The word "mask" means at least four unrelated things in deep learning — what a token can see, what counts toward the loss, what is hidden to create a task, and what is randomly dropped. One field guide to all of them, with why each exists and what breaks without it.

Jul 20, 20269 min

Intuitive Guide to LoRA: Fine-Tuning a Model by 0.2% of weights

Training Systems

You don't need a massive tech budget or a cluster of high-end GPUs to train your own AI. LoRA allows developers to fine-tune giant models right on a standard laptop. Here is the zero-jargon, first-principles explanation of the clever shortcut that leveled the playing field.

Jul 19, 202612 min

Neural Networks From Zero: From a Single Number to a Billion Parameters

Architecture

A neural network never sees a word, an image, or a sound — only a list of numbers. Starting from that one fact and a single neuron, this guide builds the whole machine: how any input becomes numbers, why weights, biases, and activations each exist, and how neurons stack into layers and layers into a model.

Jul 12, 202614 min