PLAIN ENGLISH TECHNICAL BOOK
How a Transformer Produces One Token
Follow one prompt through tokenisation, attention, inference, model formats, and a working provider implementation.
read the book
PART 2 OF 9
Attention Variants, LM Heads & Tokenisation
MQA, GQA, FlashAttention, weight tying, BPE vs SentencePiece — the implementation details.
read