Published onAugust 4, 2026Transformers & Attention (Serret 2026) — Paper Study Notespaper-studytransformersattentionllm-architectureStudy notes on Serret's intro to Transformers/attention for applied mathematicians — tokenization, kernel view of attention, MHA, encoder/decoder, KV caching, GQA, MLA.