{"id":1034,"date":"2026-07-04T16:56:11","date_gmt":"2026-07-04T16:56:11","guid":{"rendered":"https:\/\/learnerbox.net\/blog\/?p=1034"},"modified":"2026-07-05T14:59:20","modified_gmt":"2026-07-05T14:59:20","slug":"how-large-language-models-work-part-2-the-attention-mechanism","status":"publish","type":"post","link":"https:\/\/learnerbox.net\/blog\/ai-theory\/how-large-language-models-work-part-2-the-attention-mechanism\/","title":{"rendered":"How Large Language Models Work \u2014 Part 2: The Attention Mechanism"},"content":{"rendered":"\n<h6 class=\"wp-block-heading\"><em><mark style=\"background-color:rgba(0, 0, 0, 0)\" class=\"has-inline-color has-vivid-cyan-blue-color\">This is Part 2 of a five-part series on the internal mechanics of large language models. <a href=\"https:\/\/learnerbox.net\/blog\/ai-theory\/how-large-language-models-work-part-1-tokenization-and-embeddings\/\">Part 1<\/a> covered tokenization, embeddings, and positional encoding. Part 2 derives the self-attention mechanism from first principles, covers multi-head attention, and examines the KV cache.<\/mark><\/em><\/h6>\n\n\n\n<h4 class=\"wp-block-heading\">The Problem Attention Solves<\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">At the end of Part 1, we had a matrix <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>X<\/mi><mo>\u2208<\/mo><msup><mi mathvariant=\"double-struck\">R<\/mi><mrow><mi>n<\/mi><mo>\u00d7<\/mo><mi>d<\/mi><\/mrow><\/msup><\/mrow><annotation encoding=\"application\/x-tex\">X \\in \\mathbb{R}^{n \\times d}<\/annotation><\/semantics><\/math>, which was a one <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>d<\/mi><\/mrow><annotation encoding=\"application\/x-tex\">d<\/annotation><\/semantics><\/math>d-dimensional vector per token, encoding both semantic identity and position. The first transformer layer receives this matrix and must do something critical: allow each token to incorporate information from every other token in the sequence before passing its updated representation to the feed-forward network.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This is the problem that self-attention solves. Earlier sequence models such as LSTMs and GRUs processed tokens sequentially, which meant that a word at position 1 had to wait for its influence to propagate through every intermediate hidden state to reach position 512. Long-range dependencies were difficult to learn because the gradient had to flow through hundreds of recurrent steps. Attention eliminates this bottleneck entirely by allowing any token to attend directly to any other token in a single operation, regardless of distance.<\/p>\n\n\n\n<h4 class=\"wp-block-heading\">Queries, Keys, and Values<\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">The first step of self-attention is to project each token&#8217;s embedding into three separate representations using learned weight matrices:<math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\" display=\"block\"><semantics><mrow><mi>Q<\/mi><mo>=<\/mo><mi>X<\/mi><msup><mi>W<\/mi><mi>Q<\/mi><\/msup><mo separator=\"true\">,<\/mo><mspace width=\"1em\"><\/mspace><mi>K<\/mi><mo>=<\/mo><mi>X<\/mi><msup><mi>W<\/mi><mi>K<\/mi><\/msup><mo separator=\"true\">,<\/mo><mspace width=\"1em\"><\/mspace><mi>V<\/mi><mo>=<\/mo><mi>X<\/mi><msup><mi>W<\/mi><mi>V<\/mi><\/msup><\/mrow><annotation encoding=\"application\/x-tex\">Q = XW^Q, \\quad K = XW^K, \\quad V = XW^V<\/annotation><\/semantics><\/math>where <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><msup><mi>W<\/mi><mi>Q<\/mi><\/msup><mo separator=\"true\">,<\/mo><msup><mi>W<\/mi><mi>K<\/mi><\/msup><mo separator=\"true\">,<\/mo><msup><mi>W<\/mi><mi>V<\/mi><\/msup><mo>\u2208<\/mo><msup><mi mathvariant=\"double-struck\">R<\/mi><mrow><mi>d<\/mi><mo>\u00d7<\/mo><msub><mi>d<\/mi><mi>k<\/mi><\/msub><\/mrow><\/msup><\/mrow><annotation encoding=\"application\/x-tex\">W^Q, W^K, W^V \\in \\mathbb{R}^{d \\times d_k}<\/annotation><\/semantics><\/math>\u200b are learned projection matrices, and <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><msub><mi>d<\/mi><mi>k<\/mi><\/msub><\/mrow><annotation encoding=\"application\/x-tex\">d_k<\/annotation><\/semantics><\/math>\u200b is the dimension of the query and key space (typically <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><msub><mi>d<\/mi><mi>k<\/mi><\/msub><mo>=<\/mo><mi>d<\/mi><mi mathvariant=\"normal\">\/<\/mi><mi>h<\/mi><\/mrow><annotation encoding=\"application\/x-tex\">d_k = d \/ h<\/annotation><\/semantics><\/math> where <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>h<\/mi><\/mrow><annotation encoding=\"application\/x-tex\">h<\/annotation><\/semantics><\/math> is the number of attention heads, more on that shortly).<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The resulting matrices <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>Q<\/mi><mo separator=\"true\">,<\/mo><mi>K<\/mi><mo separator=\"true\">,<\/mo><mi>V<\/mi><mo>\u2208<\/mo><msup><mi mathvariant=\"double-struck\">R<\/mi><mrow><mi>n<\/mi><mo>\u00d7<\/mo><msub><mi>d<\/mi><mi>k<\/mi><\/msub><\/mrow><\/msup><\/mrow><annotation encoding=\"application\/x-tex\">Q, K, V \\in \\mathbb{R}^{n \\times d_k}<\/annotation><\/semantics><\/math> are called the <strong>Query<\/strong>, <strong>Key<\/strong>, and <strong>Value<\/strong> matrices respectively.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The intuition behind this decomposition is often explained through an analogy: think of each token&#8217;s query vector as a question it is asking (&#8220;what context do I need?&#8221;), each token&#8217;s key vector as an advertisement of its content (&#8220;here is what I contain&#8221;), and each token&#8217;s value vector as the actual information it contributes when attended to (&#8220;here is what I give you if you attend to me&#8221;). The query-key interaction determines <em>how much<\/em> each token attends to every other; the values are what actually gets aggregated.<\/p>\n\n\n\n<h4 class=\"wp-block-heading\">Scaled Dot-Product Attention<\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">Given <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>Q<\/mi><\/mrow><annotation encoding=\"application\/x-tex\">Q<\/annotation><\/semantics><\/math>, <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>K<\/mi><\/mrow><annotation encoding=\"application\/x-tex\">K<\/annotation><\/semantics><\/math>, and <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>V<\/mi><\/mrow><annotation encoding=\"application\/x-tex\">V<\/annotation><\/semantics><\/math>, the attention output is computed as:<math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\" display=\"block\"><semantics><mrow><mtext>Attention<\/mtext><mo stretchy=\"false\">(<\/mo><mi>Q<\/mi><mo separator=\"true\">,<\/mo><mi>K<\/mi><mo separator=\"true\">,<\/mo><mi>V<\/mi><mo stretchy=\"false\">)<\/mo><mo>=<\/mo><mtext>softmax<\/mtext><mtext>\u2009\u2063<\/mtext><mrow><mo fence=\"true\">(<\/mo><mfrac><mrow><mi>Q<\/mi><msup><mi>K<\/mi><mi mathvariant=\"normal\">\u22a4<\/mi><\/msup><\/mrow><msqrt><msub><mi>d<\/mi><mi>k<\/mi><\/msub><\/msqrt><\/mfrac><mo fence=\"true\">)<\/mo><\/mrow><mi>V<\/mi><\/mrow><annotation encoding=\"application\/x-tex\">\\text{Attention}(Q, K, V) = \\text{softmax}\\!\\left(\\frac{QK^\\top}{\\sqrt{d_k}}\\right) V<\/annotation><\/semantics><\/math>Let&#8217;s unpack this step by step.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Step 1 \u2014 Dot products:<\/strong> <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>Q<\/mi><msup><mi>K<\/mi><mi mathvariant=\"normal\">\u22a4<\/mi><\/msup><mo>\u2208<\/mo><msup><mi mathvariant=\"double-struck\">R<\/mi><mrow><mi>n<\/mi><mo>\u00d7<\/mo><mi>n<\/mi><\/mrow><\/msup><\/mrow><annotation encoding=\"application\/x-tex\">QK^\\top \\in \\mathbb{R}^{n \\times n}<\/annotation><\/semantics><\/math> is a matrix of raw attention scores. Entry <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mo stretchy=\"false\">(<\/mo><mi>i<\/mi><mo separator=\"true\">,<\/mo><mi>j<\/mi><mo stretchy=\"false\">)<\/mo><\/mrow><annotation encoding=\"application\/x-tex\">(i, j)<\/annotation><\/semantics><\/math> is the dot product between the query vector of token <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>i<\/mi><\/mrow><annotation encoding=\"application\/x-tex\">i<\/annotation><\/semantics><\/math> and the key vector of token <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>j<\/mi><\/mrow><annotation encoding=\"application\/x-tex\">j<\/annotation><\/semantics><\/math>:<math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\" display=\"block\"><semantics><mrow><msub><mi>s<\/mi><mrow><mi>i<\/mi><mi>j<\/mi><\/mrow><\/msub><mo>=<\/mo><msub><mi mathvariant=\"bold\">q<\/mi><mi>i<\/mi><\/msub><mo>\u22c5<\/mo><msub><mi mathvariant=\"bold\">k<\/mi><mi>j<\/mi><\/msub><\/mrow><annotation encoding=\"application\/x-tex\">s_{ij} = \\mathbf{q}_i \\cdot \\mathbf{k}_j<\/annotation><\/semantics><\/math>A high dot product means token <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>i<\/mi><\/mrow><annotation encoding=\"application\/x-tex\">i<\/annotation><\/semantics><\/math> finds token <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>j<\/mi><\/mrow><annotation encoding=\"application\/x-tex\">j<\/annotation><\/semantics><\/math> highly relevant. This is the mechanism through which, for example, a pronoun &#8220;she&#8221; can attend strongly to the noun &#8220;Alice&#8221; it refers to three sentences earlier.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Step 2 \u2014 Scaling:<\/strong> The dot products are divided by <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><msqrt><msub><mi>d<\/mi><mi>k<\/mi><\/msub><\/msqrt><\/mrow><annotation encoding=\"application\/x-tex\">\\sqrt{d_k}<\/annotation><\/semantics><\/math>\u200b\u200b. Without this scaling, when <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><msub><mi>d<\/mi><mi>k<\/mi><\/msub><\/mrow><annotation encoding=\"application\/x-tex\">d_k<\/annotation><\/semantics><\/math>\u200b is large, the dot products grow large in magnitude, pushing the softmax into regions where gradients become vanishingly small \u2014 a variant of the vanishing gradient problem. The <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><msqrt><msub><mi>d<\/mi><mi>k<\/mi><\/msub><\/msqrt><\/mrow><annotation encoding=\"application\/x-tex\">\\sqrt{d_k}<\/annotation><\/semantics><\/math>\u200b\u200b factor keeps the variance of the dot products approximately constant regardless of model dimensionality.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Step 3 \u2014 Causal masking (decoder-only models):<\/strong> In autoregressive models like GPT, token <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>i<\/mi><\/mrow><annotation encoding=\"application\/x-tex\">i<\/annotation><\/semantics><\/math> must not attend to any token <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>j<\/mi><mo>&gt;<\/mo><mi>i<\/mi><\/mrow><annotation encoding=\"application\/x-tex\">j &gt; i<\/annotation><\/semantics><\/math> meaning that it cannot look into the future. Before applying softmax, a mask is added:<math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\" display=\"block\"><semantics><mrow><msub><mi>s<\/mi><mrow><mi>i<\/mi><mi>j<\/mi><\/mrow><\/msub><mo>\u2190<\/mo><msub><mi>s<\/mi><mrow><mi>i<\/mi><mi>j<\/mi><\/mrow><\/msub><mo>+<\/mo><msub><mi>m<\/mi><mrow><mi>i<\/mi><mi>j<\/mi><\/mrow><\/msub><mo separator=\"true\">,<\/mo><mspace width=\"1em\"><\/mspace><msub><mi>m<\/mi><mrow><mi>i<\/mi><mi>j<\/mi><\/mrow><\/msub><mo>=<\/mo><mrow><mo fence=\"true\">{<\/mo><mtable rowspacing=\"0.36em\" columnalign=\"left left\" columnspacing=\"1em\"><mtr><mtd><mstyle scriptlevel=\"0\" displaystyle=\"false\"><mn>0<\/mn><\/mstyle><\/mtd><mtd><mstyle scriptlevel=\"0\" displaystyle=\"false\"><mrow><mtext>if&nbsp;<\/mtext><mi>j<\/mi><mo>\u2264<\/mo><mi>i<\/mi><\/mrow><\/mstyle><\/mtd><\/mtr><mtr><mtd><mstyle scriptlevel=\"0\" displaystyle=\"false\"><mrow><mo>\u2212<\/mo><mi mathvariant=\"normal\">\u221e<\/mi><\/mrow><\/mstyle><\/mtd><mtd><mstyle scriptlevel=\"0\" displaystyle=\"false\"><mrow><mtext>if&nbsp;<\/mtext><mi>j<\/mi><mo>&gt;<\/mo><mi>i<\/mi><\/mrow><\/mstyle><\/mtd><\/mtr><\/mtable><\/mrow><\/mrow><annotation encoding=\"application\/x-tex\">s_{ij} \\leftarrow s_{ij} + m_{ij}, \\quad m_{ij} = \\begin{cases} 0 &amp; \\text{if } j \\leq i \\\\ -\\infty &amp; \\text{if } j &gt; i \\end{cases}<\/annotation><\/semantics><\/math>The <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mo>\u2212<\/mo><mi mathvariant=\"normal\">\u221e<\/mi><\/mrow><annotation encoding=\"application\/x-tex\">-\\infty<\/annotation><\/semantics><\/math> values become zero after softmax, effectively zeroing out future positions.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Step 4 \u2014 Softmax:<\/strong> Each row of the masked score matrix is passed through softmax:<math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\" display=\"block\"><semantics><mrow><msub><mi>\u03b1<\/mi><mrow><mi>i<\/mi><mi>j<\/mi><\/mrow><\/msub><mo>=<\/mo><mfrac><mrow><mi>exp<\/mi><mo>\u2061<\/mo><mo stretchy=\"false\">(<\/mo><msub><mi>s<\/mi><mrow><mi>i<\/mi><mi>j<\/mi><\/mrow><\/msub><mo stretchy=\"false\">)<\/mo><\/mrow><mrow><munderover><mo>\u2211<\/mo><mrow><mi>k<\/mi><mo>=<\/mo><mn>1<\/mn><\/mrow><mi>n<\/mi><\/munderover><mi>exp<\/mi><mo>\u2061<\/mo><mo stretchy=\"false\">(<\/mo><msub><mi>s<\/mi><mrow><mi>i<\/mi><mi>k<\/mi><\/mrow><\/msub><mo stretchy=\"false\">)<\/mo><\/mrow><\/mfrac><\/mrow><annotation encoding=\"application\/x-tex\">\\alpha_{ij} = \\frac{\\exp(s_{ij})}{\\sum_{k=1}^{n} \\exp(s_{ik})}<\/annotation><\/semantics><\/math>The resulting matrix <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>A<\/mi><mo>\u2208<\/mo><msup><mi mathvariant=\"double-struck\">R<\/mi><mrow><mi>n<\/mi><mo>\u00d7<\/mo><mi>n<\/mi><\/mrow><\/msup><\/mrow><annotation encoding=\"application\/x-tex\">A \\in \\mathbb{R}^{n \\times n}<\/annotation><\/semantics><\/math> is the <strong>attention weight matrix<\/strong>. Each row sums to 1 and can be interpreted as a probability distribution over the sequence: the probability that token <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>i<\/mi><\/mrow><annotation encoding=\"application\/x-tex\">i<\/annotation><\/semantics><\/math> attends to each position.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Step 5 \u2014 Value aggregation:<\/strong> The output for token <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>i<\/mi><\/mrow><annotation encoding=\"application\/x-tex\">i<\/annotation><\/semantics><\/math> is a weighted sum of all value vectors:<math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\" display=\"block\"><semantics><mrow><msub><mi mathvariant=\"bold\">o<\/mi><mi>i<\/mi><\/msub><mo>=<\/mo><munderover><mo>\u2211<\/mo><mrow><mi>j<\/mi><mo>=<\/mo><mn>1<\/mn><\/mrow><mi>n<\/mi><\/munderover><msub><mi>\u03b1<\/mi><mrow><mi>i<\/mi><mi>j<\/mi><\/mrow><\/msub><msub><mi mathvariant=\"bold\">v<\/mi><mi>j<\/mi><\/msub><\/mrow><annotation encoding=\"application\/x-tex\">\\mathbf{o}_i = \\sum_{j=1}^{n} \\alpha_{ij} \\mathbf{v}_j<\/annotation><\/semantics><\/math>In matrix form: <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>O<\/mi><mo>=<\/mo><mi>A<\/mi><mi>V<\/mi><mo>\u2208<\/mo><msup><mi mathvariant=\"double-struck\">R<\/mi><mrow><mi>n<\/mi><mo>\u00d7<\/mo><msub><mi>d<\/mi><mi>k<\/mi><\/msub><\/mrow><\/msup><\/mrow><annotation encoding=\"application\/x-tex\">O = AV \\in \\mathbb{R}^{n \\times d_k}<\/annotation><\/semantics><\/math>\u200b. Each output vector is a contextualised representation of its token, so that the same word &#8220;bank&#8221; will produce a different output vector in &#8220;river bank&#8221; versus &#8220;central bank&#8221; because its attention weights will be distributed differently across the surrounding context.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"934\" src=\"https:\/\/learnerbox.net\/blog\/wp-content\/uploads\/2026\/07\/scaled_dot_product_attention-1024x934.png\" alt=\"\" class=\"wp-image-1036\" title=\"\" srcset=\"https:\/\/learnerbox.net\/blog\/wp-content\/uploads\/2026\/07\/scaled_dot_product_attention-1024x934.png 1024w, https:\/\/learnerbox.net\/blog\/wp-content\/uploads\/2026\/07\/scaled_dot_product_attention-300x274.png 300w, https:\/\/learnerbox.net\/blog\/wp-content\/uploads\/2026\/07\/scaled_dot_product_attention-768x700.png 768w, https:\/\/learnerbox.net\/blog\/wp-content\/uploads\/2026\/07\/scaled_dot_product_attention-1536x1400.png 1536w, https:\/\/learnerbox.net\/blog\/wp-content\/uploads\/2026\/07\/scaled_dot_product_attention-2048x1867.png 2048w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<h4 class=\"wp-block-heading\">Multi-Head Attention<\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">A single attention head computes one set of query-key-value interactions. But different aspects of meaning may require different attention patterns simultaneously: a token might need to attend to its syntactic head, its semantic antecedent, and its positional neighbours all at once. Multi-head attention runs <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>h<\/mi><\/mrow><annotation encoding=\"application\/x-tex\">h<\/annotation><\/semantics><\/math> attention operations in parallel:<math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\" display=\"block\"><semantics><mrow><msub><mtext>head<\/mtext><mi>i<\/mi><\/msub><mo>=<\/mo><mtext>Attention<\/mtext><mo stretchy=\"false\">(<\/mo><mi>Q<\/mi><msubsup><mi>W<\/mi><mi>i<\/mi><mi>Q<\/mi><\/msubsup><mo separator=\"true\">,<\/mo><mtext>&nbsp;<\/mtext><mi>K<\/mi><msubsup><mi>W<\/mi><mi>i<\/mi><mi>K<\/mi><\/msubsup><mo separator=\"true\">,<\/mo><mtext>&nbsp;<\/mtext><mi>V<\/mi><msubsup><mi>W<\/mi><mi>i<\/mi><mi>V<\/mi><\/msubsup><mo stretchy=\"false\">)<\/mo><\/mrow><annotation encoding=\"application\/x-tex\">\\text{head}_i = \\text{Attention}(QW_i^Q,\\ KW_i^K,\\ VW_i^V)<\/annotation><\/semantics><\/math> <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\" display=\"block\"><semantics><mrow><mtext>MultiHead<\/mtext><mo stretchy=\"false\">(<\/mo><mi>Q<\/mi><mo separator=\"true\">,<\/mo><mi>K<\/mi><mo separator=\"true\">,<\/mo><mi>V<\/mi><mo stretchy=\"false\">)<\/mo><mo>=<\/mo><mtext>Concat<\/mtext><mo stretchy=\"false\">(<\/mo><msub><mtext>head<\/mtext><mn>1<\/mn><\/msub><mo separator=\"true\">,<\/mo><mo>\u2026<\/mo><mo separator=\"true\">,<\/mo><msub><mtext>head<\/mtext><mi>h<\/mi><\/msub><mo stretchy=\"false\">)<\/mo><msup><mi>W<\/mi><mi>O<\/mi><\/msup><\/mrow><annotation encoding=\"application\/x-tex\">\\text{MultiHead}(Q, K, V) = \\text{Concat}(\\text{head}_1, \\ldots, \\text{head}_h)W^O<\/annotation><\/semantics><\/math>where <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><msubsup><mi>W<\/mi><mi>i<\/mi><mi>Q<\/mi><\/msubsup><mo separator=\"true\">,<\/mo><msubsup><mi>W<\/mi><mi>i<\/mi><mi>K<\/mi><\/msubsup><mo separator=\"true\">,<\/mo><msubsup><mi>W<\/mi><mi>i<\/mi><mi>V<\/mi><\/msubsup><mo>\u2208<\/mo><msup><mi mathvariant=\"double-struck\">R<\/mi><mrow><mi>d<\/mi><mo>\u00d7<\/mo><msub><mi>d<\/mi><mi>k<\/mi><\/msub><\/mrow><\/msup><\/mrow><annotation encoding=\"application\/x-tex\">W_i^Q, W_i^K, W_i^V \\in \\mathbb{R}^{d \\times d_k}<\/annotation><\/semantics><\/math>\u200b are per-head projection matrices, each head operates in a <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><msub><mi>d<\/mi><mi>k<\/mi><\/msub><mo>=<\/mo><mi>d<\/mi><mi mathvariant=\"normal\">\/<\/mi><mi>h<\/mi><\/mrow><annotation encoding=\"application\/x-tex\">d_k = d\/h<\/annotation><\/semantics><\/math> dimensional subspace, and <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><msup><mi>W<\/mi><mi>O<\/mi><\/msup><mo>\u2208<\/mo><msup><mi mathvariant=\"double-struck\">R<\/mi><mrow><mi>d<\/mi><mo>\u00d7<\/mo><mi>d<\/mi><\/mrow><\/msup><\/mrow><annotation encoding=\"application\/x-tex\">W^O \\in \\mathbb{R}^{d \\times d}<\/annotation><\/semantics><\/math> is a final output projection that mixes the concatenated heads back into the full <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>d<\/mi><\/mrow><annotation encoding=\"application\/x-tex\">d<\/annotation><\/semantics><\/math>-dimensional space.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">In GPT-3, <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>d<\/mi><mo>=<\/mo><mn>12288<\/mn><\/mrow><annotation encoding=\"application\/x-tex\">d = 12288<\/annotation><\/semantics><\/math> and <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>h<\/mi><mo>=<\/mo><mn>96<\/mn><\/mrow><annotation encoding=\"application\/x-tex\">h = 96<\/annotation><\/semantics><\/math>, so each head operates in a <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><msub><mi>d<\/mi><mi>k<\/mi><\/msub><mo>=<\/mo><mn>128<\/mn><\/mrow><annotation encoding=\"application\/x-tex\">d_k = 128<\/annotation><\/semantics><\/math> dimensional subspace. The 96 heads learn to specialize: empirical work in mechanistic interpretability has identified heads that track syntactic subject-verb agreement, heads that copy tokens from earlier in context, heads that attend to the most recent noun phrase, and heads that implement induction \u2014 recognizing when a pattern seen earlier in the context is repeating.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The computational cost of full self-attention is <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>O<\/mi><mo stretchy=\"false\">(<\/mo><msup><mi>n<\/mi><mn>2<\/mn><\/msup><mi>d<\/mi><mo stretchy=\"false\">)<\/mo><\/mrow><annotation encoding=\"application\/x-tex\">O(n^2 d)<\/annotation><\/semantics><\/math>, or quadratic in sequence length. For a 128K-token context window, the attention matrix alone has <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><msup><mn>128000<\/mn><mn>2<\/mn><\/msup><mo>\u2248<\/mo><mn>16<\/mn><\/mrow><annotation encoding=\"application\/x-tex\">128000^2 \\approx 16<\/annotation><\/semantics><\/math> billion entries, making naive computation prohibitively expensive. Efficient attention variants such as Flash Attention, Sparse Attention, Sliding Window Attention address this scaling problem, which we&#8217;ll cover in Part 3 alongside the full transformer block.<\/p>\n\n\n\n<h4 class=\"wp-block-heading\">The KV Cache<\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">During inference, a decoder-only model generates one token at a time. At step <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>t<\/mi><\/mrow><annotation encoding=\"application\/x-tex\">t<\/annotation><\/semantics><\/math>, the model processes the full sequence <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mo stretchy=\"false\">[<\/mo><msub><mi>t<\/mi><mn>1<\/mn><\/msub><mo separator=\"true\">,<\/mo><mo>\u2026<\/mo><mo separator=\"true\">,<\/mo><msub><mi>t<\/mi><mi>t<\/mi><\/msub><mo stretchy=\"false\">]<\/mo><\/mrow><annotation encoding=\"application\/x-tex\">[t_1, \\ldots, t_t]<\/annotation><\/semantics><\/math> and predicts <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><msub><mi>t<\/mi><mrow><mi>t<\/mi><mo>+<\/mo><mn>1<\/mn><\/mrow><\/msub><\/mrow><annotation encoding=\"application\/x-tex\">t_{t+1}<\/annotation><\/semantics><\/math>\u200b. Without optimization, this would require recomputing the key and value matrices for all previous tokens at every step so thar cost grows as <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>O<\/mi><mo stretchy=\"false\">(<\/mo><msup><mi>t<\/mi><mn>2<\/mn><\/msup><mo stretchy=\"false\">)<\/mo><\/mrow><annotation encoding=\"application\/x-tex\">O(t^2)<\/annotation><\/semantics><\/math> per generation.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The <strong>KV cache<\/strong> eliminates this redundancy. Since the key and value representations of tokens <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><msub><mi>t<\/mi><mn>1<\/mn><\/msub><mo separator=\"true\">,<\/mo><mo>\u2026<\/mo><mo separator=\"true\">,<\/mo><msub><mi>t<\/mi><mrow><mi>t<\/mi><mo>\u2212<\/mo><mn>1<\/mn><\/mrow><\/msub><\/mrow><annotation encoding=\"application\/x-tex\">t_1, \\ldots, t_{t-1}<\/annotation><\/semantics><\/math>\u200bdo not change between steps (they depend only on those tokens&#8217; positions and embeddings, which are fixed once generated), they can be cached in memory and reused. At step <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>t<\/mi><\/mrow><annotation encoding=\"application\/x-tex\">t<\/annotation><\/semantics><\/math>, only the new token&#8217;s <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>Q<\/mi><\/mrow><annotation encoding=\"application\/x-tex\">Q<\/annotation><\/semantics><\/math>, <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>K<\/mi><\/mrow><annotation encoding=\"application\/x-tex\">K<\/annotation><\/semantics><\/math>, <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>V<\/mi><\/mrow><annotation encoding=\"application\/x-tex\">V<\/annotation><\/semantics><\/math> need to be computed; the cached <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>K<\/mi><\/mrow><annotation encoding=\"application\/x-tex\">K<\/annotation><\/semantics><\/math> and <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>V<\/mi><\/mrow><annotation encoding=\"application\/x-tex\">V<\/annotation><\/semantics><\/math> matrices from previous steps are concatenated and the attention is computed against the full cached sequence.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This reduces per-step inference cost from <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>O<\/mi><mo stretchy=\"false\">(<\/mo><msup><mi>t<\/mi><mn>2<\/mn><\/msup><mo stretchy=\"false\">)<\/mo><\/mrow><annotation encoding=\"application\/x-tex\">O(t^2)<\/annotation><\/semantics><\/math> to <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>O<\/mi><mo stretchy=\"false\">(<\/mo><mi>t<\/mi><mo stretchy=\"false\">)<\/mo><\/mrow><annotation encoding=\"application\/x-tex\">O(t)<\/annotation><\/semantics><\/math> in attention computation, but at the cost of memory. A KV cache for a 128K-token context with 96 heads, 128 layers, and FP16 precision occupies on the order of tens of gigabytes. Managing KV cache memory is one of the primary challenges in deploying frontier models efficiently at scale. It is why inference hardware for long-context models requires far more VRAM than a naive parameter count would suggest.<\/p>\n\n\n\n<h4 class=\"wp-block-heading\">What Attention Actually Learns<\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">It is tempting to describe attention as a mechanism that &#8220;understands&#8221; language. It is more precise, and more useful for engineering purposes, to say that attention is a differentiable, content-based memory retrieval system. The model learns, through gradient descent on the next-token prediction objective, to configure its <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><msup><mi>W<\/mi><mi>Q<\/mi><\/msup><\/mrow><annotation encoding=\"application\/x-tex\">W^Q<\/annotation><\/semantics><\/math>, <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><msup><mi>W<\/mi><mi>K<\/mi><\/msup><\/mrow><annotation encoding=\"application\/x-tex\">W^K<\/annotation><\/semantics><\/math>, <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><msup><mi>W<\/mi><mi>V<\/mi><\/msup><\/mrow><annotation encoding=\"application\/x-tex\">W^V<\/annotation><\/semantics><\/math> matrices such that the right values get retrieved for the right queries.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The mechanism has no built-in notion of syntax, coreference, or meaning. All of that emerges from training dynamics or from the statistical regularities of billions of documents. Understanding this distinction is important for debugging model failures: when an LLM makes a coreference error or loses track of a constraint stated early in a long context, the failure mode is almost always traceable to attention weights that were distributed incorrectly, either because the training distribution underrepresented that pattern, or because the KV cache compression strategy discarded the relevant context.<\/p>\n\n\n\n<h4 class=\"wp-block-heading\">Conclusion<\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">Self-attention is the architectural innovation that made modern LLMs possible. By allowing every token to attend directly to every other token in a single parallelisable matrix operation, it solves the long-range dependency problem that defeated earlier sequential architectures. Multi-head attention extends this by learning multiple independent attention patterns simultaneously, while the KV cache makes autoregressive inference tractable at scale.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Part 3 will take the attention output <math xmlns=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><semantics><mrow><mi>O<\/mi><mo>\u2208<\/mo><msup><mi mathvariant=\"double-struck\">R<\/mi><mrow><mi>n<\/mi><mo>\u00d7<\/mo><mi>d<\/mi><\/mrow><\/msup><\/mrow><annotation encoding=\"application\/x-tex\">O \\in \\mathbb{R}^{n \\times d}<\/annotation><\/semantics><\/math> and trace it through the rest of the transformer block: the feed-forward network, layer normalisation, residual connections, and the architectural choices (decoder-only versus encoder-decoder) that distinguish GPT-style models from BERT-style ones.<\/p>\n\n\n\n<h6 class=\"wp-block-heading\"><em><mark style=\"background-color:rgba(0, 0, 0, 0)\" class=\"has-inline-color has-vivid-cyan-blue-color\">Coming next in the AI Engineering series is <a href=\"https:\/\/learnerbox.net\/blog\/ai-theory\/how-large-language-models-work-part-3-the-transformer-block-and-architecture\/\">Part 3<\/a>: The Transformer Block and Architecture.<\/mark><\/em><\/h6>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n","protected":false},"excerpt":{"rendered":"<p>This is Part 2 of a five-part series on the internal mechanics of large language models. Part 1 covered tokenization, embeddings, and positional encoding. Part 2 derives the self-attention mechanism from first principles, covers multi-head attention, and examines the KV cache. The Problem Attention Solves At the end of Part 1, we had a matrix X\u2208Rn\u00d7dX \\in \\mathbb{R}^{n \\times d}, which was a one ddd-dimensional vector per token, encoding both semantic identity and position. The first transformer layer receives this matrix and must do something critical: allow each token to incorporate information from every other token in the sequence before passing its updated representation to the feed-forward network. This is the problem that self-attention solves. Earlier sequence models such as LSTMs and GRUs processed tokens sequentially, which meant that a word at position 1 had to wait for its influence to propagate through every intermediate hidden state to reach position 512. Long-range dependencies were difficult to learn because the gradient had to flow through hundreds of recurrent steps. Attention eliminates this bottleneck entirely by allowing any token to attend directly to any other token in a single operation, regardless of distance. Queries, Keys, and Values The first step of self-attention is to project each token&#8217;s embedding into three separate representations using learned weight matrices:Q=XWQ,K=XWK,V=XWVQ = XW^Q, \\quad K = XW^K, \\quad V = XW^Vwhere WQ,WK,WV\u2208Rd\u00d7dkW^Q, W^K, W^V \\in \\mathbb{R}^{d \\times d_k}\u200b are learned projection matrices, and dkd_k\u200b is the dimension of the query and key space (typically dk=d\/hd_k = d \/ h where hh is the number of attention heads, more on that shortly). The resulting matrices Q,K,V\u2208Rn\u00d7dkQ, K, V \\in \\mathbb{R}^{n \\times d_k} are called the Query, Key, and Value matrices respectively. The intuition behind this decomposition is often explained through an analogy: think of each token&#8217;s query vector as a question it is asking (&#8220;what context do I need?&#8221;), each token&#8217;s key vector as an advertisement of its content (&#8220;here is what I contain&#8221;), and each token&#8217;s value vector as the actual information it contributes when attended to (&#8220;here is what I give you if you attend to me&#8221;). The query-key interaction determines how much each token attends to every other; the values are what actually gets aggregated. Scaled Dot-Product Attention Given QQ, KK, and VV, the attention output is computed as:Attention(Q,K,V)=softmax\u2009\u2063(QK\u22a4dk)V\\text{Attention}(Q, K, V) = \\text{softmax}\\!\\left(\\frac{QK^\\top}{\\sqrt{d_k}}\\right) VLet&#8217;s unpack this step by step. Step 1 \u2014 Dot products: QK\u22a4\u2208Rn\u00d7nQK^\\top \\in \\mathbb{R}^{n \\times n} is a matrix of raw attention scores. Entry (i,j)(i, j) is the dot product between the query vector of token ii and the key vector of token jj:sij=qi\u22c5kjs_{ij} = \\mathbf{q}_i \\cdot \\mathbf{k}_jA high dot product means token ii finds token jj highly relevant. This is the mechanism through which, for example, a pronoun &#8220;she&#8221; can attend strongly to the noun &#8220;Alice&#8221; it refers to three sentences earlier. Step 2 \u2014 Scaling: The dot products are divided by dk\\sqrt{d_k}\u200b\u200b. Without this scaling, when dkd_k\u200b is large, the dot products grow large in magnitude, pushing the softmax into regions where gradients become vanishingly small \u2014 a variant of the vanishing gradient problem. The dk\\sqrt{d_k}\u200b\u200b factor keeps the variance of the dot products approximately constant regardless of model dimensionality. Step 3 \u2014 Causal masking (decoder-only models): In autoregressive models like GPT, token ii must not attend to any token j&gt;ij &gt; i meaning that it cannot look into the future. Before applying softmax, a mask is added:sij\u2190sij+mij,mij={0if&nbsp;j\u2264i\u2212\u221eif&nbsp;j&gt;is_{ij} \\leftarrow s_{ij} + m_{ij}, \\quad m_{ij} = \\begin{cases} 0 &amp; \\text{if } j \\leq i \\\\ -\\infty &amp; \\text{if } j &gt; i \\end{cases}The \u2212\u221e-\\infty values become zero after softmax, effectively zeroing out future positions. Step 4 \u2014 Softmax: Each row of the masked score matrix is passed through softmax:\u03b1ij=exp\u2061(sij)\u2211k=1nexp\u2061(sik)\\alpha_{ij} = \\frac{\\exp(s_{ij})}{\\sum_{k=1}^{n} \\exp(s_{ik})}The resulting matrix A\u2208Rn\u00d7nA \\in \\mathbb{R}^{n \\times n} is the attention weight matrix. Each row sums to 1 and can be interpreted as a probability distribution over the sequence: the probability that token ii attends to each position. Step 5 \u2014 Value aggregation: The output for token ii is a weighted sum of all value vectors:oi=\u2211j=1n\u03b1ijvj\\mathbf{o}_i = \\sum_{j=1}^{n} \\alpha_{ij} \\mathbf{v}_jIn matrix form: O=AV\u2208Rn\u00d7dkO = AV \\in \\mathbb{R}^{n \\times d_k}\u200b. Each output vector is a contextualised representation of its token, so that the same word &#8220;bank&#8221; will produce a different output vector in &#8220;river bank&#8221; versus &#8220;central bank&#8221; because its attention weights will be distributed differently across the surrounding context. Multi-Head Attention A single attention head computes one set of query-key-value interactions. But different aspects of meaning may require different attention patterns simultaneously: a token might need to attend to its syntactic head, its semantic antecedent, and its positional neighbours all at once. Multi-head attention runs hh attention operations in parallel:headi=Attention(QWiQ,&nbsp;KWiK,&nbsp;VWiV)\\text{head}_i = \\text{Attention}(QW_i^Q,\\ KW_i^K,\\ VW_i^V) MultiHead(Q,K,V)=Concat(head1,\u2026,headh)WO\\text{MultiHead}(Q, K, V) = \\text{Concat}(\\text{head}_1, \\ldots, \\text{head}_h)W^Owhere WiQ,WiK,WiV\u2208Rd\u00d7dkW_i^Q, W_i^K, W_i^V \\in \\mathbb{R}^{d \\times d_k}\u200b are per-head projection matrices, each head operates in a dk=d\/hd_k = d\/h dimensional subspace, and WO\u2208Rd\u00d7dW^O \\in \\mathbb{R}^{d \\times d} is a final output projection that mixes the concatenated heads back into the full dd-dimensional space. In GPT-3, d=12288d = 12288 and h=96h = 96, so each head operates in a dk=128d_k = 128 dimensional subspace. The 96 heads learn to specialize: empirical work in mechanistic interpretability has identified heads that track syntactic subject-verb agreement, heads that copy tokens from earlier in context, heads that attend to the most recent noun phrase, and heads that implement induction \u2014 recognizing when a pattern seen earlier in the context is repeating. The computational cost of full self-attention is O(n2d)O(n^2 d), or quadratic in sequence length. For a 128K-token context window, the attention matrix alone has 1280002\u224816128000^2 \\approx 16 billion entries, making naive computation prohibitively expensive. Efficient attention variants such as Flash Attention, Sparse Attention, Sliding Window Attention address this scaling problem, which we&#8217;ll cover in Part 3 alongside the full transformer block. The KV Cache During inference, a decoder-only model generates one token at a time. At step tt, the model processes the full sequence [t1,\u2026,tt][t_1, \\ldots, t_t] and predicts tt+1t_{t+1}\u200b. Without optimization, this would require recomputing the key and value matrices for all previous tokens at every step so thar cost grows as O(t2)O(t^2) per generation. The KV cache eliminates this redundancy. Since the key and value representations of tokens t1,\u2026,tt\u22121t_1, \\ldots, t_{t-1}\u200bdo not change between steps (they depend only on those tokens&#8217; positions and embeddings, which are fixed once generated), they can be cached in memory and reused. At step tt, only the new token&#8217;s QQ, KK, VV need to be computed; the cached KK and VV matrices from previous steps are concatenated and the attention is computed against the full cached sequence. This reduces per-step inference cost from O(t2)O(t^2) to O(t)O(t) in attention computation, but at the cost of memory. A KV cache for a 128K-token context with 96 heads, 128 layers, and FP16 precision occupies on the order of tens of gigabytes. Managing KV cache memory is one of the primary challenges in deploying frontier models efficiently at scale. It is why inference hardware for long-context models requires far more VRAM than a naive parameter count would suggest. What Attention Actually Learns It is tempting to describe attention as a mechanism that &#8220;understands&#8221; language. It is more precise, and more useful for engineering purposes, to say that attention is a differentiable, content-based memory retrieval system. The model learns, through gradient descent on the next-token prediction objective, to configure its WQW^Q, WKW^K, WVW^V matrices such that the right values get retrieved for the right queries. The mechanism has no built-in notion of syntax, coreference, or meaning. All of that emerges from training dynamics or from the statistical regularities of billions of documents. Understanding this distinction is important for debugging model failures: when an LLM makes a coreference error or loses track of a constraint stated early in a long context, the failure mode is almost always traceable to attention weights that were distributed incorrectly, either because the training distribution underrepresented that pattern, or because the KV cache compression strategy discarded the relevant context. Conclusion Self-attention is the architectural innovation that made modern LLMs possible. By allowing every token to attend directly to every other token in a single parallelisable matrix operation, it solves the long-range dependency problem that defeated earlier sequential architectures. Multi-head attention extends this by learning multiple independent attention patterns simultaneously, while the KV cache makes autoregressive inference tractable at scale. Part 3 will take the attention output O\u2208Rn\u00d7dO \\in \\mathbb{R}^{n \\times d} and trace it through the rest of the transformer block: the feed-forward network, layer normalisation, residual connections, and the architectural choices (decoder-only versus encoder-decoder) that distinguish GPT-style models from BERT-style ones. Coming next in the AI Engineering series is Part 3: The Transformer Block and Architecture.<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[38],"tags":[],"class_list":["post-1034","post","type-post","status-publish","format-standard","hentry","category-ai-theory"],"_links":{"self":[{"href":"https:\/\/learnerbox.net\/blog\/wp-json\/wp\/v2\/posts\/1034","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/learnerbox.net\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/learnerbox.net\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/learnerbox.net\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/learnerbox.net\/blog\/wp-json\/wp\/v2\/comments?post=1034"}],"version-history":[{"count":3,"href":"https:\/\/learnerbox.net\/blog\/wp-json\/wp\/v2\/posts\/1034\/revisions"}],"predecessor-version":[{"id":1044,"href":"https:\/\/learnerbox.net\/blog\/wp-json\/wp\/v2\/posts\/1034\/revisions\/1044"}],"wp:attachment":[{"href":"https:\/\/learnerbox.net\/blog\/wp-json\/wp\/v2\/media?parent=1034"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/learnerbox.net\/blog\/wp-json\/wp\/v2\/categories?post=1034"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/learnerbox.net\/blog\/wp-json\/wp\/v2\/tags?post=1034"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}