- Published on
LLM
- Authors

- Name
- seren-wib
Contents
- 1. Language Model
- 2. LLM(Large Language Model)
- LLM basic concepts
- 1. Tokenization
- 2. Embedding
- Conversion steps: token > integer ID > vocabulary (voca) > embedding vector
- Transformer model
- Temperature
- Sampling
- 3. Transformer model
- 1. Self-attention operation
- 2. Positional embedding vector
- 3. Structure of a Transformer block
- Self-Attention
- FFN(Feed Forward Network) or MLP
- Skip Connection
- LayerNorm (normalization layer)
- Flow
- 4. mulit-head attention
- 5. Encoder and decoder
- Cross-Attention
- Encoder-decoder inference
- 4. Models built on the Transformer
- 1. BERT(Bi-directional Encoder Representations from Transformers)
- Attach a classification layer to a well-trained BERT language model to perform various NLP tasks (natural language processing tasks)
- 2. GPT(Generative Pre-trained Transformer)
- GPT flow
- 5. Training an LLM
- 1. Pre-training
- 2. SFT(supervised fine-tuning)
- 3. RLHF (reinforcement learning with human feedback)
- 6. Multimodal LLM
- How to feed images in
- Method 1. Put image vectors into the LLM like input tokens
- Method 2. Refer via Cross-Attention
- VLM(Vision Language Model)
- 1. Language Model
- 2. LLM(Large Language Model)
- 3. Transformer model
- 4. Models built on the Transformer
- 5. Training an LLM
- 6. Multimodal LLM
1. Language Model
A probabilistic model that predicts the most appropriate next word based on the context of the given text
2. LLM(Large Language Model)
A language model with a very large number of parameters, trained on a massive amount of text data
Mainly implemented as Transformer-based models
LLM basic concepts
1. Tokenization
The step that splits a sentence into smaller units called tokens
The split tokens are mapped to integer IDs according to a predefined vocabulary
2. Embedding
- Converts each token into a numeric vector representation with hundreds or thousands of dimensions
- Similar concepts point in similar directions in vector space.
- The same word gets different vector values depending on context.
Conversion steps: token > integer ID > vocabulary (voca) > embedding vector
Transformer model
The neural network architecture at the core of large language models
Processes data (the tokens' embedding vectors) in parallel and understands context in depth
Takes in the whole sentence at once, so training speed is greatly improved over earlier models
Computes the probability distribution of the next token (word, next token)
Temperature
Temperature controls the randomness of text generation by adjusting word probabilities before the model chooses a word
High temperature (≥ 1): more diverse and creative output
Low temperature (~0): more focused and predictable responses
Sampling
Top-K sampling
- Limits the word pool
- Picks probabilistically among the top K words with the highest probability
- Fixed number of candidate words; with K=5, always 5 are chosen
Top-p sampling
- Picks a word from a dynamically changing word pool, filled until the cumulative probability reaches p
- Variable number of candidate words; candidates are added until the cumulative probability reaches p.
3. Transformer model
Creates three kinds of vectors, Query, Key and Value, from each token's embedding vector, computes how related one token's Query is to every token's Key, then mixes the Values in proportion to that relevance to make a new vector
Q(Query): the information I want to find
K(Key): I am related to this kind of information
V(Value): the content actually passed on
1. Self-attention operation
input1, input2, input3 come in
Example input = [1010, 0202, 1111]Multiply each input by the trained weight matrices inside the model to make query, key, value
Example key = [[0,1,1], [4,4,0], [2,3,1]] value = [[1,2,3], [2,8,0], [2,6,3]] query = [[1,0,2], [2,2,2], [2,1,3]]query1 is compared with every key (dot product of vectors)
Example query1 = [1, 0, 2] key1 = [0, 1, 1] key2 = [4, 4, 0] key3 = [2, 3, 1] query1 · key1 = [1, 0, 2] · [0, 1, 1] = 1×0 + 0×1 + 2×1 = 2 score1 = 2 query1 · key2 = [1, 0, 2] · [4, 4, 0] = 4 + 0 + 0 = 4 score2 = 4 query1 · key3 = [1, 0, 2] · [2, 3, 1] = 2 + 0 + 2 =4 score3 = 4 score = [2, 4, 4]The comparison produces scores
Turn the scores into ratios with softmax
Example score1 = 2 score2 = 4 score3 = 4 query1 = [1,0,2] e^2 ≈ 7.389 e^4 ≈ 54.598 e^4 ≈ 54.598 denominator = 7.389 + 54.598 + 54.598 = 116.585 softmax(2) = 7.389 / 116.585 ≈ 0.063 softmax(4) = 54.598 / 116.585 ≈ 0.468 softmax(4) = 54.598 / 116.585 ≈ 0.468 softmax([2, 4, 4]) ≈ [0.063, 0.468, 0.468]Multiply every value by its ratio
Example multiplication1 0.063 × value1 = 0.063 × [1, 2, 3] = [0.063, 0.126, 0.189] multiplication2 0.468 × value2 = 0.468 × [2, 8, 0] = [0.936, 3.744, 0] multiplication3 0.468 × value3 = 0.468 × [2, 6, 3] = [0.936, 2.808, 1.404]Add the multiplied values to make output1
Example [0.063, 0.126, 0.189] + [0.936, 3.744, 0] + [0.936, 2.808, 1.404] = [1.935, 6.678, 1.593] output1 ≈ [1.935, 6.678, 1.593]query2 and query3 make output2 and output3 the same way
2. Positional embedding vector
Self-Attention computes relationships between tokens.
But without order information, the meaning of the sentence falls apart.
So a positional embedding is added to each token embedding.
As a result, the Transformer receives "word meaning + position" together.
It must have the same dimension as the token embedding.
3. Structure of a Transformer block
Self-Attention
Computes how much each token should refer to the other tokens in the sentence
FFN(Feed Forward Network) or MLP
If Self-Attention handled "the relationships between tokens", the FFN takes each token vector separately and transforms it once more
Usually made of a linear layer, an activation function, and a linear layer
Skip Connection
self-attention result + original input
Preserves information from the previous layer and makes training stable
LayerNorm (normalization layer)
Normalization that tidies vector values that get too large or uneven.
As computations pile up, the distribution of values can shift, so it keeps things stable at every layer
Flow
Input vector (token embedding + positional embedding)
↓
Self-Attention
↓
Skip Connection + LayerNorm
↓
Feed Forward Network
↓
Skip Connection + LayerNorm
↓
Output vector
The input vector comes in
Self-Attention reflects the relationships between tokens
Add the original input = Skip Connection
Tidy the values = LayerNorm
The FFN processes each token vector once more
Add the original value again = Skip Connection
Normalize again = LayerNorm
Pass it to the next Transformer block
4. mulit-head attention
A concept that came about because the self-attention approach can only focus strongly on one kind of relationship.
- Linear projection
- Multiplying a vector by a weight matrix to move it into a different vector space
- It can reduce the dimension, increase it, or just change the direction within the same dimension.
Input X
↓
Create Q1, K1, V1 for head 1
Create Q2, K2, V2 for head 2
Create Q3, K3, V3 for head 3
...
↓
Run self-attention separately in each head
↓
Concatenate the head outputs (concatenation)
↓
Linear transformation again
↓
multi-head attention output
5. Encoder and decoder
The encoder reads the input sentence and produces meaning vectors, and the decoder refers to the encoder's last layer when producing the output sentence
The way it refers here: Cross-Attention
Cross-Attention
Output of the encoder's last layer = vectors summarizing the input sentence after reading it all
Decoder Query = a request for the information needed at the position currently being output
Encoder Key = a search tag indicating what information each input token holds
Encoder Value = the input information actually passed to the decoder
Cross-Attention = the process of searching the encoder Keys with the decoder Query and fetching the matching Values
Encoder-decoder inference
Key and value vectors are created from the encoder's output for the given input
4. Models built on the Transformer
1. BERT(Bi-directional Encoder Representations from Transformers)
A Transformer Encoder-based model
Uses two sentences as input - feed in sentence A and sentence B and judge whether they are consecutive.
Attach a classification layer to a well-trained BERT language model to perform various NLP tasks (natural language processing tasks)
Tasks such as sentence classification, sentiment analysis and next sentence prediction
- Next Sentence Prediction
- Autoencoding training: restore the hidden Masked tokens
- I ate [MASK]
- Infer the MASKed part from the context
2. GPT(Generative Pre-trained Transformer)
A Transformer Decoder-based model
Auto-regressive model: generates the next token based on previous outputs
The architecture used by most LLMs
GPT flow
Input words
↓
Convert words into embedding vectors
↓
Transformer decoder processes them
↓
Predict the next word
↓
Feed the predicted word back into the input and repeat generation
5. Training an LLM
1. Pre-training
The stage that learns next-token prediction on a large corpus, i.e. text on the scale of trillions of tokens (compressing world knowledge)
Trained to predict the next word from previous words
- Learns everything from a single simple task: predicting the next token
- The labels are inherent in the data itself, so it is self-supervised
Compressing the whole internet and melting it into the parameters
At this stage it gains the basis for language structure, factual knowledge and reasoning ability
Data quality filtering is decisive
High-quality datasets go through deduplication, adult-content filters, language detection and quality classifiers
- A model that has finished this stage is called a Base Model
- It has knowledge but doesn't know well how to respond to instructions > SFT needed
2. SFT(supervised fine-tuning)
The stage where the model looks at pairs of questions and ideal responses and learns what format and attitude to answer with when given an instruction (format/attitude learning)
Data collection methods
- Human-written: professional annotators write (question, ideal response) pairs directly. Best quality, highest cost.
- Model-generated + human review: GPT-4 generates drafts → humans revise. The InstructGPT and Alpaca approach.
- Self-Instruct: the model generates diverse instructions and responses on its own. Large variance in quality
Model after SFT: follows instructions, but may also follow harmful requests as is > RLHF needed
3. RLHF (reinforcement learning with human feedback)
Adjusts the model based on the answers people prefer (human preference alignment)
- Data collection
- Generate several responses to the same question, and human raters mark preferences by pairwise comparison
- OpenAI's rater interface for InstructGPT
Reinforcement learning of the model
- Trained to produce results with high human preference while not drifting too far from the original model
6. Multimodal LLM
A model that understands and generates many data formats in one model, not just text but also images, speech and video
- Modality: data format (text, speech, image, video)
- Projection layer: an intermediate conversion layer that vectorizes another modality such as images or speech with an encoder, then shapes that vector into a form the LLM can take
- Uses a large language model as the backbone; data from other modalities is vectorized with a separate encoder and fed into the large language model
- Requires fine-tuning
How to feed images in
- Turn the chunk of pixels into image embedding vectors with an image encoder like ViT
Method 1. Put image vectors into the LLM like input tokens
[image tokens] [describe] [this] [photo]
Treats the image information as a kind of input token too, processed together with the text question that follows
The Molmo model uses this approach
Method 2. Refer via Cross-Attention
The NVLM model uses this approach
VLM(Vision Language Model)
- A model designed to understand and process visual information such as images or video together with natural language (text)
- Provides multifaceted reasoning, such as looking at an image and describing it or answering questions, as if it had "eyes"
- Translates images into tokens the LLM can read