TRANSFORMER COMPONENTS

🔹 Embeddings
Converts tokens into dense 16D vectors
🔹 Multi-Head Attention
4 attention heads compute Q·K^T to learn which tokens relate to each other
🔹 Attention Weights
Softmax over scaled scores creates probability distribution
🔹 Feed-Forward
2-layer network (ReLU activation) transforms attention output
🔹 Backpropagation
Cross-entropy loss guides gradient descent to update all weights
LEGEND:
⚪ White = Current token (focus)
🟡 Yellow = Target (training)
🟢 Green = Generated (prediction)
🔵 Cyan = Vocabulary token