How might LLMs store facts | Deep Learning Chapter 7
A video on YouTube. In Science & Engineering, a Krater category.
Watch on YouTubeSummary by Krater
This video explains the mechanics of Multilayer Perceptrons (MLPs) within transformer models, using the example of storing the fact that Michael Jordan plays basketball in high-dimensional vector space.
From the video
Answers: How do Multilayer Perceptrons store facts in large language models?
- Multilayer Perceptrons
- Transformer models
- Vector embeddings
- Matrix multiplication
- Neural network activation functions
- ReLU and GELU
- Superposition in neural networks
What it concludes
- Fact-finding research shows that facts in transformer models live inside multi-layer perceptrons.
- MLP computation essentially boils down to a pair of matrix multiplications with a nonlinear function in between.
- The Johnson-Lindenstrauss lemma explains how high-dimensional spaces can store exponentially many nearly perpendicular vectors.
- Individual neurons in large language models rarely represent a single clean feature due to superposition.
Rate it, review it and add it to your lists in Krater.
Titles and thumbnails from YouTube. Krater isn't affiliated with, endorsed by or sponsored by YouTube or Google.