But what is cross-entropy? | Compression is Intelligence Part 2
A video on YouTube. In Science & Engineering, a Krater category.
Watch on YouTubeSummary by Krater
This video explains the mathematical concept of cross-entropy, its origins in file compression and language trees, and its fundamental role as the loss function in training large language models.
From the video
Answers: What is cross-entropy and how does it relate to compression and machine learning loss functions?
- cross-entropy
- file compression
- language trees
- information theory
- large language models
- loss functions
- Kullback-Leibler divergence
What it concludes
- Cross-entropy measures the average number of bits needed to encode data from one probability distribution using a code optimized for another distribution.
- For any optimal code, the number of bits allocated to a symbol is equal to the negative log base 2 of its probability.
- The cross-entropy of distribution Q relative to P is minimized when Q equals P, and its minimum value is the entropy of P.
- In machine learning, training a language model using cross-entropy loss is equivalent to finding model outputs that match the statistical distribution of the training data.
Rate it, review it and add it to your lists in Krater.
Titles and thumbnails from YouTube. Krater isn't affiliated with, endorsed by or sponsored by YouTube or Google.