Mebin pinged me to an article called "Compression is Prediction" in a Group Chat today. While it is the intersection of the current interests I have, I did not read the article when I came across it earlier from HackerNews(!) It isn't everytime that something shows up to you multiple times, so this time I read it.
The Article itself was superbly written, tying everything down and explaining how entropy coding and ML are related, until I realised that the Entropy table does not have units!
I began to write an entire paragraph down in the gc, which began like this:-
Oh btw did you know a lot of places just don't mention a unit for entropy (including this post)
I was wondering in ML class once if the base of the logarithm depended on anything for entropy calculation
The teacher said no, but I wondered why would the base matter in calculations if all we're doing is compare the values
Turns out the base of the logarithm defines what unit it is in
... until I realised that this is my long lost TIL that I've been wanting to write since that fateful ML class.
But What's Entropy?¶
The Entropy that we all usually learn first is from Physics (Thermodynamics), where it is a measure of "randomness", or the number of possible microscopic arrangements of the particles in a system, and how they relate to the observable macroscopic properties.
However, the Entropy here (while extremely similar) is different. It's also known as Shannon Entropy, named after Claude Shannon, who made insane strides in the world of information theory.
Shannon Entropy $H(X)$ quantifies the average amount of information, uncertainty, or "surprise" inherent in the possible outcomes of a random variable $X$.
It is given by the formula:-
$$H(X) = - \sum_{i=1}^{n} P(x_i) \log_b P(x_i) = \sum_{i=1}^{n} P(x_i) \log_b \left( \frac{1}{P(x_i)} \right)$$
where:-
- $P(x_i)$: The probability of outcome $x_i$ occurring.
- $\log_b \left(\frac{1}{P(x_i)}\right)$: This term quantifies the "surprise" or self-information of a specific outcome $x_i$. If an event is extremely unlikely ($P(x_i) \to 0$), its occurrence is surprising and thus yields more information (a higher value). Conversely, a certain event ($P(x_i) = 1$) has zero surprise.
- $H(X)$: The expected value (weighted average) of the self-information across all possible outcomes.
Notice how the Thermodynamic Entropy and Shannon Entropy are very similar. However, Thermodynamic Entropy is bounded by the laws of matter, temperature and energy, while Shannon Entropy is a purely mathematical concept that can apply to any data distribution. That means you could model Thermodynamic Entropy as a special case of Shannon Entropy where you are measuring the "lack of information" about the exact positions and velocities of microscopic particles.
Lower bounds of Information Representation¶
Shannon's Entropy gives us the theoretical lower bound for data compression. This means that it is impossible to compress data to a size smaller than its entropy. That means that if you find a way to represent your data in the smallest form possible (while retaining all of it!), it's entropy would be zero!
While it's probably not because of college being relatively lax than school, I'm sure it played a direct role in Entropy not having units. The main reason comes down to the fact that Entropy is a dimensionless quantity. But does that mean it's unitless?!?
Units of Entropy¶
The units of Entropy are decided by the base of the logarithm used! The base of the logarithm talks about how many equally likely outcomes we're considering.
A few of the common units are:-
| Unit | Base |
|---|---|
| Shannon (bit) | 2 |
| Nat | e |
| Hartley (ban, dit, decit) | 10 |
Since computers rely on bits, the unit of entropy in information theory is usually Shannons. Physicists (who sometimes really want to set themselves apart) use Nats, as that helps them with their Calculus equations, while Hartleys are used in Telecommunications and Radar engineering.
Entropy is used in a lot of other places, and often it's not used directly. This is because raw entropy numbers are just harder to explain.
Thus you find units like bits per character (in linguistics/data compression), effective number of species (in biology, where it is $e^H$ ), etc.
There are extensions like Von Neumann Entropy for Quantum Physics, and Rényi Entropy for generalising the notion of information content to non-additive systems. (I genuinely dont know much about these...)
P.S. That connection between Thermodyanamic Entropy and Shannon Entropy solves a Physics Paradox called "Maxwell's Demon". Do check it out :D