|Project name|| Main project page|
Previous entry Next entry
Variational Bayes approach for the mixture of Normals
The latent variables induce dependencies between all the parameters of the model. This makes it difficult to find the parameters that maximize the likelihood. An elegant solution is to introduce a variational distribution of parameters and latent variables, which leads to a re-formulation of the classical EM algorithm. But let's show it directly in the Bayesian paradigm.
We can now introduce a distribution :
The constant is here to remind us that has the constraint of being a distribution, ie. of summing to 1, which can be enforced by a Lagrange multiplier.
We can then use the concavity of the logarithm (Jensen's inequality) to derive a lower bound of the marginal log-likelihood:
Let's call this lower bound as it is a functional, ie. a function of functions. To gain some intuition about the impact of introducing q, let's expand :
From this, it is clear that (ie. a lower-bound of the marginal log-likelihood) is the conditional log-likelihood minus the Kullback-Leibler divergence between the variational distribution q and the joint posterior of latent variables and parameters. As a side note, minimizing DKL(p | | q) is used in the expectation-propagation technique.
In practice, we have to make the following crucial assumption of independence on in order for the calculations to be analytically tractable:
This means that approximates the joint posterior, and therefore the lower-bound will be tight if and only if this approximation is exact and the KL divergence is zero.
As we ultimately aim at inferring the parameters and latent variables that maximize the marginal log-likelihood, we will use the calculus of variations to find the functions and qΘ that maximize the functional .
This naturally leads to a procedure very similar to the EM algorithm where, at the E step, we calculate the expectations of the parameters with respect to the variational distributions and qΘ, and, at the M step, we recompute the variational distributions over the parameters.
We start by writing the functional derivative of with respect to :
Then we set this functional derivative to zero. We also make use of a frequent assumption, namely that the variational distribution fully factorizes over each individual latent variables (mean-field assumption):
Recognizing the expectation and factorizing qΘ(Θ) into , we get:
Taking the exponential:
As this should be a distribution, it should sum to one, and therefore:
Interestingly, even though we haven't specified anything yet about , we can see that it is of the same form as the prior on zn, a Multinomial distribution.
We start by writing the functional derivative of with respect to qΘ:
Then, when setting this functional derivative to zero and using the factorization , we can obtain the variational distribution of each parameter.