Variational Inference and Generative Models
Author: I Am Zhang Yixian
Author’s Zhihu profile: Thinker
Variational Inference
The basic idea is simple: probabilistic models often require us to approximate distributions that are difficult to compute. Inference about any unknown quantity can be viewed as inference about a posterior probability because Bayes’ theorem allows us to construct one:
For large datasets, Markov chain Monte Carlo methods are too slow, which is where variational inference comes in.
Two kinds of variables appear repeatedly: the observed data, , and the latent variable, .
The resulting inference problem is to determine the posterior conditional distribution for the input data. Using the ELBO, we seek a tractable distribution that can approximate and stand in for the true posterior distribution .
We therefore optimize the KL divergence between them:
The KL divergence can be rewritten as follows. All expectations below are taken with respect to :
This gives us the evidence lower bound, or ELBO:
The above can, of course, be replaced by ; the equation then needs only a slight adjustment.
Whether we are working with a VAE, GAN, or NF, one inequality is especially important:
- Make and as close as possible
The objective is to make as close to 0 as possible, which in turn makes the two distributions as close as possible. Rearranging gives:
Combining the two equations, we obtain:
Because our objective is to make as close to 0 as possible, we obtain the equation from the paper:
This formula captures the key idea of variational inference. Here, is referred to as the ELBO.
VAE (Variational Autoencoder)
A VAE seeks to maximize the derived above. It does so using SGVB and the reparameterization trick.
The VAE architecture is shown below:

Put simply, it takes the real data as input and uses the generative networks to produce the desired :
It then introduces Gaussian noise and combines these quantities to construct the latent variable . This is the reparameterization trick, which makes backpropagation possible:
The decoder maps to :
When training a VAE, we must ensure that follows a normal distribution. This is a major difference between a VAE and an AE: it gives the latent space a degree of regularity, including continuity and completeness.
The VAE objective is:
The first term can be understood as the reconstruction loss and the second as the regularization loss.
Applying maximum likelihood to the first term yields:
How do we handle the second term? We can evaluate the integral directly:
(We want the distribution of to be as close as possible to ; accordingly, the distribution of has a mean of 0 and a variance of 1. For , the distribution is .)
The second term pushes as close as possible to . The first term is the MSE between the generated and ; it is the term that also appears in an AE. The second term acts as a regularizer.
To generate a completely new image, sample directly and feed it into the decoder network. Its output is the new image.
GAN (Generative Adversarial Network)
A GAN consists of a generator G and a discriminator D:

A GAN can be understood as using cross-entropy to judge the similarity between distributions:
This resembles a binary classification problem. A sample either comes from the real dataset or is produced by drawing random noise and passing it through the generator. The discriminator must then decide whether the sample is real or fake:
The discriminator answers only “right” or “wrong.” It is a binary classification network.
Suppose is the distribution of real samples. The corresponding
is then the distribution of generated samples.
represents the discriminator, so
represents the probability that it classifies a sample as real, while
corresponds to the probability that it classifies a sample as fake.
Expressed in terms of cross-entropy, this becomes:

Let the sample produced by the generator be ~, where follows the distribution of the noise fed into the generator. is a real sample point.
Extending the formulation to the continuous case, we rewrite it as an integral.
The formal procedure is as follows:
We first fix G—that is, choose an arbitrary G—and then solve for :
Differentiating the expression inside the integral gives the maximum , the point at which performs best:
At this point,
That is,
We then continue training the generator to minimize : .
Thus, when —the Nash equilibrium, where the discriminator can no longer tell which is which—we obtain the minimum:
Substituting the minimum values and rearranging gives:
When JSD is 0, and are considered equal and can no longer be distinguished. At this point, C* = -log4.
NF (Normalizing Flow)
A normalizing flow is another kind of generative network. It is based on the change-of-variables theorem.
Suppose the generative network is still G, is the latent variable with a standard normal distribution, and x is the real data.
We want the distribution of the generated x to be as close as possible to the original distribution of x.
Suppose is a sample drawn from .
We want and to be as close as possible, which gives us the objective:
Several mathematical facts matter here; among them, the volume after a linear transformation equals the determinant of the transformation matrix.
In other words, the determinant can be viewed as the local linear rate of volume change under the transformation .
The input and output sizes of an NF must be the same. This sets it apart from the other two, which can accept arbitrary inputs.
Now suppose it consists of a series of flow networks.

We modify the original network accordingly:
We seek to maximize the left-hand side—in other words, to maximize the log likelihood.
To calculate it, we can reverse the process and feed x through the network to generate z.
Because the expression above is too computationally expensive, we use the following techniques:
1. Coupling Layer: used in NICE and Real NVP

The layer maps the first 1:d dimensions directly from z to x. For the remaining d+1:D dimensions, the earlier data pass through the F and H networks to obtain and , followed by a linear combination:
After the coupling layer, the Jacobian is triangular:
`
The upper portion is copied directly, so it is 1:1, while the entries along the lower-right diagonal (D>i>d+1) are the values.
The determinant of the Jacobian can therefore be written as:
But if every network were processed this way, would the first d terms not remain unchanged?
The terms must therefore be processed in an interleaved manner. Each network randomly selects d terms, and the size of d must vary as well. Only then will stacking the networks have an effect.

- 1x1 Convolution
In addition to coupling layers, we can use a 1x1 convolution. Proposed in GLOW, it works especially well.
Suppose we are processing an image with three channels: R, G, and B. The 1x1 convolution uses a 3x3 matrix :

The size remains unchanged after the 1x1 convolution, and this matrix is in fact :
\begin{equation} % begin mathematical environment J_f = \left( % left parenthesis \begin{array}{ccc} % the matrix has three centered columns w_{11} & w_{12} & w_{13}\\ % first-row elements w_{21} & w_{22} & w_{23}\\ % second-row elements w_{31} & w_{32} & w_{33}\\ \end{array} \right) = W % right parenthesis \end{equation}
As long as W is easy to work with, the result is easy to calculate:

The result is therefore the product of the W matrices along the diagonal:
Substituting this into the formula gives:
This turns a complex computation into a simpler one.
References:
[1] Hung-yi Lee, Machine Learning
[2] Su Jianlin

