Restricted Boltzmann Machines

Figure 5.9: Restricted Boltzmann Machines.
Image fig_rbm

The turning point between shallow and deep training techniques is generally considered to have occurred in 2006, when Hinton and others at the University of Toronto introduced Deep Belief Networks (DBNs) (HOT06), an algorithm that “greedily” trains a layered structure one layer at a time using an unsupervised training algorithm. The distinctive feature of DBNs is that their layers consist of Restricted Boltzmann Machines (RBMs) (FH94,Smo86).

Let $\mathbf{v} \in \{0,1\}^{n}$ be a binary stochastic variable associated with the visible state and $\mathbf{h} \in \{0,1\}^{m}$ a binary stochastic variable associated with the hidden state. Given a state $(\mathbf{v}, \mathbf{h})$, the energy of the configuration of the visible and hidden layers is given by (Hop82)

\begin{displaymath}
E(\mathbf{v}, \mathbf{h} ) = - \sum_{i=1}^{n} a_i v_i - \su...
...1}^{m} b_j h_j - \sum_{i=1}^{n} \sum_{j=1}^{m} w_{i,j} v_i h_j
\end{displaymath} (5.109)

where $v_i$ and $h_j$ are the binary states of the visible and hidden layers, respectively, while $a_i$ and $b_j$ are the weights, and $w_{i,j}$ are the weights associated with them. A Boltzmann Machine is similar to a Hopfield network, except that all outputs are stochastic. A Boltzmann Machine can therefore be defined as a special case of an Ising model, which is itself a special case of a Markov Random Field. Similarly, RBMs can be interpreted as stochastic neural networks in which the nodes and connections correspond to neurons and synapses, respectively.

The probability of the joint configuration $(\mathbf{a}, \mathbf{b}, \mathbf{W})$ is given by the Boltzmann distribution:

\begin{displaymath}
P(\mathbf{v}, \mathbf{h} ) = \frac{1}{Z(\cdot)} e^{-E(\mathbf{v}, \mathbf{h}) }
\end{displaymath} (5.110)

where the partition function $Z$ is given by
\begin{displaymath}
Z=\sum_{\mathbf{v}, \mathbf{h}} e^{-E(\mathbf{v}, \mathbf{h}) }
\end{displaymath} (5.111)

the sum of the energies of all possible pairs of visible and hidden states.

The word restricted refers to the fact that direct interactions between units belonging to the same layer are not permitted; interactions are allowed only between adjacent layers.

Given an input $\mathbf{v}$, the binary hidden state $h_j$ is activated with probability:

\begin{displaymath}
p(h_j = 1 \vert \mathbf{v}) = \sigma \left( b_j + \sum_i v_i w_{i,j} \right)
\end{displaymath} (5.112)

where $\sigma(x)$ is the logistic function $1 / (1 + exp(-x) )$. Similarly, the visible state can readily be obtained given the hidden state:
\begin{displaymath}
p(v_i = 1 \vert \mathbf{h}) = \sigma \left( a_i + \sum_j h_i w_{i,j} \right)
\end{displaymath} (5.113)

Estimating the model parameters $(\mathbf{a}, \mathbf{b}, \mathbf{W})$ so as to correctly model the distribution of the training data is computationally expensive. However, in 2002 Hinton proposed the Contrastive Divergence (CD) algorithm, which enables much more efficient training of RBMs, finally making them suitable for large-scale applications. A detailed and practical description of RBM training can be found in (Hin12).

Paolo medici
2026-10-01