The turning point between shallow and deep training techniques is generally considered to have occurred in 2006, when Hinton and others at the University of Toronto introduced Deep Belief Networks (DBNs) (HOT06), an algorithm that “greedily” trains a layered structure by training one layer at a time using an unsupervised training algorithm. The distinctive feature of DBNs is that their layers consist of Restricted Boltzmann Machines (RBMs) (FH94,Smo86).
Let
be a binary stochastic variable associated with the visible state and
a binary stochastic variable associated with the hidden state.
Given a state
, the energy of the configuration of the visible and hidden layers is given by (Hop82)
The probability of the joint configuration of the visible states
and hidden states
, parameterized by the vectors
,
and the weight matrix
,
is given by the Boltzmann distribution:
The word restricted refers to the fact that direct interactions between units belonging to the same layer are not allowed; interactions are permitted only between adjacent layers.
Given an input , the hidden binary state
is activated with probability:
| (5.114) |
| (5.115) |
Estimating the model parameters
so as to correctly model the training-data distribution is computationally expensive.
However, in 2002 Hinton proposed the Contrastive Divergence (CD) algorithm, which enables much more efficient training of RBMs, finally making them suitable for large-scale applications.
A detailed and practical description of RBM training can be found in (Hin12).
Paolo medici