Convolutional Neural Networks (CNNs) represented a breakthrough in static-image processing, thanks to their ability to capture hierarchies of local features through convolutional filters and pooling. However, conventional CNNs are not suited to sequential data such as text, temporal signals, or video, in which the order and temporal relationship between elements are crucial. To address these problems, Recurrent Neural Networks (RNNs) were introduced. Their neurons have recurrent connections that enable them to retain a memory of the past state. RNNs have been applied to tasks such as speech recognition, machine translation, time-series analysis, and automatic image description (image captioning). However, conventional RNNs have difficulty learning long-term dependencies because of the vanishing or exploding gradient problem. To mitigate these limitations, more sophisticated architectures such as Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU) have been developed. These architectures can control which information to retain or forget through gating mechanisms.
Despite this, the first generation of deep sequential neural networks was based on an encoder-decoder paradigm, in which the information in the input sequence was compressed into a lower-dimensional tensor that had to preserve as much useful information as possible for the task. However, this approach has an inherent limitation: relevant input elements may be overlooked or attenuated during compression.
A fundamental change came with the “attention mechanism” (attention mechanism). The idea behind attention is simple but powerful: instead of compressing all information into a single vector, the model can dynamically “focus” on the most relevant parts of the input sequence. Attention weights are computed based on the relevance of each element to the others, overcoming the limitations of representations compressed into recurrent states.
Let
be the input matrix of the sequence, where
is the number of tokens (the individual units into which the algorithm divides the input sequence, varying from sequence to sequence) and
is the embedding dimension (a numerical vector representing each token, fixed for the model and capable of encoding semantic and syntactic information).
The matrices called query (), key (
), and value (
) are obtained through linear projections:
The self-attention mechanism can then be expressed in scalar form (for a single token) as:
Intuitively, attention can be viewed as a dynamic generalization of weighting methods such as Bag of Words or TF-IDF. However, whereas TF-IDF assigns static weights to terms, attention assigns contextual, task-dependent weights, allowing the model to focus selectively on the most relevant parts. Semantically, the resulting matrix has the same length as the input (
input tokens and
output tokens), but each token is enriched with information from the global context of the sequence.
The attention mechanism led directly to the development of Transformers (VSP$^+$17), now the de facto standard for sequence modeling in natural language, computer vision, and multimodal learning. In Transformers, the central operator is self-attention, which makes it possible to model the relationships among all sequence elements directly and in parallel. Compared with RNNs, Transformers offer significant advantages in terms of parallelization, numerical stability, and the ability to learn long-range dependencies.
In computer vision, the application of Transformers led to the development of Vision Transformers (ViT) (DBK$^+$20), in which an image is divided into small regions (patches) treated as a sequence, similarly to words in a text. These models have demonstrated performance competitive with or superior to that of CNNs on various classification, recognition, and segmentation tasks, especially when large amounts of data are available.
In practical applications, a single self-attention module is not sufficient to extract enough information from the input tokens. The multi-head self-attention mechanism extends the idea of self-attention by allowing the model to examine the sequence from several perspectives simultaneously. In practice:
| Model | Layers |
|
||
|---|---|---|---|---|
| Transformer (base) (VSP$^+$17) | 6 encoder + 6 decoder | 512 | 8 | 64 |
| BERT-Base | 12 encoder | 768 | 12 | 64 |
| BERT-Large | 24 encoder | 1024 | 16 | 64 |
| GPT-3 (175B) | many decoders | 12288 | 96 | 128 |
Today, RNNs and Transformers are complementary tools: the former remain useful in scenarios involving relatively short sequences or limited resources, whereas the latter form the basis of the most advanced architectures in modern deep learning. The evolution from recurrent mechanisms to attention-based mechanisms marked a paradigm shift: from the idea of compressed memory to a dynamic, contextual representation in which the model autonomously decides “what to look at” for each sequence element.
Paolo medici