Definition·
Architectures

Transformer

A neural network architecture based on the attention mechanism, foundation of modern LLMs.

Detailed explanation

Introduced in 2017 by the "Attention Is All You Need" paper, the Transformer replaces recurrent architectures with attention that weighs all positions of a sequence in parallel. It powered the rise of GPT, BERT, Claude, Gemini and most current LLMs.

Examples

GPT-4, Claude, Gemini, Mistral
BERT for text understanding
Vision Transformer (ViT) for images

Frequently asked questions

Why is it revolutionary?

Attention enables massively parallel training and long-range dependencies, unlocking very large models.

Related terms

Last updated: 7/15/2026

Talent AI

Turn theory into practice

Post a mission or join the community of top AI, Data and Machine Learning experts.