Transformer Explainer

Tokens flow through embeddings, attention heads, and a transformer block.

GPT-2 · 626 MB model download on desktop · preset examples on mobile

Adapted from work by Polo Club of Data Science and collaborators · MIT.

← All labs