A large language model can feel like one enormous black box. It becomes easier to reason about when you follow one small question at a time: how does text become input, how can earlier words affect later ones, and what exactly is the model choosing when it produces the next token?

This visual learning path is a sequence, not a promise that one animation explains every implementation detail. Open each explainer, identify the relationship it makes visible, and carry one question into the next.

1. Start with the trip a prompt takes

What Happens When You Send a Prompt is a useful map of the territory. Use it to locate tokenization, model processing, and the generated response before going deeper. Ask which steps happen in the model and which belong to the surrounding product.

2. See a neural network as a function

But what is a neural network? builds intuition from layers of simple computations. Pause when activations move between layers. The important idea is not that a network thinks like a brain; it is that learned numerical parameters transform one representation into another.

3. Watch attention connect positions

How LLMs work, 3b1b style gives a compact visual overview of transformers. Then use Induction heads in language models as a narrower interpretability example. Ask what information an attention pattern reveals—and what it does not prove about the model as a whole.

4. Connect prediction with compression

But what is cross-entropy? connects prediction quality with information and compression. You do not need to memorize the equation on a first pass. Compare a confident correct prediction, an uncertain prediction, and a confident wrong prediction. Notice how the loss changes.

5. Look under the serving layer

Training explains how parameters are learned; serving explains how a product can generate responses quickly enough to use. How DeepSeek keeps 890 bytes of KV cache per token focuses on one memory and performance technique. Treat its number as specific to the implementation described, not a universal constant for every model.

6. Finish with limits, not magic

Return to the original prompt and write down what the model had available: tokens, learned parameters, the current context, and product-level tools or retrieval if present. A visual explanation is most useful when it helps you separate the mechanism you observed from the story you might be tempted to tell about it.

Next, compare the learning path with Five ways a cleaning robot ruins an office. It shifts the question from how a model computes to how goals and assumptions can fail in a system.

Sources & further reading

Make the next question an experiment.

Browse independent creators who make big ideas visible.

Find an explainer