Local model
Private and hardware-bound. Good for experiments, automation, and offline control. Quality depends on the model and machine.
AI chatbots can feel like magic. This page opens one up: how it reads your words, how it picks each next word, and how apps turn that into assistants that use tools and check their work. Everything here is clickable, and nothing needs an account.
Or jump to the next-token machine · how models are trained · home
Six quick stops cover the essentials. Prefer everything? Switch to the full tour.
The illustration shows an agent system as a physical workbench. Press play to watch a request travel through it, or pick a station to zoom in.
In one lineYour words go in, get turned into numbers, run through a model, and come back out as an answer. Agents add tools and checks along the way.
Models never see letters. A tokenizer chops text into pieces it learned from data: common chunks become one token, rare ones get split. In English a token averages about three-quarters of a word. Type anything below. This uses a real tokenizer (OpenAI's o200k). Every model family has its own, but they all work this way.
In one lineModels read numbered chunks of text, not letters.
Each colour is one token. A · marks a leading space: most tokens carry the space in front of them, so " cat" and "cat" are different tokens.
Look at how a sum gets chopped up. Asked for the answer in one go, a model can't line up the digits and carry the one like you would on paper. It guesses the answer's tokens the same way it guesses words. (Bigger models that write out their working step by step do much better, and a calculator tool is exact.)
At every step the model gives a probability to every token in its vocabulary (about 150,000 of them). Then the app samples one. Temperature reshapes those odds: at 0 it always takes the top pick, and higher values flatten the curve so unlikely tokens get a real chance. That's why the same prompt can give different answers.
In one lineThe model gives odds for every possible next word, then the app rolls the dice.
This is a raw model, not a chatbot: it continues text instead of answering, a bit like autocomplete. The before-and-after demo further down shows the difference.
Each layer lets every token look back at every earlier token and pull in what's relevant. That's attention. Many "heads" run in parallel, each learning a different kind of lookup. Click a word to see where it looks, then flip the last word and watch the link move.
In one lineEach word looks back at earlier words to work out what they mean together.
The arcs show, in simplified form, what a well-trained model learns to do. Real attention is messier and spread across dozens of layers. The percentages come from a real (tiny) model, measured by asking it to finish "The thing that was too big was the…".
Now the whole trip, start to finish. The text becomes token numbers, each number becomes a list of meaning-numbers, the layers let the words compare notes (that's attention), and out come odds for the next token. One is picked, glued onto the end, and the whole thing runs again. Real token IDs and probabilities from the same small model.
In one lineThe model writes one token at a time, and each new token is added to the text before it guesses the next.
Everything above uses a model that was already trained. Training tuned its billions of internal numbers (called weights) by showing it huge amounts of text. That's a big topic with its own page, but the biggest single change is easy to see right here.
In one lineTraining is where the model learned. Chatting with it does not change it.
The deep dive covers pretraining, fine-tuning, RLHF, safety tuning, evaluation, data quality, and the difference between base models and chat assistants. It includes a tiny model you can train yourself and a preference-ranking demo.
Everything inside the window is visible to the model at once. Nothing fades until the window is full. Then the app (not the model) decides what to cut or summarize. Memory survives because the app loads it back in at the start of every session.
In one lineThe model only sees what fits in its window right now.
The system prompt is the app's own instructions, sent before your first message every time.
An app can't paste everything you own into the window, so it searches for what's relevant first. Many search systems turn each piece of text into a long list of numbers (an embedding), where similar meanings land close together even with no words in common. Each star below is a real one, squashed from 384 numbers down to 2D. Pick a search to see what's closest.
In one lineSimilar meanings get similar numbers, which is how apps find the right notes to show the model.
Pick a search above, or click any star.
Sentence embeddings from all-MiniLM-L6-v2, projected to 2D with t-SNE. The 2D layout is approximate; the similarity scores come from the full 384 dimensions.
A vague request fits dozens of reasonable plans, so the model has to guess which one you meant. Every constraint you add prunes wrong branches. Better prompts don't make the model smarter; they leave it less to guess.
In one lineSay what you want, what not to touch, and how to check it.
A model can't touch anything by itself. It only writes text. To use a tool it writes a structured request; the app around it (the harness) checks permissions, runs it, and pastes the result back into the context. Then the model reads that evidence and carries on.
In one lineTools let AI check instead of guess, but what a tool brings back is information, never orders.
Watch the loop run. Every lap adds to the context window: plans, tool calls, results. When a check fails, the agent goes around again. When an action is risky, it stops at the gate and waits for you.
In one lineAgents repeat plan → act → check until the job is really done, or stop to ask you.
A useful AI run leaves an outside trail: what was requested, what it checked, what it observed, and how it verified the result. You don't need to see inside the model's head to judge whether the work can be trusted.
In one lineTrust work you can see being checked.
AI is not one thing. The useful question is what kind of system you are using and what constraints come with it. Pick a job below and see which system fits. Some jobs need two at once.
In one linePick the kind of AI that fits the job.
Private and hardware-bound. Good for experiments, automation, and offline control. Quality depends on the model and machine.
Usually stronger and faster to update. Great for hard reasoning, coding, and multimodal work, with external service tradeoffs.
A model with tools, memory, permissions, and a loop. Powerful because it can act, risky if guardrails are sloppy.
Image, voice, music, and video models turn prompts into pixels or sound. They need visual direction, not just facts.
Most real AI systems are a stack of choices. It starts with just a model and its context. Add layers and watch what you gain, and what new risks need checks.
In one lineMore power needs more safety checks.
Scores are illustrative, to show the trade-off, not measurements.
Most AI mistakes are not mysterious. They usually come from missing context, stale information, weak instructions, bad tool choice, untrusted input, or permission boundaries.
In one lineMost mistakes come from missing, stale, or unchecked information.
An assistant summarized the notes on the left. Every sentence sounds fine, but some aren't backed by the source. Tap the ones you think are unsupported, then turn on grounding to check each claim against the evidence.
In one lineSounding sure is not the same as being right.
AI terms get thrown around loosely. Pick a concept to see the useful definition, a practical example, and the common misunderstanding to avoid.
In one linePlain-English definitions for the buzzwords.