Interactive AI lab

How AI Works

AI chatbots can feel like magic. This page opens one up: how it reads your words, how it picks each next word, and how apps turn that into assistants that use tools and check their work. Everything here is clickable, and nothing needs an account.

Tour ready

Six quick stops cover the essentials. Prefer everything? Switch to the full tour.

Illustrated overview

AI System Map

The illustration shows an agent system as a physical workbench. Press play to watch a request travel through it, or pick a station to zoom in.

In one lineYour words go in, get turned into numbers, run through a model, and come back out as an answer. Agents add tools and checks along the way.

Generated illustration
Illustrated AI agent system with input, token blocks, context window, model core, tools, memory, verification, and final output stations.
Act I

Inside the model

Inside, a language model does one thing: it looks at the text so far and gives odds for every possible next piece of text. Everything else is built around that.

Text becomes numbers

Token Playground

Models never see letters. A tokenizer chops text into pieces it learned from data: common chunks become one token, rare ones get split. In English a token averages about three-quarters of a word. Type anything below. This uses a real tokenizer (OpenAI's o200k). Every model family has its own, but they all work this way.

In one lineModels read numbered chunks of text, not letters.

Loading tokenizer…
Tokens0
Characters / token0
Words / token0

Each colour is one token. A · marks a leading space: most tokens carry the space in front of them, so " cat" and "cat" are different tokens.

Go deeperWhy do AI models flub maths?

Look at how a sum gets chopped up. Asked for the answer in one go, a model can't line up the digits and carry the one like you would on paper. It guesses the answer's tokens the same way it guesses words. (Bigger models that write out their working step by step do much better, and a calculator tool is exact.)

Probabilities, then a dice roll

Next-Token Machine

At every step the model gives a probability to every token in its vocabulary (about 150,000 of them). Then the app samples one. Temperature reshapes those odds: at 0 it always takes the top pick, and higher values flatten the curve so unlikely tokens get a real chance. That's why the same prompt can give different answers.

In one lineThe model gives odds for every possible next word, then the app rolls the dice.

Loading model data…

This is a raw model, not a chatbot: it continues text instead of answering, a bit like autocomplete. The before-and-after demo further down shows the difference.

0 · greedy1 · raw odds2 · chaos

spin

How words find each other

Attention

Each layer lets every token look back at every earlier token and pull in what's relevant. That's attention. Many "heads" run in parallel, each learning a different kind of lookup. Click a word to see where it looks, then flip the last word and watch the link move.

In one lineEach word looks back at earlier words to work out what they mean together.

Click a word

The arcs show, in simplified form, what a well-trained model learns to do. Real attention is messier and spread across dozens of layers. The percentages come from a real (tiny) model, measured by asking it to finish "The thing that was too big was the…".

Putting it together

The Generation Loop

Now the whole trip, start to finish. The text becomes token numbers, each number becomes a list of meaning-numbers, the layers let the words compare notes (that's attention), and out come odds for the next token. One is picked, glued onto the end, and the whole thing runs again. Real token IDs and probabilities from the same small model.

In one lineThe model writes one token at a time, and each new token is added to the text before it guesses the next.

Real model data
picked token is appended → the loop runs again
How it learned

Where the Model Came From

Everything above uses a model that was already trained. Training tuned its billions of internal numbers (called weights) by showing it huge amounts of text. That's a big topic with its own page, but the biggest single change is easy to see right here.

In one lineTraining is where the model learned. Chatting with it does not change it.

Deep dive live

Same size model, before and after training to chat

Raw base model: just continues the text
After chat training: answers you

How models are made

The deep dive covers pretraining, fine-tuning, RLHF, safety tuning, evaluation, data quality, and the difference between base models and chat assistants. It includes a tiny model you can train yourself and a preference-ranking demo.

PretrainingFine-tuningRLHFAlignmentEvalsRAG
Open the deep dive
Act II

Around the model

The model only ever sees one thing: its context window. Apps decide what goes into it: your chat, saved notes, search results, and tool results.

What the model can see

Context Window & Memory

Everything inside the window is visible to the model at once. Nothing fades until the window is full. Then the app (not the model) decides what to cut or summarize. Memory survives because the app loads it back in at the start of every session.

In one lineThe model only sees what fits in its window right now.

0% full
Window size 1K tokens ≈ 750 words
When it overflows, the app will…
Extras
system prompt memory you assistant tool result summary

The system prompt is the app's own instructions, sent before your first message every time.

Finding the right notes

Embedding Map

An app can't paste everything you own into the window, so it searches for what's relevant first. Many search systems turn each piece of text into a long list of numbers (an embedding), where similar meanings land close together even with no words in common. Each star below is a real one, squashed from 384 numbers down to 2D. Pick a search to see what's closest.

In one lineSimilar meanings get similar numbers, which is how apps find the right notes to show the model.

Real embeddings

Pick a search above, or click any star.

Sentence embeddings from all-MiniLM-L6-v2, projected to 2D with t-SNE. The 2D layout is approximate; the similarity scores come from the full 384 dimensions.

Same model, better instructions

Prompt Pruning

A vague request fits dozens of reasonable plans, so the model has to guess which one you meant. Every constraint you add prunes wrong branches. Better prompts don't make the model smarter; they leave it less to guess.

In one lineSay what you want, what not to touch, and how to check it.

12 possible plans

The prompt

Ambiguity

Guessing vs checking

Tool Use Simulator

A model can't touch anything by itself. It only writes text. To use a tool it writes a structured request; the app around it (the harness) checks permissions, runs it, and pastes the result back into the context. Then the model reads that evidence and carries on.

In one lineTools let AI check instead of guess, but what a tool brings back is information, never orders.

Model writes
Harness runs it
Result → context
Model answers
Model guesses

Agent checks

Act III

Agents in the wild

An agent is a model in a loop with tools, memory, and boundaries. The interesting part is what it checks, and when it stops.

Step through it

Agent Loop Circuit

Watch the loop run. Every lap adds to the context window: plans, tool calls, results. When a check fails, the agent goes around again. When an action is risky, it stops at the gate and waits for you.

In one lineAgents repeat plan → act → check until the job is really done, or stop to ask you.

Watch it loop
Lap 1
Context window this run0 tokens
Observable work

Trace Viewer

A useful AI run leaves an outside trail: what was requested, what it checked, what it observed, and how it verified the result. You don't need to see inside the model's head to judge whether the work can be trusted.

In one lineTrust work you can see being checked.

Audit trail
Different jobs, different tradeoffs

Local, Cloud, Agent, Media

AI is not one thing. The useful question is what kind of system you are using and what constraints come with it. Pick a job below and see which system fits. Some jobs need two at once.

In one linePick the kind of AI that fits the job.

Match the job

Local model

Private and hardware-bound. Good for experiments, automation, and offline control. Quality depends on the model and machine.

Cloud model

Usually stronger and faster to update. Great for hard reasoning, coding, and multimodal work, with external service tradeoffs.

Agent

A model with tools, memory, permissions, and a loop. Powerful because it can act, risky if guardrails are sloppy.

Media model

Image, voice, music, and video models turn prompts into pixels or sound. They need visual direction, not just facts.

Pick a job above to see which kind of system fits it best, and why the others are a worse match.
Assemble the system

Build an Agent Stack

Most real AI systems are a stack of choices. It starts with just a model and its context. Add layers and watch what you gain, and what new risks need checks.

In one lineMore power needs more safety checks.

Interactive model

Scores are illustrative, to show the trade-off, not measurements.

Capability0
Risk0
Oversight0
Where it goes wrong

Failure Modes

Most AI mistakes are not mysterious. They usually come from missing context, stale information, weak instructions, bad tool choice, untrusted input, or permission boundaries.

In one lineMost mistakes come from missing, stale, or unchecked information.

Quick checkpoint

Pick the safer approach.
A game about grounding

Spot the Hallucination

An assistant summarized the notes on the left. Every sentence sounds fine, but some aren't backed by the source. Tap the ones you think are unsupported, then turn on grounding to check each claim against the evidence.

In one lineSounding sure is not the same as being right.

Round 1 of 3
Source
    Assistant's summary: tap the unsupported claims
    Fluent is not the same as supported.
    Reference

    Glossary

    AI terms get thrown around loosely. Pick a concept to see the useful definition, a practical example, and the common misunderstanding to avoid.

    In one linePlain-English definitions for the buzzwords.

    Tap a term
    How AI Works in 60 seconds

    A silent, captioned run-through of the page. Everything in it is interactive below.