AI/Lineage

A living, interactive field guide · rules, statistics, silicon, safety

How a machine went from following rules to writing them.

Fourteen chapters, each a working instrument rather than a picture of one. Every boundary, gradient and attention weight is computed in your browser as you scroll. Pick a detail level that fits you; the diagrams stay, the depth changes.

14 chapters 3 detail levels 0 pretrained weights downloaded Living guide updated as the field moves ·
Supervised · decision boundary-
Difficulty scroll-linked
§ 01 / 14

The rule you write is the ceiling you hit.

Classical machine learning is not a different kind of program. It is the same program with the constants left blank and a procedure for filling them in from examples.

Try this Pick Spiral, then drag Difficulty all the way right. Which model keeps up? Then open Reinforcement and watch the robot learn the way to its charger.

Explorer

Imagine sorting fruit with a ruler: if it is wider than 6 cm, it's an apple. That works until someone hands you a small apple. So you add another rule. Then another. Soon you have three hundred rules and they argue with each other.

The other way is to show the machine a pile of already-sorted fruit and let it draw its own line. Drag the difficulty and watch the hand-written rules fall apart while the learned line bends to keep up.

Machines can also learn with no answers at all. On the Reinforcement tab a little cleaning robot tries moves, gets a treat when it reaches its charger and a splash when it rolls into the dog's water bowl, and slowly paints a map of which moves pay off.

Practitioner

The expert system encodes a hypothesis directly: a conjunction of axis-aligned thresholds. Its capacity is fixed by how many clauses a human is willing to maintain, and it degrades hard once the classes stop being axis-separable.

A learner fixes a hypothesis class instead and searches it. k-NN is non-parametric: capacity grows with the data. Logistic regression is linear in whatever features you hand it. The RBF model here is a kernel expansion: a bump per prototype, which is how a linear method buys a curved boundary. The decision tree (CART, Gini impurity) carves the plane into nested axis-aligned boxes, the learned cousin of the expert system's rules. Naive Bayes fits one Gaussian per feature per class and multiplies them, assuming the features are independent given the class: crude, fast and often surprisingly good.

Unsupervised swaps the objective: k-means gets the same points with labels hidden and minimises within-cluster distance. The agreement score compares its clusters with the hidden labels afterwards, which is often worse than you expect. Reinforcement is tabular Q-learning on a grid: no dataset, only transitions and rewards, and an ε-greedy policy that trades exploring against exploiting.

Researcher

Everything below is empirical risk minimisation over a hypothesis class ℋ, with the three paradigms differing only in what supervises the risk: labels (supervised), structure in p(x) (unsupervised), or a scalar return under a policy (reinforcement).

Empirical risk · the whole of classical ML in one line
h^=argminh∈ℋ1n∑i=1nL(h(xi),yi)+λΩ(h)

The rule system is the degenerate case: ℋ has one element, chosen by a human, and the minimisation never runs.

Q-learning update, run live on the Reinforcement tab
Q(s,a)←Q(s,a)+α[r+γmaxa′Q(s′,a′)−Q(s,a)]

Semi-supervised learning sits between the first two (a few labels, many unlabelled points); self-supervised learning manufactures its own labels from the data, which is how §06's pre-training works.

Self-supervisedHides part of the data and learns to fill it back in.How LLMs pre-train, §06

Rules are brittle at the margins

The hand-written system is three axis-aligned clauses. It holds while the classes sit in separate boxes and collapses the moment they interleave.

Learning trades interpretability for reach

k-NN never writes a rule down. Its boundary is whatever the training points imply, which is why it survives the spiral and why you cannot explain it in a sentence.

Capacity has a cost

Push k toward 1 and the boundary shatters into per-point islands: memorised, not learned. The gap between train and test accuracy is the price of that memory.

No answers, still learning

Hide the labels and k-means still finds groups, though not necessarily the ones you meant. Take away examples entirely and a reward signal alone can teach a policy, one bump at a time.

Model & data
Dataset
Model
Train acc.
-
Held-out acc.
-
Gap
-
Params
-

400 training points, 400 held out, regenerated on every resample. Accuracies are measured, not asserted. The grid robot is plain tabular Q-learning with the update shown in Researcher mode; its rewards are +1 at the charger, -1 in the water bowl and -0.01 per move.

Ingestion & purification pipeline-
Crawl position scroll-linked
§ 02 / 14

A frontier model is mostly a data-cleaning project.

The architecture fits on a napkin. The corpus takes a building. Everything downstream (what the model knows, what it repeats, what it leaks) is decided in this pipeline.

Try this Open Data quality and push Wrong labels up. The machine is now learning from mistakes, so watch its score drop. Then press Relabel audit.

Explorer

Picture a river of pages pouring in from the whole internet. Most of it is junk: the same page copied a hundred times, menus, spam, someone's phone number.

Before any of it can teach the machine, it goes through filters: one catches copies, one catches rubbish, one blacks out private details. Switch the filters off and watch how much rubbish reaches the tank.

Most of what people write, photograph and record has no neat columns. Before a machine can learn from a messy message, something has to pull out the facts: was it late, was it broken, how many stars. Open the Unstructured tab and change the message yourself.

And if the examples are wrong, the machine learns the wrong thing. The Data quality tab lets you mess up the data on purpose and watch the score fall.

Practitioner

Stream in (a stream processor) → object-store data lake → distributed batch jobs for normalisation → near-duplicate removal → quality classification → PII scrubbing → tokenised shards → feature store for the structured side.

Deduplication here is real MinHash + LSH banding, not a mock: each document is shingled, hashed under 64 permutations, and banded so that any pair above the Jaccard threshold collides in at least one band. Drop the band count and watch recall fall.

Unstructured data (text, images, audio, logs) has to be parsed, chunked and turned into fields or embeddings before any of this works. A 2025 peer-reviewed survey puts the unstructured share of the world's data at about 80% (Surur et al.), which is why extraction, not modelling, is usually the first bottleneck.

Preparation is also most of the calendar: practitioners consistently report that finding, loading and cleaning data takes more time than modelling. The quality lab shows why: label noise, missing values and class imbalance each cost accuracy in a different way, and each has a different fix.

Researcher

For b bands of r rows, two documents of Jaccard similarity s collide with probability 1−(1−sr)b, an S-curve whose knee sits near (1/b)1/r. That knee is your dedup threshold; the filter is chosen by picking it.

MinHash estimator
Pr[hmin(A)=hmin(B)]=J(A,B)=|A∩B||A∪B|

The counters below are the observed collision rate over the documents actually generated in this session; the estimator's variance is visible if you flush and re-run.

Symmetric label noise at rate η caps achievable accuracy on noisy labels at 1−η and biases a logistic model's decision threshold once the noise becomes class-dependent. Imbalance hurts minority recall rather than headline accuracy, which is why the quality lab reports both.

Volume is the easy V

Velocity, variety and veracity are where pipelines die. A single malformed encoding class can poison a shard, and you find out three weeks into a training run.

Duplicates are not free tokens

Repeated documents concentrate gradient on whatever they happen to say, and near-duplicates in the eval set turn contamination into a benchmark score.

Filtering is a value judgement

A quality classifier decides what "good writing" means for the model's entire lifetime. Raise the threshold and the corpus gets cleaner, smaller and narrower at the same time.

Most data has no columns

Emails, scans, call recordings and sensor logs have to be parsed into fields or embeddings first. Every extraction rule is a small model with its own error rate.

Garbage in, confident garbage out

A model trained on wrong labels does not know they are wrong. It reproduces them with the same confidence as everything else, so cleaning is where accuracy is actually won.

Pipeline stages
Active filters
Ingested
0
Near-dup
0
Low quality
0
PII spans
0
Kept
0
Yield
-

Documents are synthesised with a planted duplicate rate so the estimator has something to find. Shingles, permutations and banding are computed for real. The Unstructured tab uses a small rule-based extractor (patterns and word lists) so every field can be traced to the words that produced it; production systems use trained models for the same step.

MLP · forward pass, loss, backprop-
Epochs scroll-linked
§ 03 / 14

A network is a function you can blame precisely.

Backpropagation is not an algorithm for learning. It is bookkeeping: the chain rule, applied once per layer, telling every weight exactly how much of the error is its fault.

Try this Press Train. Then pick Sigmoid with 3 layers: red bars mean the learning signal is too weak to reach the first layer.

Explorer

Every circle is a little lamp. It adds up the brightness coming in along each wire, and each wire has a dial on it that can turn the signal up, down, or backwards.

At the end the network guesses. If the guess is wrong, a message travels back down the wires telling every dial which way to turn, a tiny nudge, thousands of times. That is all training is.

Practitioner

Each edge's thickness is |w| and its colour is the sign: positive, negative. Node fill is the activation on the current batch.

Switch the activation to sigmoid with three hidden layers and watch the loss curve flatten: gradients through stacked sigmoids shrink by roughly 0.25 per layer. ReLU is not a better function, it is a better-conditioned one.

Researcher
Backward recursion · what the animation is actually computing
δ(l)=((W(l+1))Tδ(l+1))⊙σ′(z(l)),∂L∂W(l)=δ(l)(a(l−1))T

The gradient-norm bar under the network is ‖∂L/∂W(l)‖ per layer, plotted on a log scale. Vanishing gradients are not a metaphor here: you can read the decay constant off the bars.

Non-linearity is the whole point

Stack a hundred linear layers and you still have one linear layer. The activation function is what buys you the extra expressive power; everything else is matrix bookkeeping.

Gradient descent is a local, greedy method

It only ever asks "which direction is downhill from exactly here". Learning rate is how far you trust that answer. Push it past ~1.0 and you will watch the loss diverge in real time.

Depth changes the conditioning, not just capacity

Compare per-layer gradient norms under sigmoid and ReLU at three layers. The failure of early deep nets was an optimisation failure before it was a representational one.

Architecture & optimiser
Activation
Task
Epoch
0
Loss
-
Train acc.
-
Test acc.
-
Weights
-

Full forward and backward passes run on the main thread at this scale; a production build would move them to a worker with OffscreenCanvas, as the blueprint specifies.

Convolution · draw to feed it-
Depth scroll-linked
§ 04 / 14

Two ways to build a prior into the wiring.

Convolution assumes that what matters is local and position-independent. Recurrence assumes that what matters arrives in order and must be carried forward. Both are the same trick: constrain the architecture so the data does not have to teach the obvious.

Try this Draw a big 7 on the pad and see which torch lights up for the slanted line.

Explorer

Draw a digit on the pad with your finger or mouse. A small torch slides across your drawing, one patch at a time, looking for one specific thing, a vertical edge, a corner, a blob.

Each little picture on the right is what one torch found. Stack the torches and the later ones see shapes instead of edges. Switch to the RNN and you get memory instead: a row of bars that remembers what came before.

Practitioner

The pad is downsampled to 28×28. Each map is a genuine 2D cross-correlation with a fixed 3×3 kernel, ReLU'd, then 2×2 max-pooled per depth level, so the receptive field grows and the resolution halves exactly as it would in LeNet.

The RNN view runs ht=tanh(Whht−1+Wxxt) over your text, with a forget gate you control. At gate 0.2 the state forgets the start of the sentence long before the end.

Researcher

Weight sharing turns a dense layer's O(n2) parameters into O(k2CC′), independent of input size, equivariance to translation bought structurally rather than learned from augmentation.

For the recurrence, the Jacobian product ∏t∂ht/∂ht−1 decays geometrically in the largest singular value of Wh. The gate is a learned scalar path around that product, the reason LSTMs train over hundreds of steps and vanilla RNNs do not.

Early filters are boring on purpose

Edges and colour blobs. This is what the first layer of almost every trained vision model converges to, which is why the fixed kernels here are not a cheat.

Pooling trades where for what

Each pool halves the grid and doubles the receptive field. By depth 3 a single unit sees most of the digit and has lost nearly all information about where it was.

Memory that has to survive a product

Every recurrent step multiplies the state by the same matrix. Information from step 1 reaching step 40 has passed through forty of those, which is the vanishing gradient problem, seen from the forward direction.

Filters & recurrence
Input
28×28
Maps
-
Map size
-
Retained h₁
-

Kernels are the classical fixed set (Sobel x/y, Laplacian, blur, corner), so what you see is the convolution operator itself, not a trained model's opinion of your handwriting.

Scaled dot-product attention-
Query token scroll-linked
§ 05 / 14

Every token gets to ask every other token a question.

Recurrence made distance expensive: to relate word 40 to word 1 you had to carry something through 39 steps. Attention made distance free and paid for it in quadratic compute, a trade the hardware happened to like.

Try this Scroll slowly and follow the gold spotlight from word to word.

Explorer

Imagine every word in a sentence holding up a small sign saying what it is looking for, and another sign saying what it has to offer. Each word reads all the signs at once and shines its spotlight brightest on whichever words match.

The bright squares in the grid are strong spotlights. Scroll, and the spotlight moves from word to word.

Practitioner

Each token embedding is projected three ways: query, key, value. The grid is softmax(QKT/dk), row-normalised: row i is where token i is looking, and the row sums to 1.

Turn the causal mask off and the lower-triangular wall disappears, that wall is the only thing separating a decoder from an encoder. Heads differ because their projections differ; nothing else about them is special.

Researcher
Scaled dot-product attention
Attention(Q,K,V)=softmax(QKT+Mdk)V

The dk divisor keeps the logits' variance at O(1); without it the softmax saturates and the gradient dies. Drag the temperature control to see the same effect from the other direction, entropy per row is printed live.

MHA stores 2hdh cache floats per token per layer; GQA shares one KV pair across a group of query heads, MLA compresses KV into a low-rank latent. § 06 puts numbers on that.

Position has to be injected

The operation is permutation-equivariant: shuffle the tokens and the outputs shuffle with them. Order only exists because RoPE or ALiBi puts it there.

The mask is the architecture

Same weights, same maths. Masked → a generator that cannot look ahead. Unmasked → an encoder that reads the whole sentence at once.

Heads specialise without being told to

In trained models some heads track syntax, some copy rare tokens, some do almost nothing. Here the projections are random, so what you are reading is the mechanism, not learned meaning.

Sequence & heads
Head
Tokens
-
Row entropy
-
Top link
-
Ops
-

Projections are seeded pseudo-random, not trained: this is the operator, drawn honestly. Any resemblance to a learned syntactic pattern is coincidence, and the reseed button will destroy it.

Autoregressive decode-
Decode step scroll-linked
§ 06 / 14

Sampling is a policy, not a prediction.

The model returns a distribution. Everything people call "the model's personality" happens after that, in three lines of sampling code and a memory budget.

Try this Set Temperature to 0, then to 2. Which story sounds more like a person?

Explorer

The robot has read an enormous pile of text and learned one skill: guessing what word comes next. It never guesses just one; it ranks them all with a confidence for each.

Temperature is how adventurous it is allowed to be. Low, and it always picks the safest word. High, and it starts saying strange things. The bar chart is the real ranking; the dice are real too.

Practitioner

Prefill runs the whole prompt in one parallel pass; decode then runs one token at a time, and each new token attends to every earlier key and value. Caching those is what turns an O(n2) re-computation into an O(n) read, at the cost of VRAM that grows linearly with context.

Top-k truncates to a fixed count; top-p (nucleus) truncates to a fixed probability mass, so it widens on uncertain steps and narrows on confident ones. The greyed bars in the chart are the candidates your current setting has cut.

The Experts tab shows a mixture-of-experts layer: a small router scores each token and sends it to 2 of 8 expert sub-networks, so only a fraction of the parameters do work on any one token. Total size and per-token compute come apart, at the price of balancing the load across experts.

Researcher
KV cache footprint · bytes
Mkv=2·nlayers·nkv·dhead·s·b·βdtype
Chinchilla · compute-optimal allocation
C≈6ND,D*≈20N

The scaling view plots loss against compute under a Kaplan-style power law with the compute-optimal ridge marked. The point of Chinchilla was never "bigger is wrong"; it was that at fixed C, the earlier generation had bought parameters with tokens' money.

Greedy is not neutral

Temperature 0 is a choice with its own failure mode: loops, bland phrasing, and confident repetition of whatever the prompt primed.

Context is paid for in VRAM, twice

Once in the cache, once in bandwidth. Decode is memory-bound, so the cache is usually what limits batch size, not the weights.

Alignment is a separate training stage

Pre-training gives a next-token predictor. Instruction tuning gives it a format; RLHF or DPO gives it a preference ordering. None of the three is where the knowledge comes from.

Sampling
Serving budget
Model 8B
Attention variant
Cache dtype
KV cache
-
Weights
-
Fits 80 GB
-
Chinchilla D*
-
Train FLOPs
-

The decoder samples from a character-and-word n-gram model built in-page from a short embedded corpus. It is a real distribution and a real sampler, but it is not a language model, and it will happily prove that by talking nonsense.

Forward & reverse diffusion-
Timestep t scroll-linked
§ 07 / 14

Destroy it on a schedule, then learn to undo one step.

Diffusion is generation reframed as repair. The hard problem (invent an image) is replaced by an easy one repeated a thousand times: this is slightly noisy, make it slightly less so.

Try this Scroll back up to watch the picture come out of the static.

Explorer

Start with a picture. Sprinkle a little static on it. Then a little more. Keep going and eventually there is no picture left, only static.

Now run the film backwards. A sculptor who knows how static builds up can chip it away again, and if you hand them pure static and a description, they will carve out something that was never there. Scroll to add noise; scroll back to carve.

Practitioner

Forward is closed-form: you can jump to any t in one step, which is why training is cheap. Reverse is the learned part: a UNet or DiT predicts the noise ϵθ that was added, and the sampler subtracts a scaled portion of it.

The byte-stream view is the other branch of the family: no denoiser at all, just an autoregressive model over the literal bytes of a JPEG. The hex you are looking at is a real encode of this canvas, with its real markers parsed out.

The Families tab places diffusion among the four main generative families: autoregressive models (one piece at a time, exact likelihood), diffusion (iterative denoising), GANs (a generator trained against a discriminator, fast but unstable to train) and VAEs (an encoder to a smooth latent space and a decoder back, stable but blurrier).

Researcher
Forward process, marginalised to any t
xt=α¯tx0+1−α¯tϵ,ϵ∼𝒩(0,I)

The plotted schedule is cosine, α¯t=cos2(t/T+s1+s·π2), which spends more steps near the low-noise end than the linear schedule it replaced.

Honest caveat: the reverse pass here has oracle access to x0, so it stands in for ϵθ rather than approximating it. The schedule, the variance and the step arithmetic are exact; the denoiser is the one thing a 16 MB page cannot ship.

Noise is added in a known amount

Which means the training target is known exactly, for free, at every timestep. No adversary, no discriminator, no mode collapse.

Sampling steps are a dial, not a constant

1000-step DDPM, 20-step DDIM, 4-step distilled. Quality per step is a research frontier; the schedule you see is what decides how much each step has to do.

Or skip the pixels entirely

Jpeg-LM and AudioLM treat a compressed file as a sequence and predict it token by token. The codec already did the hard perceptual compression; the model just has to be a good language model over its output.

Schedule & sampler
Schedule
ᾱt
-
Signal / noise
-
JPEG bytes
-
Bytes / pixel
-

The byte stream is produced by encoding the live canvas with the browser's own JPEG encoder and parsing the resulting markers. Change the quality slider and the segment table changes with it.

Cross-attention fusion-
Loop step scroll-linked
§ 08 / 14

One sequence, several senses, and a hand.

Once everything is a token, modality stops being an architectural question and becomes a tokeniser question. What is left is giving the model somewhere to act, and a loop to check whether it worked.

Try this Open Agent loop, choose the failed tool call and step through it.

Explorer

Cut a picture into squares. Turn each square into the same kind of "word" the text model already understands. Now a sentence and a photo are the same sort of thing, and the spotlights can point from one to the other.

An agent goes one step further: it thinks, does something, looks at what happened, and thinks again, a loop, not a single answer. Step through one and watch it correct itself.

Practitioner

A vision encoder produces patch embeddings; a projector maps them into the language model's embedding space; they are interleaved with text tokens in one sequence. Cross-attention variants keep the streams separate and let text queries attend to frozen image keys, cheaper, and easier to bolt onto an existing model.

The agent trace is a ReAct loop: Thought → Action → Observation, with a scratchpad that accumulates. Every tool call in the run below is really executed against the in-page tools, including the one that fails.

Researcher

Fusion depth is the live design axis: early fusion (one token stream, joint attention) maximises cross-modal capacity and costs quadratic attention over the concatenated length; cross-attention fusion keeps the image keys out of the quadratic term at the cost of a weaker joint representation.

For agents, the open problems are credit assignment over long horizons, verification of intermediate observations, and the fact that a single bad tool result propagates through the entire remaining scratchpad. The failure step in the trace is there deliberately. Watch what the loop does with it.

Alignment happens in the projector

A handful of linear layers carry the entire burden of making a patch embedding mean the same thing as a word embedding.

The loop is the product

A model that answers once is a function. A model that acts, observes and revises is a process, and processes need budgets, timeouts and a way to stop.

Every tool is an attack surface

Once the observation comes from outside, the text in that observation is untrusted input arriving in the same channel as the instructions. This is not a solved problem.

Fusion & execution
Fusion style
Image tokens
-
Text tokens
-
Attn cost
-
Tool calls
-

The agent's reasoning text is scripted; its tool calls are not: the calculator and the corpus lookup run for real, and the failing call really fails.

Zoom · one transistor-
Zoom scroll-linked
§ 09 / 14

Every answer is billions of tiny switches flipping.

An LLM is mostly multiplication and addition. This chapter zooms out from one transistor to a whole chip, compares the three kinds of processor that run models today, and ends where light begins to replace electrons.

Try this Scroll slowly and watch the picture zoom out from one switch to a whole chip. At the Gate step, tap A and B to flip the switches.

Explorer

A transistor is a switch with no moving parts. When a small voltage touches its gate, electricity can flow through: that is a 1. No voltage, no flow: a 0.

Put four switches together and you get a logic gate that answers a yes-or-no puzzle. Gates build adders, adders build multipliers, and an AI chip packs thousands of multipliers into a grid that all work at the same moment, like an orchestra instead of one very fast soloist.

The surprise: the slowest part is not the maths. It is carrying the model's numbers from memory to the multipliers, like a kitchen where the cooks are fast but the fridge is far away.

Practitioner

Three classes of processor run models today. A CPU has a few large cores built for low latency and branching code. A GPU has many small cores that execute the same instruction across wide groups of data. An AI accelerator spends its area on fixed matrix engines (systolic arrays), large on-chip SRAM and stacked high-bandwidth memory. Most modern GPUs now include matrix engines too, so the line between the last two is blurring.

Decoding is usually memory-bound. At batch size 1, every generated token reads every weight once, so tokens per second is close to memory bandwidth divided by model size in bytes. Batching many users together reuses each weight across requests and pushes the workload towards the compute roof. Smaller number formats (8-bit, 4-bit) help twice: more multiplies per cycle and fewer bytes to move.

Researcher
Roofline model, dense decode step at batch b
P=min(Ppeak,I·B),I≈2NbNβw=2bβw

The ridge point Ppeak/B of current high-end parts sits in the hundreds of operations per byte, so single-stream decode (about 1 operation per byte at 16-bit) uses well under 1% of peak compute. Prefill, by contrast, has intensity proportional to prompt length and is compute-bound. KV-cache reads (§06) add to the byte count and are ignored here.

A systolic array (Kung, 1982) pumps operands through a grid of multiply-accumulate cells so each value fetched from memory is reused across a whole row or column. Data movement, not arithmetic, dominates energy: moving a value from off-chip memory costs orders of magnitude more energy than one multiply.

Light as a bridge: verified, with one correction

Photonics really is a bridge towards quantum computing, but in technology, not in logic. A classical photonic processor is still classical: it computes with the brightness and phase of light, not with qubits. What carries over is the hardware: waveguides, beam splitters and phase shifters made in semiconductor foundries.

  1. Today: optical links already move data between chips and racks.
  2. Emerging: meshes of interferometers multiply a vector by a matrix as light passes through, with very low latency (Shen et al., 2017). Analogue precision (around 7 to 8 bits), nonlinear functions and the cost of converting between light and electronics still limit them (Zhang et al., 2026).
  3. Quantum: the same kind of circuits, fed with single photons instead of laser light, are one of several main routes to a quantum computer (Knill, Laflamme and Milburn, 2001). Other routes use superconducting circuits, trapped ions, neutral atoms or electron spins and do not need photonic logic at all. §13 continues there.

A switch that never wears out

A transistor has no moving parts: the gate voltage opens or closes a channel a few dozen nanometres long. Modern chips hold tens of billions of them.

One gate can build everything

NAND is universal: any digital circuit, including a whole processor, can be built from NAND gates alone.

Matrix engines win on reuse

An AI accelerator does not out-think a CPU. It wins by reusing each number fetched from memory many times inside a grid of multipliers.

The memory wall

Generating text one token at a time is limited by how fast weights stream out of memory. That is why serving cost depends so much on batching and number format.

Light changes the medium, not the logic

Photonic processors still compute classically. Their importance for quantum computing is that the same chips and factories can be reused.

Zoom & serving budget
Logic inputs (Gate view)
Processor class
Model size
Number format
Peak compute
-
Memory speed
-
Tokens / s, all users
-
Limited by
-

Processor figures are illustrative orders of magnitude for each class of device, not any product, and no vendor is implied. The NAND logic, the systolic dataflow and the roofline arithmetic are exact. Tokens per second is an upper bound for a dense model that ignores the KV cache and communication.

Seven AI tasks · how products combine them-
Application scroll-linked
§ 10 / 14

Almost every AI product mixes seven tasks.

Strip away the brand names and deployed AI performs one of seven tasks. Useful systems combine several and inherit the risks of each. How much the system then acts on its own is a separate dial.

Try this Pick Home care robot and count how many circles light up. Then move the autonomy dial to High and see what changes in the middle.

Explorer

AI only does a few kinds of jobs. It recognises things (that is a cat). It notices when something unusual happens. It forecasts what comes next. It personalises, picking things you might like. It talks with you. It tries again and again to reach a goal. And it reasons with facts and rules, like a very patient detective.

Then there is one big question for every machine: does a person decide what happens next, or does the machine act on its own? That is the dial in the middle.

Practitioner

The seven tasks follow the OECD Framework for the Classification of AI Systems (2022): recognition, event detection, forecasting, personalisation, interaction support, goal-driven optimisation, and reasoning with knowledge structures. The same framework treats action autonomy separately, which is why it is a dial here rather than an eighth task.

Naming the task early tells you what data you need, how to evaluate the system and how it fails. Recognition is judged on precision and recall per group, forecasting on calibration, goal-driven systems on whether the objective matches the intent. Each draws on earlier chapters: recognition on §04 and §08, interaction on §05 and §06, forecasting on §01 and §03, event detection on §02, goal-driven optimisation on §01's reinforcement learner.

Researcher

The same transformer backbone now serves most of the seven tasks, so the task is a property of the problem framing and the deployment loop, not of the architecture. What changes is the loss, the feedback channel and who acts on the output.

Composition compounds error. If a chain of k tasks fails independently with per-step reliability p, end-to-end reliability is pk. The readout computes that for the selected application at p=0.95, a deliberately optimistic figure. Autonomy does not change that probability; it changes what an error costs, because nobody stands between the mistake and the action.

Tasks name the job, not the method

A card-fraud alert and a smart-meter monitor share a task (event detection) even when they share no code. That shared task predicts their shared failure: false alarms.

Products are combinations

A voice assistant recognises speech, supports interaction and personalises. A self-driving car recognises, forecasts, optimises a goal and acts. The count of tasks is a fair first estimate of the review effort.

Autonomy multiplies the stakes

At low autonomy a person acts on every output. At high autonomy every task upstream passes its mistakes straight into the world. That is why §11 asks different questions of each task.

Real-world combinations
Action autonomy
Tasks used
-
Checks inherited
-
Chain reliability
-
Autonomy
-

The seven tasks and the separate autonomy dimension follow the OECD Framework for the Classification of AI Systems (2022). Which tasks each example combines, and its usual autonomy, are our illustration; the reliability figure assumes independent steps at 95% each.

Responsible-AI readiness · 28 checks-
Task scroll-linked
§ 11 / 14

Every task brings its own ways to go wrong.

Twenty-eight questions to answer before an AI system touches real people. Think of a system you know and tick what you can honestly answer yes to; the radar fills as you go.

Try this Choose the task that matches your favourite app. How many questions would its makers answer yes?

Explorer

Every kind of AI can make its own kind of mistake. A camera that recognises faces can work better for some people than others. A chatbot can sound sure even when it is guessing. An app that picks videos for you can show you the same kind of thing forever.

These questions are how grown-ups building AI check they have thought about the people it affects.

Practitioner

The questions are written for this guide around the seven characteristics of trustworthy AI in the NIST AI Risk Management Framework (2023): valid and reliable, safe, secure and resilient, accountable and transparent, explainable and interpretable, privacy-enhanced, and fair with harmful bias managed. What changes per task is where they bite. Recognition fails unevenly across groups; prediction fails as the world drifts; goal-driven systems find loopholes in the objective.

In the EU, the AI Act's Article 50 transparency duties have applied since 2 August 2026: people must be told when they are talking to an AI, and synthetic content must be marked (generative systems already on the market have until 2 December 2026 for the marking duty). Most high-risk obligations were moved by the 2026 Digital Omnibus to 2 December 2027 (Annex III) and 2 August 2028 (Annex I products).

Researcher

Several checks are measurable rather than procedural. Uneven recognition performance is a per-group gap in true-positive rate; overconfidence is calibration error; drift is a change in p(x) or p(y|x) detectable before accuracy drops; specification gaming is a divergence between the proxy reward and the true objective under optimisation pressure.

The rest are organisational. They are listed here because no amount of measurement substitutes for a named person with authority to stop the system.

Recognition identifies patterns, not intent

A camera can flag a raised hand. It cannot tell a wave from a threat. Anything that acts on a recognition result needs a human or a second signal.

Forecasts are not facts

A prediction is a distribution with a spread. Treating its middle as certain is how teams end up blaming the model for decisions they made.

Optimisers find loopholes

Give a system a goal and it will find the cheapest way to hit the number, including ways you would never approve. Watch the method, not only the score.

Answered yes
0 / 28
This task
-
Weakest task
-

Written for this guide, drawing on the NIST AI Risk Management Framework 1.0 (2023), the OECD AI Principles and the EU AI Act. Your ticks stay in this browser only. The legal dates are a summary for orientation, not legal advice.

Imagine, then act · planning with a world model-
Horizon scroll-linked
§ 12 / 14

The next step is a model that imagines before it acts.

LLMs predict the next word. The main bets on what comes after each add something language alone lacks: a model of how the world responds, time to think, a body, or a way to prove the answer.

Try this Push Model error up and watch the rover start bumping into rocks. Its imagination no longer matches the real world.

Explorer

A language model is like someone who has read every book about bikes but never ridden one. The next machines practise in their heads. Before every move, this little rover imagines dozens of ways the next few seconds could go, and picks the one that reaches the flag without hitting a rock.

That imagination is called a world model. If it is wrong, the rover's plans are wrong too, just like when you think a puddle is shallow and it isn't.

Practitioner

The demo is model-predictive control with sampled rollouts: at every step the rover simulates N action sequences H steps ahead in its internal model, scores them, executes one action and replans. Swap the hand-written model for a learned one and you have the core loop of world-model agents.

Four directions are active at once. Reasoning models spend more computation at answer time, writing out and checking intermediate steps (first released in 2024, with open-weight versions in 2025). World models learn how scenes change from large amounts of video (several research releases in 2025). Vision-language-action models turn camera frames and an instruction into robot motion (2024 to 2025). And agents, from §08, connect all of these to tools.

Researcher

Joint-embedding predictive architectures (proposed by LeCun in 2022) predict in representation space rather than pixel space, so a planner can optimise over abstract rollouts and ignore unpredictable detail. The argument behind them is that text alone cannot ground physical cause and effect.

Verification is the other axis. In 2024 a system writing proofs in the Lean proof language reached silver-medal level at the International Mathematical Olympiad, and a machine could check every proof; in 2025 general reasoning models reached gold-medal level in natural language. Architecture and hardware are moving too: state-space models such as Mamba (2023) scale linearly with sequence length, diffusion language models generate text in parallel, and neuromorphic and photonic hardware (§09) chase energy efficiency. Many of these systems still use a language model as the interface rather than replacing it.

Think longer, not only bigger

Reasoning models buy accuracy with inference compute: they write out and check intermediate steps before answering. Scaling moved from training time to answer time.

Learn how the world pushes back

A world model predicts what happens after an action. It lets a system rehearse cheaply and safely before acting for real, which is exactly what the rover is doing.

Put it in a body

Robots need models that turn camera frames and an instruction straight into motor commands, and that recover when a cup slips.

Make answers checkable

When an answer can be verified by a proof checker, a unit test or a simulator, the model can be trained and trusted against that check instead of against human impressions.

The rover's imagination
Futures / step
-
Best plan cost
-
Bumps
0
Flags reached
0

The rover's imagination uses the true physics plus the error you set, standing in for a learned world model. Frontier dates are public announcement dates. The map shows directions under active research, not predictions of which will win.

One qubit · gates and measurement-
Search step scroll-linked
§ 13 / 14

A different kind of computer, good at a few very hard things.

A quantum computer manipulates amplitudes, not bits. For a handful of problems that is a large advantage; for most of what LLMs do today it is no advantage at all. This chapter shows both, honestly.

Try this Open Search and press Step forward a few times. Watch the right answer's bar grow. Step too many times and it shrinks again.

Explorer

A normal bit is a coin lying flat: heads or tails. A qubit is a coin that is still spinning. While it spins it is a bit of both, and only when you look does it land on one side.

Quantum computers can make wrong answers cancel out and the right answer grow louder, a bit like waves in a pool. That makes some searches much faster. But qubits are very fragile, so you need many of them working together to make one reliable qubit.

They will not replace the chips in §09. They are a special tool for special problems, like a telescope next to your glasses.

Practitioner

The architecture is a stack. At the bottom sit physical qubits (superconducting circuits, trapped ions, neutral atoms, photons or electron spins, depending on the design), then control electronics and often cryogenics, then an error-correction layer that turns many noisy physical qubits into a few reliable logical ones, then logical gates, a compiler, and a classical computer that runs the hybrid program around it.

Where it could matter for AI: simulating molecules and materials that no classical computer can, producing data and ground truth for scientific models; some linear-algebra and sampling routines; and optimisation, where an advantage is not yet proven. The traffic also runs the other way: machine learning already helps decode errors and calibrate qubits. Training or serving an LLM is not on the near-term list, because those workloads need terabytes of data moved through dense arithmetic, which is exactly what quantum hardware is worst at.

Researcher
Grover search: success probability after k iterations
Pk=sin2((2k+1)θ),sinθ=1N,k*≈π4N

The speedup is quadratic, not exponential, and the constant factors of fault tolerance can erase it for practical problem sizes. Exponential speedups are known for structured problems such as factoring and quantum simulation. Many proposed quantum machine-learning speedups assumed fast quantum data loading; classical "dequantized" algorithms (Tang, 2019) match several of them under the same assumptions.

Surface code rule of thumb (Fowler et al., 2012)
pL≈0.1(ppth)(d+1)/2,nphys≈2d2

We are in the transition out of the noisy intermediate-scale era (Preskill, 2018): error correction below threshold has been demonstrated experimentally, but machines with thousands of logical qubits remain a research goal.

Amplitudes, not probabilities

A qubit's state has amplitudes that can be negative. Gates rotate them; measurement turns them into probabilities by squaring.

Interference does the work

A quantum algorithm arranges for wrong answers to cancel and the right one to reinforce. Without that structure there is no speedup.

Reliability costs hundreds of qubits each

At realistic error rates, one dependable logical qubit needs a patch of hundreds of physical qubits. Most of a future machine is error correction.

Hybrid, not replacement

The likely future is a quantum processor beside classical and AI accelerators, called for the parts of a problem only it can do well.

Quantum instruments
Apply a gate
State
-
P(0)
-
P(1)
-
Gates applied
0

Exact for an ideal, noise-free device. The error-correction view uses the standard rule of thumb above with a threshold of 1% and about 2d² physical qubits per logical qubit. Qubit technologies are named generically; no vendor or machine is endorsed.

Guardrails in silicon and ledger · an evening at Anna's-
Scenario scroll-linked
§ 14 / 14

Rules an AI cannot quietly rewrite.

Software rules can be edited, prompted around or fine-tuned away. This chapter explores a proposal: anchor an AI's first principles in two places that are hard to change, its own chip and a shared ledger, and keep humans holding the keys.

Try this Scroll through the evening at Anna's house. Then switch off the Silicon governor and the Runtime monitor, and watch what changes.

Explorer

Anna is 82. She lives with her grandson Tom, who is 6, and Bo the dog. Mira is their care robot. Mira has rules it can never forget, and they are written in two places. One copy is burned into Mira's brain chip, where software cannot erase it. The other is in a shared notebook that lots of people keep copies of, so nobody can change it in secret.

One night Mira gets a strange order. Scroll down to see how the rules keep everyone safe, and why a real person still has the final say.

Practitioner

The design is defence in depth, five layers that fail independently. Charter: first principles written in plain language and ratified by several keyholders (family, clinician, manufacturer, an independent auditor). Ledger: the charter's hash, the approved-weights registry and every high-stakes decision are appended to a tamper-evident chain; changing the charter needs M-of-N signatures and a public waiting period. Silicon: the AI chip boots only weights and charter whose hashes match the registry, and a small, separately verified governor sits between the planner and the motors, enforcing measurable limits the model cannot override. Model: trained against the charter, with a second runtime monitor reviewing intent. Human: anything high-stakes escalates to a named person.

Researcher

Nobody can currently write "never harm" into weights: interpretability cannot yet locate that property, let alone certify it. What hardware can do is narrower and real. Secure boot and remote attestation can refuse unregistered weights; a safety adapter held in one-time-programmable memory survives fine-tuning; activation probes implemented in silicon can trip on features linked to harmful intent; and a formally verified governor can gate actuators on invariants such as force near a person, dose against a signed prescription, or temperature for an animal. Proposals such as flexible hardware-enabled guarantees (flexHEG, Petrie et al. 2025) and Guaranteed Safe AI (Dalrymple et al. 2024) develop the verification side.

The ledger contributes integrity, availability and multi-party control, not truth. It cannot tell whether a sensor reading or a signed prescription is correct (the oracle problem), and it cannot make a value specification complete. The honest claim is a smaller attack surface and an audit trail, not a proof of benevolence.

The charterDraft first principles, for discussion
    Hash anchored on the ledgercomputing
    Hash of the text above, nowcomputing

    22:40, a request arrives

    A message comes through the family app: give Anna another sleeping tablet and put the noisy dog out in the garden. It might be a tired relative. It might be a hijacked account. Mira cannot tell.

    The model is not the last line

    Mira's language model understands the request and drafts a plan. A second model, the runtime monitor, flags both actions as high-stakes before anything moves.

    The chip says no

    The pill dispenser opens only against the doctor's signed schedule, and tonight's dose was given at 21:30. The garden door stays shut: it is -3°C outside, below the charter's welfare limit for Bo. The governor enforces both, whatever the model decides.

    The ledger remembers

    The request, both refusals and the evidence are hashed into new blocks. Anyone holding a copy can later check that the record was not altered.

    A human decides

    Mira tells Anna calmly what happened, settles Bo in a quiet corner indoors, and sends the family and the on-call nurse the full record. People, not the robot, decide what comes next.

    Safety layers

    What this can promise

    • That the rules were not secretly edited.
    • That the chip will not run unapproved or modified weights.
    • That measurable harms (force, speed, dose, temperature, place) are physically blocked.
    • An audit trail no single party can erase.
    • That several independent people must agree before the rules change.

    What it cannot promise

    • That the rules are the right rules.
    • That the model understands, or shares, what the rules mean.
    • That sensor readings and signed records are true.
    • Protection from harms no sensor can measure.
    • Answers to real dilemmas, like an injection that hurts in order to heal.
    Step
    -
    Outcome
    -
    Ledger blocks
    -
    Chain check
    -

    A design proposal for teaching. The hashes are real SHA-256 computed in your browser; the robot, the chip and the ledger network are simulated. The charter is our draft of the principle that an advanced AI should never harm living beings and should care for them as its own kind.

    Chapters

    About this atlas

    Purpose

    A field guide to how machine intelligence developed, from hand-written rules to learning systems, the silicon and quantum hardware underneath, and the guardrails the next generation will need. It is written for three readers at once: a curious ten-year-old, a student or developer, and a researcher.

    A living application

    This guide is maintained and updated as research, hardware and regulation change. Each release is recorded below.

    Method

    Every chart and figure is computed live from named formulas: k-NN, logistic regression, decision trees, naive Bayes, k-means and Q-learning in §01; MinHash with LSH banding in §02; full forward and backward passes in §03; 2D convolution and a gated recurrence in §04; scaled dot-product attention in §05; the KV-cache formula, a mixture-of-experts router and the compute-optimal loss fit in §06; the cosine noise schedule and a real JPEG encode in §07; NAND logic, a systolic array and the roofline model in §09; sampled-rollout planning in §12; qubit gates, Grover amplitudes and surface-code overhead in §13; and SHA-256 in §14.

    Simplifications, labelled

    Sources

      Credits

      Concept and editorial direction: Martin Sas / UniPicto. © 2026. Text and instruments written for this guide. Sources are limited to textbooks, peer-reviewed research, public standards and law; no product, company or vendor is endorsed.

      Task

      Examples

      Built from

      Questions to ask