← Lab 01 · slide 01

Denoise

A diffusion model on a training set small enough to see. Images are 64 by 64 greyscale. Noise is added on a schedule; a denoiser estimates the clean image from the noisy one; generation runs the denoiser from pure noise. The training set is ten images from a Stable Diffusion finetune Ariel made, plus anything you upload or draw.

The training set

0 / 12

Draw one

Forward: adding noise

clean image

with noise · 50%

denoiser’s estimate of the clean image

At low noise the estimate is the selected image. At high noise the noisy input could have come from any image in the set, so the estimate is their average.

Reverse: generating from noise

not started

Seed
Steps
Scheduler

The seed sets the starting noise; the same seed always gives the same image. DDIM adds no noise after the first draw, so similar training images pull it to the same result. DDPM adds fresh noise at every step, so the seed matters more. Steps: fewer steps, coarser path.

Step by step: from noise to an image

seed 1 · 25 steps · DDIM

The generation above, one step per slide. It uses the seed, steps and scheduler set there, so the last slide is the image Generate produces. Each slide shows three things the model has at that moment.

Loading the training set…

From words to numbers: how a prompt gets in

The model on this page has no text encoder; it only denoises. Latent diffusion adds one. These steps follow a prompt from words to the numbers that steer each denoising step. Steps one to three are illustrations with Stable Diffusion's real figures beside them. Step five runs for real on the images above.

  1. 1. The prompt is cut into tokens

    Pieces from a fixed vocabulary, not whole words. Common words are one token; longer ones split. </w> marks the end of a word, and every prompt is wrapped in start and end tokens. Stable Diffusion's text encoder reads at most 77 tokens; anything past that is dropped.

    <start>a</w>pizza</w>on</w>a</w>wooden</w>board</w>,</w>glowing</w><end>

    11 of 77 tokens used

    Simplified: the real tokenizer is byte-pair encoding over 49,408 pieces.

  2. 2. Each token becomes a number

    Its position in the vocabulary. Start and end are 49406 and 49407, as in CLIP; the others here are made up.

    <start> 49406a 40724pizza 10235on 17170a 40724wooden 36889board 8239, 43699glow 40070ing 34271<end> 49407
  3. 3. Each number becomes a list of numbers

    A vector: 768 numbers per token in Stable Diffusion 1.x, twelve shown here. The looked-up vector is the same wherever the word appears. The text encoder then mixes each token with the others, so after it runs the same word carries different numbers in different prompts. Bars up are positive, bars down negative; the values are illustrative.

    <start>
    a
    pizza
    on
    a
    wooden
    board
    ,
    glow
    ing
    <end>

    top row: looked up · bottom row: after the text encoder

  4. 4. The vectors steer every denoising step

    This is cross-attention. At each step, every position in the image grid asks which tokens matter to it and mixes in their vectors. The questions come from the image, the answers from the words:

    Attention(Q, K, V) = softmax(QKT / √d) V

    Q comes from the image, K and V from the prompt's vectors. The prompt never paints a pixel. It changes what the denoiser predicts the noise to be, at every one of the 25 steps.

  5. 5. Guidance decides how hard to follow it

    This part is real. Mark the training images that fit your prompt. The conditional denoiser only considers those; the unconditional one considers all of them. Each step uses unconditional + s × (conditional − unconditional), where s is the guidance scale. Same seed for every result.

    Mark at least one image to see guidance.

What is running here

Forward process: xt = √ᾱt x0 + √(1−ᾱt) ε, with ε gaussian and ᾱt on a cosine schedule from 1 down to 0.

Denoiser: the expected clean image given the noisy one. For a training set of a few images, that is exactly a weighted average of the set, each image weighted by how likely it is to have produced the input at this noise level. The bars show those weights. A trained network approximates this same quantity for a training set of billions of images it cannot store.

Reverse process: at each step the denoiser estimates x0, the noise is re-estimated from it, and the pair is recombined at the next lower noise level. DDIM adds no randomness after the first draw. DDPM adds a fresh gaussian at each step, scaled to the schedule. Both are reproducible from the seed.

With this few images the process lands on one of them exactly. That is what an ideal denoiser does on a small set. Back to Lab 01.

Pandaemonium Architecture 6.0 — ATEK-639/439 — Fall 2026 · QR code