← Back to the lab

Gaussian sampling

A Gaussian, or normal distribution, is the familiar bell curve. It describes how values are distributed around a center. Sampling produces one value at a time. Draw enough values and their collective shape starts to resemble the curve.

Draw from the curve

The mean μ moves the center. The standard deviation σ controls the spread. Values near the center are denser; values far out in the tails are less common. A sample can land on either side.

Theoretical densityObserved histogram
Gaussian density with mean 0.0 and standard deviation 1.0; 0 sampled valuesprobability density0.000.240.48-8-4048sample value x

Draws0

Latest value—

Sample mean—

Sample SD—

Changing μ or σ clears the samples. The horizontal scale stays fixed; the vertical scale fits the curve and bars. The Gaussian continues beyond the visible window: off-screen draws are counted, never clipped into the histogram.

Start with a single draw, then twenty, then a thousand. The histogram is what happened. The curve is the distribution we sampled from. An uneven batch does not mean the distribution changed.

Does a value below the mean make a value above it due?

No. These draws are independent and the distribution stays fixed until you move a control. Sampling does not compensate for previous outcomes. The sample mean tends toward μ over many draws, but it need not move closer after every new draw.

Probability is area

A bell curve shows probability density. To ask for a probability, choose an interval: what is the chance that a draw falls between these two values? The answer is the area under the curve between them.

Standard normal / μ = 0, σ = 1

Standard Gaussian with the interval -1.0 to 1.0 shaded; probability 68.27 percentprobability density0.000.240.48-4-2024sample value x

P(-1.0 ≤ X ≤ 1.0)68.27%of the total area

About 68% of this distribution lies within one standard deviation of the mean; about 95% within two. These are long-run proportions, not a promise about the next twenty draws.

What about the chance of exactly one value?

In an ideal continuous Gaussian, a single exact value has probability zero: it occupies no width. An interval can have positive probability. The displayed sample is rounded, so a displayed value such as 0.100 stands for a small range of values. Computer-generated numbers have finite precision too.

The entire area is 1. A narrow curve is taller so that its area stays the same. Density can exceed 1; probability cannot. Not every real-world distribution is Gaussian—we chose this one so its rule is visible.

Monte Carlo

Monte Carlo methods use repeated random samples to estimate a quantity. We already know the area between −1 and 1 under a standard Gaussian. Now estimate it by drawing values and counting how many fall inside that interval.

Estimated probability = draws inside the interval ÷ total draws.

Standard normal / fixed interval −1 to 1

Monte Carlo estimate of Gaussian interval probability: 0 of 0 draws inside −1 to 1probability density0.000.240.48-4-2024sample value x

0 inside / 0 total draws

Estimate—

Calculated area68.27%

Absolute error—

Trials accumulate until you reset. Error is in percentage points. Every draw counts, including those outside the plot.

Try ten trials, then a thousand. Reset and repeat. The estimate changes from run to run. Larger samples usually make it more stable, but each new batch can move it closer to or farther from the calculated area. For independent trials like these, typical error shrinks in proportion to 1/√N: about four times as many trials to halve it.

Why estimate an area we can calculate?

Here the calculated area lets us check the method. Monte Carlo is useful when a direct calculation is difficult but we can simulate outcomes. The Monty Hall simulation is another example: simulate games and count wins to estimate a strategy’s win rate.

A sample produces one outcome. A Monte Carlo estimate combines many outcomes to answer a question. More trials reduce random error; they cannot correct the wrong simulation rules or a distribution that poorly represents the situation.

Change the spread

Keep the center at zero. Sample a standard Gaussian value z, then reuse it as the curve widens or narrows. This lets us change the distribution without also changing the random input.

Original σ = 1 · dashedAdjusted σ = 1.000 · solid
Gaussian temperature comparison: original standard deviation 1, adjusted standard deviation 1.000. Temperature 1.00.probability density0.000.240.48-8-4048sample value x

Original draw z—

Adjusted draw √T · z—

Always pick the peak0.000

The blue dot is the original draw; the gold dot is the adjusted draw. Values beyond the plotted range still appear in the readout. Replaying a seed reproduces z. Moving temperature reuses that same z.

Lower temperature concentrates the Gaussian around its mean. Higher temperature spreads it out. Always choosing the peak gives zero every time, which does not reproduce the bell-shaped distribution.

Why √T rather than T?

For this experiment, temperature means raising a density to the power 1/T and renormalizing it. For a Gaussian this leaves the mean unchanged and multiplies the variance by T. Standard deviation is the square root of variance, so the new spread is σ√T. We keep T positive; T = 0 is not a Gaussian density with positive width.

Language-model temperature uses the same idea of reshaping relative likelihoods, but its distribution is over discrete tokens, not a bell curve over numbers. Diffusion models use Gaussian noise directly. Not every control called “temperature” in an image tool has this exact meaning.

From numbers to tokens

A Gaussian makes the distinction visible: the curve can stay exactly the same while every draw is different. A language model has many possible next tokens instead of a continuous axis. Its selected token becomes part of the context for the next prediction, so the distribution can change after every step.

In diffusion, Gaussian samples supply noise. In Nekhen’s The Next Word, the class constructs a distribution over candidate words. The shared question is: what possibilities were available before this one outcome appeared?

Explore Gaussian noise in Denoise →

Pandaemonium Architecture 6.0 — ATEK-639/439 — Fall 2026 · QR code