# Why Path Tracing Is Noisy: The Rendering Equation, Monte Carlo Sampling, and the Real-Time Sample Budget

Whitted's rays answer specific questions cheaply; the rendering equation asks about all of the light at once, and answering it with random samples is exactly where the grain comes from.

> global illumination · path tracing · real-time ray tracing performance · denoising · About 11 min · Oct 6

## Key points

1. Whitted's shader only keeps the parts of light transport where light comes straight from a light source or bounces in exactly one direction, so it misses bounced light, color bleeding, soft shadows from area lights, and blurred reflections.
2. The rendering equation states that light leaving a point equals what the point emits plus the integral over all incoming directions of arriving light times the material term times a cosine factor; because the arriving light at one point is the outgoing light of another, the equation refers to itself and its solution is an infinite sum of bounces.
3. Path tracing solves that equation by Monte Carlo estimation: follow one random path per sample and average many samples, which is correct on average but random in any single sample.
4. Noise is the spread of that random estimate, and its typical error shrinks like σ/√N, so halving the noise needs four times the samples; a clean offline image can need thousands of samples per pixel, while a game can afford about one.
5. Variance is not uniform across the screen: matte surfaces with direct light are nearly noise-free, while glossy reflections and indirectly lit surfaces have huge variance because most random paths contribute almost nothing and rare ones contribute a lot, producing speckle and bright fireflies.
6. Real-time rendering copes with better sampling (aiming rays at lights, ReSTIR-style sample reuse worth roughly 6×–100× efficiency), temporal accumulation across frames using motion vectors, spatial denoisers such as SVGF that blur where estimated variance is high, and trained networks like Ray Reconstruction.
7. At one sample per pixel the displayed image is a reconstruction rather than a computation, so its characteristic weaknesses are smeared fine detail, ghosting in motion, and noise wherever temporal history is lost.
8. Judging the trade-off means checking whether the game ray-traces individual effects or the whole lighting (path tracing), remembering that ray-traced cost scales roughly quadratically with render resolution so upscaling is the biggest lever, and looking for evidence of the sample budget in motion rather than in still screenshots.

---

Last time we followed a single ray from the eye: it found the nearest surface, then fired a shadow ray at each light, a reflection ray for mirrors, and a refraction ray for glass. That shader gets shadows, reflections and glass right because geometry answers the question. So why does turning on “path tracing” in a modern game produce a grainy, shimmering image that looks like television static? The grain is not a bug. It is the visible signature of a much bigger job than Whitted's shader ever attempted.

## What the ray tree quietly leaves out

Whitted's shader makes two assumptions. First, light reaches a point along a straight line from a light source — one shadow ray per light, yes or no. Second, reflections and refractions are perfect: a mirror sends light in one exact direction, glass bends it in one exact direction.

Walk into a real room and both assumptions fail, in ways you have probably noticed without naming them:

- A room lit by one window is not only bright near the window. The wall opposite the window glows, and the floor near it picks up light. Light has bounced off the wall before reaching your eye.
- A white wall beside a red wall picks up a pink tint. The bounced light carries the color of whatever it bounced off.
- A ceiling lamp is a physical object with size, so shadow edges are soft gradients, not knife edges.
- Brushed metal and painted car bodies reflect a blurred image, not a sharp mirror copy.

None of these need new physics. They need *more directions* of light. Every point on a surface receives light not just from lamps, but from every other surface in view — and those other surfaces are themselves lit by everything they can see. Each point's brightness depends on the brightness of every other point. That is the bootstrap problem Whitted's shader sidesteps.

## One equation that says “all of it”

In 1986 James Kajiya wrote down the statement of that bootstrap problem in a single line, now called the **rendering equation** [1]. In its familiar form it looks like this:

$$L_o(x,\omega_o) = L_e(x,\omega_o) + \int_{\Omega} f_r(x,\omega_i,\omega_o)\, L_i(x,\omega_i)\, (\omega_i \cdot n)\, d\omega_i$$

Take it apart piece by piece, because every term is doing something you need to follow later.

- $L_o(x,\omega_o)$ is how much light leaves the point $x$ in the direction $\omega_o$ — the quantity the camera ultimately wants to know.
- $L_e$ is light the point emits itself. For most surfaces this is zero; for a lamp or a glowing sign it is the whole story.
- $f_r$ is the material term: given light arriving from some direction, how much of it does this surface send back out in the viewing direction? A mirror's $f_r$ picks out one direction; a rough surface spreads it out; a red surface suppresses red-absorbing colors.
- $L_i(x,\omega_i)$ is the light *arriving* at $x$ from one particular direction $\omega_i$.
- The integral $\int_\Omega \ldots d\omega_i$ says: don't pick one incoming direction, add up contributions from *all* directions over the hemisphere above the point.
- The $(\omega_i \cdot n)$ factor is the cosine term: light arriving at a grazing angle delivers less energy per unit of surface area, so it contributes less.

The self-reference is in $L_i$. Light arriving at $x$ from direction $\omega_i$ is just light *leaving* the surface that lies that way — in other words, the same quantity $L_o$ evaluated at a different point. So the equation is built on itself: the unknown appears on both sides. Physically that means the answer is a nested sum. Direct light is bounce one. Light that bounced off another wall before reaching our point is bounce two, and it depends on light that had bounced off something else, and so on, infinitely.

Whitted's shader is, in this language, a drastically pruned version. It keeps the $L_e$ term and the parts of the integral where $f_r$ is a spike pointing at exactly one direction (perfect mirror, perfect glass). The smooth, spread-out parts of the integral — the ones that create glow, color bleeding and soft shadows — are simply dropped, which is why game engines spent decades faking them with hand-authored “ambient” terms.

## Path tracing: guessing the integral with random samples

You cannot evaluate an integral over a whole hemisphere by hand, and for a whole scene you cannot evaluate it exactly at all. So you guess — but you guess in a way that is right *on average*. This is **Monte Carlo estimation**: rather than computing the average of a function over all directions, draw $N$ random directions, evaluate the function at each, and average the results.

Write the estimator as

$$\hat{L} = \frac{1}{N}\sum_{i=1}^{N} \text{(light arriving from random direction } \omega_i\text{)} \times \frac{f_r}{p(\omega_i)}$$

where $p(\omega_i)$ is the probability with which you chose that direction. Dividing by $p$ is what keeps the average fair: directions you sample rarely must count more when they do appear. If you sample every direction with equal probability, the weight is just a constant.

**Path tracing** turns this into a way of following rays [1]. Instead of building a branching tree of rays at every hit point like Whitted's shader does, you follow a *single* random path: the eye ray hits a surface, you pick one random new direction weighted by the material, follow that ray to wherever it lands, pick another random direction there, and continue until the path dies out or hits a light. One path is one sample of the integral. The pixel's color is the average over many paths.

Two details matter for what comes later. First, at each surface the path also fires one ray straight at a light (chosen randomly if there are many lights). That single aimed ray is the surviving descendant of Whitted's shadow ray. Second, following one path instead of a full tree is not just cheaper, it is statistically better: in Whitted's tree almost all the rays are deep in the recursion, where they contribute the least to the image, while a path spends its one ray per bounce on exactly the light that matters [1].

In offline film rendering, $N$ is often thousands of paths per pixel. In a game running at 60 frames per second, $N$ is usually **one**. That gap is the entire story of the noise.

## Where the noise comes from

Here is the crucial mental shift: with Monte Carlo, the pixel value is not a computed number, it is a *random number* whose average over infinitely many samples would be correct. One sample, one draw. So a pixel rendered with $N$ samples has an error whose typical size shrinks like

$$\text{error} \sim \frac{\sigma}{\sqrt{N}}$$

where $\sigma$ is how widely a single sample's value varies. Halving the noise therefore needs four times the samples; cutting it tenfold needs a hundred times the samples. This square-root law is why the Vulkan documentation describes 1 sample per pixel as a mess of black and white dots — the “salt and pepper” look — and notes that a clean image might take 1000+ rays per pixel, i.e. seconds or minutes per frame [2].

Two more facts turn that abstract $\sigma$ into the specific speckle you see:

**Different pixels have very different $\sigma$.** A pixel on a matte wall under direct sunlight is nearly deterministic — every sample comes back with a similar value, so 1 sample per pixel is already fine. A pixel looking at glossy metal, or at a surface lit only indirectly, has enormous $\sigma$: most random paths wander off and hit something dark, contributing almost nothing, while an occasional path lands exactly on a bright light and contributes a lot. Averaging a hundred near-zero samples and one huge one gives a wildly unreliable answer. Rare samples with very large values also produce **fireflies**: isolated, extremely bright pixels that stand out because they sit far above their neighbors.

**Noise in the *indirect* term is the hard part.** Whitted's shadow ray produced crisp, noise-free shadows because it deliberately aimed at a light — a good sample distribution. Path-traced indirect lighting has no such target; the bounce ray goes wherever the material says, and the light it eventually finds may be small. That is why the shimmer in a path-traced game shows up first in corners, under furniture, and inside the tinted glow that one wall casts onto another.

Now multiply. A 1920×1080 frame is about 2.07 million pixels. At 60 fps with one sample per pixel, you are averaging roughly 124 million random paths per second, each of which is several bounce rays. Film renderers spend thousands of samples per pixel and take seconds to minutes per frame; a game has about 16.7 milliseconds, including everything else it must do. There is no version of “just take more samples” that fits.

## What games actually do about it

Real-time path tracing is therefore a sampling-efficiency contest, and modern engines attack it on four fronts.

**Sample smarter.** Instead of picking directions blindly, sample the ones that matter: aim shadow rays at lights, sample glossy reflections in the lobe the material actually reflects, and — the big modern win — *reuse* samples. ReSTIR resamples a pool of candidate light samples and then shares good samples between neighbouring pixels and between frames; the original 2020 paper reports equal-error results 6×–60× faster than the previous state of the art while tracing at most 8 rays per pixel [3], and a 2023 SIGGRAPH course puts highly optimized variants at up to roughly 100× efficiency gains, with the blunt warning that in production engines “even one ray or path per pixel may only be feasible on the highest-end systems” [4].

**Reuse across frames.** Because the camera moves, you can find where each pixel's content was in the previous frame (motion vectors) and blend the old noisy value with the new one. If nothing moved, hundreds of frames accumulate into what is effectively hundreds of samples. The moment something new appears on screen — a disocclusion behind a moving character — the history is gone and the pixel is back to one sample, which is exactly where you see sudden noise when you whip the camera around.

**Denoise.** A denoiser looks at the noisy image and tries to reconstruct what the missing samples would have said, guided by how noisy each region is estimated to be. The classical recipe, SVGF, has three stages: temporal accumulation, variance estimation, and a spatial filter that blurs harder where variance is high and backs off where the image has become stable [5]. In the Quake II RTX research project this full denoiser ran at about 3.5 ms at 1440p, roughly the second-largest cost after tracing itself, with the tracer operating at a single sample per pixel [5].

**Learn it.** Newer titles replace several hand-tuned denoisers with one trained network. Cyberpunk 2077's path-traced mode has used as many as five separate denoisers per frame, and DLSS Ray Reconstruction collapses them into a single AI model [6]; NVIDIA reports that the full DLSS stack (upscaling, ray reconstruction, frame generation) is on average 4.9× faster than native 4K rendering with full ray tracing on [7].

The consequence for image quality is worth stating plainly: at one sample per pixel the picture on screen is a *reconstruction*, not a calculation. The denoiser has to guess, and its hardest guess is distinguishing genuine fine detail — a chain-link fence, a face in a reflection — from noise that looks structurally similar. When it guesses wrong, the result is smeared detail, a slightly plastic look, or ghostly trails that linger for a frame after a light moves.

## How to read the trade-off yourself

With that background, the graphics menu stops being a list of magic words.

**Check what “ray tracing” means in this game.** The usual “RT” toggle renders the frame the old way and then ray-traces a few effects on top — reflections, shadows, sometimes one bounce of global illumination. “Path tracing”, “full ray tracing”, or a preset like Overdrive means the whole lighting solution is path traced and then denoised. The costs are in completely different leagues, and so are the artifacts: effect-level ray tracing is mostly crisp, while path tracing lives or dies by its reconstruction.

**Know the biggest performance lever.** The cost of ray-traced effects grows roughly quadratically with render resolution [6], so rendering at a lower resolution and upscaling buys far more than nudging a quality slider. This is why path-traced modes are practically designed around upscaling; in 2023, an RTX 4090 could not stay above 30 fps at native 4K with path tracing in Cyberpunk 2077 without DLSS [6].

**Look for the sample budget in motion, not in screenshots.** Screenshots favour ray tracing because a still camera gives the accumulation time to clean up. In motion, watch for: grain and shimmer that appears when you turn quickly or step past an edge (history lost, one sample per pixel exposed); soft, waxy textures in reflective surfaces (denoiser eating detail); trails behind moving lights or characters (temporal reuse that should have been rejected); and brightness that leaks or lags briefly around newly visible objects.

**Use the artifacts as a quality signal.** If reflections on rough metal are crisp where they should be blurry, the reflection is probably screen-space and only contains what the camera can already see. If the corners of a room stay flatly dark instead of picking up glow from a bright wall, indirect light is still being faked. If everything is correct but crawls with a static-like texture, you are looking at honest path tracing being paid for with a very small sample count.

In short: Whitted's rays buy correct *answers to specific questions* — is this point lit, what does the mirror show — at a fixed cost. Path tracing buys the whole lighting solution, but only as an average of random samples, so its price is paid in variance, and variance is bought back with clever sampling, sample reuse across space and time, and denoising. When you feel a quality-versus-frame-rate dial in a game, that dial is really a sample budget being spent and a filter being asked to cover the shortfall. Two related topics this article skips: how a renderer decides when to stop a path (Russian roulette and other termination rules), and the sampling machinery for effects that Monte Carlo handles poorly by default, such as caustics.

## Sources

1. [Kajiya, "The Rendering Equation" (SIGGRAPH 1986) — the integral equation, its Monte Carlo solution, and why one path beats a branching ray tree](https://www.cs.princeton.edu/courses/archive/fall03/cs526/papers/kajiya86.pdf)
2. [Vulkan documentation: Ray Tracing Optimization — Denoising and Adaptive Sampling — 1 spp "salt and pepper" variance, the √N rule, and 1000+ rays per pixel for a clean image](https://docs.vulkan.org/tutorial/latest/ML_Inference/Desktop_Applications/07_ray_tracing_optimization.html)
3. [Bitterli et al., "Spatiotemporal reservoir resampling for real-time ray tracing with dynamic direct lighting" (SIGGRAPH 2020) — ReSTIR, 6×–60× equal-error speedups at ≤8 rays per pixel](https://dl.acm.org/doi/10.1145/3386569.3392481)
4. ["A Gentle Introduction to ReSTIR Path Reuse in Real-Time" (SIGGRAPH 2023 course) — up to ~100× efficiency gains, and one ray per pixel being feasible only on high-end systems](https://dl.acm.org/doi/10.1145/3587423.3595511)
5. [NVIDIA: Path Tracing Quake II in Two Months — 1 spp path tracing, SVGF's temporal accumulation / variance estimation / à-trous filter, and denoiser cost at 1440p](https://developer.nvidia.com/blog/path-tracing-quake-ii/)
6. [HotHardware on Cyberpunk 2077 Ray Reconstruction — as many as five denoisers per frame, quadratic cost scaling with render resolution, 4K path tracing below 30 fps without DLSS](https://hothardware.com/news/dlss-rr-tested-cyberpunk)
7. [NVIDIA: DLSS 3.5 and Ray Reconstruction in Cyberpunk 2077's full ray tracing mode — replacing hand-tuned denoisers with one AI model, 4.9× average speedup vs native 4K](https://www.nvidia.com/en-us/geforce/news/dlss-3-5-cyberpunk-2077-phantom-liberty-available-now/)

---

Original article: https://eulore.ai/articles/why-path-tracing-is-noisy-c0c4c779

> **Eulore** · Learn a little. Understand a lot.
>
> Eulore is an AI learning tool that turns what you want to learn into a continuing series. Share a topic, and it gets to know your starting point before creating articles you can read in 5–10 minutes. Ask as you read, and shape what comes next.This article was created in the same way.
>
> Start your own series → https://eulore.ai
