# From Hours to 16.7 ms: BVH, RT Cores, Denoising, and How to Read a Ray Tracing Settings Menu

Three machines close the thousand-to-one sample gap between an offline render and a game frame: a tree of boxes that makes each ray cheap, reconstruction that guesses the samples you never traced, and upscaling that buys rays by rendering fewer pixels.

> real-time ray tracing performance · denoising · BVH · RT cores · upscaling · About 11 min · Oct 6

## Key points

1. A film frame takes minutes to hours and a 60 fps game frame takes 16.7 ms; almost the whole gap is sample count, not per-ray speed, because noise falls only as σ/√N while a game traces about one sample per pixel.
2. The BVH is a tree of nested boxes that lets a ray reject most of a scene with a few box tests, cutting triangle tests from O(N) to about O(log N) on average; it must still be built or refitted every frame, and its performance depends heavily on whether rays are coherent.
3. Camera rays are coherent and fast; the bounce and shadow rays a path tracer fires afterwards are incoherent, which is where real-time ray tracing time actually goes and why peak ray rates are not achieved in games.
4. RT cores are dedicated hardware units that perform BVH traversal and ray-triangle intersection, freeing shader cores; NVIDIA quotes about 1.1 Giga rays/sec in software on the prior generation versus 10+ Giga rays/sec with them, but they accelerate only intersection, not shading or sample count.
5. Because halving noise costs four times the rays, the answer is not more rays but reconstruction: denoisers re-use nearby pixels (edge- and variance-aware spatial filters) and previous frames (reprojection through motion vectors, verified by depth, normals and object ID, blended as an exponential moving average).
6. Temporal reuse creates all the characteristic artefacts — ghosting when history is stale, noise flashes on disocclusion when history does not exist, and lag from every frame of history kept — which is what the denoiser quality setting really trades.
7. Upscaling buys rays rather than saving them: rendering 1080p for a 4K output is a quarter of the pixels, so about four times the rays per pixel at the same frame time; DLSS Quality at 4K is 1440p internal (2.25× fewer pixels) and Performance is 1080p (4×).
8. Ray Reconstruction replaces separate hand-tuned denoising and upscaling passes with one neural network that takes the noisy low-resolution frame plus engine guide buffers (motion vectors, depth, normals, roughness, albedo, specular hit distance) and outputs a clean full-resolution image.
9. Frame generation inserts AI frames between rendered ones: it raises the displayed rate without reducing the rays spent in simulated frames and without improving input latency, so it is a smoothness lever rather than an image-quality one.
10. A practical setting order is upscaler first, then the best available denoiser, then only the two RT effects you notice, then resolution preset, and frame generation last once the simulated rate is already comfortable.

---

Last time we found out why a path-traced frame is grainy: the renderer has to answer the rendering equation, it answers it with random samples, and a game can only afford about **one sample per pixel**. The grain is that one sample showing through.

So the real-time problem is not "trace more rays." It is: make one sample per pixel look like a picture, and do it inside 16.7 milliseconds — the frame time at 60 fps. Every part of the RT pipeline you meet in a graphics menu is one of three machines built for that job. One makes each ray cheaper. One guesses the samples you didn't trace. One lets you buy more rays by rendering fewer pixels.

## The size of the gap

An offline path-traced frame for film takes minutes to hours on a render farm; complex shots can run for days. A 60 fps game frame is 16.7 ms. That is roughly a four-orders-of-magnitude gap, and almost all of it is sample count. Offline renderers use thousands of samples per pixel because the estimate's error shrinks like $\sigma/\sqrt{N}$: to halve the noise you need four times the samples. A game uses roughly one.

You cannot close a thousand-to-one sample gap with faster hardware alone. You can close maybe ten-to-one of it. The rest has to come from not computing the samples in the first place.

## A tree of boxes: how one ray avoids testing ten million triangles

Start with the cheapest possible ray query: a camera ray leaving the eye, and you want to know which triangle it hits first. A modern game scene can hold tens of millions of triangles. Testing the ray against all of them, for every pixel, every frame, is hopeless.

The fix is the **bounding volume hierarchy**, or BVH. You wrap groups of nearby triangles in boxes, wrap groups of boxes in bigger boxes, and keep going until one box contains the whole scene. Now tracing a ray is a walk down the tree. Test the root box; if the ray misses it, the entire scene is rejected in one test. If it hits, test the two boxes inside, descend only into the ones the ray actually passes through, and when you reach a leaf, test the ray against the handful of triangles stored there. Because a ray is a straight line through space, it visits only a thin slice of the tree, and on average the number of triangle tests falls from $O(N)$ to about $O(\log N)$ [1][2].

![Bounding volume hierarchy for a small scene: primitives wrapped in nested boxes, and the corresponding tree](https://pbr-book.org/4ed/Primitives_and_Intersection_Acceleration/pha07f03.svg)

Two properties of this structure matter for everything that follows.

First, the BVH has to exist *before* any ray is traced, and scenes change. Objects move, doors open, foliage sways. Each frame the game either rebuilds the tree or "refits" the existing one with updated boxes, and both cost real time out of the same 16.7 ms.

Second, not all rays are equally friendly. The camera rays leaving your eye are **coherent** — neighbouring pixels shoot nearly parallel rays, which hit the same boxes in the same order and behave well on hardware designed to process many rays in parallel. The rays a path tracer fires *after* the first hit — bounce rays, shadow rays — scatter in random directions. Those are **incoherent**: each thread in a GPU executes a different walk through the tree, so the hardware sits idle waiting on memory [1]. This is where real-time ray tracing time actually goes, and it is why the peak ray rates you see in marketing are not what a game achieves. NVIDIA's own numbers for the Turing generation list roughly 1.1 Giga rays/sec doing traversal in software on the previous architecture, versus 10+ Giga rays/sec using dedicated hardware [3]. Treat the peak as an upper bound that coherent, simple work approaches and incoherent path tracing does not.

## RT cores: moving the walk out of the shader

Before 2018, the BVH walk was ordinary software: the GPU's shader cores ran the box tests and triangle tests themselves, at a cost of thousands of instruction slots per ray [3]. Those cores are also the ones doing everything else — shading, particles, post-processing — so ray tracing competed directly with the rest of the frame.

**RT cores** are dedicated units that do the walk instead. Each one contains hardware for the box tests that steer traversal and hardware for ray–triangle intersection; the shader core just launches a ray and gets back a hit or a miss, and stays free for other work [3][4]. The 2018 Turing generation shipped the first of these with DirectX Raytracing; AMD's equivalent units handle the box and triangle tests while shader code drives the traversal [2].

That is what makes real-time ray tracing possible at all, but notice exactly what it accelerates: *finding intersections*. It does not shade, and it does not reduce the sample count. A game with RT cores still traces about one sample per pixel, which — as the last article showed — is a noisy image. The hardware also has limits you can see in a game: geometry marked as needing an alpha-test shader, like a chain-link fence, interrupts the hardware walk and drags performance down, which is why developers are told to flag solid geometry as opaque [8].

## The wall: samples cost more than they look

Here is the arithmetic that forces everything else. Suppose you want half the noise. From $\sigma/\sqrt{N}$ you need four times the samples, meaning four times the rays per frame at the same frame rate — or the same number of rays at a quarter of the frame rate. Quality and frame rate trade off quadratically in ray count, not linearly.

That is a bad exchange rate for a settings slider. So instead of tracing the missing samples, real-time renderers **guess** them, and are surprisingly good at it because neighbouring information is enormously redundant.

## Borrowing samples: denoising

A denoiser exploits two kinds of redundancy.

**Spatially**, adjacent pixels usually see the same surface under nearly the same lighting, so an average over a small neighbourhood is a better estimate than any one pixel. A plain blur would destroy texture and edges, so the filter is *edge-aware* and, crucially, *noise-aware*: it blurs hard where the signal is unreliable and barely at all where it looks clean. The classic method, SVGF, estimates per-pixel variance by accumulating the first and second moments of the luminance over time, and uses that variance to drive a multi-scale wavelet filter. As published, it reconstructed a path-traced image from one path per pixel in about 10 ms at 1080p [5].

**Temporally**, the previous frame already contains an answer for the same surface point. If the game knows where each pixel moved — its **motion vector** — it can reproject last frame's result onto this frame, check that the two really are the same surface by comparing depth, surface normal and object ID, and blend the old result with the new one. A blend that keeps a fraction of history each frame is an exponential moving average; eight frames of history at one sample each act like eight samples [4][6]. This is where almost all the apparent quality at 1080p comes from: not more rays this frame, but rays from previous frames, re-used.

Temporal reuse is also where the artefacts come from, and they are all the same shape — history that is *wrong* still being trusted.

- **Ghosting.** A shadow moves or an object is animated in a way the motion vectors don't describe. Reprojection drags stale pixels along and you see a smear trailing the object. The standard defence is *neighbourhood clamping*: if the reprojected history value is far outside the range of values in the current frame's local neighbourhood, it is rejected as stale [6].
- **Disocclusion.** When the camera turns, pixels that were hidden are revealed and have no valid history at all. Those regions flash noise until history rebuilds, so the denoiser temporarily boosts spatial blur there [6].
- **Lag.** Every frame of history you keep buys smoothness and costs responsiveness, because the image is partly a picture of the past. That trade-off is what the "denoiser quality" setting is really choosing.

## Upscaling: buying rays by rendering fewer pixels

Denoising decides how good one sample looks. Upscaling decides how many pixels you have to sample — and because rays per frame is pixels × samples, cutting pixels is the same as multiplying your ray budget.

The numbers are worth having in your head. 1080p is 2.1 million pixels; 4K is 8.3 million, exactly four times as many. Rendering at 1080p and outputting 4K means a quarter of the pixels, so for the same frame time you can afford about four times the rays per pixel. That is what an upscaler does: it renders at a lower internal resolution, then reconstructs the full-resolution image using motion vectors, depth, and history. DLSS Quality at 4K renders internally at 1440p, which is 2.25× fewer pixels; Performance renders at 1080p, which is 4× fewer [7]. The jittered sampling and history reuse mean the upscaler is also feeding the denoiser extra samples, which is why upscaling and denoising are not independent choices in an RT game — they are two halves of the same budget.

**Ray Reconstruction** merges them. Instead of running a hand-tuned denoiser and then an upscaler as separate passes, one neural network takes the noisy low-resolution ray-traced frame plus a set of **guide buffers** the engine provides — motion vectors, depth, surface normals, roughness, diffuse and specular albedo, and for reflections a specular hit distance — and outputs a denoised, full-resolution image in a single pass [7][9][10]. Introduced with DLSS 3.5 in 2023, it replaces heuristics that had to be hand-tuned per effect with a model trained to recognise what reflections, shadows and indirect light look like, and it has been updated since. The practical consequence for you: when a game offers it, it is usually both better *and* cheaper than the hand-tuned denoiser it replaces, so it should be on.

**Frame generation** is a different thing and easy to confuse with the above. It inserts AI-constructed frames between genuinely rendered ones. It raises the number on your frame counter without reducing the rays spent in the frames that were actually simulated, and it does not reduce input latency. It is a smoothness lever, not an image-quality one. If the simulated rate is already marginal, generating frames on top will look smoother and still feel sluggish.

## Reading the settings menu

With that machine in mind, the menu stops being a list of mysteries.

**The currency is internal resolution × samples per pixel × rays per sample (bounces), times the cost of each effect.** Every setting spends one of those.

Ray-traced effects are offered separately because their costs differ enormously. RT shadows and RT ambient occlusion are the cheap end — they replace shadow maps and screen-space AO, and their ray counts are modest. RT reflections sit in the middle and give the most visible payoff per frame, because they fix the specific failure of the old technique: screen-space reflections can only reflect what is already on screen, so a puddle stops reflecting a building the moment you look away from it. RT global illumination brings colour bleeding and bounced light. A full path-tracing mode — sometimes sold as an "Overdrive" preset — does all of it at once with one or two bounces, and in practice only runs with a denoiser, an upscaler, and often frame generation carrying it.

A workable order of operations:

1. **Turn the upscaler on first.** Quality or Balanced. This either frees up a large share of the frame or lets you spend it, and nothing else you decide afterwards is meaningful until it's set.
2. **Enable Ray Reconstruction if the game has it**, or the game's best denoiser. It is usually cheaper than the alternative.
3. **Pick the two RT effects you actually notice.** For most players that's reflections and shadows. Ray-traced global illumination is beautiful and the first thing to drop.
4. **Only then adjust the upscaler preset or internal resolution** to hit your target frame rate.
5. **Add frame generation last**, and only once the simulated rate is already comfortable.

Two rules of thumb. First, target slightly more frame rate than you think you need, because temporal accumulation is measured in wall-clock time: eight frames of history is 0.13 s at 60 fps but 0.27 s at 30 fps, so at low frame rates ghosting and reveal artefacts hang around roughly twice as long. Second, prefer a higher upscaler preset over a lower RT preset — a slightly softer image from a good upscaler is far less distracting than reflections dissolving into noise.

## What the artefacts tell you

The image itself is a diagnostic, once you know which machine each failure belongs to.

- **Holds still, still grainy.** Not enough samples or history. Raise RT quality, or turn the denoiser on.
- **Smears and trails behind moving objects.** The denoiser is trusting stale history. Lower its aggressiveness, or raise RT quality so it has less to fix.
- **Sparkles and throbs on edges revealed when you turn the camera.** Disocclusion. It is inherent to temporal reuse; it shrinks with more rays or a better denoiser.
- **Shimmering on fences, railings and thin geometry at distance.** The upscaler, not the rays. Raise the preset, or use DLAA if you have headroom.
- **Reflections that appear only when the reflected object is on screen.** That is screen-space reflections, not ray tracing. The setting is not doing what you think.

None of this is a "realism" slider. It is a set of knobs on three machines — the intersection hardware, the reconstruction network, and the resolution you choose to pay for — and the settings menu is only asking you how to divide 16.7 ms between them.

## Sources

1. [Physically Based Rendering: Bounding Volume Hierarchies — BVH structure, traversal by culling, the surface area heuristic, and coherent vs. incoherent rays](https://pbr-book.org/4ed/Primitives_and_Intersection_Acceleration/Bounding_Volume_Hierarchies)
2. [Bounding volume hierarchy — logarithmic test counts, and the RT core / Ray Accelerator descriptions for NVIDIA and AMD hardware](https://en.wikipedia.org/wiki/Bounding_volume_hierarchy)
3. [NVIDIA Turing GPU Architecture whitepaper — RT core design, thousands of instruction slots per ray in software, and the ~1.1 vs 10+ Giga rays/sec comparison](https://images.nvidia.cn/aem-dam/en-zz/Solutions/design-visualization/technologies/turing-architecture/NVIDIA-Turing-Architecture-Whitepaper.pdf)
4. [NVIDIA Developer: Ray Tracing — ray casting, BVH building and refitting, and real-time denoising with tensor cores](https://developer.nvidia.com/discover/ray-tracing)
5. [Spatiotemporal Variance-Guided Filtering — temporal accumulation, variance estimation, and ~10 ms at 1080p from one path per pixel](https://cwyman.org/papers/hpg17_svgf.pdf)
6. [FidelityFX Denoiser documentation — reprojection, disocclusion handling, variance boost, and neighbourhood clamping against ghosting](https://gpuopen.com/manuals/fidelityfx_sdk/techniques/denoiser/)
7. [NVIDIA Technical Blog: DLSS 3.5 Ray Reconstruction — what the neural denoiser replaces and what it is trained on](https://developer.nvidia.com/blog/generate-groundbreaking-ray-traced-images-with-next-generation-nvidia-dlss/)
8. [Best Practices for Using NVIDIA RTX Ray Tracing — why opaque geometry and avoiding any-hit shaders matter for hardware traversal performance](https://developer.nvidia.com/blog/best-practices-for-using-nvidia-rtx-ray-tracing-updated/)
9. [NVIDIA: Decoding AI-Powered DLSS 3.5 Ray Reconstruction — plain-language account of temporal and spatial denoising, and how frame generation differs from super resolution](https://blogs.nvidia.com/blog/ai-decoded-ray-reconstruction/)
10. [Streamline DLSS Ray Reconstruction programming guide — the exact guide buffers RR expects, including specular motion vectors and specular hit distance](https://github.com/NVIDIAGameWorks/Streamline/blob/main/docs/ProgrammingGuideDLSS_RR.md)

---

Original article: https://eulore.ai/articles/real-time-ray-tracing-bvh-rt-cores-denoising-e51cc94c

> **Eulore** · Learn a little. Understand a lot.
>
> Eulore is an AI learning tool that turns what you want to learn into a continuing series. Share a topic, and it gets to know your starting point before creating articles you can read in 5–10 minutes. Ask as you read, and shape what comes next.This article was created in the same way.
>
> Start your own series → https://eulore.ai
