# The Journey of a Single Ray: How Whitted-Style Ray Tracing Gets Reflections, Shadows and Glass Right

Following one ray from the eye through shadow, reflection and refraction rays — and counting what that recursion costs per frame.

> whitted ray tracing · reflections and refraction · real-time ray tracing performance · About 9 min · Oct 6

## Key points

1. Ray tracing reverses the rasterization question: instead of pushing triangles to pixels, it fires one ray per pixel from the eye and searches for the nearest surface, so visibility is resolved per pixel without a depth buffer.
2. At each hit point, extra rays answer scene-wide questions: a shadow ray toward the light (a cheap any-hit test that ends at the first blocker), a reflection ray along $R = D - 2(D \cdot N)N$, and a refraction ray bent by Snell's law using the material's index of refraction.
3. Shadows, reflections, and refraction look right because their correctness comes from scene geometry rather than from a shadow map's resolution or from pixels that happen to be on screen; reflected objects may sit entirely outside the camera's view, which screen-space techniques can never capture.
4. Whitted's shader builds a ray tree per pixel, and the final color is assembled bottom-up: local lighting plus reflectance times the reflection ray's color plus transmittance times the refraction ray's color, truncated by a maximum depth (often 3) or by a decaying weight.
5. Cost multiplies: primary rays equal the pixel count, and each hit can spawn further rays, so $b$ branches over $d$ levels means up to about $b^d$ rays per pixel. BVH acceleration structures and hardware RT cores (roughly 10 gigarays/s on Turing versus about 1.1 gigarays/s in software) make this tractable but not free.
6. The frame budget is what limits quality: 16.7 ms at 60 fps, 6.9 ms at 144 fps, and cost scales roughly with pixels times rays per pixel, so high resolution with several traced effects runs out first.
7. Games stay hybrid — rasterize primary visibility, trace secondary effects at reduced resolution with few rays per pixel, then reconstruct with spatial and temporal filtering (EA SEED's Pica Pica: about a quarter ray per pixel for reflections, roughly 2.25 rays per pixel total).
8. Whitted-style tracing gives hard shadows, mirror-sharp reflections, and clean glass, but no soft shadows from area lights, no blurred glossy reflections, and no diffuse color bleeding; those require many samples per pixel (path tracing) and show up as noise instead.

---

Last time we looked at how games draw: geometry goes in, pixels come out, and a depth buffer quietly decides which surface wins. That pipeline never asks what is actually visible along a direction — so anything that depends on other parts of the scene (shadows, reflections, light bouncing) has to be faked in extra passes. This article follows the opposite approach, one ray at a time, and shows why the results suddenly look right — and where the cost lands.

## One pixel, one ray: asking the question backwards

A ray is just an origin $O$ and a direction $D$: $P(t) = O + tD$, with $t > 0$ meaning "in front of me". To render, the camera does this for every pixel: build a ray starting at the eye and passing through that pixel's center, then ask a single question — **what is the nearest surface this ray hits?**

Answering it means solving for the smallest valid $t$ among all objects. For a sphere that is a quadratic equation; for a triangle you first intersect the ray with the triangle's plane, then check whether the hit point lies inside the triangle [1]. In his 1980 paper, Turner Whitted wrapped objects in bounding spheres so a ray that misses a sphere can skip everything inside it [2].

The nearest hit gives you a point, a surface normal, and a material. Notice that visibility is now resolved *per pixel, by construction*: no depth buffer, no dependence on draw order, no interpolation shortcuts. But also notice the reversal. Rasterization is a production line for triangles; ray tracing is a search per pixel — like firing a laser pointer through every pixel and seeing what lights up. That is slower, and it is the reason the correct amount of light becomes available.

## Step two: the shadow ray replaces a grid with one question

Once you know a surface point, you can do the ordinary local shading — diffuse plus specular from the lights. But is the light *actually reaching* this point? So you spawn a second ray from the hit point toward the light. If anything is intersected at a distance shorter than the distance to the light ($0 < t < t_{\max}$), the point is in shadow for that light and the light's contribution is dropped [1][3].

Two details matter. First, this query only needs "does anything block the way?" — not "what is the closest thing". The ray can stop at the very first hit, which makes shadow rays cheaper than primary rays. That distinction is why ray tracing hardware has separate *any-hit* and *closest-hit* queries [4]. Second, the shadow's shape is determined by geometry alone. There is no shadow map resolution to make edges blocky, no depth bias to cause acne, no peter-panning, and no shimmering when the light's projection changes as you move. The correctness comes free with the method, not from tuning.

One honest caveat: game lights are treated as points, so a surface is either lit or not. That gives razor-sharp shadow edges. Real light sources have area, and a soft penumbra requires sampling many directions toward the light and averaging them — which is where noise enters the picture, and where Whitted's model stops.

## Step three: the reflection ray looks at what your screen cannot see

For a shiny surface, the incoming direction reflects about the surface normal:

$$R = D - 2(D \cdot N)\,N$$

Spawn a new ray from the hit point along $R$, and whatever it hits is what you see reflected. The color it returns *is* the light arriving from that mirrored direction, so the final color of a point is assembled from the bottom up: its own local lighting from the lights, plus the reflectance coefficient times the color returned by the reflection ray, plus the transmittance coefficient times the color returned by the refraction ray.

This is the heart of what Whitted did. The paper describes the shading information as "a tree of rays" extending from the viewer to the first surface, and from there to other surfaces and to the light sources, created per pixel and handed to the shader [1]. The recursive part is what previous algorithms lacked: visibility calculations did not stop at the nearest intersection; each visible intersection spawns more rays, and the process repeats until no new ray hits anything [1][3].

Here is why that fixes the failure you already know. The object you see in a reflection lies along the reflected direction, which is usually not along your line of sight to that pixel — it can be off to the side, behind a wall, or entirely outside the camera's view. Screen-space reflections, the technique from the previous article, can only reuse pixels that are already on screen, which is exactly why reflections slide out of existence at screen edges. A traced reflection ray leaves the screen's memory behind and searches the actual scene. Mirrors facing mirrors also come out right, because the recursion simply continues: the reflected ray hits a second mirror, which spawns its own reflection ray.

## Step four: refraction makes glass look solid

For transparent materials, the ray does not bounce — it bends. Its new direction obeys Snell's law, $\eta_1 \sin\theta_1 = \eta_2 \sin\theta_2$, where $\eta$ is the index of refraction: about 1.0 for air, 1.33 for water, 1.45–1.65 for glass, 2.42 for diamond [3]. You trace the transmission ray in that bent direction and take the color it returns, weighted by the material's transmission coefficient.

Because the bending depends on the surface normal at each point, a curved pane of glass shifts and magnifies what is behind it — the ray even travels *through* the object and bends again when it exits, so both surfaces matter. Total internal reflection appears on its own: past a certain angle the bend is geometrically impossible, and the ray can only reflect. Whitted's paper treats the reflection and transmission coefficients as constants, but explicitly notes that for accuracy they should follow the Fresnel law — varying with incidence angle, which is why real glass grows more reflective at grazing angles [2]. Note also that games usually use one index for all three color channels; using different values per channel is what would split light into a rainbow, so dispersion stays a luxury effect.

## The bill arrives: a tree, not a ray

Now count the rays. Primary rays equal the number of pixels: 1,920 × 1,080 is about 2.1 million, 4K is about 8.3 million. Add one shadow ray per light, one reflection ray, one refraction ray per hit. If every ray spawns $b$ children and recursion runs $d$ levels deep, the worst case per pixel is roughly $b^d$ rays — depth 3 with three branches is already 27. In practice rays terminate when they miss geometry or when the material's reflectivity is negligible; lecture notes put a depth of 3 as often sufficient, with a weight that decays each bounce as another stopping rule [3].

Intersecting every triangle with every ray would be hopeless. The fix is an acceleration structure: a bounding volume hierarchy (BVH), a tree of nested boxes where a ray that misses a box skips everything inside it. Whitted's bounding spheres were the same idea in 1980 [2]. Traversal cost per ray drops from "number of triangles" to roughly the depth of that tree.

Hardware closed much of the rest of the gap. NVIDIA's Turing architecture (2018) added RT Cores that perform BVH traversal and ray–triangle intersection in fixed-function units, cutting thousands of software instructions per ray down to a hardware probe. The claimed figure is on the order of 10 gigarays per second versus about 1.1 gigarays per second in software on the previous generation — around ten times faster [4][5]. Notably, NVIDIA's own guidance says: "Do not expect hundreds of rays cast per pixel in real-time" [4].

Do the budget math for yourself. At 60 fps you have about 16.7 ms per frame; at 144 fps, about 6.9 ms. Take 1080p with one shadow ray and one reflection ray per pixel: roughly three rays per pixel, about 6.2 million rays per frame, about 373 million rays per second at 60 fps. Against 10 gigarays per second that sounds affordable — and this is why *limited* ray-traced effects became viable. But: 4K multiplies the pixel count by four; neighboring pixels' reflection rays scatter to completely different places, so they share almost no cached data; every hit still needs shading; rays that miss still pay traversal; and each additional effect or bounce adds unique rays. That is the real performance wall — not a single expensive ray, but hundreds of millions of *unpredictable* rays inside a frame that also has to do everything else.

## How games actually spend the rays today

Almost no game ray-traces primary visibility. Rasterization still finds the visible surfaces cheaply, and the depth/normal/material buffer it produces becomes the launch point for rays that only handle the effects that need scene-wide information. EA SEED's Pica Pica demo rasterizes first, then traces reflections at half resolution — a quarter ray per pixel — plus a quarter ray per pixel for shadow rays at reflection hits, reconstructing full resolution from both spatial and temporal filtering, for roughly 2.25 rays per pixel total [6][7].

That reconstruction step is the escape hatch. Instead of shooting 16 rays per pixel to get a clean result, you shoot one and let neighboring pixels' ray hits fill in the gaps. So the quality knob in a game's settings is not simply "ray tracing on or off". It is: which effects receive rays, how many rays per pixel, at what resolution the ray buffer runs, how many bounces, and how aggressive the denoiser is. Turning ray tracing off usually removes an effect outright; lowering its quality usually trades sharpness and noise instead of correctness.

## Judging what you see

The next time something looks off, the artifact tells you the technique. A reflection that dissolves as it approaches the screen edge, or that omits something you know is off-screen, is screen space. A shadow edge that is blocky on a fixed grid, or that crawls and shimmers as you move, is a shadow map. A reflection that is geometrically correct but slightly soft or grainy, sharpening when you hold still, is traced and temporally denoised — noisy rather than missing is the signature of a real ray. And because frame cost scales roughly with pixel count times rays per pixel, the crunches happen at high resolution with several traced effects, which is why upscaling and ray tracing are usually discussed together.

The boundary of Whitted-style ray tracing is worth stating plainly, because his own paper does: it handles mirror-like reflection, hard shadows, and clean refraction, but it does not provide for diffuse reflection from distributed light sources, so there is no soft shadow, no blurred glossy reflection, and no color bleeding from one matte surface onto another [2]. Those need many samples per pixel averaged together, which is path tracing — and its characteristic artifact is noise traded against frame rate. That is the next step.

<details>
<summary>Answers to a quick self-check</summary>

1. Why can shadow rays stop at the first hit while reflection rays cannot? A shadow ray only needs to know whether anything blocks the path to the light, so any hit is decisive. A reflection ray needs the color of the surface actually seen, so it needs the *closest* hit.
2. A 1440p image (2,560 × 1,440, about 3.7 million pixels) with one reflection ray per pixel at full resolution needs roughly 3.7 million reflection rays per frame — about 221 million per second at 60 fps.
3. Why does turning down ray tracing quality often produce blur instead of a missing object? Fewer rays are cast, and the denoiser reconstructs the missing detail from neighboring pixels and previous frames, so error appears as softness, noise, or smearing rather than as a structurally wrong image.
</details>

## Sources

1. [An Improved Illumination Model for Shaded Display (Whitted, CACM 1980) — ACM entry for the original recursive ray tracing paper](https://dl.acm.org/doi/10.1145/358876.358882)
2. [Whitted 1980 full paper (PDF) — ray tree, shadow rays, bounding volumes, and the model's stated limits](https://cseweb.ucsd.edu/~viscomp/classes/cse274/fa21/readings/whitted.pdf)
3. [Ray tracing lecture notes, Lund University — trace/directIllumination recursion, termination, refraction indices](https://fileadmin.cs.lth.se/cs/Education/EDAN30/lectures/L2-rt.pdf)
4. [NVIDIA Turing Architecture In-Depth — RT Cores, BVH traversal, gigarays per second, rays-per-pixel expectations](https://developer.nvidia.com/blog/nvidia-turing-architecture-in-depth/)
5. [NVIDIA GeForce: ray tracing questions answered — denoising and per-pixel ray cost for different effects](https://www.nvidia.com/en-us/geforce/news/geforce-gtx-dxr-ray-tracing-available-now/)
6. [EA SEED, Hybrid Rendering for Real-Time Ray Tracing — half-resolution reflection rays and reconstruction](https://media.contentapi.ea.com/content/dam/ea/seed/presentations/2019-ray-tracing-gems-chapter-25-barre-brisebois-et-al.pdf)
7. [EA SEED, Shiny Pixels and Beyond (GDC 2018) — ray budget accounting, roughly 2.25 rays per pixel](https://media.contentapi.ea.com/content/dam/ea/seed/presentations/gdc2018-seed-shiny-pixels-and-beyond-real-time-raytracing-at-seed.pdf)

---

Original article: https://eulore.ai/articles/whitted-ray-tracing-single-ray-journey-19d77c00

> **Eulore** · Learn a little. Understand a lot.
>
> Eulore is an AI learning tool that turns what you want to learn into a continuing series. Share a topic, and it gets to know your starting point before creating articles you can read in 5–10 minutes. Ask as you read, and shape what comes next.This article was created in the same way.
>
> Start your own series → https://eulore.ai
