Last time we looked at how games draw: geometry goes in, pixels come out, and a depth buffer quietly decides which surface wins. That pipeline never asks what is actually visible along a direction — so anything that depends on other parts of the scene (shadows, reflections, light bouncing) has to be faked in extra passes. This article follows the opposite approach, one ray at a time, and shows why the results suddenly look right — and where the cost lands.
One pixel, one ray: asking the question backwards
A ray is just an origin and a direction : , with meaning "in front of me". To render, the camera does this for every pixel: build a ray starting at the eye and passing through that pixel's center, then ask a single question — what is the nearest surface this ray hits?
Answering it means solving for the smallest valid among all objects. For a sphere that is a quadratic equation; for a triangle you first intersect the ray with the triangle's plane, then check whether the hit point lies inside the triangle 1. In his 1980 paper, Turner Whitted wrapped objects in bounding spheres so a ray that misses a sphere can skip everything inside it 2.
The nearest hit gives you a point, a surface normal, and a material. Notice that visibility is now resolved per pixel, by construction: no depth buffer, no dependence on draw order, no interpolation shortcuts. But also notice the reversal. Rasterization is a production line for triangles; ray tracing is a search per pixel — like firing a laser pointer through every pixel and seeing what lights up. That is slower, and it is the reason the correct amount of light becomes available.
Step two: the shadow ray replaces a grid with one question
Once you know a surface point, you can do the ordinary local shading — diffuse plus specular from the lights. But is the light actually reaching this point? So you spawn a second ray from the hit point toward the light. If anything is intersected at a distance shorter than the distance to the light (), the point is in shadow for that light and the light's contribution is dropped 13.
Two details matter. First, this query only needs "does anything block the way?" — not "what is the closest thing". The ray can stop at the very first hit, which makes shadow rays cheaper than primary rays. That distinction is why ray tracing hardware has separate any-hit and closest-hit queries 4. Second, the shadow's shape is determined by geometry alone. There is no shadow map resolution to make edges blocky, no depth bias to cause acne, no peter-panning, and no shimmering when the light's projection changes as you move. The correctness comes free with the method, not from tuning.
One honest caveat: game lights are treated as points, so a surface is either lit or not. That gives razor-sharp shadow edges. Real light sources have area, and a soft penumbra requires sampling many directions toward the light and averaging them — which is where noise enters the picture, and where Whitted's model stops.
Step three: the reflection ray looks at what your screen cannot see
For a shiny surface, the incoming direction reflects about the surface normal:
Spawn a new ray from the hit point along , and whatever it hits is what you see reflected. The color it returns is the light arriving from that mirrored direction, so the final color of a point is assembled from the bottom up: its own local lighting from the lights, plus the reflectance coefficient times the color returned by the reflection ray, plus the transmittance coefficient times the color returned by the refraction ray.
This is the heart of what Whitted did. The paper describes the shading information as "a tree of rays" extending from the viewer to the first surface, and from there to other surfaces and to the light sources, created per pixel and handed to the shader 1. The recursive part is what previous algorithms lacked: visibility calculations did not stop at the nearest intersection; each visible intersection spawns more rays, and the process repeats until no new ray hits anything 13.
Here is why that fixes the failure you already know. The object you see in a reflection lies along the reflected direction, which is usually not along your line of sight to that pixel — it can be off to the side, behind a wall, or entirely outside the camera's view. Screen-space reflections, the technique from the previous article, can only reuse pixels that are already on screen, which is exactly why reflections slide out of existence at screen edges. A traced reflection ray leaves the screen's memory behind and searches the actual scene. Mirrors facing mirrors also come out right, because the recursion simply continues: the reflected ray hits a second mirror, which spawns its own reflection ray.
Step four: refraction makes glass look solid
For transparent materials, the ray does not bounce — it bends. Its new direction obeys Snell's law, , where is the index of refraction: about 1.0 for air, 1.33 for water, 1.45–1.65 for glass, 2.42 for diamond 3. You trace the transmission ray in that bent direction and take the color it returns, weighted by the material's transmission coefficient.
Because the bending depends on the surface normal at each point, a curved pane of glass shifts and magnifies what is behind it — the ray even travels through the object and bends again when it exits, so both surfaces matter. Total internal reflection appears on its own: past a certain angle the bend is geometrically impossible, and the ray can only reflect. Whitted's paper treats the reflection and transmission coefficients as constants, but explicitly notes that for accuracy they should follow the Fresnel law — varying with incidence angle, which is why real glass grows more reflective at grazing angles 2. Note also that games usually use one index for all three color channels; using different values per channel is what would split light into a rainbow, so dispersion stays a luxury effect.
The bill arrives: a tree, not a ray
Now count the rays. Primary rays equal the number of pixels: 1,920 × 1,080 is about 2.1 million, 4K is about 8.3 million. Add one shadow ray per light, one reflection ray, one refraction ray per hit. If every ray spawns children and recursion runs levels deep, the worst case per pixel is roughly rays — depth 3 with three branches is already 27. In practice rays terminate when they miss geometry or when the material's reflectivity is negligible; lecture notes put a depth of 3 as often sufficient, with a weight that decays each bounce as another stopping rule 3.
Intersecting every triangle with every ray would be hopeless. The fix is an acceleration structure: a bounding volume hierarchy (BVH), a tree of nested boxes where a ray that misses a box skips everything inside it. Whitted's bounding spheres were the same idea in 1980 2. Traversal cost per ray drops from "number of triangles" to roughly the depth of that tree.
Hardware closed much of the rest of the gap. NVIDIA's Turing architecture (2018) added RT Cores that perform BVH traversal and ray–triangle intersection in fixed-function units, cutting thousands of software instructions per ray down to a hardware probe. The claimed figure is on the order of 10 gigarays per second versus about 1.1 gigarays per second in software on the previous generation — around ten times faster 45. Notably, NVIDIA's own guidance says: "Do not expect hundreds of rays cast per pixel in real-time" 4.
Do the budget math for yourself. At 60 fps you have about 16.7 ms per frame; at 144 fps, about 6.9 ms. Take 1080p with one shadow ray and one reflection ray per pixel: roughly three rays per pixel, about 6.2 million rays per frame, about 373 million rays per second at 60 fps. Against 10 gigarays per second that sounds affordable — and this is why limited ray-traced effects became viable. But: 4K multiplies the pixel count by four; neighboring pixels' reflection rays scatter to completely different places, so they share almost no cached data; every hit still needs shading; rays that miss still pay traversal; and each additional effect or bounce adds unique rays. That is the real performance wall — not a single expensive ray, but hundreds of millions of unpredictable rays inside a frame that also has to do everything else.
How games actually spend the rays today
Almost no game ray-traces primary visibility. Rasterization still finds the visible surfaces cheaply, and the depth/normal/material buffer it produces becomes the launch point for rays that only handle the effects that need scene-wide information. EA SEED's Pica Pica demo rasterizes first, then traces reflections at half resolution — a quarter ray per pixel — plus a quarter ray per pixel for shadow rays at reflection hits, reconstructing full resolution from both spatial and temporal filtering, for roughly 2.25 rays per pixel total 67.
That reconstruction step is the escape hatch. Instead of shooting 16 rays per pixel to get a clean result, you shoot one and let neighboring pixels' ray hits fill in the gaps. So the quality knob in a game's settings is not simply "ray tracing on or off". It is: which effects receive rays, how many rays per pixel, at what resolution the ray buffer runs, how many bounces, and how aggressive the denoiser is. Turning ray tracing off usually removes an effect outright; lowering its quality usually trades sharpness and noise instead of correctness.
Judging what you see
The next time something looks off, the artifact tells you the technique. A reflection that dissolves as it approaches the screen edge, or that omits something you know is off-screen, is screen space. A shadow edge that is blocky on a fixed grid, or that crawls and shimmers as you move, is a shadow map. A reflection that is geometrically correct but slightly soft or grainy, sharpening when you hold still, is traced and temporally denoised — noisy rather than missing is the signature of a real ray. And because frame cost scales roughly with pixel count times rays per pixel, the crunches happen at high resolution with several traced effects, which is why upscaling and ray tracing are usually discussed together.
The boundary of Whitted-style ray tracing is worth stating plainly, because his own paper does: it handles mirror-like reflection, hard shadows, and clean refraction, but it does not provide for diffuse reflection from distributed light sources, so there is no soft shadow, no blurred glossy reflection, and no color bleeding from one matte surface onto another 2. Those need many samples per pixel averaged together, which is path tracing — and its characteristic artifact is noise traded against frame rate. That is the next step.
Answers to a quick self-check
- Why can shadow rays stop at the first hit while reflection rays cannot? A shadow ray only needs to know whether anything blocks the path to the light, so any hit is decisive. A reflection ray needs the color of the surface actually seen, so it needs the closest hit.
- A 1440p image (2,560 × 1,440, about 3.7 million pixels) with one reflection ray per pixel at full resolution needs roughly 3.7 million reflection rays per frame — about 221 million per second at 60 fps.
- Why does turning down ray tracing quality often produce blur instead of a missing object? Fewer rays are cast, and the denoiser reconstructs the missing detail from neighboring pixels and previous frames, so error appears as softness, noise, or smearing rather than as a structurally wrong image.