← Back

raytrace-on-screen: where a traced frame goes

A case study of one framebuffer demo, to learn what makes emitted Codex slow in Roc. raytrace-on-screen traces codex/test/raytracer-test's scene, two spheres over a floor, at 160 x 120 with Cobblestone's Raytracer, and copies the trace onto the screen. On the page it takes about 105 ms a frame.

The instruments

What the page's frame is

what share of the page's frame
rt-intersect-obj: the sphere and plane tests 22%
rt-closest: the walk over the scene for a ray 17%
geo-sqrt: a square root, written in Codex as Newton's method 11%
the demo's copy onto the screen, one poke-32 a pixel 11%
rt-shade 6%
the small vector functions and List.get 11%
the zig host about 1%

It is Raytracer's own Roc. The host's memory and screen are not where the time goes.

Step 1: the square root is the machine's

Geometry's geo-sqrt and Quaternion's quat-real-sqrt are the same loop: guess, refine, stop when two guesses are within ~, four ULPs. rocemit now writes both as Roc's sqrt behind the same guard (zero for an argument that is not positive), keyed by chapter and name, beside its existing rule that writes DeviceMath's real-sqrt as F64.sqrt. A test runs the loop against the instruction from a millionth to a trillion and on the first 2,000 perfect squares: they never differ by more than four ULPs.

30 frames, native trace render checksum
the Newton loop 0.86 to 1.12 s 1.26 to 1.35 s unchanged
Roc's sqrt 0.57 to 0.64 s 0.82 to 0.88 s unchanged

A third of the frame, and not one pixel moves. It has landed (rust-codex-compiler ad24db4) with every gate green: safari's 54 units, the ladder at 779 of 1,032 with its ledger unchanged, the gpu and games smoke tests, and verify.sh with all 18 hashes the same. The page's wasm now takes 61 to 69 ms a frame in Node, where it took 96 to 101.

Step 2: what rt-closest carries

rt-closest walks the scene's objects and keeps the best hit so far, and a hit is a whole RtHit: a flag, a distance, a position and a normal (two 3-vectors), and the material (a colour and three integers). Every object, for every ray, returns one, and the walk compares it and carries one forward. In the LIR, each step of the loop takes the best hit apart into its twelve scalars and builds it again. A sphere's hit also normalizes its normal, a square root, even when the sphere is not the closest.

Two hand-edited variants of the emitted Raytracer test what that costs. Both find the nearest object by its distance alone, then build one RtHit, for that object; the comparison is the original's, so the same object wins every time.

30 frames, native, fastest of five, interleaved trace render
as emitted, with Roc's sqrt 0.600 s 0.859 s
dist-calls 0.527 s 0.767 s
dist-inline 0.549 s 0.778 s

The checksums are the emitted build's in all three. Carrying distances instead of hits saves about 12%. Inlining the vector functions saves nothing that five runs can tell apart from noise.

Step 3: what is left

After both changes a ray costs about 900 nanoseconds against three objects, which is slow for native code. perf over dist-inline's trace finds no copies through libc and under 1% in reference counting. The time is in the functions' own code, and the disassembly says what that code is:

function share of trace instructions moves to or from a stack slot f64 arithmetic
roc__proc_343, the nearest-object walk with the distance tests folded in 48% 1,749 1,503 54
roc__proc_321, the bench's pixel loop, with rt-trace folded in 12% 1,936 1,712 0
roc__proc_342, called by the walk 11% 233 181 0

Eighty-six percent of the hottest function's instructions move a value to or from the stack. The dev backend gives a value a stack slot, not a register, and a record moves field by field from slot to slot. The hottest instructions in perf annotate are movsd to a slot. That is the dev backend's code, not Raytracer's shape: no rewrite of the program removes it.

Step 4: the same bench under LLVM

Steve allows LLVM for a bench that builds in under a minute; this one builds in 2 s. Fastest of five, interleaved, every checksum the same:

30 frames, native dev trace dev render LLVM trace LLVM render
as emitted, with Roc's sqrt 0.592 s 0.804 s 0.036 s 0.082 s
distances first, the distance computed in the walk (dist-calls) 0.520 s 0.753 s 0.032 s 0.081 s
distances first, read off Geometry's own ray3-sphere and ray3-plane (dist-rayhit) 0.555 s 0.788 s 0.032 s 0.084 s

LLVM is sixteen times the dev backend on the walk and ten times on the render. Under it, distances first still take a tenth off the walk, and nothing measurable off the whole render, where shading is the rest.

The upstream change is dist-rayhit's shape: it stays inside Raytracer and writes no arithmetic twice. Under LLVM it equals the variant that duplicates Geometry's; under dev it gets half as much.

What this leaves to decide

  1. The sqrt rule has landed (step 1).
  2. Distances before hits in rt-closest is Raytracer's change, not the emitter's: about 12% natively, the same images, and it would help every backend. It could go upstream as a PR.
  3. The rest is the dev backend. An optimizing build (--opt=speed, LLVM) is what keeps values in registers. By the standing rule that is for a finished page, not for exploring, so measuring it here is your call.
  4. For Zulip, this adds a second concrete question to the profiling one: whether the dev backend's register allocation is planned, since for floating-point code like this, stack moves are most of what it emits.
  5. Renderer3D is the scene demos' version of this question, and the same minimal-program approach applies to r3d_scan_cols_sh!.