The Roc machine now draws Cobblestone's screens: scene-on-screen's 3D
scene, the four GPU widget tests, and gop-padded-stride. It draws them
slowly: 3 to 4 seconds for a 320x240 frame in the page's wasm build. This
note is for deciding how the screen should be built before any more
profiling. It sets out what a pixel costs, why, and the choices, with a
recommended order. The questions at the end are Steve's.
Measured in the page's own build (wasm, Roc's dev backend) unless it says
otherwise. machine/batch/PERF.md has the details.
| what | cost |
|---|---|
gop-padded-stride: 123,000 pixel writes, 46,000 reads, one render |
3.0 s |
scene-on-screen: one 320x240 3D render |
3.4 to 4.1 s |
the same frame's pixels drawn by MachineGpu directly (native, dev) |
0.09 s |
a whole 640x480 GPU frame in MachineGpu (native, dev) |
0.37 s |
MachineGpu shows what Roc can do with a flat buffer written in place:
about a microsecond a pixel, including a depth test and a colour
interpolation. The emitted programs pay far more per pixel than that, and
the difference is not in the memory itself.
A Codex program that touches memory or a device is emitted with the machine
threaded through it: every such function takes the machine and hands it
back. One poke-32 in a loop becomes this Roc:
fill_fb! = |machine, base, i, n| (if (i >= n) { (machine, 0) } else { ({
(machine1, _d) = Machine.store(machine, base, (i * 4), sentinel, 4)
fill_fb!(machine1, base, (i + 1), n)
}) })
Machine.store finds what backs the address and returns
{ ..m, gpu: ... }: a new machine record. That record is flat, and wide.
It has 24 fields. The e1000, the NE2000 card, the HPET, the APICs, the PCI
table, the disks and the GPU are records inside it, and a Roc record inside a
record is stored inline. Only lists and dictionaries sit behind a reference.
So every door call copies the whole machine: roughly a kilobyte, an estimate
from the field list, since Roc does not report sizes. It does that twice for
every pixel a program writes and reads back.
BASIC measured this exact effect. Its step's cost grew with the machine record's width:
| added to BASIC's machine record | P134's run time |
|---|---|
| 15 lists, as loose fields | +35% |
| 15 integers, as loose fields | +13% |
| 15 lists, in a nested inline record | the same as loose |
| 15 lists, behind one reference (a one-element list) | about 0 |
BASIC's next step was to split its machine into hot state and a cold part behind one reference. It was planned, not built.
There is a second, older problem: Roc copies a list it can still reach
(reference_roc_list_copying). A value threaded down a recursion and handed
back, or a field taken out of a record in a separate statement, makes the
next write copy. MachineGpu avoids it by construction, measured flat at
every size; any new shape has to be measured the same way.
Steve's three modes for the screen, and what each one needs:
machine/native, whose C allocator already runs gop-padded-stride in
3.1 s against 22.4 s on Roc's default platform), checking the image, for
instance against a hash. This needs a native build and something to compare
the image with.A. Split the machine into hot and cold. The machine keeps the fields a
pixel touches (memory, the GPU's planes, the clock) and puts everything else
behind one reference:
Machine : { mem, gpu, clock, cold : List(Cold) }. A memory or GPU door then
copies a few dozen bytes instead of a kilobyte. A device door takes the cold
part out and puts it back inside its own record update, the shape BASIC's
devices proved.
Machine.roc's device doors, nothing in
rocemit.mmap count at three sizes, as
MachineGpu was.B. Memory owned by the host, as effects. peek and poke become hosted
calls into a byte array the platform owns, so the machine record stops
carrying memory at all.
C. The image code directly (mode 3). Emit the drawing chapters as a library, with memory as one flat, bump-allocated arena rather than the machine's tree. A browser app renders a frame per animation step into a framebuffer it hands the page.
D. An LLVM build for the finished page. Steve has agreed to LLVM for a finished product. BASIC's LLVM build ran 2.7 times faster, and struct copies are the kind of cost an optimizer can remove.
E. Report the copying to Roc. roc-apps/findings/ holds reduced programs
for three of the copying rules, and none has been reported. Steve's call.
scene-on-screen's scene turning, using the
code the machine already runs.B changes what the machine is, and I would not start it without a reason A leaves behind.
Steve chose A, and it is built (roc-apps 31e7fd0). The machine record is
{ mem, gpu, clock, devices }, and the other devices are one Devices record
in a list of one. A door that changes a device takes it out of the list,
writes it, and puts it back.
First, the question under it, measured: does updating one field copy an inline record beside it? On Roc's dev backend, yes. 20 million updates of one field took 0.13 s alone, 2.19 s beside a 512-byte record, and 0.21 s with that record in a list of one.
| what | before A | after A |
|---|---|---|
| gop-padded-stride, the page's wasm build | 1,506 ms | 716 ms |
| scene-on-screen, the page's wasm build | 2,662 ms | 1,413 ms |
| e1000-tx-deadline, natively on the ladder's platform | 10.18 s | 3.92 s |
The ladder is unchanged, and no unit measured allocates more.
A list of one, or a Box? Carried untouched, a Box costs the same as the
list. But writing through a Box means boxing again, which allocates on every
write: 200,000 device writes made 200,004 allocations, against 4 through the
list, and ran ten times slower. Roc's own documentation calls box and unbox
"expensive". So the devices stay in the list of one, and a Box is for a part
of the machine that is never written after it is made.
The numbers and the programs behind them are in roc-apps/machine/batch/PERF.md
and machine/batch/probes/.