Skip to content

P8X Project Backlog

Add ideas as they come; move items between sections as they progress. Last updated: 2026-10-02

How to use

  • NEXT — committed, in rough priority order
  • IDEAS — captured, not yet committed
  • VERIFY — open questions / checks before trusting something
  • WONT-DO / SUPERSEDED — settled decisions NOT to do something, kept so they are not re-litigated. Check here before starting anything that looks obvious.
  • Completed work lives in BACKLOG-DONE.md — the project log, plus every finished item lifted out of the sections above. Nothing here is done; if it is in this file, it is still live.

[~] marks a partially-done item: the finished part is described inline, the remainder is why it is still here.


NEXT

  • [ ] CF drive 1 port base: firmware + emulator vs the cf-card rev B (2026-09-19). The routed cf-card rev B decodes drive 1 at its own $FF18-$FF1F (header J5, own buffer + strobes); the firmware CF driver (CFSEL/DRVSEL ORed into CFHEAD) and the emulator's -c2 model select drive 1 with the ATA device bit on the shared $FF10-$FF17 task file. Drive 0 works as is; drive 1 on this card needs a drive-selectable port base in CFSETL/CFINIT (and the rest of the CF driver) plus the emulator's port decode -- data-integrity code, so test-validated (the /d1 mount tests). Or decide the card goes back to one port.

  • [ ] FPGA RTL write-protects $0000-$1FFF, the map says ROM ends at $17FF (found 2026-10-02). fpga/tang-nano-20k/rtl/p8x_top.v (the cpu/lcd board builds) and fpga/rtl/p8x_soc.v (the co-sim SoC) write RAM only for mem_addr >= $2000; since the 2026-09-14 ROM shrink the emulator (RAMBASE $1800), the firmware and the OS keep scratch in $1800-$1FFF (IBUF, SBUF $1D00, the BIOS block at $1F00 incl. the monitor line buffer). Move both boundaries to RAMBASE ($1800), then re-run the co-sims and rebuild + reflash the board. Related: fpga/sim/run.sh does not elaborate today -- p8x_soc.v instantiates sdram_model/sdram_arb/gfx_mem, which are not in its iverilog file list -- so the co-sim harness needs its sources brought up to date first.

  • [~] Hand-asm: from-scratch redesigns on the Tier A ISA (2026-09-12). tools/tierA_rewrite.py only covered the idioms it could prove safe; the goal is genuine rewrites that use the ISA's shape (word variables, (Pn+d) records, LEAW, PHW (Pn+d), table dispatch) — order: assembler, compiler, BASIC, then the OS/WM kernel and monitor last; the asm twins of the C commands are left alone (C versions will be compared against them later, and asm retired where C wins).

    • [x] apps/p8xasm.asm DONE 2026-09-12 — drop-in rewrite: chained-hash symbol table (256 buckets, 16-byte entries at $8000–$C5FF, 1,120 symbols), first-letter opcode index, DISPTAB operand dispatch, generic two-operand path (no MOVW/LDPn special cases), .org pads itself. 5,178 → 4,065 B; cover source 2.4× faster, self-host 5.7×, a 1,000-symbol source > 30× (the old scan did not finish in 900 M cycles). Byte-identical to the host on every test; error messages unchanged (a backward .org is now reported on the .org line). Layout note: code + OPCTAB must stay below $8000; INCBUF moved to $CC00, the BIOS dir-scan page to $CE00, path buffers to $D000. 2026-09-13: .ascii/.asciiz decode \n \t \r \0 \\ \" as the host assembler does (it copied bytes verbatim and a \" ended the string — the on-board cc emits C escapes raw); asm.c too; 4,116 B.
    • [x] basic/p8xbasic.asm DONE 2026-09-12 (moved ahead of the compiler) — drop-in rewrite: token-indexed STMTTAB/FACTAB dispatch (CHECKLINE reads STMTTAB for legal leaders), PHW/PLW around the evaluator, one CMPW + a relation mask for every compare, (P1+d) variable / FOR / GOSUB records, early-exit sorted line search, keyword matching only on letters. 11,151 → 9,124 B; arithmetic loop 1.6×, GOSUB+variables 2.0×, strings 2.8× faster. Verified by all 12 make test-basic tests plus a differential scripted session against the old binary (the only diffs are fixes: no stray ? before a lowercase-keyword error, no ?SYNTAX ERROR IN 0 after RUN, immediate FOR prints). FOR nests 3 deep (was 2, unchecked); a 4th GOSUB / 33rd variable / 17th string variable is ?SYNTAX ERROR; division stays UNSIGNED as before. Same -D build knobs; PROG moved to BASRAM+$580 unchanged.
    • [x] apps/p8xcc.asm DONE 2026-09-13 — drop-in rewrite that GENERATES the same code (differential compile of pwd/wc/grep/vi on the machine: text-identical apart from the tab indent): one arena/first-letter- chain name-table mechanism for all seven tables with (P1+d) access and length-first rejection, keyword codes from the lexer, word ops for slots/literals/decimal, single-pointer emit, tables at $B000 (code-only binary). cc.bin 20,915 → 10,075 B (code 13.3 → 10.1 KB); compiles 1.2–1.6× faster (pwd 2.60 → 2.10 M, wc 30.8 → 20.6 M, grep 55.8 → 36.1 M, vi 48.8 → 30.0 M cycles); emitted text ~35% smaller. Found and fixed an OLD bug: a char array declared after an int array got word elements (grep.c collect). Calls to undeclared functions now emit the name; syntax errors bail. The frame model (P3 frames instead of static slots) stays a separate, later item — it changes the generated code and needs its own tests. Later the same day, from the C twin's differential: the right operand of && leaves condition mode (if (a && b == c) fell into the body when a was false), 16-bit label numbers (a byte counter wrapped at 256: grep/vi got duplicate labels), 16-bit local-array sizing (char b[300] got 21 slots); caps 250 functions / 250 macros, an 11.5 KB name arena at $A000. 10,182 B. cc_c_test.sh guards all three.
    • [ ] os/p8xos.asm + os/wmkernel_body.asm, firmware/p8xmon.asm (last). Measure with p8xemu -L cycle stamps (scratch copies with STA $FF02 at entry and before the final message); each module with its own tests.
  • [~] Generate the ABI/scratch includes from the single source — C HALF DONE (2026-09-14). gen_memmap.py now also emits os/commands/lib_mem.c + os/commands-asm/lib_mem.inc (COMMAND_SYMS: the scratch/graphics/TPABASE addresses commands name). The C consumers are converted: the graphics commands (screen/term/desk/finder/paint/wdesk/write + lib_gfx), basic.c, and apps/cc.c + apps/asm.c //#use mem instead of hand-//#defineing; a memory-map move of those is now one python3 generators/gen_memmap.py. REMAINING (documented as intentionally-hardcoded in reference_p8x_memmap_singlesource, NOT started): the 28 asm command .org $5900, ~111 test/run.sh --base, the asm-app equates (p8xasm/ p8xbasic/p8xedit), and the fpga loaders -- ;#use mem there hits the ;#use(4)/.include(1) limits or has no build-context path, and the TPABASE literals are a uniform greppable sweep anyway. Also still open (nice to have): generate the memory-map DOC and the p8xos/p8xmon header layout comments from the table, and derive the test size caps (symtab − TPABASE).

  • [ ] Port the frame model into apps/p8xcc.asm — the deferred half (2026-09-13). apps/cc.c is frame-model but apps/p8xcc.asm (the default /bin/cc) is still static-slot, so the two compilers no longer emit byte-identical text and the cc_c_test twin DIFF is SUSPENDED (it now verifies cc.c by running its output). Porting the same frame + leaf convention into p8xcc.asm's ~3,900 lines (rewrite EMSLOT/EM_SAVESLOTS/EM_RESTSLOTS/EM_POPPARAMS/FUNCDEF prologue + add the leaf lookahead) restores the twin diff AND gives the DEFAULT on-board compiler the recursion-correctness + ~13% size win. Large, differential-tested against cc.c's output. Not on the self-host path.

THE BOARD HAS TWO STALENESS SURFACES; A FEATURE MAY NEED BOTH. The BITSTREAM carries the CPU, microcode, monitor ROM and graphics RTL (build.sh lcd load); the SD CARD carries the OS, /bin and BASIC (tools/imgsend.py). The ellipse spanned RTL and BASIC, and updating only the bitstream left the hardware understanding a command nothing could issue -- ?SYNTAX ERROR from a BASIC that could not parse it. Check both.

And check p8x_cpu.fs's MTIME after a build. This has hidden two separate failures (a 48/46 placement failure, and a multiply-driven st net). build.sh itself now stops correctly -- set -euo pipefail plus an explicit exit 1 on each of synthesise / P&R / pack, so load is unreachable after a failed build.

  • Finder desktop (two-mode P4) — the rest of the app frame. finder.c shipped a full-screen file browser + full-screen launch (SYS_EXEC). Still to do:
  • Mouse support (via lib_ptr, keyboard-only today). (The Apps menu -- press a -- and the FILE menu -- press f: rename / duplicate / move / new folder / delete -- both shipped 2026-09-10. The file ops delegate to mv/cp/del/rmdir/mkdir through the launch-and-return chain, so P8XFS needs no rename/rmdir primitive; c_finder_fileops_test.sh.)
  • Real pull-down menus. The APPS and FILE menus are dropdowns picked by a letter, and the top bar is a key-hint strip -- not Mac-style press-drag-release pull-downs from named bar titles (the old tiled desk had those via lib_wm). Apps should likewise take over the bar with a File/Quit pull-down. Pairs with the mouse work above.
  • (Kermit shipped 2026-09-10 — kermit send|recv /path over the P3 second port ($FF08/$FF09), P5's last app; P5 is now complete. Term, Write, and the Paint/Image frame adaptation had already shipped. See the asm-twin note below.)
  • Interactive serial terminal command (gap vs the original two-mode notes, 2026-09-10). The notes call for a pass-through terminal on the 2nd port — keystrokes out $FF09, port-2 bytes onto the console — for talking to another machine from inside Term, with file-transfer apps like kermit running over it. Only the transfer half shipped (kermit); there is no dumb-terminal command yet. Needs a poll loop over both ACIAs (console RDRF + port-2 RDRF) and an escape key to exit; C + asm twins per the /bin rule.
  • Real Kermit protocol interop. kermit is a fire-and-forward SEQ LEN data CHK stream with no ACK/NAK, so it only talks to itself (P8X↔P8X, or the emulator's -2i/-2o loopback). Talking to a host Kermit (Mac kermit, C-Kermit) needs the real protocol: SOH-framed packets, the S/F/D/Z/B packet types, ACK/NAK + retransmit, the char-encoding (tochar/ctl/unchar) and the init-parameter negotiation. Decide whether that interop is wanted before building it; the current framing is fine for P8X-to-P8X.
  • Finder File menu. The bar has an Apps dropdown (a) but no File menu; the notes want one (open / duplicate / rename / move / quit) — pairs with the rename/duplicate/move file ops above. Apps likewise should take over the bar with a real File/Quit dropdown rather than the key-hint strip Write/Term draw.
  • Retire the tiled desk/wdesk once Finder covers their use; their FILES/ launch logic carried forward, the tiling did not.
  • Real-serial arrow-key timing: finder's ESC-sequence decode uses a bounded keyrdy() spin; over a slow link a lone ESC vs an arrow may still race -- revisit if it misbehaves on hardware.

  • Glass TTY (two-mode P2) — screen/kermit asm twins. The glass TTY behind BIOS CONOUT is ALWAYS-ON (2026-09-10) and is now a hardware text OVERLAY (a char-gen plane composited over the GL bitmap at scanout, gtxt.v): proper scrollback, per-cell erase, and the per-glyph-GTEXT-program speed cut all shipped with it — see BACKLOG-DONE "Glass-TTY text overlay". screen off still disables the mirror for a session. Still open:

  • screen and kermit command asm twins. os/commands/screen.c and os/commands/kermit.c both shipped C-only; per the /bin dual-twin rule each needs an os/commands-asm/*.asm twin (and the run.sh build lists updated) — see feedback_p8x_new_command_dual. kermit's twin needs the 2nd-ACIA poll ($FF08/$FF09) plus the FS wrappers it already uses.
  • VERIFY (RTL fit): the overlay char RAM is 80×34×8 = 2,720 bytes ≈ 2 Gowin BSRAMs, and sdram_video's ax/6 / ax%6 are constant-divisor divides; confirm placement + timing on the next card/lcd synth run. The co-sim is clean (c_gl_ovl_rtl_test is byte-identical to the emulator), and the ROM's per-glyph GTEXT machinery is gone, so there is headroom on both sides.

  • [~] FPGA build (Tang Nano 20K) — MILESTONES 0-4 DONE (2026-08-12); clock-up and IRQ remain. A standalone FPGA P8X running the same microcode and the same unmodified monitor/OS/toolchain. Parallel track to the TTL build, not a replacement. See fpga/README.md.

    • Done: first light (UART echo); the CPU core co-simulated against the emulator cycle-for-cycle across all 88 opcodes; ACIA + driven console with console output diffed; the core on real hardware with the full 64K map; SD-over-SPI behind the $FF10..$FF17 CF task file, with P8X/OS booting from a microSD and running the whole /bin toolchain. P8X is in the board's flash, so it comes up standalone on power.
    • Milestone 5 — clock up. Currently 9 MHz: the fabric runs at 27 MHz and a microcycle takes three phases, because it needs two dependent block-RAM reads (the microcode word first, since its PSEL field picks the pointer that drives mem_addr, then the memory byte). Fmax is ~50 MHz, so there is a lot on the table. Options: overlap the two reads by pipelining the microcode fetch a cycle ahead; drop to two phases; or raise the fabric clock with a PLL. Any change must still diff clean against the emulator (fpga/sim/run.sh x3) — that is the regression test.
    • Milestone 5 — IRQ. irq_set is currently tied low in fpga/tang-nano-20k/rtl/p8x_top.v. The core already implements the rev-C forcing-buffer entry ($08 injection, vector $0808, EI/DI/RTI) and isa_test.asm exercises it in simulation; it just needs a real source wired up (timer and/or the ACIA).
    • Milestone 6 — graphics display for BASIC. A 4.3" Sipeed 480x272 RGB panel, driven only by new BASIC statements (LINE, COLOR, BOX fill/nofill) — NOT a text console, so no font ROM, no PUTC hook and no OS changes; serial stays the console.
      • Geometry is forced by block RAM. 6 spare blocks = 12288 bytes; 480x272 needs 16320 at even 1 bpp, so the panel resolution does not fit at any depth. The framebuffer is 240x136 at 2 bpp (8160 bytes, 4 blocks, 2 spare), pixel-doubled to fill the panel with square pixels. Four pens index a 12-bit RGB palette. 8 colours (3 bpp) would cost all 6 remaining blocks and straddle byte boundaries; 16 colours at full resolution is an SDRAM project.
      • The drawing engine is in the DEVICE, not in software: BASIC loads registers and writes a command byte. Spends the resource there is spare (12.7k LUT4) instead of the one there is not. A full-screen fill is ~1 ms instead of ~180 ms, and BASIC never has to mask sub-byte pixels.
      • DONE: the emulator models it ($FF20-$FF26, p8xemu -g/-G, make test-gfx), which is the golden model the RTL gets written against. Two rules are load-bearing and pinned by test/gfx_test.sh: endpoints are INCLUSIVE, and off-screen pixels are DISCARDED rather than clipped (coordinates are bytes, and y*60 + (x>>2) would otherwise fold x>=240 onto the next row).
      • The bus card is the SAME device (2026-08-14). The planned physical card is a Tang Nano 20K plus this same 4.3" panel, so resolution, command set, RTL core and golden model are shared; only the front-end differs (internal CPU bus vs. an external bus interface). That kills the earlier worry that a smart engine was affordable on the FPGA but not in TTL — there is no TTL engine to build. Command set is settled and modelled: PLOT/LINE/BOX/BOXFILL/CLS/SETPAL/CIRCLE/CIRCLEFILL/POINT plus SELFTEST/RESET/IDENT, with a "PG" presence signature and an IDENT record carrying the geometry. Coordinates are 16-bit pairs (a low-byte write clears its high byte) so 480x272-over-SDRAM stays reachable without a protocol change.
      • Card hardware, still open: the P8X bus is 5 V TTL and the Nano is 3.3 V, so the interface needs level translation on D0-D7 (bidirectional) plus the address/control inputs — 74LVC245-class parts. Address decode is the standard I/O-page detect from docs/p8x-card-standards.md plus A7..A4 = 0010. The bus write strobe is asynchronous to the Nano's 27 MHz, so it needs synchronising, and per-IC 100nF decoupling applies as on every card.
      • DONE: the BASIC statements (2026-08-14). COLOR pen, CLS, LINE x0,y0,x1,y1, BOX x0,y0,x1,y1[,FILL|,NOFILL] — tokens $A5-$AA, covered by emulator/test/basic_gfx_test.sh. Three things that are not obvious: NOFILL HAS to be a real keyword (with FILL tokenised and NOFILL not, CRUNCH matches FILL inside the word and an outline silently comes out solid); CLS needs the GPEN RAM shadow because GCOL is write-only in the device, so the pen cannot be read back and restored; and FILL/NOFILL had to be added to CKLEAD's blacklist or a bare FILL line would be accepted as a statement.
      • DONE: the RTL (2026-08-14). fpga/rtl/gfx.v (registers, drawing engine, framebuffer, palette) + fpga/rtl/video_rgb.v (480x272 timing, 2x-doubled scanout). fpga/sim/gfx.sh byte-compares the frame the RTL produces against p8xemu -g for both payloads: identical. Fits the board: BSRAM 44/46, Fmax 49 MHz, and NO PLL (9.009 MHz wanted, 27/3 = 9.000 delivered by the divider the CPU already uses).
      • The BUSY contract, learned the hard way. The emulator draws instantaneously; the RTL takes roughly 9 ms for a full fill, and a command written while another runs ABORTS it. Software MUST poll GSTAT bit 7. Code written against the emulator alone looks perfect there and draws a few scattered pixels on the RTL -- which is exactly what the first frame diff showed. BASIC now has GWAIT/GEXEC and the payloads call GWAIT before every command; the poll is free when BUSY is never set, so one binary is correct on both. This is also why graphics cannot be CYCLE-diffed: a program polling GSTAT legitimately reads different values on the two models, so the framebuffer is the thing that must agree.
      • DONE ON HARDWARE (2026-08-16). build.sh lcd load, then I/B/ basic, and BASIC draws on the panel. Pinout and timings verified from Sipeed's own 480x272 example (CLK 77, DEN 48, R 38-42, G 32-37, B 27-31; 560x297 at 9 MHz = 54.11 Hz; DE-only, no HSYNC/VSYNC). BSRAM 44/46, Fmax ~48 MHz, no PLL.
      • Two display bugs, both invisible in simulation:
      • fb_data >> ((2'd3 - ax[2:1]) << 1) -- a shift AMOUNT is self-determined in Verilog, so it evaluated in TWO bits and gave shifts of 2,0,2,0 instead of 6,4,2,0. Every pixel in the left half of a byte was invisible and every pixel in the right half drawn twice. Same class as the px_row truncation.
      • the framebuffer inferred as TRUE dual port, which halves a Gowin block's depth: 8 blocks instead of 4, 48/46, would not place. It now shares ONE port with the engine holding one cycle in three. Writing the shared read with two destination registers is NOT synthesisable as block RAM (falls back to 1020 RAM16SDP4); one read register feeding both is, at the cost of a pipeline stage.
      • A STALE p8x_cpu.fs can be reprogrammed without anyone noticing -- that hid the 48/46 failure for two rounds. The script's own guards are now correct (see the note at the top of this file); what is left is the shared bitstream FILENAME between the cpu and lcd targets. Check p8x_cpu.fs's mtime if a fix appears to do nothing.
      • Test gap that let bug 1 through, now closed: tb_video checked frame shape and gfx.sh checked framebuffer contents; nothing checked the MAPPING between them. sim/tb_scanout.v does.
      • FIXED (2026-08-17): CPU register writes were gated by the scanout hold. The if (sel && wr) block sat inside the engine's else -- the branch that does not run while the scanout owns the framebuffer port -- so any write landing on a scanout cycle was silently dropped: one in four in simulation, one in three on the board. That is the whole of the "RTL misses SETPAL and BOXFILL" mystery: their GCMD write happened to collide with a hold and the command never arrived, while every command that did not collide went through. It also explains the earlier clue that only commands where GWAIT had to SPIN were skipped -- a spin puts the following write at a different, unluckier phase. Register writes now live in their own always block, ungated; they touch no framebuffer port, so there was never anything to gate. gfx.sh passes on all three payloads.
      • OPEN: the co-sim exercises the shared port (irregular LFSR hold), but reintroducing the pending-write bug did NOT make it fail. The contention coverage is therefore unproven and worth understanding.
      • DONE (2026-08-17): ELLIPSE ($0A) / ELLIPSEFILL ($0B) in the RTL, matching the emulator pixel for pixel; test_gfx3.asm covers a wide outline, a tall fill and a near-circle and is part of gfx.sh. Two bugs the frame diff caught, neither visible by reading: the region-1 initialiser is 4ry2 - 4rx2ry + rx2 and the middle term was written ry2ry, so the walk never started; and a FILL must load its first span through the span-init state, because the circle begins its walk at x=r (seeding cx to ccx-r is right there) while the ellipse begins region 1 at x=0, where the span is the single pixel ccx.
      • CLOSED by retirement (2026-09-01): SELFTEST ($F0) was emulator-only (the RTL always rejected it), and the single-interface migration removed it with the rest of the CPU door -- there is no register to poke it through any more. Its "prove a card with no software" role belongs to the bridge PING + GLID probe and the monitor's wake-up console (its blank screen + banner on the LCD, which replaced the splash).
    • SD error paths are now tested (fpga/tang-nano-20k/sim/tb_sd_spi.v with sd_model.v +sdfail=1|2); that found and fixed two lockups. Still unexercised: CRC failure, a card that reports write-protect, and card removal mid-transfer.
    • Console newlines: FIXED (2026-08-13). CONOUT/PUTC now expands a bare LF into CR LF, so P8X no longer depends on a host tty doing it. Files and pipes are untouched (they route through OUTCH to a file or capture buffer and never reach CONOUT), and an existing CR LF is not doubled (TTYLST). TTYRAW ($60A1) disables it for binary transfers. BASIC's private PUTC/GETC — a leftover from the retired standalone build — now tail-call the BIOS, so it inherits the same behaviour. The standalone BASIC build (BASORG=$0000, BASIC as the whole ROM) is now genuinely dead — it would fail at the BIOS call — and is marked RETIRED in basic/README.md. See the separate NEXT item for BASIC's CWD bug.
    • Not done: nothing uses the board's 64 Mbit SDRAM — it turned out to be unnecessary once the microcode ROM was compacted (see fpga/tang-nano-20k/mk_compact_ucode.py), but it is there if a future build wants more than 64K.
  • [~] Second CF drive — FULL DUAL-VOLUME core DONE (2026-06-27); cross-drive single-command copy deferred. Landed (emulator + firmware + OS, hardware card deferred): two CF cards as equal read/write P8XFS volumes, each with its own current directory, 0:/1: prefixes, a switchable current drive, drive 0 = boot/default.

    • Emulator (a886a47): struct cf_state cf[2], -c2 <img>, ATA DEV-bit routing (CFHEAD bit 0), absent-device safe.
    • Firmware (30f388b): DRVSEL ORed into CFHEAD by CFSETL/CFINIT; CFSEL ($0148) / CFCURDRV ($014B) jump-table entries; bounded CFWAIT/CFDRQ (~4096 polls) so an absent drive times out, not hangs. cf2_test.sh.
    • OS (dd7beb4): CURDRIVE + a drive-1 CWD backing block; SWITCHDRV (swap working↔backing CWD, lazy CFINIT via a DRVINIT bitmask); PARSEDRIVE in RV_START (one-shot N: prefix → resolve from that drive's root); bare 0:/1: switch (CKDRIVESW), CD N:/dir, prompt shows the drive; SYS_SETDRIVE/SYS_GETDRIVE. os_dualvol_test.sh (prefix
      • switch + isolation). Single-drive behavior byte-identical; full suite green. Bulk cross-drive copy DONE — IMPORT N:/dir built-in. Provisions a fresh card from a "master": walks the source directory (drive N) collecting its files, then for each one reads the whole file into a RAM buffer on the source drive and writes it into the CWD on the destination drive — flipping the ATA device bit between the read and the write of every file. os_import_test.sh (build a boot volume + a master with /BIN/{ALPHA,BETA}, IMPORT 1:/BIN, host-verify both land on drive 0 with exact content). Root-cause fix that unblocked cross-drive I/O (emulator): the CF model stored the LBA/feature task-file registers per device and routed writes to the currently-selected device. But CFSETL writes CFLBAx before it writes CFHEAD (the device-select), so on a drive switch the LBA landed on the old device and the newly-selected one executed with a stale LBA — the dev=1 lba=0 / wrong-refill-LBA symptom that had blocked cross-drive copy. Real ATA has a shared task-file bus (both drives latch LBA writes; the DEV bit picks who runs the command), so the emulator now mirrors feature/LBA writes to both devices and routes only data/command to the active one. This is the same defect that stalled single-command CP 1:/X 0:/Y. Firmware foundation (152a51a): CFSEL lazy-CFINITs a drive on first select (CFIMASK); BIOS read/write streams carry their own drive (ROSDRV/WOSDRV, captured by FOPEN/FWOPEN, re-asserted by FG_FILL/FW_FLUSH/FCLOSE). Full suite green; single-drive byte-identical. SUPERSEDED (2026-07-08) by the Unix-style mount migration (branch mount-drives, see docs/mount-drives-design.md). The 0:/1: prefix model below was replaced by mounting drive 1 at /D1 in one namespace: the drive decision moved into a single FRESOLVE/RV_START mount redirect, so lib_drive.c and all per-command prefix code were deleted and commands are drive-unaware (CAT /D1/X, cross-mount CP /D1/A /B). This retires the filter-tool limitation — grep/wc/sed/… reach /D1 for free (they build absolute paths → FRESOLVE) with zero code growth, because no command parses a drive. The historical record of the N:-prefix work is kept below.

    Inline N: prefix on /BIN commands — DONE for cat/dir/cp/mv/diff (2026-07-07). New shared lib_drive.c (hasdrive/pdrive/seldrive over BIOS CFSEL) wired into the self-contained openers catpath (cat) and dir, and into abspath (cp/mv/diff). A path may carry a 0:/1: prefix (CAT 1:/X, DIR 1:/BIN, DIR 1:/*.C), and cp/mv/diff take a prefix on either path — so single-command cross-drive CP 1:/A 0:/B works: each stream keeps its own drive (ROSDRV/WOSDRV) and the shell's per-command SYNCDRV means routing to the other card never leaks. The earlier revert failed only because of the emulator's per-device task-file bug (fixed with the shared-bus model); with that gone the per-stream approach is correct. os_binprefix_test.sh verifies CAT/DIR/CP across drives (and that CP doesn't leak onto the wrong card). Deliberately excluded: the stdin-filter tools (grep/wc/head/tail/ more/sort/uniq/sed) share lib_stdin/lib_globx; the largest, grep (which also carries the -r tree walk), sits right at the TPA ceiling — its globals already overlap the $EA00 FSDIRBUF page and it works only at its exact current size, so the ~1.2 KB of prefix code pushed it into its own $FA00/$FC00 I/O buffers and corrupted -r/glob. Those tools follow the current drive (switch-then-run). Would need a size cut (or a per-command lib) to include them. Also open: drive-scoped PACK/FORMAT/FSCK act on the current drive via DRVSEL (no dedicated test); find/tree walk the current drive's CWD only (no path arg to prefix).

    Original scope note (superseded — we went full dual-volume, not read-only): keep a "master" CF holding core files (e.g. /BIN binaries) and, in the field with no host, provision a fresh card by copying from it — FORMAT, insert master, IMPORT 1:/BIN, done. Scope was to be read-from- drive-1 only, NOT full dual-volume: the working/boot volume stays drive 0 with the normal CWD; drive 1 is just a source you read/copy from. This avoids the heavy FS refactor (no per-drive CWD, no mounting). - HW: a second CF port at its own decode (e.g. $FF18–$FF1F) — one more '138 term + buffers + socket. (Master/slave on one channel is too unreliable for True-IDE CF; the driver also hardwires $E0=drive 0 today.) - BIOS: make sector I/O drive-aware — CFRDSEC/CFWRSEC select the drive (port base / DEV bit) per transfer; the read stream carries its source drive and the write stream its dest drive, so FGETB(drive 1) and FPUTB(drive 0) interleave in one copy loop. Add a select call / per-init IDENTIFY+SET FEATURES for card 1. - OS: honor a leading N: drive prefix on source paths in the resolve/FOPEN path (unprefixed = drive 0, with CWD). Then CP 1:/BIN/X /BIN/X works as-is (cp reads src/writes dst). Add a bulk IMPORT 1:/BIN (walk drive 1 with the find/dir-R recursion, copy each file to drive 0). - Emulator: a 2nd image (-c2 disk2.img) modelling the 2nd device. p8xfs.py is already per-image (build the master with it). Bonus: this also solves the post-FORMAT bootstrap (repopulate /BIN with no host), which unblocks the minimal-kernel split (DIR/PWD→/BIN) below. Effort: HW small, BIOS moderate/low-risk (additive), OS N:+IMPORT is the real work but far less than general dual-volume.

  • P8XFS v2 — remaining loose ends (the hierarchy itself is DONE; see DONE):

    • on-target FORMAT — DONE (2026-06-22, see DONE). Added the FORMAT command; it fit once the OS moved to $4000 (rev D).
    • OS code size — 16 KB ceiling (rev E). The boot loader (CMD_B) loads the OS to $2000 upward. The firmware/BIOS scratch sits at $6000 (monitor line buffer $6000, param/state block $6040, SBUF $6100), so the OS image must end below $6000 — i.e. 16 KB of RAM ($2000–$5FFF). This now matches the on-disk OS region (LBA 1–32 = 16 KB) exactly, so RAM and disk impose the same cap. The OS is ~9.5 KB today → ~6.5 KB headroom. (rev E dropped the OS to $2000 and the scratch/TPA −$1000, growing the TPA to ~37.9 KB.) BIOS scratch/SBUF/OS vars: LBA $6047, SBUF $6100, OS vars $6300.
  • [ ] History persistence (optional). The history ring is RAM-only (cleared at cold start). If cross-session history is wanted, add explicit history -w [file] / history -r [file] (dump/load the whole ring in one FCREATE/FOPEN) rather than a per-command append — P8XFS is contiguous one-extent-per-file, so appending each command would rewrite+reallocate the file and churn the disk.

  • [ ] memmap: build-time regeneration. generators/gen_memmap.py emits the committed memmap.{inc,h,py}. They must be re-run by hand after editing the canonical MAP. Follow-up: have run.sh / the Makefiles invoke gen_memmap.py before assembling/compiling so the generated files can never be stale (the "option 2" deferred when this landed). Low urgency — the table changes rarely.

  • [ ] memmap: fold the file-local temps in (full flat map). TMP/TMP2/CNT are kept out of memmap.inc because firmware and the OS each define them at different addresses (a name collision). Forcing them in means renaming the OS side (TMP alone = 97 refs, 131 total) to unique names. Deferred as high-churn / low-value (they're working temps, not layout).

  • [ ] memmap: auto-single-source the compiler-emitted .org. apps/p8xcc.asm and compiler/p8cc.c emit .org $5900 (= TPABASE) as literal text — they can't .include memmap.inc (their TPA buffers reuse OS-scratch names like NAMEBUF) nor interpolate a symbol into emitted text. p8cc.py already reads memmap.TPABASE. Options: rename the internal buffers to avoid the clash then .include, or teach the emit path to substitute the value. Pointer comments mark the coupling meanwhile.

  • [ ] Multi-stage pipes (a | b | c). The shell's pipe state machine (PIPEF/PIPESCAN/PIPE_RHS) handles exactly two stages: it splits on the first |, runs the left into PIPE.TMP, then re-dispatches the right. The re-dispatch jumps to DISPATCH without re-scanning for |, so a third stage is swallowed as args of the second command. To support N stages, PIPE_RHS would need to re-run PIPESCAN on the remaining line (chaining temp files), or the splitter could iterate left-to-right. Until then, CAT f | GREP x | WC silently drops the | WC.

  • [ ] Verify DIN 41612 footprints against physical connectors in stock (row A/C orientation when mated, mounting holes, press-fit vs solder)

  • [ ] Order backplane PCB first as the cheap validation article

  • [ ] SYS_OPEN — open-by-name as a syscall (2026-07-16). Commands each repeat the same four steps: SYS_GETCWD -> build an absolute path in their OWN path[80] -> FRESOLVE -> FOPEN. 16 of 25 commands call FOPEN directly; 15 build their own CWD-prefixed path. That per-command buffer is exactly where Wave 2 found overflows in dir/cmp/mv/cat/find — five buffers, five bounds to get wrong. Fold it into one OS call with ONE bounded buffer. No ABI change: commands still receive a raw string, they just make one call instead of four. lib_apath.c's abspath() stays — cp/mv/diff need the path STRING, not just an open. Also kills the FSDIRBUF/SBUF footgun (see the cat fix, eda2b7f): any command that opens a file while its stdout is a redirect must currently remember to FSDIRBUF its dir scan off SBUF, or the scan overwrites the redirect's buffered output. That belongs in the OS once, not in every command forever — and pipes make it systematically likelier, since every pipe stage is a redirect-writer. Scoped 2026-07-16 — the implementation is THIN, the OS already has the primitive: RESOLVE (p8xos.asm:1704) takes a path at P2 and yields SDIR = parent dir + NAMEBUF = leaf, already handling CWD-relative resolution and the mount/drive redirect via RV_START. SYS_MKDIR (:1704 area) is the model for a path-taking syscall: it just moves P1->P2 and calls a CORE routine. So SYS_OPEN is roughly P2 = P1; JSR RESOLVE; bridge SDIR/NAMEBUF -> the BIOS FNAME/DIRLBA; JSR FOPEN. See FINDP2 (:1484) for the existing resolve-then-find bridge — reuse it rather than reinvent. Free syscall slots: $2024 and $2027 (the table ends at SYS_MKDIR $2021). Open design decision — no free 512-byte page in OS scratch for SYS_OPEN to own a private dir buffer ($6000-$62FF is BIOS scratch + SBUF $6100; $6300-$69FF is fully allocated: RUNPATH $6740, PATHBUF, APBUF $6800...). Three routes: 1. bounded OS-side path buffer only; caller still supplies the scan page. Kills the overflow class; cheapest; leaves FSDIRBUF with the caller. RECOMMENDED — most of the win, least risk. 2. flush the write stream before the dir scan — no page needed and kills the footgun outright, but partial-sector flush-then-append is FS surgery. 3. reclaim a page from $63xx-$69xx. Pairs with shell-side glob+argv (IDEAS): together they are the whole "stop making commands resolve paths and expand globs" thesis. This is the cheap half.

  • [ ] CWDPATH is 48 bytes — deep paths truncate (2026-07-16). SETPATH now bounds appends (b41f5f1), so a path deeper than the buffer clamps instead of overflowing into INMODE/INARM/CWDLH — but it still truncates. A real tree deeper than 47 chars leaves CWDPATH short, and since SYS_GETCWD hands that string to programs to resolve relative paths, a command run from a deep directory can resolve against the wrong (truncated) path. CWDL/CWDN stay exact, so the OS itself is fine; only the string is short. Bounded, not solved. Options: (a) grow CWDPATH — needs space in OS scratch ($6300-$69FF is fully allocated, same wall as SYS_OPEN's dir buffer); (b) make cd refuse a path that would not fit, which is honest but makes a legal directory unreachable; (c) store the CWD as LBAs and render the text on demand by walking parents, which removes the buffer as a limit but costs a directory walk per prompt. Depth of ~5 components with 8-char names is the practical ceiling today. Nobody has hit this in normal use — /src/os-bios is 12 chars — so it is recorded, not urgent.

Bus test card (USB bring-up controller)

Design is settled; the board is routed in KiCad (2026-09-18, the uniform 280×140 card, 0 unconnected; first designed at 27 parts on a 160×100 Eurocard). See hardware/bustest-card/p8x-bustest-card-design.md. The thesis: a control card you can type at, driving the backplane one microcycle at a time over a USB serial line, with the emulator as the reference model. Nothing below has been built or measured.

  • [ ] Make buscon.c actually build (2026-07-22). The firmware (hardware/bustest-card/firmware/buscon.c, 241 lines) has the whole protocol — field lookup against the microcode's own DOE/DLD/PSEL names, ownership groups, the rest→A→B→rest phasing, the 5V-present interlock — but there is no CMakeLists.txt, so it has never been compiled, and mcp_write/mcp_read are empty shells marked TODO(hw). Needs: pico-sdk build wiring, the real MCP23S17 SPI transaction, and SPI timing chosen against the datasheet (clock rate, CS setup/hold). Do this before ordering the board (the copper is already routed). It is the cheapest way to find a design error: the pin map, the field→(chip,bit) allocation and the netlist all have to agree, and a compile plus a host-side harness catches a swapped chip or an off-by-one bit while it is still a text edit.

  • [ ] Generate the field→(chip,bit) map instead of asserting it (2026-07-22). FIELDS[] in buscon.c maps each microcode field to an expander chip and bit, and the comment says it "MUST match the netlist allocation" in gen_eagle.py. That agreement is currently maintained by hand and checked by nobody — the classic way to lose an afternoon on the bench. gen_eagle already computes the allocation (alloc), so it can emit a buscon_pins.h the firmware includes, making the two structurally one source. Same pattern as gen_memmap.py.

  • [ ] Resolve the design doc's own open items before fab (2026-07-22). Carried in §10: where CLK parks when halted (affects listen-mode sampling only); the eight status LEDs in §5.4 are a guess and want a second opinion; and contention margin — dropping the bus series R means a wrong drive is limited only by device R_on (~25–50 mA, abs-max-safe but not indefinite), accepted for a careful bench tool and reversible by adding 100 Ω. Also blocked on the project-wide DIN 41612 mating-orientation check already listed above, which bites this card as much as any other.


IDEAS

  • [ ] Off-screen content: scrolling / scroll bars (2026-09-17). How to handle content that doesn't fit one 480x272 screen -- the immediate case is a DIRECTORY in finder with more entries than the icon grid holds, but the same problem hits the spreadsheet (sheets past A1..H12, deferred at its first cut) and any long list (man/scrollback already scroll their own way). Options to weigh:

    • a scroll BAR (a thumb on the right/bottom edge, draggable with the mouse, page on click above/below the thumb) -- the most discoverable, and the mouse + following cursor now exist to drive it;
    • keyboard paging (PgUp/PgDn, or arrows past the edge auto-scroll) -- cheap, already how finder's selection could grow;
    • a viewport model shared by finder + sheet: a first-visible-row/col offset, clamp the selection to scroll the window, redraw the visible slice. Both apps draw a fixed grid today, so this is a real refactor (draw cell/icon at (index - offset)). Decide the interaction once and apply it to finder AND sheet so they feel the same. See finder.c (icon grid) and sheet.c (cell grid).
  • [ ] BASIC graphics cursor (2026-09-17). A software crosshair/pointer for BASIC programs, matching what finder/paint already draw client-side. Three primitives to add (statements + maybe function forms):

    • turn the cursor ON / OFF -- e.g. CURSORON / CURSOROFF, an XOR/ complement crosshair the interpreter tracks and redraws (self-inverse, no read-back, the finder cur_xdraw idiom -- draw single LINES per arm, NOT degenerate rectlines, or you get the "four compass dots" bug just fixed in finder).
    • read the cursor position -- CURSORX / CURSORY functions (or one CURSOR(0/1)), returning the current window coords so a program can react to where the pointer is.
    • Move source: the pointer already arrives as xterm SGR on the console (lib_ptr), and on hardware via lib_ps2 / the PS/2 mouse. BASIC would run a poll/event hook (a MOUSE-style read, or fold into INPUT) to advance the cursor. Decide whether motion is polled by the program or tracked live by the interpreter's input loop.
    • Both twins (basic.c + p8xbasic.asm) share the token ABI; free tokens after the layer keywords are below $FC (the GL verbs fill $B4..$F5, GLRD $FB, TEXTON..GRAPHICSOFF $FC..$FF) -- so this needs a token-space plan (reclaim an unassigned slot, or a MOUSE/CURSOR sub-keyword scheme). See the GXEN layer work for the pattern.
  • [ ] write: rich text -- multiple fonts / sizes / colours (2026-09-17). The write editor currently lays one font at one size in one colour. Let a document carry runs with different fonts, sizes, and colours. The GL text layer already has the mechanisms: TSIZE scales glyphs, the pen (GCOL/pen colour) sets colour, and the stroke font streams from /FONT.GL -- so size and colour are nearly free; multiple fonts is the real work (either additional /*.GL stroke files selectable per run, or the chargen bitmap font as a second face). Design questions to settle:

    • Document model: store runs as (text, font, size, colour) spans, not a flat char buffer -- pick a representation the editor can edit in place and re-flow (insert/delete inside a run splits/merges it).
    • Line layout with mixed sizes: a line's height/baseline follows its tallest run; word-wrap and the cursor must walk variable-width glyphs (GTEXT/TSIZE give advance widths). Today write assumes a fixed cell grid.
    • UI to set attributes: a menu or key/combo to change the font/size/ colour of the selection or the next-typed run (mouse selection now exists).
    • Save format: extend the on-disk file to record the spans (or a simple markup) so a reload restores the styling; keep plain .txt readable. Start with size + colour (cheap, high payoff) and add font faces after the font-file story is decided. See write.c, lib_gfx (TSIZE/pen/GTEXT), and the GTEXT 2D-text notes; the md panel-renderer idea wants the same size/colour plumbing.
  • [ ] Shared app frame / consistent UI chrome (2026-09-17). paint, write, sheet, term (and finder) each hand-draw their own chrome -- menu bar, title, close/quit box, mouse cursor, status line -- with subtly different geometry and behaviour (the recent cursor + menu fixes had to be applied app by app). Factor a common full-screen app frame into a lib so every app gets the same look and input handling for free:

    • a standard menu bar (title at left, menu items, consistent height/ colours) with mouse hit-testing + keyboard access;
    • the window frame / title and a close box in the same spot every time (the paint red-X quit convention);
    • one mouse cursor implementation (the XOR crosshair cur_xdraw idiom -- single LINES per arm, not degenerate rectlines) shared, not re-coded;
    • an optional status line and, once the scrolling idea lands, standard scroll bars -- so all of this composes. Design it as a small retained-mode helper (draw frame, register menus + handlers, run the event loop, call back into the app for the content area) so an app supplies only its canvas + commands. Big consistency payoff and it collapses the per-app chrome duplication. Pairs with the scrolling/scroll-bar and BASIC-cursor ideas above. See paint.c, write.c, sheet.c, term.c, finder.c, and lib_gfx; the resident WM kernel already does chrome for its windows -- decide whether the full-screen apps share that or a lighter lib.
  • [ ] imgsend: VERIFY pass (2026-08-21, from a real corruption). A clone delivered trit.bin with the right SIZE but corrupt content — "acked every sector, finished with 'K'" certifies transport, not bytes — and the corrupt program wild-jumped the machine to the monitor while the identical image ran perfectly in the emulator. A re-clone fixed it. Fix: per-sector checksum in the protocol, or a read-back verify pass after the clone (loader-side CRC of the whole image vs host). Until then: a board program that crashes impossibly while emulator-clean is PRESUMED CORRUPT — re-clone before debugging logic.

  • [ ] md: a panel (Tier B) renderer (2026-08-28). The console md command's parser, re-targeted at the 480x272 panel via GTEXT: size-2 colour headings, green code, 80x34 grid, keypress paging. ~200 extra lines; GTEXT paints a page in a second or two, fine for reading. Gets genuinely good after stage 10h vector text (proportional sizes). The parser is already structured for a second back end (esc()/nl()/spaces() are the only output paths).
  • [ ] p8cc: block-local declarations in NESTED blocks miscompile (2026-08-28, found building md). Locals declared in a block nested deeper than function level (e.g. char *t; char *u; inside an else-arm inside a while) silently corrupt: a branch testing those pointers took the wrong path while the identical shape at function scope worked, and the same shape in a tiny standalone program ALSO worked -- it needs surrounding function locals to collide with. Workaround (applied in md.c): declare every local at function top, C89-style. Fix: p8cc block-scope allocator; add a compiler test with nested-block locals beside live function locals.
  • [ ] Faster image transfer (2026-08-29; REGRESSION noted 2026-08-31). Moving pixels is the slowest thing the machine does: host->board rides the 115200 bridge (a full-screen P8I is ~256KB = ~22 s of line time) and on-target IMAGE draws pixel-by-pixel through the register window (563 cycles/px asm; the mandrill ~1.4 s). OBSERVED on the board: the current BASIC and C IMAGE are noticeably SLOWER than they used to be. MEASURED 2026-09-01 (emulator cycle brackets, no board needed): BASIC IMAGE of the mandrill is 39.09M cycles TODAY and 39.09M cycles at 66fe59e (pre-migration) -- byte-for-byte the same CPU cost (~596 c/px), so the CODE did not regress. The regression is ARCHITECTURAL: the card arc moved the CPU off the chip, so what used to be on-chip register writes now cross the 115200 bridge -- ~8-10 wire bytes per pixel (colour pair, x, GCMD) PLUS a GSTAT poll round-trip per pixel = tens of seconds for a full-screen image where the on-chip era took ~1.4 s. The fix is therefore exactly the rungs below -- and BLIT LANDED 2026-09-01: BASIC's IMAGE is one GL BLIT per row, the P8I bytes streamed verbatim, ~14 wire bytes + a GWAIT round trip per pixel down to 2 burst-streamed bytes. The remaining floor is the UART itself (262KB at 115200 is ~23 s however framed); the 2 Mbaud raise below is the next lever. MEASURED ON THE BOARD 2026-09-01: the mandrill via BASIC IMAGE fell 393.5 s -> 37.7 s (10.4x). The C image command still walks the device door until the C-library rung. Benchmark hook added for this: the emulator's -L LED trace is cycle-stamped now, so POKE 65282,n brackets time any code span. Candidate rungs, mostly independent: (a) raise the UART -- the BL616 USB-serial on the Tang Nano runs 2 Mbaud+; bridge DIV is one parameter on each side and protocol v1 is rate-agnostic; (b) a card-side BLIT command: set a rect, then stream raw RGB565 bytes into the span filler (the burst writer exists -- this is the P8I inner loop moved into fabric, turning IMAGE into FGETB+poke at wire speed); (c) cheap RLE in P8I v2 for flat-colour art (photos won't compress, UI will). Measure (a) first; it may make (b) moot for the SD path.
  • [ ] Restore a 16-line shell history by relocating the C commands' high scratch pages (2026-09-07). The command-history ring shrank 32 -> 16 -> 8 lines when the WM kernel was folded into the OS image: the only free block above CSTACKTOP is $F800..$F9FF (512 B), because $FA00..$FBFF is the FSDIRBUF dir/glob sector page (dir, cat, glob_expand: bios(FSDIRBUF,0,0xFA)) and $FC00..$FDFF is RDBUF, the shared file-read buffer -- both //#defines in os/commands/lib_*.c, NOT memmap anchors (which is how a 16-line ring briefly overlapped FSDIRBUF). To get 16 lines (1 KB) back: move those two pages (e.g. both down 512 B with CSTACKTOP lowered to match, costing the TPA 512 B), touch every command that names page $FA / RDBUF, update HISTN + HISTRING in gen_memmap.py and the monitor-doc memory map, then run the full suite -- and ADD the missing test: history recall after a globbing command. The OS region's spare 907 B is not 1 KB. Judged invasive for the gain on 2026-09-07; 8 lines shipped instead.
  • [ ] wdesk VIEW: cache the picture instead of re-streaming per repaint (2026-09-07). view_body() re-opens the .p8i and streams every row from disk (FGETB) on EVERY kernel repaint while VIEW is focused, exactly as desk's drawview does -- there is no framebuffer. So dragging the VIEW window (each drag step is a repaint) re-reads the whole file from CF, which is slow for a large image. Options: keep a dirty flag and redraw content only when it actually changed (not on a bare move), or decode the image once into a RAM/card scratch and BLIT from there. Cosmetic (correctness is fine); revisit if VIEW gets real use. See docs/p8x-wm-design.md rung 13.
  • [ ] PS/2 keyboard + mouse card (2026-09-04; the sketch). A TTL bus card giving the machine native human input -- and, with the LCD-as-a-terminal entry, a fully HEAD-DOWN P8X: panel, keyboard, mouse, no Mac. Philosophy: hardware receives, software understands -- the card is two dumb PS/2 receivers; scan-code decode, mouse-packet assembly and device init all live in a shipped library, the P8X way. WINDOW $FF58-$FF5F (free, adjacent to the GL port -- the human-interface corner of the I/O page): $FF58 PSADAT r: port A (keyboard) byte, ready-flag cleared on read -- raw Set-2 scan codes, no translation $FF59 PSAST r: bit0 ready, bit1 OVERRUN (byte arrived while one waited -- 1-byte holding register, the ACIA/CF precedent; PS/2 is ~1 ms/byte, the CPU laps it), bit2 parity error w: bit0 = force CLOCK low, bit1 = drive DATA low (open collector) -- the host-to-device transmit is BIT-BANGED by software, hardware does only the receive shift $FF5A PSBDAT r: port B (mouse) byte } same shape, $FF5B PSBST rw: as PSAST } second port $FF5C PSLINE r: live line states (Aclk Adat Bclk Bdat) for the bit-banged transmit's polling $FF5E PSID r: 'K' ($4B) -- presence, the single-byte GLID convention (an absent card floats $FF) RECEIVE PATH per port, ~5 TTL ICs: 74HC164 shifter clocked by the device's falling clock edges, 74HC161 bit counter to 11 (start + 8 data + parity + stop), 74HC574 latch + ready FF, 7407 open-collector drivers for the two lines. Two ports + bus decode + presence ~= a dozen through-hole ICs, one 100nF per IC (house rule). 5V logic throughout -- PS/2 is native TTL. TRANSMIT (needed once: $F4 enable-streaming to the mouse; keyboards talk unasked): software holds clock low >100 us via PSAST bit0, pulls data, releases clock, then feeds bits as the DEVICE clocks them -- pure polling against PSLINE, no timing the CPU cannot make. SOFTWARE: lib_ps2.c/.inc -- Set-2 make/break -> ASCII with shift/caps state, 3-byte mouse packets -> (dx, dy, buttons), the $F4 init dance; then the pointer abstraction above, with paint's event loop as the first client and the OS console (CONIN from the keyboard port when present) as the second -- that is the standalone-machine door. IRQ is optional sugar ($FF06 convention) -- polling suffices at PS/2 rates. EMULATOR: model the window (PSID 'K', script-fed FIFOs) so lib_ps2 and its tests run before any solder melts -- the golden-model discipline, as ever. PACKAGING -- three homes, weighed 2026-09-16: 1. Standalone TTL backplane card -- RECOMMENDED for the eventual TTL machine. The receive path is ~12 ICs (a full Eurocard), too much to graft onto the working IO card without a risky re-place/re-route, and the 8-slot backplane has ~2 free slots (6 cards + the planned IRQ card). One card per function. 2. Fold into the IO card -- only argument is slot economy, which we do not need; rejected (respin risk + no room for a dozen ICs beside the two ACIAs + CF buffers). 3. On the GRAPHICS FPGA CARD -- the FASTEST path to a working mouse + keyboard, because that card and its bridge already exist. Two PS/2 receivers in fabric (a shift register clocked by the device clock + a small FIFO -- the RTL mirror of the 164/161/574) on spare GPIO, level-shifted 3.3V<->5V (PS/2 is 5V open-collector). Fit is the caveat: the card sat at 18,537/20,736 LUT4 (89%) near the ~19,150 placement cliff, so two receivers (~a few hundred LUT4 + FIFOs) must be measured and may cost a GL feature. Payoff: the runcard daily-driver (CPU in the emulator, card = display) becomes a COMPLETE head-down machine on hardware we already have -- display + mouse + keyboard on one card. WIRING/PARTS (option 3, 2026-09-16): the breadboard interface is designed -- two TXS0102 auto-direction level translators (one per port, 5V<->3.3V, internal pull-ups, no externals), the four lines to free Nano GPIO (example 76/75/74/73). Full pinout/BOM/.cst/open-drain rule in fpga/tang-nano-20k/PS2-INTERFACE.md (+ .pdf diagram). Hardware not wired yet (a few days out); software continues on the emulator/lib_ps2 path meanwhile. BRIDGE REVERSE CHANNEL (option 3 / the card personality): the bridge is today mostly emulator->card (GL writes) + the PING. Add card->emulator EVENT packets -- the fabric receiver captures a PS/2 byte and ships it back tagged port A/B; the emulator drops it into the matching $FF58 FIFO. The SAME $FF58 model then serves the lcd personality (on-chip register read straight off the bus) AND the card personality (fed over the wire), and lib_ps2 is written once against the emulator regardless of where the bytes originate. MOUSE -> POINTER: no app change -- Finder is ALREADY pointer-driven via lib_ptr (ptr_ev/ptr_x/ptr_y, click + right-click menu; d9f8280). lib_ps2 decodes the 3-byte mouse packet to (dx,dy,buttons) and feeds the same pointer abstraction. Decide: grow lib_ptr a second backend (PS/2 packets vs console SGR), or have lib_ps2 synthesise the same events. Either way the app layer is untouched -- this is the native replacement for the backed-out serial-mouse shim (which died on adapter power, not on the software). EMULATOR PLAN (golden-model-first, the concrete first step): a. Add the $FF58-$FF5F window to p8xemu: PSADAT/PSAST/PSBDAT/ PSBST/PSLINE/PSID('K'), two device FIFOs, the ready + overrun + parity bits, and the write-side CLOCK/DATA drive bits for the bit-banged transmit. b. Feed the FIFOs three ways, all landing in the same window: - SCRIPTED (tests): -ps2a/-ps2b raw-byte files (or a console escape) -- the minimum for lib_ps2 unit tests: Set-2 make/break, the $F4 mouse init dance, 3-byte assembly, overrun. - HOST-TERMINAL (interactively "attach" devices): translate the emulator TTY's keystrokes -> Set-2 scan codes -> port A, and its xterm mouse SGR -> PS/2 3-byte packets -> port B. Drives the EMULATED PS/2 ports from the Mac's own keyboard/mouse -- the inverse of the old shim, exercising the REAL native path. - CARD-BRIDGED: the reverse-channel bytes above, once the fabric receiver exists. c. Tests: c_ps2_kbd / c_ps2_mouse (scripted), then Finder driven by the host-terminal mouse through lib_ps2 instead of lib_ptr/SGR.
  • [x] LCD as a terminal (2026-08-29) — SHIPPED as the text overlay (2026-09-15). Both missing pieces landed: the console state machine (the glass TTY, always-on 2026-09-10) and SCROLL. Scroll took a fourth option beyond the (a)/(b)/(c) below — a char-gen overlay plane (gtxt.v): the console is a grid of ASCII cells composited over the bitmap at scanout, and TXSCR scrolls the cells (a char-RAM move on the card), so no scanout base-offset register and no fabric blitter were needed. See BACKLOG-DONE "Glass-TTY text overlay". Original notes kept for the record: mirror BIOS CONOUT to the display so the machine is usable head-down, serial only for file transfer. Scroll options once weighed: (a) redraw the whole screen from a line ring (TEXT is fast enough for a demo, not for dir spam); (b) a scanout base-offset register -- vertical scroll becomes one register write, the classic terminal trick, ~30 LUT in sdram_video + a wrap rule; (c) a fabric copy-rect (a real blitter rung, also what image GRAB wants). Keyboard stays the serial RX. Fits the FPGA-CPU era (idea 2): CPU and console on one board.
  • [ ] Single-interface card: what remains after the BASIC migration (2026-08-31). DONE: the category-2 statements (LINE, BOX, CIRCLE, CLS, PIXELW) emit GL -- chosen WINDOW-space semantics (y up), full-screen window established by BASIC at cold start and after the native RESETF (the raw port's is DEGENERATE), PRMFIL shadowed at PRMSH so BOX/CIRCLE restore it, everything records inside CLBEG/CLEND, bridge-proven (BURST frames on the wire). BASIC's device door is down to the DMA gap: PIXELR() read, IMAGE write, GTEXT's rasterizer. REMAINING for the single-interface end state: (a) the C gfx library (lib_gfx C+asm twins) still drives $FF20 -- same migration decision there; (b) a GL pixel-read verb (IMAGER/PIXRD through the RB FIFO, ~50-80 LUT); (c) the card-side blit (the faster-image- transfer item) for IMAGE/GTEXT-class raw speed; then (d) the $FF20 window can close for ~100-150 LUT net of the read verb.
  • [ ] Burn the font into the card (successor board; 2026-09-01). Today the OS streams /FONT.GL to the glyph bank at boot (FONTLD) -- zero fabric cost, user-replaceable font, but the card has no text STANDALONE (bare monitor, or a different host machine driving the card). The PGC-authentic alternative: a BSRAM ROM holding the font stream plus a power-up init FSM that replays it into the glyph slots (SDRAM powers up undefined, so the bitstream cannot hold the glyphs directly). BSRAM is abundant (41/46 blocks free); the copier FSM is LUTs we do NOT have at today's ~19,150 placement cliff -- hence successor board. Keep /FONT.GL as the OVERRIDE path either way (a file swap = a new typeface; TDEFIN = custom glyphs).
  • [ ] Restore AREAPT, ARC/SECTOR and CLMOD on a successor board (removals decided 2026-08-30/09-01). CLMOD (opcode 78, the one-byte in-place list patch) went 2026-09-01: its measured 342 LUT4 funded BLIT at the placement cliff; zero ecosystem users beyond its own test, and CLRD-out + re-record is the workaround. The original entry: Two cuts bought placement headroom against the chip's PRACTICAL cliff (~19,150-19,250 LUT4, well under the nominal 20,736 -- see STAGE10-DESIGN.md "round four"): AREAPT (opcode E7, the patterned fill mask; -569 measured) and ARC/SECTOR (3C/3D, the 4-degree polyline walk + fan fill). CIRCLE/ELIPSE, LINPAT and the trig ROM stayed. The as-built designs are fully documented in STAGE10-DESIGN.md and in git history at a518415..HEAD; re-adding either is a revert plus keyword regeneration. Until then: arcs = short DRAW chains at 4-degree steps, patterned fills = software span masks.
  • [ ] GETLN drops input past 63 chars SILENTLY (2026-08-26, found via a long gl one-liner). LINEBUF is 64 bytes; GETLN just stops storing (and echoing) at 63 — no beep, no error, no truncation marker. The dropped tail cost an afternoon: gl ... CLRUN 2 lost its 2, which left the GL decoder legally WAITING for CLRUN's parameter (no error — a partial command just waits), and the next gl invocation's stale-error drain silently consumed the downstream evidence. Fix candidates: BEL on the dropped char (one JSR), a bigger LINEBUF (the history ring pairs 64-byte slots — HISTLEN moves with it), or both. Until then: long GL content goes in a file (gl FILE.GL streams it), never a one-liner.
  • [ ] GL power-up viewport is degenerate; the PGC's is full-screen (2026-08-26, found replaying the manual's HOUSE example). par[17..20] power up 0,0,0,0, so a faithful PGC stream that sets WINDOW but never VWPORT (legal on the PGC — its power-on viewport is the whole screen) maps every vertex into one pixel. Workaround: lead with VWPORT 0 479 0 271. Real fix to consider: power-up + RESETF default par[17..20] = 0,479,0,271 in emulator + RTL (small, and it is what the manual's own examples assume). Check no test relies on the degenerate default before changing it.
  • [ ] cube.bin is 161 bytes below the C-stack top (2026-08-21). Stage-9 library growth pushed cube.bin (36,191 B from $6A00) to $F79F against CSTACKTOP $F800 — a deep call chain will collide. Options: shrink the E3MAX pool (512 records is generous for a demo), split lib_g3d so LINE-only clients skip the TRI machinery, or the long-standing p8cc codegen shrink. Any new g3d client must check its map.
  • [ ] 800x480 panel support (5"/7") — a stage, not a flag (2026-08-21). DE-only panels cannot be auto-detected (one-way interface, no EDID): selection = a strap pin read at config, or an SD config byte the monitor reads at boot. Both timing sets in one bitstream; software asks IDENT, as designed. The real work: ~33 MHz pixel clock (PLL; the 27/3 CPU symmetry breaks), stride 1024 -> 2048 (a line spans two SDRAM rows; scanout pays 2 activations -- headroom exists), 768 KB framebuffer pages (page-flip bit moves), ~4x scanout bandwidth. Own design doc when a panel is actually in hand.
  • [ ] Stage 9 candidates for the geometry engine (2026-08-20). The 8b engine (STAGE8B-DESIGN.md) deliberately left rungs: colour per edge (the engine draws white-only; a per-edge or per-list pen), a list-base register / multiple lists (one fixed list at $100000 today), indexed meshes (shared vertices instead of 12 bytes per edge), camera helpers (sin/cos stays software by design — but a matrix-compose helper could live in lib_g3d), and BASIC statements over the engine (needs arrays or a statement-level world builder). Filled faces / hidden lines are a different algorithm class — their own design first.
  • [ ] BASIC could use the MDU (2026-08-20). Interpreter multiply/divide still runs the software loops; routing them through $FF30 (with the probe-and-fallback idiom from lib_g3d) would speed every arithmetic program. Measure first: interpreter overhead may dominate the way p8cc poke overhead did.
  • [ ] Board successor scouting (2026-08-20). If the BSRAM wall (42/46) arrives: ULX3S (ECP5-85F, SDR SDRAM, mature open flow) for continuity, or Colorlight i5/i9 + hand-built carrier for the hardware route; avoid DDR3-only boards (obsoletes the proven SDR controller). The emulator- as-golden-model discipline makes a port mostly pinout + video backend.
  • [ ] Emulator bus server — run one script against the card AND the reference model (2026-07-22). The bus test card's design doc (§1.2) makes the emulator the reference model, and the firmware deliberately uses the microcode's own field names (DOE, DLD, PSEL, ALUS) so a script means the same thing on both sides. But nothing in emulator/ or tools/ speaks the card's ASCII protocol, so that equivalence is a claim, not a test. A server that accepts the same w FIELD=VAL / step / r lines and drives the emulated machine would let a single script run against silicon and model and diff the answers — turning "the card behaves correctly" from a judgement call into a check. This is what would make the card trustworthy, and it is a host-side program, so it can be built before any board exists.

  • [ ] Shell-side glob + an argv ABI (2026-07-16). The biggest structural fix available, and the one real architectural drift from Unix: the shell does not expand globs — the commands do. cat *.LOG hands cat the literal string "*.LOG" and cat expands it itself, carrying lib_glob + lib_globx (~150 lines) into its binary; cat, cp, mv and lib_stdin each call glob_expand independently. In Unix the shell expands and passes argv, and the command never sees a *. Blocker: the program-arg ABI is a raw string (P2 = arg tail), so the shell could not pass expanded matches even if it wanted to. Fixing it means an argv ABI — a breaking change across all 25 commands x 2 twins. Payoff: lib_glob/lib_globx leave EVERY command binary; one glob implementation instead of four; the per-command pattern-buffer overflow class (e.g. dir's gpat[16], Wave 2) disappears. Trigger: do this when the TPA size ceiling actually bites. That is the real motivation — vi.c compiles to 32.7 KB of the 38 KB TPA, and grep sits close enough that a single //#use pushed its SOURCE over p8cc.c's buffer (b19ae24). Until something misses the TPA, the cost/benefit does not clear. Pairs with SYS_OPEN in NEXT: together they are the whole "stop making commands resolve paths and expand globs" thesis; SYS_OPEN is the cheap half and needs no ABI change, this is the expensive half.

  • [ ] CODE_REVIEW.md — the remaining findings (2026-07-16). All 6 HIGH-severity items were re-checked on 2026-08-12 and are already fixed; see the status block at the top of the file. The medium/low items below remain unverified, and nothing under fpga/ has ever been reviewed. ~417 items across 68 files from the fresh-eyes review. The 6 high-severity are all fixed; two mechanical/logic sweeps landed (59c46f7 Wave 1, bbad538 Wave 2) plus the hot-path bounds (ac6f414) and BASIC STEP (f0888da). What is left is mostly efficiency and docs findings, and a tail of medium ones. The review is PLAUSIBLE CLAIMS, NOT VERIFIED DEFECTS. Measured on the two sweeps: 215 claims rejected vs 85 applied — more than twice as many wrong as right. Several were already fixed but still listed; several are right about the abstract rule and wrong as an action here (acting on the signed-compare finding shipped a buffer overflow the 87-test suite passed — see WONT-DO). So this is NOT a checklist to grind: verify each item against current code, and treat rejecting one as a success. What worked: fan out one agent per command owning BOTH twins, scoped to a named category, told explicitly the review may be wrong. Keep the hot paths (p8xasm/p8xcc/p8xos/p8xmon/p8xbasic/p8xedit/p8cc/p8lib) OUT of any fan-out — a confident-but-wrong flag/carry edit there breaks everything. Given the hit rate, "finish the review" is probably not worth doing as a project; mine it for the real bugs when touching a file anyway.

  • [ ] ln command — symbolic links (2026-07-12). Wanted: ln /bin/dir /bin/ls so ls runs dir (command aliasing, and general path aliasing). Design decided after analysis: implement as a symlink, NOT a hard link. The P8XFS v2 entry (name12|startLBA4|len4|load2|exec2|flag1|spare7, flags $00/$01/$02/$FF) has no reference counting, and PACK relocates extents by rewriting the single owning directory entry — so a hard link (a 2nd entry sharing an extent) dangles the moment either name is deleted or the volume is packed. Making hard links safe = adding per-extent refcounts to the on-disk format + delete/PACK changes (a format change; rejected). Symlink plan: a new flag $03 whose file content is the target path string. Work: (1) firmware FRESOLVE detects a $03 leaf, reads its stored path, and re-resolves it (absolute, or relative to the link's directory) behind a depth counter (≤8) to break cycles — decide whether FFIND/FOPEN follow; this is the boot-critical, highest-risk piece. (2) tools/p8xfs.py: recognize $03, ls prints ls -> dir, optional host ln subcommand. (3) OS DIR shows -> target; RUN follows for free (loads via FOPEN). (4) new ln command C and asm twins (reuse FWOPEN/FPUTB/FCLOSE, set flag $03), man page, os_ln_test, /src tree + an ln target in the /src/commands/* Makefiles. NB: cp /bin/dir /bin/ls already aliases a command today, safely, at the cost of a duplicate binary — so ln is convenience, not a capability gap.

  • [ ] Move tools/clib.py -> compiler/clib.py. clib.py is a C-toolchain preprocessing pass (the //#use lib_*.c splicer) — conceptually a sibling of p8cc.py/p8cc.c and the prototype of the future native CPP pass, not a disk/ROM utility like the rest of tools/. Grouping it under compiler/ makes the toolchain legible. The lib_*.c files STAY in os/commands/: they're command-specific helpers and clib resolves //#use NAME to lib_NAME.c relative to the source dir. Low-risk but mechanical — update the $ROOT/tools/clib.py refs in os/run.sh and the ~8 c_*/os_* test scripts, plus doc mentions, then re-run the suite. Deferred (2026-06-26).

  • [ ] ISA additions to shrink program size (2026-06-26). p8cc codegen is bulky partly because the ISA lacks ops the compiler emits constantly. The approach: histogram the generated asm, find the most frequent multi-instruction idioms, collapse each to one opcode — but only where it shrinks real programs (each opcode costs microcode/emulator/assembler + a genucode.OPC entry). Complements the codegen-improvement and asm-rewrite items — an ISA-level win that helps every compiled program at once.

    Compiler-only wins DONE first (2026-09-11, p8cc.py, full suite green): before touching the ISA, the two biggest codegen idioms were fixed in the compiler alone — (a) leaf operands load straight into __t (gen_operands/gen_leaf_t: a constant/global/scalar-local on one side of a binary op no longer costs a PHW/PLW spill pair), (b) conditions branch directly on compare flags (gen_cond + __cmp16: one JC/JNC/JZ/JNZ per relation in if/while/for/&&/||/!, no 0/1 materialised then re-tested; global scalars tested in place). Measured with the new tools/p8cc_sizes.sh over ALL 45 /bin C commands: 627,172 → 532,728 bytes (−94,444, −15.1%), best del/awk −23.6%, finder −18%. Compares stay UNSIGNED on purpose (see docs/memory). Remaining: port both levers to the self-hosting compilers (compiler/p8cc.c, apps/p8xcc.asm) so on-target cc builds shrink too; the ISA sketch that follows is in docs/p8x-isa-c-extensions.md.

    Tier A ISA DONE on the emulator (2026-09-11, 24 pure-microcode opcodes, 88 → 112 defined): LDPn #imm16 ($38–$3A, a real 3-byte op in BOTH assemblers — every LDPn site in the monitor/OS/apps shrank a byte), ADDP3/SUBP3 #imm8, LDA/STA (Pn+d), LDW a,(Pn+d)/STW (Pn+d),a, LDW a,#imm8/#imm16, ADDW/SUBW/CMPW a,b, INCW/DECW a. Carry propagation runs through the condition planes (ALU step latches C → next step routes it → C=0/C=1 plane pair). Host assembler: new shapes, imm8/imm16 chosen from the operand TEXT (pass-stable); byte stream = address word first, then disp/imm, for both LDW and STW. Disassembler (C + asm twin) decodes all shapes (codes 10–21); test_isa.asm C1–D1 prove each op on its carry-plane case; ISA card + programmer's guide regenerated.

    p8cc emitters for Tier A DONE (2026-09-11): LDW a,#n for every 16-bit constant / string / global address; statement-level g = g ± k on a global word → INCW/DECW/ADDW/SUBW in place; orderings of a global word vs a leaf in a condition → CMPW g,__t + one branch; LPW1 for the bios()/puts() pointer setup; LDW __t,#k ; ADDW __ax,__t for member offsets, -e, ~e and the post-call argument drop. Measured over all 45 /bin C commands: 532,728 → 428,320 bytes (−104,408, −19.6%); from the pre-campaign 627,172 that is −31.7%; finder 32,630 → 16,311 (−50%). Frames on P3 DONE (2026-09-11): locals/args moved from the software C-stack onto the hardware stack — SUBP3 #L prologue, LDW/STW/LEAW (P3+d) for every local (scalars laid out first so they stay inside the 255-byte window; a far path for bigger displacements), args PHW'd and dropped with ADDP3, the whole __push/__enter/__entf/__leave/ __lea/__ldw/__ldb/__stw/__stb runtime deleted. The program saves the OS's P3 in __sp0, runs on P3 = CSTACKTOP-1 (where the C-stack was) and restores it with the new LPW3. Char scalars keep a zero-high 2-byte slot (stores write it; char params zeroed at entry) so they load with one LDW. Microcode added for it: ADDW/SUBW/CMPW a,#imm8 ($A0–$A2), LEAW a,(Pn+d) ($A4–$A6), LPW3 ($79), and PHW now pushes hi-first so a pushed word is little-endian on the stack (PHW/PLW only ever used as pairs). Startup relocates P3 to CSTACKTOP-1 ONLY when the inherited P3 is above CSTACKTOP; a nested launch (shell running a script on a C program's stack, Finder's auto-return) keeps its P3 — relocating up trampled the shell's return addresses. 428,320 → 374,672 (−12.5%); 627,172 → 374,672 = −40.3% overall; finder 32,630 → 13,909 (−57%). (2026-09-12: the self-hosting compilers now emit Tier A too — p8cc.c rewritten, p8xcc.asm's templates ported, native asm parses the shapes; only the TTL EPROM reburn remains from the list below.) Still open (historical): the self-hosting compilers (p8cc.c, p8xcc.asm) emitting PARKED (2026-09-11): monitor + OS rewrite for the new ISA. The hand-written monitor/OS/apps were only RE-ASSEMBLED for Tier A (gain: the 3-byte LDPn, ~156 bytes). A measured idiom count shows the easy substitutions are small (monitor ~50 sites / 150–250 B, OS ~67 sites / 250–350 B, p8xcc.asm ~96 sites; LPWn/INCW/ADDW/MOVW/LDW candidates) because hand asm never built the compiler's costly idioms — and each site needs reading, since the Tier A memory ops clobber A + flags unlike the byte sequences they replace. Direction: consider SCRATCH rewrites designed around the new ISA (frames on P3, 16-bit word ops, displacement addressing), starting with the easy replacements; the OS is the one that matters (16 KB ceiling). Revisit after the software-only compiler list. 2026-09-12: the os-rewrite branch (in-place rewrite with -ref copies, step 0 = native assembler two-operand/(Pn+d) shapes, WM kernel pass 1: 14,681 → 13,965 B) was DROPPED — the rewrite will restart from scratch. Its tip is kept as the tag archive/os-rewrite-2026-09-11; the native-assembler step 0 (commit 3e0e3c8 there) is still worth cherry-picking when the on-target toolchain is next touched. Tier A / the P3 frame model (they still use the software C-stack and work, but their output is ~40% larger), the native assembler parsing the compiler-only shapes, EPROM reburn for the TTL build.

    Relative branches DONE (2026-09-11): JMP/BZ/BNZ/BCP/JNC/BLT/BGE/BLE/ BGT rel8 ($A8–$B0, 128 opcodes), 2 bytes, signed d8 from the next instruction; the taken path pushes A and saves FLAGS in T2, then restores both (A is the ALU's only input — the first cut clobbered it and broke LDA #0 / JNC skip / LDA #1 in __add; caught by c_disasm), so they are drop-in for the absolute forms (14 steps; 2 when not taken). Assembler: .relax directive + iterative shrink-only relaxation (converges because sizes only shrink) and a forced MNEMONIC.R form; the compiler writes .relax at the top of its output, hand sources stay byte-identical with the native assembler (verified). Disassembler (C + asm twin) prints the resolved target (shape 22). 374,672 → 369,209 (−1.5%; only branches within ±127 bytes shrink, about half). 627,172 → 369,209 = −41.1% overall. Narrow values + peephole DONE (2026-09-11): is_narrow (char loads, byte constants) → gen_byte_a straight into A for putchar, bios()'s A operand, byte stores (through pointers and to char variables), truth tests and 8-bit CMP compares in conditions (same C/Z sense as __cmp16 on zero-extended values); __b scratch byte when the source needs the address scratch. Peephole: drop reload-after-store (STW/LDW, MOVW, STA/LDA unless a branch follows), jump-to-next-line, LDW #k+LDA __ax → LDA #k. 369,209 → 341,137 (−7.6%; the estimate was ~3-5%). 627,172 → 341,137 = −45.6% overall. Inline word ops + flag conditions + dead functions DONE (2026-09-11): the arithmetic helpers are retired — + - & | ^ are ADDW/SUBW/ANDW/ORW/ XORW on __ax (immediate right side, pointer scale folded, ±1 = INCW/DECW, leaf into __t, k - x parks x with MOVW), -x/~x are XORW #65535 (+ INCW). Twelve more pure-microcode opcodes for it: ADDW/SUBW/CMPW a,#imm16 ($B1–$B3), ANDW/ORW/XORW a,b / a,# / a,#w ($B4–$BC); every immediate form has a full 16-bit Z (0/1 marker of the low byte's Z in T2, re-latched via the Z plane when the high byte is 0; 14 steps — the a,b forms have no room and stay high-byte-only). So a condition is one CMPW + one branch: orderings normalised to C = (L>=R) (a>b is b<a, k=k+1), x == k → CMPW #k ; JZ, if (x & m) on ANDW's Z, if (x) → CMPW #0, globals compared in place; only var==var keeps __cmp16. Relops/!/&&/|| as values materialise 0/1 via gen_cond. Functions main() never reaches are not compiled (dead //#use library code). 341,137 → 293,890 (−13.9%); 627,172 → 293,890 = −53.1% overall; finder 9,508. Speed: an int add/compare is one 14-step instruction instead of a JSR into a 15-instruction loop. 140 opcodes. Tests: test_isa E1–E5, c_compile WORD-OK + dead-function check, c_disasm decodes ANDW a,# / ADDW a,#w. PHW (Pn+d) + first argument in __ax DONE (2026-09-11): PHW (Pn+d) $BD–$BF (2 bytes, 10 steps, pushes hi-first like PHW a; d measured before the push; A!) — push_arg pushes a scalar local/param straight from its slot and a global word with PHW label; argument 0 is evaluated last into __ax and never pushed (one-arg calls: no push, no ADDP3); the callee stores __ax into the slot at P3+1 in its prologue (STW (P3+1),__ax), skipped when the body never names the parameter; params 1.. sit above the return address at L+3+2(i-1). 293,890 → 285,072 (−3.0%); 627,172 → 285,072 = −54.5% overall; finder 8,903; 143 opcodes. Tests: test_isa E6, c_compile ARG-OK, c_disasm decodes PHW (P3+d). Microstep audit of control flow DONE (2026-09-11, minimal testing by choice: test-isa + test-quick, no full suite): two datapath facts shorten every absolute transfer — a pointer can be loaded FROM THE BYTE IT ADDRESSES in one step (doe=MEM, dld=PTRH, psel=0 reads mem[P0] into P0.hi; same timing as the fetch), and a pop's SP++ ; read merges into one post-increment read. JMP a 4→3 steps, absolute Jcc 4→3 taken / 3→2 not taken, JSR a 12→9 (return pushed through T2 while the target low byte waits in T; P0 steps back onto the high operand byte and loads P0.hi from it), RTS 6→5, RTI 9→7, PLW 11→9, PHW 9→8, PHW (Pn+d) 10→9. JSR (P1) (8) and the relative branches have no slack without new hardware. test_isa 4,615 → 4,382 cycles (−5%); OS boot to a program's first instruction −5.3%. Finding — the relative branches are the speed cost in compiled code: a taken Jcc rel8 is 14 steps (push A, save FLAGS, sign-extend, add, carry, restore) against 3 for the absolute form, so .relax trades 1 byte for 11 cycles on every taken branch. A compiled benchmark (poke(65282,n) LED stamps under p8xemu -L; fib(12) + 200 string sums) ran only −0.4% faster from the audit but −6.5% with .relax removed (+8 bytes of 287). Options, to decide: (a) compiler emits always-taken JMP (loop back-edges, else-skips) as absolute via a new forced .A suffix, relaxing only conditional branches; (b) ISA: drop the A/flags preservation from the taken relative path (14 → 8 steps) and free the compiler of the four idioms that rely on it (LDA #0 / JNC skip / LDA #1 → LDA #0 / ROL, branch-free __cmp16); (c) both. Also fixed here: test_isa.asm's body had grown past the IRQ vector at $0808 (the E-series), ran into the handler bytes and hit an RTI on garbage; it "passed" only because the stray RTI landed on a path that ended with A=00. The body now starts above the vector (JMP start), and all E-tests were re-verified under both the previous and the new microcode (both HALT clean with A=00). Relative-branch decision DONE (2026-09-12, both (a) and (b)): the taken relative path no longer saves/restores A and the flags (14 → 8 steps; B kept; not-taken still 2 and untouched) — an ISA contract change, documented on the card/guide/design doc; test DA now asserts B + stack + not-taken flags. The compiler's four dependent idioms went branch-free (bios() carry → LDA #0 / ROL / STA, __mul carry the same, __divmod borrow = ROL / XOR #1, __cmp16 = SUB/SUB/OR with no branch), and every unconditional jump it emits is JMP.A — a new forced-absolute assembler suffix (3 steps vs 8 taken; +1 byte). Only conditional branches stay relaxed. Benchmark workload 783,637 → 735,142 cycles (−6.2%, within 0.3% of an all-absolute build) for +3 bytes; /bin total 285,072 → 284,835 (the ROL idioms are shorter than the old skip-branches). Compiled code is now ~6% faster than before the audit and OS/monitor code ~5%. Next software-only levers: the self-hosting compilers (the on-target cc/asm still build old-ISA binaries); the OS-resident shared runtime is PARKED (see the item below).

  • [ ] OS-resident shared C runtime — PARKED (2026-09-12). After the inlining work each compiled program carries only __mul, __div/__mod/ __divmod, __shl/__shr and __cmp16, ~150 bytes, ≈6.7 KB across the 45 /bin commands (2.4%). Moving them into the OS would need a fixed entry table (programs bind by address), a version byte so a stale binary fails cleanly, and OS space (1,703 B free on graphics-card, ~2,400 on os-rewrite); speed is unchanged (a JSR into the OS costs the same). Judged the smallest lever on the list; revisit if the OS budget allows or if more helpers return (e.g. a 32-bit or signed runtime).

    Data-driven priority (measured on 5 compiled commands, 19,897 instrs): - Done — the move idioms (the big win): PHW/PLW + LPW1/LPW2 (pure microcode, −15.3%) and now MOVW (the mem→mem move, adds the PT2 scratch pointer — software done 2026-06-27, −3.6% more; hardware regen pending). See the PROTOTYPED note below and the MOVW item. 67% of instructions were LDA/STA (16-bit data shuffled byte-by-byte through A); the rev-D ops together are now ~−18–19% off compiled command size. - DROPPED — generic memory inc/dec & 16-bit ALU (INCM/DECM/ADDW). Originally floated, but arithmetic (ADD/SUB/INC/DEC) is only 1.1% of instructions — new opcodes here would shrink real programs by a fraction of a percent. Not worth it for size; would only help speed in arithmetic-heavy code, which the text-tool workload isn't. - PROMOTED — frame-relative addressing for local access (best remaining microcode-only lever). (Historical: DONE 2026-09-11 as Tier A LDW/STW/LEAW (P3+d) with frames on P3 — __csp no longer exists; the measurement below is the 2026-06 baseline.) LDA __csp appears 192× — every local-var access has the compiler compute local_addr = __csp + offset inline (LDA __csp … LDA __csp+1, then load a pointer) right after each JSR __enter. Far more frequent than arithmetic. A targeted op — LPW1 __csp,#off (load P1 = the word at __csp+immediate offset), or a general (Pn+disp) load/store mode — would collapse these. Likely pure microcode (compute base+offset through the ALU into PT, then access — the ALU is free mid-instruction; no 2nd scratch pointer). This is now the best remaining microcode-only lever — MOVW (below) is done, so frame-relative addressing is the next size win to prototype. - Hardware: MOVW+PT2 (separate item below) — the biggest single idiom, the only one needing a chip. Software done 2026-06-27 (microcode/emulator/assembler/p8cc.py); register-bank schematic regen pending.

    PROTOTYPED & MEASURED (2026-06-26): an empirical histogram of 5 compiled commands showed 67% of all instructions are LDA/STA — 16-bit data moved one byte at a time through A. Three pure-microcode 16-bit ops were added and validated (full suite green, both compilers): - PHW a / PLW a (0x74/0x75) — 16-bit push/pop of a memory word, replace the compiler's LDA/PHA/LDA/PHA & PLA/STA/PLA/STA (push_ax/pop_t). - LPW1 a / LPW2 a (0x76/0x77) — load a 16-bit pointer from a memory word, replace LDA a/TAP1L/LDA a+1/TAP1H (ax_to_p1). Read-via-PT / write-to-Pn is sequential, so no 2nd scratch pointer needed. Measured on sed/sort/grep/dir/cat/wc: PHW/PLW alone −12.1%, plus LPW1 (wired only into the central ax_to_p1 site) −15.3% total (sed −16.9%, dir −17.2%). All pure microcode: opcodes + emulator runs the regenerated u*.bin directly + assembler gets them via genucode.OPC; only genucode.py + p8cc.py + p8cc.c changed. Remaining to finish this op set: convert p8cc.c's LPW1 sites (only p8cc.py's central one is done) + the remaining inline/once-per-program LPW1 sites in both, then update the ISA docs (opcode table in docs/p8x-monitor.md, the programmer's-guide PDF via gen_progguide.py).

  • [ ] Userland text tools — the size-blocked remainder (2026-07-08, trimmed 2026-07-16). Most of this item SHIPPED — vi, find, regex +/?, and a field-oriented awk are all in /BIN with C+asm twins and man pages; see DONE. What is left is blocked on program size, not design:

    • Full awk with an expression evaluator — blocked. The shipped awk.c (242 lines) is the "field tool" this item proposed as the fits-now alternative: [/re/] { print items }, $0..$NF, NF, NR, -F c, string literals, stdin or file. It deliberately has no arithmetic, variables, or if. A real awk needs a recursive eval(), and that was measured at >64 KB — over the whole address space, ~2.3x the ~37.9 KB TPA (15,584 lines of asm from ~400 lines of C), because p8cc's non-optimizing codegen expands 16-bit ops byte-by-byte. grep is already ~32.5 KB at the ceiling, so any expression language is over budget. Unblocked by the codegen/ISA shrink work, a p8cc peephole/temp-reuse pass, or a larger TPA (overlays / bank switching). Revisit once compiled command size drops materially. (The old note here said to dodge p8cc's lack of mutual recursion by writing one self-recursive eval(minbp). That limit is gone — see below — so the parser shape is now free; only the SIZE blocks this.)
    • Regex character classes [a-z] / [^..] and \ escapes — blocked. The highest-value remaining regex feature. Even minimal buffers overflowed once variable-length atoms (atomlen/atomone) were added: it is blocked on grep's host (p8cc.c) build size, where grep's globals already collide with the $EA00 (-r) / $FA00 (glob) FNEXT pages. Unblocked by the same codegen-size work, or by restructuring grep (e.g. splitting -r content search into its own command to free grep's globals). vi's / search is literal and would also benefit. 2026-09-12: the size blocker is gone — the p8cc.c host build of grep is 12,704 bytes (p8cc.py 10,953) after the Tier A codegen rewrite, roughly half the old figure, so the classes can be attempted whenever regex is next touched.
    • find enhancements — small, in-budget, just not done. -type f|d, -name, -exec, and a path argument: find and tree are the two commands that take no path at all (FIND pattern / TREE walk the CWD only), so searching elsewhere means cd-ing there first. Note this is no longer a drive limitation — the N: prefix model was superseded by the /D1 mount, and commands are drive-unaware now, so a path arg would reach /D1 for free via FRESOLVE.

    p8cc subset gaps confirmed while writing awk (bite any ambitious command; note in the compiler docs): no break/continue — still true; restructure loops with a flag / condition (p8cc.c itself does this in register_struct).

    ~~no forward declarations or mutual recursion~~ — NO LONGER TRUE (c57da4e). Forward prototypes parse and mutual recursion works; verified on the machine with both compilers (is_even/is_odd round-trip: EVEN-OK/ODD-OK/MUTUAL-OK), and p8cc.c now relies on a forward prototype itself (int toobig(char *);). (Calls to an undeclared function still default to int, so a prototype is only needed to keep the host cc build warning-clean.)

  • [~] MOVW dst,src — 16-bit memory→memory move — SOFTWARE DONE (2026-06-27); register-bank hardware regen PENDING. Opcode $78, shape a,a (5 bytes: opcode + dst16 + src16), 12 microcode steps. Landed: microcode (genucode.py MOVW + PT2=5), emulator (P[6], PT2 init), assembler (MOVW dst,src two-operand encoder), and p8cc.py (a mov16 helper at the two hot absolute→absolute sites — assignment-yields-value and int var-load). Measured on the /BIN commands: sort −833 B, grep −1106 B, sed −931 B (~3–5% each, on top of the −15.3% from PHW/PLW/LPW). Tested: isa_wordops_test.sh gained a MOVW round-trip (mem→mem via both scratch pointers); full suite green. Docs synced (ISA card + programmer's guide PDFs regen'd → 88 opcodes, system-design §9, bus-definition, regbank theory). Two findings while doing it: 1. No backplane change — PSEL is already 3 bits (PSEL0–2 on C20/C21/C27) and U33 already decodes select 5. Confirmed against the regbank generator's exported bus set. 2. The register-bank hardware needs more than "+2 chips", and there's a latent gap: MOVW increments BOTH scratch pointers, and the already-shipped PHW/PLW/LPW also PINC PT — but the current PT is load-only 74377 latches ("PT does not count"). So PT must become four 74169 counters, PT2 is another four 74169s + a 74244 buffer pair, and the count/load decoders (U39/U40) must produce -CNT4/-CNT5 + -LDL5/-LDH5. ~10 chips, all on the regbank card, no backplane/control-card change. Documented in the regbank theory as rev-D PENDING. The remaining work is the gen_eagle.py schematic regen + BOM + placement PDFs + DRC — deliberately its own pass (canon generator, no DRC backstop; the project warns against rushing bus-facing schematic edits). 3. p8cc.c does NOT yet emit MOVW — its codegen is P1-indirect (address-through-__ax), so it lacks the absolute→absolute idioms MOVW targets; benefiting would need a codegen refactor. p8cc.py (which builds every /BIN command) gets the win; p8cc.c adoption is a follow-up. Original analysis below.

  • [ ] MOVW register-bank hardware regen (rev D). See the MOVW item above and the regbank theory rev-D note: turn PT into 74169 counters, add PT2 (PSEL=5) counters + -SEL5 buffers, extend U39/U40 count/load decode. Change the regbank netlist in gen_eagle.py (CARDS["regbank-card"]), then rebuild the KiCad board (generators/build.sh regbank-card: ERC + gate-sim + DRC) and gen_bom.py; check one-hot pointer-bus drive. No backplane/control- card change. TTL build-blocker (2026-10-02): the routed KiCad regbank is still the rev B netlist, and compiled code uses PHW/PLW (C argument pushes and pops), which PINC PT. Also teach the on-target assembler (apps/p8xasm.asm + gen_p8xopc.py) the two-operand shape if MOVW is ever to be assembled on the target (host-only today; the opcode table and cover-test currently skip two-operand shapes).

  • [ ] (orig) MOVW dst,src — 16-bit memory→memory move (needs a 2nd scratch pointer = HARDWARE) (2026-06-26). The single largest idiom from the histogram: LDA src/STA dst/LDA src+1/STA dst+1 (~3,335 sites across the 5 commands), 12 bytes that MOVW dst,src would collapse to 5. Deferred from the PHW/PLW/LPW prototype because, unlike those, it can't be done in pure microcode: a mem→mem move needs two addresses live at once (read src / write dst), but there's only one hidden scratch pointer (PT); P1/P2 belong to the running program. And the pure-microcode shortcut (LDAX/STAX on a fixed pseudo-accumulator) is impossible because __ax is a per-program label, not a fixed address microcode could name. So MOVW requires: - Hardware: a second hidden scratch pointer PT2 (PSEL=5) — one more 74169 counter pair + extend the register-bank PSEL decode (74138 U33, currently decodes 0–4) to output 5. ~2 chips; PSEL is already 3 bits so the control word needs no change. PT2 would also enable future two-address ops. Nothing's fab'd, so this is free to design now. - Emulator: widen P[] to 6 entries (P[5]=PT2); psel is already 3-bit. - Assembler: a two-operand absolute shape MOVW dst,src (the parser currently handles a single a operand). - Compiler: a movw(dst,src) helper replacing the scattered inline LDA/STA/LDA/STA mem→mem moves, mirrored in p8cc.py + p8cc.c. Microcode sketch (12 steps, fits the 16-step budget): load dst→PT2, src→PT (4 steps each via the _ld_pt pattern), then 2× (read mem[PT]→T, PT++; write T→mem[PT2], PT2++). Projected to stack on top of the −15.3% already measured, plausibly reaching the 25–40% total from the original analysis. The one item in the ISA-shrink program that needs hardware — do it as a deliberate hardware decision, after the pure-microcode ops above are banked.

  • [ ] Emulator: optional real-clock-pace mode (2026-06-26). p8xemu currently runs as fast as the host (bounded only by -l cycle cap or TTY blocking). Add a mode that throttles execution to the hardware's actual clock rate so timing, I/O pacing, and the interactive feel match the real machine — useful for sanity-checking that programs aren't relying on host speed, demoing at realistic speed, and validating any future timing-sensitive I/O. Needs the target clock frequency as a parameter and a host-time pacing loop (sleep to align cycle count to wall-clock); keep it opt-in so the test suite stays fast. Cross-check against the planned hardware clock once a board exists.

  • [ ] Hardware changes enabling "deep" (multi-step) instructions to assist the ISA (2026-06-26). The current microcoded engine sequences each instruction through control-store steps; richer instructions (block moves, deeper addressing modes, the ISA additions above) may need more microcode steps or wider control than the present sequencer/step-counter and control-word budget allow. Explore the hardware mods that would unlock longer/deeper microprograms: a wider step counter, more control-word bits (spare backplane lines exist — see SPARE12–SPARE23), or a second-level sequencer. Since no board is built yet, this is free to design now. Tie the scope to whichever ISA additions (item above) prove worthwhile, so the hardware serves real instructions rather than speculative ones.

  • [ ] Disassembler (reverse assembler): point it at an address block, get assembler back. A tool that walks a memory/file region and decodes each byte stream back into P8X mnemonics + operands — the inverse of p8xasm. Counterpart to the on-target ASM (#43); together they round-trip code on the machine (DUMP shows hex, this shows instructions). Sketch:

    • Core: a single opcode→(mnemonic, addressing-mode, length) table — ideally generated from the same source the assembler/emulator use so it can't drift (the ISA table is canon; don't hand-maintain a second copy). Linear sweep from the start address: read opcode, look up length, format the operand per its mode (#imm, $abs, $zp, relative-branch target as $abs so output re-assembles), advance, repeat.
    • Output: re-assemblable text — emit a leading .org, optional addr: bytes columns (like a listing), and resolve branch displacements to absolute targets. Round-trip test: disasm a known .bin, re-p8xasm it, assert byte-identical (the strongest correctness check).
    • Where: start as a host tool (tools/p8xdis.py, fast to iterate + easy round-trip test in make test), then optionally an on-target /BIN program (disasm.c) once the table can be shared — pairs with DUMP for on-machine reverse-engineering. Caveat: pure linear sweep mis-decodes data interleaved with code (no control-flow tracing) and can't recover labels/comments — acceptable for a v1; a later pass could follow branches to mark code vs. data.
  • [ ] Optimize monitor/OS/BASIC hot paths with the rev-C T-operand ALU ops (LDT/ADDT/SUBT/CMPT/etc.): these let you compute A := A ⟨op⟩ T without first shuffling the operand through B, so spots that currently do "save B, load operand into B, ALU, restore B" can collapse. Purely an optimization — the firmware is already correct as-is (the new ops are additive; existing code is byte-identical). Do this only after the ALU card is built and the B-mux is verified in hardware — code using the T-operand ops won't run on bare metal until the 74157 mux (U32/U33) is actually populated, so until then it would only work in the emulator.

  • The OS stdio stream model and pipes are DONE (see DONE — OS stream syscalls, program </> redirection, and |). Remaining sugar, if wanted: useful filter commands to pipe through (a MORE/WC/GREP as C programs — now writable over getchar/putchar + the BIOS), and a separate stderr stream (errors currently go to the console via the BIOS directly, which is the desired behaviour, just not a distinct syscall).

  • [~] Offload OS commands to loadable programs — TREE/DUMP/DEP DONE; FSCK and PACK remain resident. Now that the monitor publishes a shared filesystem API ($0118 FFIND / $011B FCREATE, see DONE — BASIC SAVE/LOAD), the heavy/self-contained OS commands can move OUT of the resident OS image into .COM-style programs loaded into the TPA and RUN — shrinking the kernel and freeing boot-ceiling space. Good candidates: PACK (~1 KB), FSCK (~0.5 KB), TREE, DUMP, DEP — anything that mostly needs sector/file access rather than live shell state. Two enablers: (a) widen the ROM FS API beyond flat root files to what these need — directory iteration, delete/tombstone, free- pointer read/write, ideally path resolution (or each program re-walks via FFIND); (b) a stable program ABI for args (the OS already passes a command tail). Net effect: the OS keeps only the shell, parser, path layer, and thin built-ins; everything else lives on disk and shares one ROM FS layer with BASIC and any user program. Sequence after the FS API grows those few calls; pairs with the on-target assembler/editor ideas below. Progress (2026-07-09): TREE was already a /bin command. DUMP and DEP offloaded — they need only peek/poke + a console key (no FS state), so they became os/commands/{dump,dep}.c + byte-identical hand-asm twins, run by bare name via PATH; the kernel shrank ~400 bytes. FSCK (~392 asm lines, read-only via $010C CFREAD) and PACK (~832 asm lines, filesystem-mutating via $010F CFWRITE) remain resident: both are feasible on the current BIOS (raw sector I/O is exposed) but each would be a large C + hand-asm reimplementation, and PACK is delicate enough (a bug corrupts a card) that kernel residence is defensible. Revisit FSCK next (safe, read-only); treat PACK as opt-in.

  • [~] /src/os-bios/asm — OS + monitor sources on-card (partial, 2026-07-11). The OS + BIOS-monitor asm sources now ship under /src/os-bios/asm with a bin output dir, and make os-bios (script /src/mk/os-bios) assembles both. Browsing works; blocker #2 (subdir-source read) is now FIXED, so a subdir source of any size streams correctly (the 58 KB monitor assembles on-target, standalone or under sh/make). Both original blockers are now fixed: 1. >64 KB files — FIXED (2026-07-12, branch feature/24bit-filesize). The BIOS file length is now 24-bit (FLEN/FSAV/ROREM/ROCNT/ WOTOT/FLAREM), matching ROLBA — max file 16 MB. The FS scratch block was reflowed for 3-byte length fields (all callers + sh read- stream save/restore updated), and the read AND write math widened (FSCAN/FNEXT/FOPEN/FGETB/FG_FILL/FLOADAT; FWOPEN/FPUTB/ FCLOSE/FCOM_CORE, 16-bit sector count for >255-sector files). No on-disk format change (the entry already stored 4 length bytes). Verified: 66–70 KB files read + write round-trip byte-identical on-target (os_bigfile_test), and an on-target assemble resolves a label living past the 64 KB source mark (was ?undefined). So the 121 KB os/p8xos.asm now assembles on-target — though it's slow in the emulator (~wall-time bound, minutes). PACK/FSCK are 24-bit too (OS SECCOUNT → 16-bit SECCNT:SECCH; del+pack relocates a 66 KB file byte-intact in os_bigfile_test). dir's size column and wc's counts are 24-bit too (byte-wise divmod10 in both twins). See firmware/WIDE_FILELEN.md. Nothing about 24-bit file lengths remains open. 2. Subdirectory source read empty / over-read — FIXED (2026-07-12). ROOT CAUSE (found via emulator memory-watch on FLEN/ROREM + an LBA trace, apps/p8xasm.asm): the size/sh/LINEBUF theories were all wrong. The startup sequence resolved the source path (FRESOLVE → DIRLBA=parent, FNAME=leaf), then called FFIND to confirm it exists — but FFIND ends in FRESET, reverting DIRLBA to the root. SAVESRC ran after FFIND, so it recorded the source's directory as root, not the real parent. Each pass, PASSINIT → RESTSRC restored DIRLBA=root and FOPEN→FFIND scanned root for the leaf, found nothing, left ROREM=0, and streamed a 0-byte output. A root-level source worked only because its parent is root. Under sh the empty read then over-read one sector into the adjacent script file, surfacing as ?syntax: pwd — which mislabeled the bug as size/sh-dependent. FIX: call SAVESRC before FFIND so it captures the resolved parent dir. Guard: os_asm_test.sh check (4) assembles a source in /src/os-bios/asm/ and requires byte-identical output to the same source at root.

  • [ ] Housekeeping (from 2026-06 consistency audit; not yet decided):

    • Tracked generated binaries: microcode/u0-u3.bin are committed but regenerate byte-identically from genucode.py. Consider gitignoring them and letting make build them. (Lean keep — project frames them as the canonical EPROM images, burned and interpreted.) NB the .hex question is RESOLVED: the burnable Intel HEX now lives only in rom/ (see DONE), and microcode/u?.hex were untracked + gitignored — microcode/ holds just the .bin the emulator/tests load.
    • busnet() is duplicated in gen_eagle.py, gen_bus_pdf.py, and gen_bus_card.py (kept in sync by hand; drift risk). De-dup is now feasible: gen_eagle's file-writing is gated behind EMIT = (__name__ == "__main__") (DONE), so it is importable WITHOUT side effects — the other scripts could import its busnet instead of keeping their own copies. (The import-scatters-board-files footgun itself is fixed.)
    • Smoke tests test1-3.asm overlap test_isa.asm (per-opcode). They give higher-level scenario coverage (banner, JSR/RTS, countdown); keep as complementary unless trimming.
  • [ ] Native toolchain follow-ups (EDIT + ASM landed — see DONE). Remaining polish on the on-target assembler/editor, none blocking: - Tools write to the flat root only. EDIT W and ASM output go to the P8XFS root via the BIOS FFIND/FCREATE layer, so they can't save into /BIN etc. Folds into the "make the BIOS file routines hierarchy-aware" item above — once that lands, the tools inherit paths. - ASM capacity — mostly lifted (2026-06-23). Source + output are now streamed to/from disk (bounded by the disk, not RAM); symbol table is ~850 entries. Remaining caps: 12-char names, 127-char source lines, single .org (backward .org rejected). Multiple .org would need per-region output rather than one monotonic stream. - ASM features not yet supported: .equ NAME,expr form (only NAME = expr), string escapes in .ascii (raw chars only), and macros/conditional assembly (the host has none either). - Self-host check — DONE (2026-06-23). ASM assembles its own ~37 KB source on-target to a binary byte-identical to the host build (make test-asm-selfhost). - EDIT: 8-bit line count (≤255 lines), whole-file rewrite on W (orphans sectors until PACK), no search/replace or block ops.

  • [ ] BASIC variable limits are tunable — names are significant to 6 chars (NAMLEN) and capped at 32 variables (NVARS, 8-byte entries in the 256-byte VARTAB at $x100). Both are constants in p8xbasic.asm; bump them if programs need longer names or more variables (grows the symbol table and may require nudging VARTAB/PROG placement). Also: names longer than 6 chars silently alias on their first 6 — could warn/error instead.

  • [ ] BASIC: name the line on the other runtime errors too. A runtime ?SYNTAX ERROR now reports its line (?SYNTAX ERROR IN 100) via the RUNNING flag + CURLINE. Extend the same to ?UNDEF'D LINE (run_undef) and ?RETURN WITHOUT GOSUB — both have their own handlers that print a bare message and would benefit from IN <line>. Small, mechanical: reuse the SYNERR pattern (check RUNNING, read the line number from CURLINE, PRDECU).

  • [ ] Tiny BASIC port (after Forth? Forth kernel is smaller and self-hosting)

  • [ ] Forth kernel — pointer bank makes NEXT 4 cycles; arguably the native language of this machine

  • [ ] FAT16 read-only support in P8X/OS (v3; Mac-side tool covers interchange until then)

  • [ ] RESIZE for growable directories (P8XFS v3)

  • [ ] FAT-style cluster allocation to eliminate PACK (P8XFS v3, entry format already compatible)

  • [ ] DS1302 RTC on I/O card → file timestamps. Footprints provisioned (see DONE): DS1302 (U16) + 32.768kHz crystal (X3) + coin cell (BT1) + a 3-wire breakout header (J3), all DNP. Remaining: connect the 3-wire to a CPU port (reserved $FF08 / PORT DEC U2 Y3) — jumper J3 to spare port bits or add a small latch/buffer — write the bit-banged DS1302 driver, and VERIFY the crystal + coin-cell land patterns against the real parts (placeholder THT footprints used).

  • [ ] Interrupt support — HARDWARE CONTROLLER WIRING (architecture done + footprints provisioned DNP, see DONE). The microcode/emulator/ISA side is implemented and tested (EI/DI/RTI, $08 IRQ entry, vector $0808, $FF06 raises IRQ in the emulator). The control card now carries DNP footprints U20 (74244 forcing buffer) + U21 (7474 IE/pending FF) and B29 = IRQ is a reserved bus line; the safe connections are wired (buffer inputs = $08, outputs forced high-Z, IRQ -> FF). What remains is the BUS-CRITICAL wiring, to design with DRC/breadboard before populating: - connect U20 outputs (Y1-8) onto the data bus (currently unwired) - opcode decode for EI/DI/RTI (drives the IE FF) + a fetch/step-0 detector - service sequencer so the buffer enable (!G) asserts at the injected fetch AND during the two PTR-load steps (DOE=idle) -> P0=$0808, and is off otherwise (currently !G is tied high = permanently disabled) - SUPPRESS the memory read during the injected fetch (cross-card: gate the memory card's -RD/-OE with the IRQ-service signal) so the buffer isn't fighting the EEPROM on the bus RISK: it drives the shared data bus; a wiring error = bus contention = dead machine, and there's no DRC backstop in the generator. Recommend designing it deliberately (breadboard/DRC, or a small daughtercard). Monitor needs an ORG $0808 stub (JMP to a handler / RAM trampoline) once the hardware exists.

  • [ ] p8x.pretty KiCad footprint lib if ever returning to KiCad round-trip

  • [ ] Front-panel bus-monitor LED card (passive, address + data, great demo)

  • [ ] Faster clock experiments once stable: 74F/74AHCT in critical paths, measure where it breaks


VERIFY

  • Register bank: address bus floats for PSEL = 6, 7 (2026-06 review; updated rev D). U33 (74138) is always enabled; in rev B/C only PSEL 0-4 are populated (P0-P3 + PT), and the address drivers U25/U26 are always on, so unpopulated codes drive an undefined value onto A0-15. Rev D adds PT2 at PSEL = 5 (MOVW), so 5 is now a real driver — only 6, 7 remain undriven. Safe ONLY if microcode never emits PSEL > 5 (PT2 = 5 is the max in rev D). Confirm the constraint, or add a default-select / pull so the bus can't float.

  • System-wide data-bus arbitration is one-hot (2026-06 review). Bus drivers are distributed: ALU U20 decodes DOE 1-6 (reg/ALU/flags); at DOE = 7 exactly one of memory/IO/CF should drive based on address decode. No check enforces "no DOE/address combination enables two drivers" across cards. In particular confirm the memory card is fully silent in the $FF00-$FFFF I/O page (via -IOPG) so it can't fight the I/O / CF cards on a read. (Backplane RN1 10k pull-ups hold the bus at $FF when nothing drives, so a no-driver case is defined.)

  • Control card single-step circuit (7474 one-pulse + self-clear NAND): verify one-clock-per-press behaviour at bring-up; refine debounce RC if needed.

  • I/O card SEL LED is source-driven from a gate output (deviation from the sink-drive standard) - noted on schematic; confirm brightness acceptable.

  • [ ] Final pinout confirmation against physical datasheets before fab. A knowledge-based audit was done (see DONE) and fixed the 74260; still worth eyeballing the actual datasheets for the parts you'll buy — at minimum the 74260 (odd input/output split) and the wide DIPs (74181, 28C64, 62256, 6850) — since manufacturer/variant pinouts can differ.

  • [ ] CF card 8-bit mode support — buy 2–3 candidates (SanDisk/industrial), test SET FEATURES $EF/$01 early. Fallback latch footprint provisioned DNP (see DONE): U9 (74374) with the CF high data byte D8-15 wired to its inputs, output high-Z and clock grounded. Only populate if a card refuses 8-bit mode; then wire the Q outputs onto D0-7 + a decoded read/latch-clock (design with DRC — it drives the data bus).

  • [ ] Backplane CLK at far slot on scope after bring-up → decide whether to populate RC terminators (R2/C13, R3/C14 shipped DNP)

  • [ ] PSU sizing — measure actual draw at bring-up vs the 4–5 A budget. ESTIMATE (~130 HCT chips + ~52 LEDs): HCT dynamic draw at a few MHz is a handful of mA/chip → ~1 A logic; LEDs (bus-monitor arrays via 330R + status LEDs via 1K) ~0.3–0.4 A; memory/ACIA ~0.1 A ⇒ ~1.5 A typical, ~2 A worst case — comfortable margin under 4–5 A. Confirm with a meter at bring-up.


WONT-DO / SUPERSEDED

Decisions already made and deliberately not being revisited. Each was reached with real analysis; the full reasoning is in BACKLOG-DONE.md under the entry named.

  • Do NOT convert the commands into pure stdin/stdout filters (2026-07-16). Sounds like the clean Unix answer; it is not the fix. The 25 commands split three ways and only one third could convert: ~10 pure filters (wc/head/tail/uniq/sort/sed/more/cat/awk/grep-no-r) — these ALREADY work as filters via the shell's <, > and |. Converting them gains nothing; the work is done. 2 multi-input (cmp, diff) — need two file handles; stdin gives one. Blocked without an fd model. ~12 FS manipulators (cp/mv/del/touch/mkdir/dir/tree/find/pack/fsck) — inherently need the filesystem. Routing their calls through the OS RELOCATES code, it does not remove it: dir/tree/find genuinely must iterate directories. Also note the pipe-able commands are already filters, so multi-stage pipes (a | b | c) need ZERO command changes — that blocker is a loop in the shell's splitter (PIPE_RHS re-scan), unrelated to this. The real drift is that commands resolve paths and expand globs; see SYS_OPEN (NEXT) and shell-side glob+argv (IDEAS) — those are the fix.

  • Bus test card resistor packaging policy (2026-07-22, supersedes the earlier "discrete resistors" decision). The banks were briefly all-discrete; that was reverted. The rule now, project-wide:

    • DIP-16 isolated networks (RNISO8D) for 8-way isolated banks — LED current-limiting and any other bank where both ends of each leg differ. On bustest: RN1 (probe-LED 330R), RN2 (status-LED 330R), RN3 (probe series 1k). One package per bank instead of eight parts.
    • SIP-8 / 9-pin bussed (SIP9) for pull-ups / pull-downs, where one side is a shared node. Already true of RN1/RNP on backplane/cf/io.
    • Discrete only for 1s and 2s (dividers, single pull-ups) — e.g. bustest R5–R8 (5V-sense + MISO dividers). A bussed SIP-8 CANNOT substitute for an isolated bank (its shared COM would short all eight legs), and 8 isolated resistors do not fit a 9-pin part, so the isolated banks are DIP-16, not SIP. Do not "unify" the two network types.
  • Bus test card indicators are 16 INDIVIDUAL labelled LEDs, not bar arrays (2026-07-22). LPR/LST (LEDARR8 / DIP-16) became LED1–16 (LED1–8 probes, LED9–16 status), each with a silkscreen label — the point being a bench tool you read by glancing needs a printed name at every indicator. Status meanings map to ST0–7 = Pico GP7–GP14: 5V-OK, ARMED, LISTEN, CLK, CLKB, -RES, ERR, USB-ACT. Colors are provisional (tied to the §10 status-set open item). This is another deliberate part-count increase (48 → 62); same standing as the discrete resistors above — not to be re-arrayed without a reason beating the labelled, through-hole clarity.

  • Do NOT re-add the -full.brd companion boards (2026-07-22, 3ab0ed9). Every card used to emit a second board with the auto-flow placement left ON the outline, as a "starting layout". It was never any use: the placer walks parts in dictionary order, not signal flow, so U1 sat beside U2 because of its name — nothing it produced was worth dragging into shape rather than placing from the ratsnest. They also shipped their own silkscreen collisions (regbank alone had 117 labels over a neighbouring part) which read as real defects in every audit and had to be explained away each time. The flow placement itself is still computed: it orders the parked parts and answers "do these parts fit?". It is just not emitted as if it were a layout.

  • SUPERSEDED (2026-09-18): the KiCad flow made every plug-in card, this one included, a uniform 280×140 mm (hardware/KICAD-BOARDS.md); the decision below is the Eagle-era record. Do NOT widen the bus test card past 160×100 (2026-07-22, 5576bcb). It was scoped at W=200 when it was ~45 parts. After the cuts it is 27 and auto-flow fits them in two rows at 31 % area with 10 mm of slack. The reason to stay standard is mechanical, not spatial: 200 mm cantilevered off the DIN connector is carried by just the two mounting holes at y=±45, and this is the card that gets handled most — every grabber clip and USB insertion puts a moment through the connector. At 160 mm it sits in the card guides like everything else. Fab also gets one panel size across the set. Nothing is lost: J1 still hugs the left edge with parts flowing +x, so the USB socket, probe header and LED bank stay at the OUTER (reachable) end — that came from the flow direction, not the extra width.

  • Do NOT add bus pull-downs, bus series resistors, or D0-7/A0-15 monitor LED arrays to the bus test card (2026-07-22). All three were in the first cut and all three were removed after being challenged; the card went 45 → 27 parts. Pull-downs (RPD1-6): the scenario they defended does not need them — the firmware holds CLK low from init, so nothing latches while the bus floats (design doc §3.1a). Bus series R: MCP23S17 is 5 V tolerant, so the level-shift argument does not apply; dropping it is a deliberate tradeoff recorded in §3.2 (a wrong drive is then limited only by device R_on) and is reversible with 100 Ω if it proves too sharp in use. Monitor LED arrays: 10 ICs → 7 by cutting them; the probe LEDs already cover what you actually watch. Re-adding any of these needs a NEW argument, not the original one.

  • Do NOT make p8cc's < > / % signed. It looks like a bug and is not. p8cc has no unsigned type, so the codebase uses int AS an unsigned 16-bit value for every size/offset/count and depends on the unsigned compare/divide. Acting on this (from a CODE_REVIEW finding) shipped a BUFFER OVERFLOW that the full 87-test suite passed — cmp.c counts past its 8K buffer deliberately, and a signed < made a 40000-byte file write b1[40000] into an 8192-byte array (reverted, 88ba592). The real fix is adding unsigned to the subset and migrating every size onto it — a language feature, not a bug fix. >> is separately fine to leave: every shipped use is masked ((v >> 8) & 255).

  • Do NOT chase the last ~3.4% of native-cc code size (temp reuse, peephole fusion such as MOVW __ax,V + PHW __ax -> PHW V). It needs lookahead / buffering in a single-pass emit-as-you-parse compiler, which grows cc.bin (already 22.6 KB of the ~37 KB TPA, and it must fit WHILE compiling) and risks the miscompile class the code review found. The cheap wins are all taken — wc.c is 4617 instructions vs p8cc.py's 4467. Revisit only if a real command misses the TPA by a few KB. See "C compiler — Milestone A/B".

  • Rejected paths for the on-target codegen wall (from the Milestone B analysis): sharding the existing optimizing codegen — the shared infra defeats it; a bigger flat memory region — tops out ~46 KB, still short of the ~82 KB needed, so it would require banking (major firmware/OS/hardware work). The answer was a NEW deliberately-small codegen (apps/p8xcc.asm), which exists. The earlier cpp|lex|cc1 split front end is deprecated and no longer shipped.

  • Milestone A is self-ACCEPT; p8cc.c now ALSO self-compiles on the host, but it is still NOT self-hosting (updated 2026-07-16 — the earlier wording here was overtaken by that day's fixes and every number in it is now wrong). Three different properties, kept straight:

    • self-accept — the subset accepts its own source (p8cc.py p8cc.c). This is what Milestone A built and tested. TRUE, and the c_selfhost test guards it.
    • self-compile — the p8cc.c bootstrap compiles p8cc.c and emits complete, correct asm. NOW TRUE (c57da4e fixed a forward-prototype bug that silently dropped 22 of 90 function bodies; b19ae24 raised src 32K->64K and 32 other host-side tables). 90/90 bodies, zero undefined.
    • self-host — the compiler RUNS on the P8X. FALSE and staying that way: the self-compiled output stops at the assembler with "address past 64K" because the compiled p8cc exceeds the machine's entire address space. That wall is exactly why Milestone B went the from-scratch apps/p8xcc.asm route, and that is the compiler you use on-target. So self-compilation is a correctness result — the compiler is good enough to reproduce itself — not a step toward running it on the machine. Do not chase self-hosting p8cc.c; apps/p8xcc.asm already is the native compiler.