Skip to content

p8cc — C cross-compiler for the P8X

A tiny C compiler that runs on the host and emits P8X assembly for assembler/p8xasm.py. Output targets the OS transient program area ($5900), so a compiled program is a RUNnable .BIN.

python3 compiler/p8cc.py prog.c -o prog.asm
python3 assembler/p8xasm.py prog.asm -o prog.bin --base 0x5900
python3 tools/p8xfs.py put disk.img prog.bin --name /PROG.BIN --load 0x5900 --exec 0x5900
# then on the P8X:  RUN /PROG.BIN

This host cross-compiler is the primary build tool — every /bin OS command is compiled with it. For compiling C on the machine, see the from-scratch native compiler apps/p8xcc.asm (/bin/cc), which does the whole compile on-target (Milestone B, achieved through v0.28 — functions, recursion, pointers, arrays, structs, a //#use splicer, and the standard builtins). (An earlier path-A front end — cpp | lex | cc1, front half on the P8X with host-side codegen — was retired in favor of cc on 2026-07-14; see git history.)

Two compilers: p8cc.py and p8cc.c

There are two implementations of the same compiler:

  • p8cc.py — the original, written in Python. The everyday tool and the reference oracle. It is the bootstrap and is never removed.
  • p8cc.c — the compiler rewritten in p8cc's own small-C subset (Milestone A, done). It is simultaneously valid standard C and valid p8cc-subset C, so it builds two ways:
cc compiler/p8cc.c -o p8cc_host        # native bootstrap: ./p8cc_host < prog.c > prog.asm
python3 compiler/p8cc.py compiler/p8cc.c -o p8cc.asm   # the self-compile proof

It reads C from stdin and writes assembly to stdout (EOF is 0 from the P8X console or -1 from host getchar). p8cc.py compiling p8cc.c cleanly is the proof that the subset is self-sufficient — "small C written in small C". Correctness is checked by a differential test (emulator/test/c_selfhost_test.sh): a sample compiled by both p8cc.c and p8cc.py runs to identical output on the P8X. (The two emit behaviourally equivalent asm — same program output — not byte-identical text; they differ in label names and argument order.)

Same ISA from both compilers (2026-09-12). p8cc.c's code generator was rewritten to emit exactly the Tier A model described below: frames on P3 with SUBP3/LDW/STW/LEAW (P3+d) and PHW pushes, ADDW/SUBW/ ANDW/ORW/XORW on the accumulator word, CMPW plus one branch for conditions, JMP.A for always-taken jumps, .relax for the rest, the same runtime helpers (__mul, __divmod, __shl/__shr, the branch-free __cmp16). Because it is single-pass — it emits while it parses — two things differ from the Python compiler: a leaf operand (constant, variable, string, array address) is deferred until its use is known and the right operand of a binary op is classified by a one-token peek, so x + 3 still becomes LDW __ax,(P3+d) / ADDW __ax,#3; and arguments are parsed left to right, so the last argument travels in __ax (the Python compiler passes the first) with the others pushed left to right — each compiler's callees match its own callers, so programs are self-consistent. Not ported: the narrow 8-bit paths beyond putchar/bios/byte stores, and dead-function elimination (a single pass cannot know reachability). Over the 45 /bin commands the C-written compiler produces 342,372 bytes against the Python compiler's 284,835 (+20%), all on the new ISA; the compiler test program passes every case under both. The same rewrite closed three old gaps in p8cc.c: brace initializer lists and string initializers for globals (int tab[] = {…}, char *s = "…", char b[8] = "…"), plain #define NAME value, and function return-type tracking (an int * result now scales pointer arithmetic), so every /bin command compiles with it.

As a fast native tool. Build p8cc.c with the host cc for a fast (~no startup) alternative to the Python tool — it is literally the C codebase compiled for the host:

cc -O2 -w compiler/p8cc.c -o p8cc-host   # a stdin->stdout filter (the in-subset
p8cc-host < prog.c > prog.asm            #   source can only use getchar/putchar)

The test suite builds it exactly this way to differentially cross-check every command against p8cc.py (see emulator/test/c_* and c_selfhost_test).

Milestone B (a C compiler that runs on the P8X and compiles its own source) is ACHIEVED — by the lean twin, not this compiler. p8cc.c here is the two-pass HOST bootstrap: it compiles to ~82 KB, larger than the machine's entire 64 KB address space, so it cannot run on-board and its back end stays on the host. The on-board self-host was reached instead by apps/cc.c — the single-pass twin of apps/p8xcc.asm, written in the subset both accept — once it got the P3-stack-frame codegen (~13% smaller output) and the +2 KB of TPA from the ROM reclaim (TPABASE $5900). It self-compiles to 30,843 bytes, assembles and runs on the machine, and reproduces itself byte-for-byte (a fixed point; emulator/test/cc_selfhost_test.sh). A back end small enough for this compiler to run on-board would still need a new, deliberately-small code generator. See BACKLOG.md.

p8cc.c is single-pass and so requires declare-before-use (function prototypes for mutual recursion, globals/structs before reference); p8cc.py is two-pass and more lenient. Both accept the same subset otherwise.

Execution model

The P8X has no 16-bit accumulator, so expression results live in a 16-bit pseudo-accumulator AX (the memory word __ax). The hardware stack (P3) holds expression temporaries (PHW/PLW) and call return addresses (JSR/RTS). + - & | ^ and every comparison are one Tier A word instruction on AX (ADDW/SUBW/ANDW/ORW/XORW/CMPW, see the last stage below); only multiply, divide, the shifts and the 16-bit equality of two variables still go through small runtime helpers (__mul, __divmod, __shl, __shr, __cmp16), and only the helpers a program actually uses are emitted. Two size levers (2026-09-11, −15.1% across /bin): a leaf operand (constant, string, global, scalar local) on one side of a binary op is loaded straight into the helper's __t input instead of being spilled through the stack, and conditions branch directly on the flags of a 16-bit compare (__cmp16, unsigned like the value helpers) rather than materialising a 0/1 and re-testing it. Measure any codegen change with sh tools/p8cc_sizes.sh.

Tier A instructions (2026-09-11, a further −19.6% across /bin; −31.7% from the pre-campaign baseline). With the Tier A opcodes in the ISA (see docs/p8x-isa-c-extensions.md) the generator emits: LDW a,#n for every 16-bit constant, string or global address (4–5 bytes, was 10 — the single biggest item); INCW/DECW/ADDW/SUBW in place for a statement-level g = g ± k on a global word (no load, helper or store; pointers scale k by the element size); CMPW g,__t for an ordering of a global word against a leaf in a condition (not ==/!=, whose Z would be high-byte-only); LPW1 for the pointer setup in bios() and puts(); LDW __t,#off ; ADDW __ax,__t for member offsets, negation, bitwise NOT and the argument drop after a call. Frames on P3 (same day, a further −12.5%; −40.3% from the pre-campaign baseline, finder −57%): the frame model above, plus ADDW/SUBW/CMPW a,#imm8 and LEAW a,(Pn+d) in the ISA for it.

Relative branches (same day, −1.5% more; −41.1% overall): the compiler's output starts with .relax, so the assembler encodes every Jcc whose target lies within ±127 bytes as the 2-byte relative opcode (shrink-only iterative relaxation, so both passes agree). Hand-written sources carry no .relax and stay byte-identical with the native assembler's output. Speed note (2026-09-12): a taken relative branch costs 8 microsteps against 3 for the absolute form (it used to cost 14, saving and restoring A and the flags; that preservation was dropped in the speed audit, so code after a taken relaxed branch must not read A or the flags — the compiler's own four such idioms became branch-free: LDA #0 / ROL turns the carry into 0/1, and __cmp16 computes its Z without branching). Every unconditional jump the compiler emits — loop back-edges, else-skips, return, the 0/1 materialisation — is always taken, so it is written JMP.A and stays the 3-byte absolute form: one byte more per jump, five cycles less every time it runs. Conditional branches stay relaxed.

Narrow values + peephole (same day, −7.6% more; −45.6% overall). A char load or a byte-sized constant is a narrow value: when its consumer only wants a byte — putchar, the A operand of bios(), a byte store (buf[i] = c, *p = c, c = ...), a truth test (while (*s)), or a compare against another narrow value — it goes straight into A from a 1–3 byte load and skips the 16-bit zero-extension, the spill and the helper: if (c == 'x') is now LDA (P3+d) / LDB #120 / CMP / JNZ. An 8-bit unsigned compare of two zero-extended bytes orders exactly like the 16-bit one, so the branch sense is unchanged. A small peephole pass then removes a reload right after the matching store, a jump to the very next line, and shortens a byte reload of a constant just written; it only ever looks at adjacent lines with no label between them.

Inline word ops, flag conditions, dead-function elimination (2026-09-11, −13.9% more; −53.1% overall: 627,172 → 293,890 bytes, finder 32,630 → 9,508). The arithmetic helpers are gone: + - & | ^ compile to ADDW/SUBW/ANDW/ORW/XORW on __ax — with an immediate when the right side is a constant (folded with the pointer scale: p + 2 on an int * is ADDW __ax,#4; x + 1 is INCW __ax), a direct __t load when it is a leaf, the stack spill only in the general case; k - x computes x first and parks it in __t with MOVW. -x is XORW __ax,#65535 ; INCW __ax, ~x the XORW alone. That needed twelve more pure-microcode opcodes: ADDW/SUBW/ CMPW a,#imm16 and ANDW/ORW/XORW in the a,b, a,#imm8, a,#imm16 shapes (140 opcodes now). Every immediate form has a full 16-bit Z (the microcode keeps a 0/1 marker of the low byte's Z in T2 and re-latches Z from it when the high byte comes out zero), so conditions are one CMPW and one branch: orderings are normalised to L < R / L >= R with C = (L >= R) (a > b is b < a; k < x is x >= k+1), x == k is CMPW __ax,#k ; JZ, if (x & m) branches straight on the ANDW's Z, a plain if (x) is CMPW __ax,#0, and a global word is compared in place (CMPW g,#k). Only x == y of two variables keeps __cmp16, because the a,b forms' Z is still high-byte only (no room in 15 microsteps). Relational and logical operators used as values (ok = a < b) materialise 0/1 through the same condition code instead of the old __lt/__eq/__not helpers. Finally, functions main never reaches through a call are not compiled at all (//#use splices whole library files, so small programs carried unused directory helpers); the output notes each dropped one. Speed: an int add or compare is now one instruction (14 microsteps) instead of a JSR into a 15-instruction loop.

PHW (Pn+d) and the first argument in AX (2026-09-11, −3.0% more; −54.5% overall: 293,890 → 285,072 bytes, finder 8,903). One more pure-microcode opcode group, PHW (Pn+d) ($BD–$BF, 2 bytes, 10 steps): push the word at Pn+d, so a local or parameter goes onto the stack straight from its slot instead of LDW __ax,(P3+d) ; PHW __ax (7 bytes); a global word is PHW label. And the first argument is no longer pushed at all: the caller evaluates it last into AX, so a one-argument call is just the value and a JSR with no push and no ADDP3. The callee stores AX into the first local slot in its prologue (STW (P3+1),__ax), or not at all when the body never names that parameter. Fewer stack round-trips per call as well as fewer bytes.

Calling convention / frames (2026-09-11: on the hardware stack). Call frames live on P3. A caller pushes arguments 1…n−1 right-to-left with PHW (each word lies little-endian at P3+1), leaves argument 0 in AX, JSRs, then drops the pushed ones with ADDP3 #2(n−1); the callee reserves its locals with SUBP3 #L and stores AX into parameter 0's slot. Everything is then a small positive displacement from P3, read and written with one LDW/STW/ LEAW (P3+d) instruction: parameter 0 at P3+1, the other locals at P3+3 … P3+L (scalars first, arrays and structs above them), the return address at P3+L+1,+2, parameter i ≥ 1 at P3+L+3+2(i−1). The compiler adds whatever it has pushed itself (spills, pending arguments) to every displacement, so a temporary on the stack never moves a local; a displacement over 255 takes a slower computed-address path. One frame per call, so functions are reentrant and recursion works. char scalars occupy a 2-byte slot whose high byte is kept zero (stores write it, char parameters are zeroed at entry), so they load with one LDW like an int. Globals keep static storage. On entry the program saves the caller's P3 in __sp0 and, if that P3 is above CSTACKTOP (a normal launch from the OS's small stack), sets P3 = CSTACKTOP-1, so frames grow down from $F800 into the free TPA exactly where the old software C-stack lived. A P3 already below CSTACKTOP means a nested launch — the shell running a script on a C program's stack, as Finder's auto-return chain does — and then P3 is kept, so the new frames grow beneath the caller's pending return addresses instead of over them. On exit LPW3 __sp0 restores P3 before the final RTS. (Before this, frames were a software C-stack __csp/ __fp in RAM with a __ldw/__stw/__entf/__leave runtime — see git history.)

Supported subset

area supported
types int (16-bit), char (8-bit), pointers T *, arrays T a[N], struct/union (nestable)
top level struct/union definitions, function definitions with parameters, global variable declarations
statements { }, declarations, if/else, while, for (e; e; e), return [e];, expr;, ;
operators = \|\| && \| ^ & == != < > <= >= << >> + - * / % unary - ! ~ & * member . ->
functions parameters, stack locals, recursion, return value in AX
pointers &lvalue, *ptr (load/store), pointer +/- scaled by element size, a[i]
primaries int / char / string literals, identifiers, calls, ( )
builtins console: getchar() putchar(e) puts(e) (OS SYS_GETC/SYS_PUTC/SYS_PUTS); memory: peek(addr) poke(addr,v); general: bios(constaddr, p1, a)

Library functions are written in C. The console builtins are thin wrappers over the OS stream syscalls ($200C/$2009/$200F), not the raw BIOS — so a program's output is redirectable by the shell: RUN PROG >FILE streams its putchar/puts to a file with no source change, and RUN PROG <FILE binds its getchar to a file (which returns -1 at end of file). Both combine — RUN CAT.BIN <IN >OUT copies a file (see os/commands/cat.c). (Directory iteration and the write stream default to the same BIOS SBUF; a program that does both — DIR — calls FSDIRBUF to move iteration onto its own buffer so it can stream output while iterating; see os/commands/dir.c.) Everything else a program needs (strlen, getline, strcmp, …) is ordinary C compiled alongside it, now that pointers, arrays, and char work. See the strlen in the test below for the pattern.

Reaching the rest of the BIOS. peek(addr)/poke(addr, v) do byte memory access (e.g. the switch/LED ports at $FF00/$FF02), and bios(addr, p1, a) calls any monitor routine: it sets P1 = p1 and A = a, JSRs the (constant) address, and returns the routine's A in the low byte with the carry flag in bit 8 (result & 256). So a status-returning call like FNEXT is usable from C — while ((bios(0x013C, 0, 0) & 256) == 0) loops until end-of-directory. The whole jump table is reachable: bios(0x0112, str, 0) is puts without the newline, and the file API (FOPEN/FGETB/FLOADAT/…) is driven by poke-ing its RAM ABI variables and bios-ing the call. (bios's address must be a literal — it becomes the JSR target.) argstr() returns the program's command tail (the RUN argument in P2) as a char *. Both compilers support all of these — see os/commands/dir.c, the OS DIR command written in C.

p8lib.c — a C-source standard library

p8lib.c is a small library written in the subset over those builtins: strlen/strcpy/strcmp, getline/putdec, and file helpers loadfile(name, dest) / savefile(name, data, len) (which drive FNORM/FFIND/FLOADAT and FWOPEN/FPUTB/FCLOSE). There is no #include or linker, so you use it by prepending it to your program:

Object-like #define. Both compilers (p8cc.py and the native cc) support #define NAME value where value is a decimal or 0x hex integer — a compile-time textual substitution (no function-like macros). This is how a command names raw BIOS/OS addresses without any code cost: //#use abi splices os/commands/lib_abi.c (#define FOPEN 0x0124, …) so the source reads bios(FOPEN, RDBUF, 0) and compiles byte-for-byte identically to the old bios(0x0124, 0xFC00, 0). (On-target, cc writes it //#define, matching its //#use directive style.)

cat compiler/p8lib.c prog.c > all.c
python3 compiler/p8cc.py all.c -o all.asm      # (or build p8cc.c with cc and run: p8cc-host < all.c > all.asm)

Caveats it documents: bios() doesn't surface the carry flag, so loadfile/savefile don't detect a missing file or a full disk; and loadfile reads whole sectors, so its destination must be sector-sized (≥ 512 bytes for any file ≤ 512). Exercised end-to-end by emulator/test/c_libfile_test.sh (a savefile→loadfile round-trip, differential across both compilers).

int is 16-bit (comparisons and / % are unsigned 16-bit); char is 8-bit. The compiler tracks types so a dereference loads/stores the right width (int/pointer = 2 bytes, char = 1) and pointer arithmetic scales by element size. Scalar locals/params occupy a 2-byte slot; arrays occupy count * elemsize. String literals are pooled and evaluate to their address.

Current limitations (next phases)

  • struct/union are used by pointer: no by-value struct parameters, returns, or whole-struct assignment (assign individual members, or pass a pointer). Union members all share offset 0; no bitfields; no sizeof() operator yet. Members are laid out with no padding (byte-addressed machine).
  • Global initializers are supported and must be compile-time constants: scalar int/char, a string for a char * or char[] (length inferable from []), and brace lists for arrays — including string tables (char *t[] = {"a","b"}). Not yet: &global address constants, nested-aggregate braces, or initialized locals beyond a scalar expression.
  • Function return types are tracked (a T *-returning call participates correctly in pointer arithmetic and dereference); a call to an undeclared function still defaults to int.
  • for-init is an expression, not a declaration: locals are function-scoped, so declare the loop variable before the loop (int i; for (i = 0; ...)).
  • Locals are function-scoped (no per-block shadowing); the C-stack and the hardware return stack are both in the TPA, so deep recursion is bounded by RAM. See the project backlog.

Testing

emulator/test/c_compile_test.sh (make test-c) compiles a C program, assembles it, RUNs it under P8X/OS, and checks the output: a while loop printing 12345, then recursive fact(5) → FACT-OK, a two-arg add → ADD-OK, and FOR-OK/LOG-OK/BIT-OK/SHIFT-OK covering for, short-circuit &&/||, bitwise & | ^ ~, and shifts << >>.

emulator/test/c_libc_test.sh exercises input: a program reads a line with getchar(), upper-cases it into a char buffer, puts() it, and prints its length using a strlen() written in C — end-to-end proof that the I/O builtins and C-source library functions work together.

emulator/test/c_struct_test.sh covers struct/union: a nested struct Rect of struct Points, ./-> access, pointer-to-struct, an array member, and a union — checking the rendered output 796A.

emulator/test/c_global_test.sh covers global initializers: a scalar, a char * string, an int[] list, a char *[] string table, and an inferred char[] — output 7HI / 6CY / YO.

emulator/test/c_selfhost_test.sh is the Milestone-A check for p8cc.c: it builds the host bootstrap with cc, confirms p8cc.py self-compiles p8cc.c, and runs a feature-spanning sample compiled by both compilers, asserting byte-identical P8X output (12345678Y120AZ5QRSTG).