The definitive reference is the BRAD_DEVICES_COMPLETE_SPEC.md compendium — 1,733 lines across the whole ecosystem, written and kept current by me. It consolidates 78 per-subsystem documents and opens today — it is the paper trail everything on this site points to.
The one book. Parts 0–10 + appendices: ecosystem blueprint, CPU, GPU, NPU, memory & storage, OS, compute/software, console & products, company, roadmap — with honest Gen 1 / Gen 2 / Vision tagging and a source-of-truth directory tree.
Every PDF now wears the brand — gold-dark rules, the official wordmark, and register-level diagrams drawn in one shared TikZ style. Download the set:
The GPU architecture — three pipelines drawn as stage diagrams.
The ISA — real bit-field figures for RRR / RI / BR / JMP.
The vector ISA v2 — 16-bit compressed layout + 256-bit VSET lanes.
BradXon Gen 1 platform — address-layout bands, MMIO, ANF fabric governor.
The full product line — brand-hierarchy tree + software portfolio.
Whitespace vs the giants — the software moat line is shipped.
Core-by-core vs the incumbent — benchmark tables, honest rows.

GPU pipeline. The six stages every BradVector core runs — drawn as numbered stage diagrams.

NPU topology. The 2×4 NCC mesh, HB-AIM, scheduler and fabric — §3.2 of the book.

Memory. BradRAM + SPMP — the two-tier unified pool, §4.1 of the book.
Every subsystem has a deep-dive. A selection:
The BradVector instruction set — the ISA every Torox GPU shares.
Torox microarchitecture, die naming and the Torox G1 die registry.
Compute model — kernels, warps, scheduling, shared memory.
The stable host-facing surface: compile → pack → session → launch → run → inspect → breakpoints. One C header, one implementation, exposed 1:1 to the browser via site/js/bradvector.js.
The software platform & strategy — why the toolchain is the moat.
SoC platforms, boards, cooling, memory modules.
Memory die, unified pool, power/security/timing deep-dives.
The operating system across Mobile–Compute editions.
The console — a Torox GPU in a box, with the Nexus controller.
The Python-first AI weapon — import tifa: one-line optimize(), AutoNeuroStream partition (BNV/BVX/BFX), MoE expert spooler, exact FP8/FP4/INT4 digit formats. Stdlib-only, runnable today.
The cloud-service developer surface — REST spec plus stdlib-only reference server: BradID auth, BradMail, BradHub, BradPay ledger, BradCloud, BradNotify, BradGeo.
The reference developer CLI — braddev cvt / run / dbg / gpu info / gpu top / timeline / test built on the real shipped stack (bradc → .bvbc → BVRT → BradTimeline/bradgdb) plus the SPMP/EROE/Fabric drivers. All self-checks pass; saxpy & fib run breathable kernels end-to-end.
The reconstruction engine for the Torox GPU. The hardware (Neuro-Stream units, semantic fabric channel, frame generation) is designed and unfabbed — so the honest move is to ship the software reference first and let the numbers argue. M0 is done; here is exactly what it proves and what it does not.
Real pixels, not a stub. bradsense_reconstruct() used to zero-fill its output buffer and book performance stats; it now upscales the frame you supply. Bilinear and Catmull-Rom bicubic at arbitrary ratio, driven by the quality-mode table, plus a luma-adaptive sharpen with an anti-ringing clamp and a flat-region gate so it does not amplify noise. It returns -ENODATA rather than inventing a frame.
2× super-resolution PSNR against ground truth: 39.43 dB bilinear, 39.44 dB bicubic, 39.57 dB bicubic + sharpen — all asserted in the test suite, along with byte-determinism, exact corner mapping, an untouched flat field, a crisped step edge, and a per-pixel anti-ringing bound. The gate caught a real bug: the luma mean was divided by 9 in border neighbourhoods with less coverage, sharpening every frame edge.
M1 temporal core: history buffer, motion-vector reprojection, responsive masks so disocclusions do not ghost. M2 the metrics harness that makes it a repeatable receipt. M3 expressing the temporal core as BradVector kernels — the first proof the ISA can run the pipeline the Neuro-Stream units will eventually accelerate. M4 the trained per-object semantic experts. None of these are claimed as working.
"We use AI to upscale" is the floor, not a differentiator — FSR 4, DLSS 4 and XeSS 2 are all machine-learned. The claim worth making is the object-aware semantic pipeline: geometry and object identity feeding the reconstruction, not a filter over a finished frame. M0 is the baseline that claim is measured against. The algorithm is FSR-3-derived but reimplemented from published behaviour (AMD's FidelityFX SDK is MIT), never cargo-copied and never branded FSR.