Proton NPUCerebral ChipsEvery machine should think.
Block 03 / packed dot products

INT8 matrix engine

Sixteen processing elements compute small tiles, while a C library assembles those tiles into larger matrix products.

INT8 × INT8INT32 accumulationCustom-1 commands

A small, explicit matrix engine

The engine combines open-source integer compute RTL with our tile buffers, scheduling controller, AXI target and custom command path. Each PE computes a packed dot product of four signed byte pairs and adds an INT32 partial sum.

4 by 4 packed INT8 PE mesh
Mesh rows represent K groups, not output matrix rows. The controller streams the four A rows with staggered timing and gathers the results into a separate C tile buffer. Open full-size diagram ↗
PropertyCurrent implementation
Array4 × 4 = 16 physical PEs
Per PEFour INT8 products + INT32 accumulation
OperationC[4][4] += A[4][16] × transpose(BT[4][16])
StorageA: 64 bytes; BT: 64 bytes; C: 64 bytes
Output layoutRow-major INT32, modulo 2^32
Scheduling10 mesh advances, then final capture; 12 observed issue-to-done cycles
mzeroClear all sixteen C values
mmaccAccumulate one K=16 tile; return scalar status 0
Larger dimensionsC library tiles M/N by 4 and K by 16; zero-pads incomplete tiles

Instruction and memory paths

The CPU writes operands to the uncached AXI window at 0xE0000000. A starts at offset 0x100, BT at 0x140, and C at 0x180. The engine performs no DMA. Software copies results back to ordinary RAM for subsequent scalar or RVV work.

.insn r 0x2b, 0, 0, rd, x0, x0   # mzero
.insn r 0x2b, 0, 1, rd, x0, x0   # mmacc

LLVM’s assembler already supports .insn, so the application needs no compiler modification. The hardware recognizes these two encodings; reserved encodings trap. This is a custom ISA contract.

Understand and verify the mesh

Step through a small example

The interactive PE explorer explains operand movement and partial sums. It is a teaching model, distinct from recorded RTL evidence.

Open PE explorer →

Inspect actual RTL

The matrix-wave command produces FST, a compact VCD and per-command arithmetic checks from real simulation signals.

Waveform commands ↗