Cerebral Chips / Proton NPU
An interactive hardware lesson

Inside the matrix engine.

Follow a packet of four numbers. Watch its sum grow. See how sixteen cells work together.

4 × 4 physical PEs4 signed INT8 products per PEINT32 partial sumsK = 16 capacity per operation
Integrated datapath · teaching model
Pipeline steps are illustrative advances, not measured RTL clock cycles. The integrated RTL is verified separately; see the Proton NPU guide.
Step 0 / 14Tile 0
0

The PE array

Click a cell to inspect it.
PE = processing element.
Real input positionsPadded positionsTracked result

A packets travel right →. INT32 partial sums travel down ↓. Each cell has its own B packet. Hardware rows select groups of k; they are not rows of C.

Follow one output

Where the operands come from

Blue = selected A packet · Purple = selected B packet

C in application RAM

Updated only when the scalar copy-back step runs.

— = not copied yet. Green = copied and checked.

Independent reference

Ordinary scalar multiplication: sum A[i,k] × B[k,j].

This reference does not use the PE pipeline.

Correctness

application outputs copied and verified

  • Every valid PE pairs the same output identity from left and above.
  • Completed tile values are checked against scalar multiplication.
  • Previous / Next restores the complete snapshot.
Read the model: software, hardware, and what a step means

One program, three kinds of work

The scalar core stages A and transposed B with scalar memory accesses. A custom matrix command starts the array. The scalar core then copies the local C tile to RAM, where the vector unit can consume it with RVV instructions. This page animates the matrix part; it does not execute scalar or RVV instructions.

for each 4×4 output tile:
    pack A and BT into local buffers
    clear the 16 INT32 result values
    run one K=16 matrix operation
    copy the 4×4 result tile to RAM

K below 16 is zero-padded. K=4 uses one arithmetic row; K=8 uses two; K=16 uses all four. Each PE sums four products, so a 4×4 result tile does not limit K to four. The four-tile example runs operations serially.

What the visualization models

Each pipeline advance reads the previous snapshot, then updates all 16 cells together. A row i enters hardware row r at step i+r+1. B packets are held fixed during the operation. Cells with no valid work are shown as idle; unspecified stale register bits are omitted.

Operand copies, weight loading, clearing, and result copying are condensed into named steps. The 10 array advances show aligned, stall-free dataflow. Real RTL adds setup, handshakes, skew/deskew logic, and memory delays. This is not an RTL waveform, measured latency, or proof that the integration works.

Local C is shown receiving each result when it exits the bottom row. The upstream output alignment and register-file write schedule are abstracted. The selected output remains highlighted through its four PE stages and the copy to RAM.

Inputs are signed INT8. Each sum uses 32-bit wraparound; these examples stay far from overflow. Displayed packed words use byte 0 in the least-significant eight bits.

Arithmetic and wiring references: Integer MACPE registersMeshOperation specification