Inside the matrix engine.
Follow a packet of four numbers. Watch its sum grow. See how sixteen cells work together.
Pipeline steps are illustrative advances, not measured RTL clock cycles. The integrated RTL is verified separately; see the Proton NPU guide.
The PE array
Click a cell to inspect it.PE = processing element.
A packets travel right →. INT32 partial sums travel down ↓. Each cell has its own B packet. Hardware rows select groups of k; they are not rows of C.
Follow one output
Where the operands come from
Blue = selected A packet · Purple = selected B packetC in application RAM
Updated only when the scalar copy-back step runs.
— = not copied yet. Green = copied and checked.
Independent reference
Ordinary scalar multiplication: sum A[i,k] × B[k,j].
This reference does not use the PE pipeline.
Correctness
application outputs copied and verified
- Every valid PE pairs the same output identity from left and above.
- Completed tile values are checked against scalar multiplication.
- Previous / Next restores the complete snapshot.
Read the model: software, hardware, and what a step means
One program, three kinds of work
The scalar core stages A and transposed B with scalar memory accesses. A custom matrix command starts the array. The scalar core then copies the local C tile to RAM, where the vector unit can consume it with RVV instructions. This page animates the matrix part; it does not execute scalar or RVV instructions.
for each 4×4 output tile:
pack A and BT into local buffers
clear the 16 INT32 result values
run one K=16 matrix operation
copy the 4×4 result tile to RAMK below 16 is zero-padded. K=4 uses one arithmetic row; K=8 uses two; K=16 uses all four. Each PE sums four products, so a 4×4 result tile does not limit K to four. The four-tile example runs operations serially.
What the visualization models
Each pipeline advance reads the previous snapshot, then updates all 16 cells together. A row i enters hardware row r at step i+r+1. B packets are held fixed during the operation. Cells with no valid work are shown as idle; unspecified stale register bits are omitted.
Operand copies, weight loading, clearing, and result copying are condensed into named steps. The 10 array advances show aligned, stall-free dataflow. Real RTL adds setup, handshakes, skew/deskew logic, and memory delays. This is not an RTL waveform, measured latency, or proof that the integration works.
Local C is shown receiving each result when it exits the bottom row. The upstream output alignment and register-file write schedule are abstracted. The selected output remains highlighted through its four PE stages and the copy to RAM.
Inputs are signed INT8. Each sum uses 32-bit wraparound; these examples stay far from overflow. Displayed packed words use byte 0 in the least-significant eight bits.