Proton NPUCerebral ChipsEvery machine should think.
Block 02 / data-parallel execution

RVV 1.0 vector unit

Two lanes execute vector instructions from the scalar core using wide architectural registers and an AXI path to shared memory.

RVV 1.0VLEN 2,048 bitsELEN 64 bits

RISC-V vector execution

The vector unit implements the RISC-V Vector 1.0 instruction set. A vector register is much wider than a scalar register, and its storage and processing are spread across the lanes. VLEN describes architectural register capacity; the lane count describes physical execution resources.

Two-lane RVV 1.0 vector architecture
The drawing summarizes control, register storage, arithmetic and vector memory access. Shared mask, slide and reduction resources are omitted for clarity. Open full-size diagram ↗
PropertyCurrent implementation
ArchitectureRISC-V Vector 1.0
Lanes2 physical execution lanes; these are not CPU cores
Architectural vector registers32 × 2,048 bits = 8 KiB total
Maximum element widthELEN = 64 bits
Lane datapath64 bits per lane
Memory interface64-bit AXI in this two-lane system
ProgrammingLLVM RVV intrinsics; no compiler fork
Interaction with matrixShared RAM through software copies; no direct vector-register-to-matrix port

From C to vector hardware

vint8m1_t x = __riscv_vle8_v_i8m1(src, vl);
x = __riscv_vadd_vx_i8m1(x, 1, vl);
__riscv_vse8_v_i8m1(dst, x, vl);

The example loads signed bytes, adds one to each active element and writes the result to RAM. The matrix library later packs those bytes into its local A buffer. After matrix computation, a second RVV loop processes the INT32 outputs.

Contributor starting points

Use ./scripts/ara smoke for the baseline or ./scripts/ara matrix-smoke for the matrix-enabled system.