Proton NPUCerebral ChipsEvery machine should think.
System architecture / milestone 01

Proton NPU

An open hardware foundation for Cerebral Chips: one RISC-V CPU, two vector lanes, and an INT8 matrix engine working under one program.

RV64 + RVV 1.04 × 4 matrix meshBare-metal RTL simulation

Our first working hardware foundation. One bare-metal program executes scalar, vector and custom matrix instructions on real scalar + vector + matrix RTL. “Proton NPU” is the project name; the current architecture is a single-core RISC-V system with attached vector and matrix execution units.

1RV64 CPU core
2RVV vector lanes
16packed INT8 matrix PEs
1 ELFone architectural program

The complete system

System block diagram: RISC-V core, RVV vector unit, matrix engine and AXI
Blue routes instructions and completions. Teal carries memory traffic. The dashed amber path carries matrix commands. The matrix is an AXI target, not a memory-fetching DMA engine. Open full-size diagram ↗
01

64-bit RISC-V core

Fetches instructions, executes scalar C and retires all instructions in order.

Explore the block →
02

RVV 1.0 vector unit

Executes RVV operations across two lanes and accesses shared RAM through AXI.

Explore the block →
03

INT8 matrix engine

Computes signed INT8 tiled products using sixteen packed-dot-product PEs.

Explore the block →

Follow one workload

  1. Scalar setup: The scalar core prepares dimensions and pointers in ordinary C.
  2. Vector preparation: explicit RVV intrinsics generate input-processing instructions for the vector unit.
  3. Matrix computation: the C library packs A and transposed B through scalar MMIO, then issues mzero and mmacc.
  4. Vector finishing: results copied to RAM are processed by the vector unit.
  5. Verification: scalar C compares every result with independently generated expectations; host tools reconcile instruction and hardware traces.

What is verified

PropertyCurrent implementation
PE arithmetic262,144 signed-byte / lane checks
Tile arithmetic256 directed and randomized tile cases
Combined program8 workloads; 115,650 simulated cycles
Execution evidence69,007 scalar + 64 RVV + 52 mzero + 34 mmacc retired
Waveform arithmeticAll 86 matrix command results checked from actual RTL FST
Regression9/9 scalar/vector tests with matrix on; 9/9 with matrix off
Failure handlingInjected wrong result and forced timeout rejected

This is a recorded functional milestone. A simulation cycle is not a measured chip clock frequency, and these tests do not establish full ISA conformance or application throughput. See the committed verification summary for source and artifact hashes.

Start contributing

git clone --branch hardware https://github.com/cerebralchips/proton-npu.git
cd proton-npu
./scripts/ara vm-start  # Apple Silicon
./scripts/ara provision
./scripts/ara doctor
./scripts/ara matrix
./scripts/ara matrix-wave

Contributor map and verification gates →