Bandwidth-first hardware for local LLMs

AI boards built for
memory bandwidth.

Small computers for robots, home inference and local language models. On small hardware, token speed is set by memory bandwidth, not by TOPS, so that is where the budget goes: a 64-bit LPDDR5 bus on every host, and 3D-stacked DRAM beside the accelerator where the price allows.

Edge Tower
Live 3D · drag to spin
Edge Tower · 4 nodesFour Edge Nodes, one thermal core, one fanUp to 40 GB of VRAM; Qwen3-32B at about 13 tok/s (estimate)
Edge Node board render
Edge NodeRK3588 + RK1828 with 5 GB stacked DRAM61 tok/s on Qwen3-8B, measured
Edge Nano board render
Edge NanoA 64-bit LPDDR5 bus on a 69 x 53 mm boardFor robots and vision, about $95 BOM
Edge Tower cutaway with thermal core and copper pedestals
Edge TowerThe core is the heatsinkCopper pedestals press every hot chip onto one finned core
3D diagram of the Edge Node memory tiers
Memory hierarchyWeights live where the bandwidth isStacked DRAM beside the accelerator, about 10x the host LPDDR
Annotated top view of the Edge Node
Edge Node, topEvery part on the board, labelled160 x 100 mm, two 2.5GbE LINK ports
01 / 06
61 tok/sQwen3-8B on one Edge Node
measured on the RK1828
~3xthe 8B decode speed of a Jetson Orin Nano Super
40 GBVRAM in a four-node Edge Tower
with M.2 cards
$95Edge Nano BOM, 2 GB
10k units, Q3 2026 DRAM

The thesis

Decode is a bandwidth equation.

A language model generates one token at a time. For each token the accelerator reads every weight once, plus the KV cache built so far, at about two operations per byte. Speed follows memory bandwidth:

tokens/s ~ effective_bandwidth / (weight_bytes + kv_bytes x context)

On the RK3588, llama.cpp runs a 1 GB model at 22.6 tokens/s and a 4.7 GB model at 5.4 tokens/s. Both work out to 22 to 25 GB/s. More TOPS does not change these numbers; more bandwidth does.

DRAM prices rose sharply through 2026, to about $13 per GB. Capacity picks the model; bandwidth picks the speed.

Decode speed per $100 for 8B-class models: Cognitora Edge Node against an RK3588S board with an RK1828 card, Jetson Orin Nano Super and a four-board Raspberry Pi 5 cluster
8B-class models, 4-bit, batch 1. Cognitora is BOM; the others are list or retail prices.

Product family

One architecture, three sizes.

Every board shares the same software stack, the same memory hierarchy and the same real-time MCU for robotics I/O.

Edge Nano board with RK3588S, heatsink and expansion headers

Edge Nano

68.85 x 53.34 mm, for robots and vision
  • RK3588S, 6 TOPS NPU, 2 or 4 GB LPDDR5 on a 64-bit bus
  • Real-time MCU, 2.54 mm expansion headers, Qwiic, MIPI CSI
  • Wi-Fi 5, Bluetooth, USB-C power
  • Qwen3-1.7B at about 23 tok/s (estimate)
BOM, 10k unitsabout $95
Edge Node board

Edge Node

160 x 100 mm, for home inference
  • RK3588 host plus an on-board RK1828: 20 TOPS INT8 and 5 GB of stacked DRAM as VRAM
  • M.2 slot for a second RK1828
  • Two 2.5GbE ports to chain nodes without a switch, CAN-FD
  • Qwen3-8B at 61 tok/s (measured on the RK1828)
BOM, 10k unitsabout $305
Edge Tower: black cylindrical enclosure with a fan in the top

Edge Tower

182 mm across, 264 mm tall
  • Four Edge Nodes on one finned thermal core
  • One 160 mm fan for the whole cluster
  • Up to 40 GB of VRAM with M.2 cards
  • Qwen3-32B at about 13 tok/s (estimate)
BOM, 10k unitsabout $1,220

BOM includes assembly at 10k units with LPDDR at about $13.4 per GB. Board specifications are design targets; the whitepaper marks every figure as measured or estimated.

Edge Tower

The core is the heatsink.

Instead of a heatsink and a small fan on every chip, every hot chip in the tower presses onto one shared finned core, and one slow fan moves all the air.

Edge Tower closed, with the top fan and the air gap under the shell labelled Edge Tower cutaway: thermal core, copper pedestals and a board swung open on its standoffs

Unified thermal core

A 100 mm square aluminium extrusion with 28 radial fins runs up the middle of the tower.

Boards chip side in

Each Edge Node mounts on one face of the core on 16 mm standoffs, components facing the core.

Copper pedestals

The RK3588, the on-board RK1828 and the M.2 RK1828 press onto pedestals through 0.5 mm thermal pads.

One slow fan

A 160 mm fan pulls air in under the shell and up through the fins: about 140 W average and 284 W peak for four nodes.

Hardware status

From concept to EVT boards.

The first hardware is a carrier for a bought-in compute module, for each board. The module brings a working SoC, DRAM and accelerator. That lets the first spin test everything the carrier adds: power, I/O, the board controller, the accelerator power path, mechanics and thermals.

Edge Node EVTEdge Nano EVT
Host moduleTuring RK1 (RK3588) on a 260-pin SO-DIMMRadxa CM5 (RK3588S2), CM4-format connector
Accelerator2x DFRobot DFR1263 (RK1828) on M.2, switched 12 V per card
Board160 x 100 mm, 6 layers, 314 parts68.85 x 53.34 mm, 4 layers, two-sided
Board controllerSTM32G474: power sequencing, fan, temperature, CAN-FDSTM32L552: expansion headers, Qwiic
EVT unit costabout $994 with modules, $117 for the carrierabout $122 with module, $42 for the carrier
Done

Schematics

Generated KiCad projects with pinouts and sub-circuits from open reference designs. ERC clean, netlists verified.

Done

Layout and BOM

Placement DRC clean, first-pass routing, every part with an MPN and a distributor price, manufacturing outputs.

Next

Signal integrity, then bring-up

Hand-route the PCIe, USB 3, CSI-2 and Ethernet pairs, order the boards, bring them up, then build a four-node tower.