How to Run LLM Models on RK3588
Convert DeepSeek-R1-Distill-Qwen-1.5B to RKLLM on an Ubuntu x86_64 PC, deploy it to RK3588 over SSH/SCP, and run the official C++ example.
Quick start: Run LLM models on RK3588
| Item | Environment used for this verification |
|---|---|
| PC | Ubuntu 20.04.6 LTS, x86_64, Python 3.11, NVIDIA GTX 1070 8 GB |
| Board | reComputer RK3588, Debian 12 / Armbian 26.08.0-trunk |
| Board kernel | 6.1.115-vendor-seeed-rk3588 |
| RKNPU driver | v0.9.8 |
| SDK | RKLLM-Toolkit / Runtime 1.3.0, repository commit 878f936 |
| Example account and address | BOARD_USER@BOARD_IP |
Run in the board's local terminal:
sudo apt update
sudo apt install -y openssh-server git
sudo systemctl enable --now sshRun on the PC:
sudo apt update
sudo apt install -y git wget openssh-client
ssh BOARD_USER@BOARD_IP#1. Fixing a missing RKNPU driver
Run on the board:
uname -m
uname -r
mountpoint -q /sys/kernel/debug || sudo mount -t debugfs debugfs /sys/kernel/debug
sudo cat /sys/kernel/debug/rknpu/version
sudo cat /sys/kernel/debug/rknpu/load
grep "^CONFIG_ROCKCHIP_RKNPU=" /boot/config-$(uname -r)
sudo dmesg | grep -i rknpu | tail -n 20
for node in /sys/class/drm/renderD*; do
printf '%s -> ' "$node"
readlink -f "$node/device/driver"
doneExcerpt from the actual board output:
RKNPU driver: v0.9.8
NPU load: Core0: 0%, Core1: 0%, Core2: 0%
/sys/class/drm/renderD130 -> /sys/bus/platform/drivers/RKNPU
CONFIG_ROCKCHIP_RKNPU=yrenderD130 is the node on this board, not a fixed number. SDK 1.3.0 requires driver v0.9.8 or later; CONFIG_ROCKCHIP_RKNPU=y means the driver is built into the kernel, so its absence from lsmod is normal. If the version file does not exist, also check dmesg and the driver binding.
On the reComputer RK3588 Seeed/Armbian Debian system used here, first inspect the matching kernel and device-tree packages:
sudo apt update
apt-cache policy linux-image-vendor-seeed-rk3588 linux-dtb-vendor-seeed-rk3588After confirming that the candidate versions of both packages come from this board's package source, install/reinstall the matching packages if the driver is missing or older than 0.9.8:
sudo apt install --reinstall linux-image-vendor-seeed-rk3588 linux-dtb-vendor-seeed-rk3588
sudo reboot#2. Install and verify RKLLM Runtime and RKLLM-Toolkit
| Component | Installed on | Purpose |
|---|---|---|
| RKNPU kernel driver | Board kernel | Access the NPU hardware |
RKLLM Runtime librkllmrt.so | Board aarch64 | Load .rkllm and provide a C/C++ inference API |
RKLLM-Toolkit rkllm_toolkit | Ubuntu x86_64 PC | Convert, quantize, and export .rkllm |
#2.1 Get the same SDK version on both devices
Run on the PC:
mkdir -p ~/RKLLM_Project
cd ~/RKLLM_Project
git clone --branch release-v1.3.0 --depth 1 https://github.com/airockchip/rknn-llm.gitRun on the board: This example checks out only the files required for building and running, reducing board storage use.
mkdir -p ~/RKLLM_Project
cd ~/RKLLM_Project
git clone --filter=blob:none --sparse --depth 1 --branch release-v1.3.0 \
https://github.com/airockchip/rknn-llm.git
cd rknn-llm
git sparse-checkout set examples/rkllm_api_demo/deploy rkllm-runtime/Linux#2.2 Verify Runtime on the board
Run on the board:
cd ~/RKLLM_Project/rknn-llm
file rkllm-runtime/Linux/librkllm_api/aarch64/librkllmrt.so
ldd rkllm-runtime/Linux/librkllm_api/aarch64/librkllmrt.sofile should report an AArch64 ELF file, and ldd should not show not found. Runtime does not require pip install. When Section 4 builds the example, CMake copies the matching librkllmrt.so to the deployment directory's lib/; set LD_LIBRARY_PATH before running it. On this board, ldd found dependencies such as libgomp.so.1. If your output says not found, install the missing board-side libraries first.
#2.3 Install Toolkit on the PC
The following uses Ubuntu x86_64 and Python 3.11 in an isolated Conda environment. If Miniforge is installed, start at the source line.
Run on the PC:
cd ~/RKLLM_Project
wget -O Miniforge3-Linux-x86_64.sh \
https://github.com/conda-forge/miniforge/releases/latest/download/Miniforge3-Linux-x86_64.sh
bash Miniforge3-Linux-x86_64.sh -b -p "$HOME/miniforge3"
source "$HOME/miniforge3/etc/profile.d/conda.sh"
conda create -n rkllm-1.3.0 python=3.11 -y
conda activate rkllm-1.3.0
python -m pip install --upgrade pip
python -m pip install ./rknn-llm/rkllm-toolkit/packages/rkllm_toolkit-1.3.0-cp311-cp311-linux_x86_64.whl
python -m pip install huggingface_hub
python -m pip check
python -c 'from rkllm.api import RKLLM; print("RKLLM-Toolkit import OK")'
nvidia-smi --query-gpu=name,memory.total --format=csv,noheader
python -c 'import torch; print("CUDA available:", torch.cuda.is_available())'In the wheel filename, cp311 means Python 3.11 and linux_x86_64 means the PC architecture; do not install this wheel on the board. This test used an existing ~/miniconda3 installation and reported No broken requirements found., RKLLM-Toolkit import OK, and CUDA available: True. The fresh-install commands above use Miniforge. Section 3 also requires a working NVIDIA driver on the PC.
#3. Convert the LLM model on an Ubuntu PC
#3.1 Download the original model
Run on the PC:
source "$HOME/miniforge3/etc/profile.d/conda.sh"
conda activate rkllm-1.3.0
mkdir -p ~/RKLLM_Project/models
hf download deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B \
config.json tokenizer.json tokenizer_config.json generation_config.json model.safetensors \
--local-dir ~/RKLLM_Project/models/DeepSeek-R1-Distill-Qwen-1.5B
ls -lh ~/RKLLM_Project/models/DeepSeek-R1-Distill-Qwen-1.5B/model.safetensorsThis downloads the original Hugging Face model, which cannot be passed directly to the board Runtime. Conversion uses data_quant.json from the official example directory as a demonstration calibration set; for production, replace it with question-and-answer data representing the target use cases.
#3.2 Generate the RK3588 model
Run on the PC: This test used a CUDA environment with an NVIDIA GTX 1070. Create the conversion script with the following parameters.
cat > ~/RKLLM_Project/convert_rk3588.py <<'PY'
from pathlib import Path
from rkllm.api import RKLLM
root = Path.home() / 'RKLLM_Project'
model = root / 'models/DeepSeek-R1-Distill-Qwen-1.5B'
dataset = root / 'rknn-llm/examples/rkllm_api_demo/export/data_quant.json'
output = root / 'models/DeepSeek-R1-Distill-Qwen-1.5B_W8A8_rk3588.rkllm'
llm = RKLLM()
ret = llm.load_huggingface(model=str(model), model_lora=None,
device='cuda', dtype='float16',
custom_config=None, load_weight=True)
if ret != 0:
raise RuntimeError(f'load_huggingface failed: {ret}')
ret = llm.build(do_quantization=True, optimization_level=1,
quantized_dtype='W8A8', quantized_algorithm='normal',
target_platform='RK3588', num_npu_core=3,
extra_qparams=None, dataset=str(dataset),
hybrid_rate=0, max_context=4096)
if ret != 0:
raise RuntimeError(f'build failed: {ret}')
ret = llm.export_rkllm(str(output))
if ret != 0:
raise RuntimeError(f'export_rkllm failed: {ret}')
print(output)
PY
python ~/RKLLM_Project/convert_rk3588.py
ls -lh ~/RKLLM_Project/models/DeepSeek-R1-Distill-Qwen-1.5B_W8A8_rk3588.rkllmRK3588 uses W8A8 quantization, the normal algorithm, and 3 NPU cores; this is the combination in the official example. max_context=4096 matches the run-time argument below. PC conversion succeeded in this test, producing a W8A8 model of about 2.0 GB. Put the model file on the corresponding RK3588 board.
#4. Deploy and run the C++ example on the board
#4.1 Build the official llm_demo on the board
Run on the board: Use the board's native aarch64 compiler instead of the PC cross-compiler path hard-coded in the official build-linux.sh.
sudo apt update
sudo apt install -y cmake g++ make
cd ~/RKLLM_Project/rknn-llm
cmake -S examples/rkllm_api_demo/deploy \
-B examples/rkllm_api_demo/deploy/build/native \
-DCMAKE_BUILD_TYPE=Release
cmake --build examples/rkllm_api_demo/deploy/build/native -j2
cmake --install examples/rkllm_api_demo/deploy/build/native
ls -lh examples/rkllm_api_demo/deploy/install/demo_Linux_aarch64/llm_demo \
examples/rkllm_api_demo/deploy/install/demo_Linux_aarch64/lib/librkllmrt.so#4.2 Transfer and run the model
Run on the PC:
scp ~/RKLLM_Project/models/DeepSeek-R1-Distill-Qwen-1.5B_W8A8_rk3588.rkllm \
BOARD_USER@BOARD_IP:~/RKLLM_Project/rknn-llm/examples/rkllm_api_demo/deploy/install/demo_Linux_aarch64/
sha256sum ~/RKLLM_Project/models/DeepSeek-R1-Distill-Qwen-1.5B_W8A8_rk3588.rkllm
ssh BOARD_USER@BOARD_IPRun on the board:
cd ~/RKLLM_Project/rknn-llm/examples/rkllm_api_demo/deploy/install/demo_Linux_aarch64
sha256sum ./DeepSeek-R1-Distill-Qwen-1.5B_W8A8_rk3588.rkllm
export LD_LIBRARY_PATH="$PWD/lib${LD_LIBRARY_PATH:+:$LD_LIBRARY_PATH}"
ldd ./llm_demo | grep -E 'librkllmrt|not found'
export RKLLM_LOG_LEVEL=1
./llm_demo ./DeepSeek-R1-Distill-Qwen-1.5B_W8A8_rk3588.rkllm 128 4096The PC and board SHA-256 values should match; both were 95fb13cf8287c2fa0e6338c550b0ceeea9eb2dfe752c55a2d0b519035bff78ee in this test. At the user: prompt, enter 1+1等于多少?请只回答数字。. Excerpt from the actual RK3588 output:
rkllm init start
I rkllm: rkllm-runtime version: 1.3.0, rknpu driver version: 0.9.8, platform: RK3588
I rkllm: rkllm-toolkit version: 1.3.0, max_context_limit: 4096, npu_core_num: 3, target_platform: RK3588, model_dtype: W8A8
rkllm init success
...
因此,**1 + 1 等于** \(\boxed{2}\)。The model's exact wording may vary. 128 is the maximum number of generated tokens for this run, and 4096 is the context limit; enter exit to quit.
If ldd shows not found, repair the missing shared library first. If initialization fails, check the driver version, model integrity, platform parameters, and Toolkit/Runtime versions. .rkllm files for RK3576 and RK3588 are not interchangeable.