Convert YOLO to RKNN and Run It on the RK3588 NPU
Export YOLO to ONNX, quantise and convert it to .rknn with RKNN-Toolkit2 on an x86 host, then run it on the board with RKNN-Toolkit-Lite2, using all three RK3588 NPU cores.
This is the standard path for getting a YOLO detector onto a Rockchip NPU. The same flow applies to RK3576, RK3568/RK3566 and RV1106; only target_platform changes.
Overview
PyTorch (.pt) ──export──► ONNX ──RKNN-Toolkit2 (x86 host)──► model.rknn
│
board: librknnrt + RKNN-Toolkit-Lite2 / C API
1. Export to ONNX (the NPU-friendly way)
Stock YOLO exports include post-processing (DFL, box decode, concat) that the NPU runs poorly or not at all. Rockchip's RKNN Model Zoo documents export settings (and links a modified Ultralytics fork) that end the graph at the raw heads, leaving decode and NMS to the CPU. Follow the export instructions for your YOLO version in the model zoo's examples/yolov8 (or v5/v10/v11) directory.
2. Set up RKNN-Toolkit2 on an x86 Linux host
python3 -m venv rknn && source rknn/bin/activate
git clone https://github.com/airockchip/rknn-toolkit2
pip install -r rknn-toolkit2/rknn-toolkit2/packages/x86_64/requirements_cpXX-*.txt
pip install rknn-toolkit2/rknn-toolkit2/packages/x86_64/rknn_toolkit2-*-cpXX-*.whl
Pick the wheel matching your Python version (cpXX). The toolkit runs on x86_64 Linux; an ARM64 build of the full toolkit is also provided for converting on-device.
3. Convert and quantise
Create dataset.txt listing 20–200 representative images (one path per line, taken from your real camera if possible). Then:
from rknn.api import RKNN
rknn = RKNN(verbose=True)
rknn.config(
mean_values=[[0, 0, 0]],
std_values=[[255, 255, 255]], # YOLO expects 0..1 input
target_platform='rk3588', # rk3576 / rk3568 / rk3566 / rv1106 ...
)
assert rknn.load_onnx(model='yolov8n.onnx') == 0
assert rknn.build(do_quantization=True, dataset='./dataset.txt') == 0
assert rknn.export_rknn('yolov8n.rknn') == 0
rknn.release()
Read the verbose log: it lists any operators that fall back to the CPU. If accuracy drops after INT8 quantisation, run rknn.accuracy_analysis() to find the offending layers, or try hybrid quantisation.
4. Prepare the board
# NPU driver present?
sudo cat /sys/kernel/debug/rknpu/version
# Install the runtime + Lite2 matching your toolkit version
sudo cp librknnrt.so /usr/lib/
pip install rknn_toolkit_lite2-*-cpXX-*-linux_aarch64.whl
The runtime version on the board must be compatible with the toolkit version used for conversion. Mismatches cause load failures or wrong results.
5. Run inference
import cv2
import numpy as np
from rknnlite.api import RKNNLite
rknn = RKNNLite()
rknn.load_rknn('yolov8n.rknn')
rknn.init_runtime(core_mask=RKNNLite.NPU_CORE_0_1_2) # all 3 RK3588 cores
img = cv2.imread('bus.jpg')
img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)
img = cv2.resize(img, (640, 640)) # use letterbox in production
outputs = rknn.inference(inputs=[np.expand_dims(img, 0)])
# outputs: raw head tensors → decode boxes + NMS on CPU
# (reuse the post-processing from rknn_model_zoo/examples/yolov8/python)
rknn.release()
6. Performance tips
- Multi-core:
NPU_CORE_0_1_2spreads one model across cores; for multiple camera streams, run one model instance per core (NPU_CORE_0,_1,_2) in separate threads for higher total throughput. - Pre-processing: use RGA for resize and colour conversion instead of OpenCV on the CPU.
- Watch load:
sudo cat /sys/kernel/debug/rknpu/loadwhile running. - C API: for production, the C API (
rknn_api.h) with zero-copy input buffers is significantly faster than Python.
¿Encontraste un error o quieres añadir algo? Sugerir una mejora.