Parker Hu2026-09-20

RK3588 VPU and NPU Hardware Acceleration Benchmark

Reproducible MPP benchmarks for NV12-to-MJPG/H.264/H.265 hardware encoding, MJPG-to-YUV420 hardware decoding, and RKNN acceleration of YOLOv8n, YOLOv8s, and YOLOv8m inference on the reComputer RK3588.

vpunpubenchmarkDownload

This project quantitatively tests three core capabilities of the RK3588 platform:

  1. MPP video encoding
  2. MPP video decoding
  3. RKNN inference

Test coverage:

  • Encoding: NV12 -> MJPG / H.264 / H.265
  • Decoding: MJPG -> YUV420
  • Inference: 640x640 input with YOLOv8n / YOLOv8s / YOLOv8m
  • Resolutions: 720p / 1080p / 2K / 4K / 8K

#RK3588 Benchmark Summary

  • Generated at: 2026-09-20 17:09:50
  • Codec frames: warm-up 20 + measured 120
  • Inference frames: warm-up 20 + measured 200
  • CPU governor: ondemand
  • Temperature before/after: approximately 35–36°C / 36–37°C

#Video Encode Benchmark (NV12 -> MJPG / H.264 / H.265)

ResolutionSizeCodecBitrate (Mbps)Pure FPSPipeline FPSPure Avg (ms)Pure P95 (ms)Output (MB)Status
720p1280×720MJPG20.0128.094.97.88.217.2OK
720p1280×720H.2644.0397.3233.72.52.81.8OK
720p1280×720H.2654.0400.0285.12.52.61.1OK
1080p1920×1080MJPG30.0203.095.94.95.338.6OK
1080p1920×1080H.2648.0218.0121.74.65.13.6OK
1080p1920×1080H.2658.0213.4111.14.74.82.4OK
2K2560×1440MJPG45.0120.745.68.38.568.4OK
2K2560×1440H.26416.0137.092.77.37.47.3OK
2K2560×1440H.26516.0138.593.87.27.34.2OK
4K3840×2160MJPG60.055.920.317.919.3153.9OK
4K3840×2160H.26432.064.035.415.616.014.6OK
4K3840×2160H.26532.064.835.615.415.79.6OK
8K7680×4320MJPG90.015.610.264.265.4615.2OK
8K7680×4320H.26464.016.711.359.960.727.5OK
8K7680×4320H.26564.032.418.230.831.231.0OK

#Video Decode Benchmark (MJPG -> YUV420)

ResolutionSizeFramesPure FPSPipeline FPSPure Avg (ms)Pure P95 (ms)Status
720p1280×720120/120497.4477.82.02.4OK
1080p1920×1080120/120311.7290.33.23.4OK
2K2560×1440120/120208.8198.84.85.2OK
4K3840×2160120/120106.1100.69.410.1OK
8K7680×4320120/12030.428.832.934.4OK

#YOLOv8 640x640 Inference Benchmark

ModelInference FPSAvg (ms)P95 (ms)Preprocess (ms)NPU (ms)Status
YOLOv8n33.928.633.10.919.6OK
YOLOv8s16.759.264.60.837.8OK
YOLOv8m9.9100.8105.80.881.8OK

#1. How to Run the Benchmark Project

Run the benchmark on the reComputer RK3588 board.

bash
sudo apt update && sudo apt install unzip -y
wget https://files.seeedstudio.com/RK3576/rk3588_mpp_benchmark_install.zip -O install.zip && unzip install.zip
cd ./install && export LD_LIBRARY_PATH="$(pwd)/lib:${LD_LIBRARY_PATH}" && chmod +x ./bin/rk3588_benchmark
sudo ./bin/rk3588_benchmark

The test results are saved in:

bash
ls ./benchmark_results
benchmark_summary.md  video_codec_benchmark.csv  yolov8_inference_benchmark.csv

Common parameters:

bash
./bin/rk3588_benchmark \
  --output-dir ./benchmark_results \
  --resolutions 720p,1080p,2k,4k,8k \
  --models yolov8n,yolov8s,yolov8m \
  --codec-warmup 20 \
  --codec-frames 120 \
  --infer-warmup 20 \
  --infer-frames 200

The wrapper script points LD_LIBRARY_PATH to the packaged lib/ directory. If the current user cannot access /dev/mpp_service or /dev/rga, it reruns through sudo.

#2. Project Capabilities

The project automatically generates test frames, runs multi-resolution encoding, MJPG hardware decoding, and 640x640 YOLOv8 inference tests, and exports a Markdown summary with CSV raw data.

#3. Test Process

  1. Generate dynamic YUV420SP (NV12) frames at the selected resolution.
  2. Run the MPP MJPG, H.264, and H.265 encoders separately.
  3. Retain the generated MJPG elementary stream and decode it to YUV420 through the MPP advanced task API.
  4. Generate 640x640 NV12 frames and preprocess them through RGA.
  5. Run the RK3588-specific YOLOv8n/s/m RKNN models.
  6. Aggregate FPS, average latency, P95 latency, and output size.
  7. Generate Markdown and CSV reports.

#4. Technical Approach

#4.1 Encoding

The encoder uses Rockchip MPP with dynamic YUV420SP (NV12) input. MJPG uses Q=80; H.264 and H.265 use CBR at the listed bitrates. Each group warms up for 20 frames and measures 120 frames.

  • Pure FPS: measured frames divided by accumulated time in MPP encode_put_frame and encode_get_packet
  • Pipeline FPS: full loop including frame generation, input writes, and output copies
  • Pure Avg/P95: per-frame latency of the measured MPP interval
  • Output: total encoded size during the 120 measured frames

Pure FPS is a serial-submission service rate. It excludes periods when the VPU may be idle while the CPU prepares the next frame, so it is not the final frame rate of a camera, network, display, or storage pipeline.

#4.2 Decoding

The decoder reuses the MJPG stream at each resolution and runs MJPG -> YUV420 through the RK3588 MPP advanced task API. Each group warms up for 20 frames and measures 120 frames.

  • Pure Decode FPS: measured after compressed-packet allocation and copying until MPP returns the output frame
  • Pipeline FPS: also includes packet preparation and task recycling

#4.3 Inference

The inference path generates 640x640 NV12 frames, preprocesses them with RGA, and sends them to RKNN Runtime. It tests:

  • yolov8n_rk3588.rknn
  • yolov8s_rk3588.rknn
  • yolov8m_rk3588.rknn

Each model warms up for 20 frames and measures 200 frames. The report records complete Infer, RGA preprocessing, and RKNN inference-stage latency.

#5. Explanation of Result Metrics

#5.1 Encoding Results

Use Pure FPS to compare VPU service capability and Pipeline FPS to assess the unoptimized complete loop. Their gap indicates the cost of frame generation and memory copying.

#5.2 Decoding Results

Pure Avg/P95 and Pure FPS describe hardware decoding after packet preparation. Pipeline FPS covers the complete packet-processing loop.

#5.3 Inference Results

Inference FPS comes from the complete serial Infer call. Preprocess isolates RGA work; NPU reports the RKNN stage. Camera capture, networking, display, and application queues are excluded.

#6. Objective Assessment Based on the Current Results

After excluding CPU-side NV12 generation, input-memory writes, and output copying, pure 4K H.264/H.265 encoding on the RK3588 reaches approximately 64 FPS, pure 4K MJPG decoding reaches 106.1 FPS, and YOLOv8n inference reaches 33.9 FPS. The lower end-to-end encoding figures mainly reflect frame preparation and memory movement in this benchmark pipeline rather than the VPU itself. Although pure 8K H.265 encoding reaches 32.4 FPS, the complete loop reaches only 18.2 FPS; sustained end-to-end 8K@30 still requires zero-copy, parallel submission, and optimization of the complete media path.

  1. Use YOLOv8n by default. For YOLOv8s/m, use frame sampling and bounded queues.
  2. Production 4K H.264/H.265 pipelines should use zero-copy or shared capture buffers and reserve memory bandwidth for other workloads.
  3. Do not use 32.4 Pure FPS to claim end-to-end 8K@30; the current pipeline reaches only 18.2 FPS.
  4. Leave engineering margin for 8K MJPG decoding because P95 exceeds one 30-FPS frame interval.
  5. Before publishing formal results, fix clocks where appropriate, run at least three rounds, and report median, P95, and temperature.