CV / PaddleOCR

PaddleOCR v5 Mobile Detection

Text-region detection for document images on reComputer R Series with Hailo-8 or Hailo-10H.

24 Downloads
Größe
5.49 MB
Speicher
4GB+
Präzision
Hailo HEF / HailoRT

Wähle das Gerät, das du verwendest. Die Einrichtungsanleitung und Dokumentation werden entsprechend aktualisiert.

Erste Schritte

Bereitstellen
sudo docker run --rm \
  --name cm5-hailo8-paddle-ocr-v5-mobile-detection \
  --privileged \
  --net=host \
  -e PYTHONUNBUFFERED=1 \
  --device /dev/hailo0:/dev/hailo0 \
  -v /usr/lib/libhailort.so.4.23.0:/usr/lib/libhailort.so.4.23.0:ro \
  -v /usr/lib/libhailort.so:/usr/lib/libhailort.so:ro \
  ghcr.io/seeed-projects/recomputer-hailo8-cv/paddle_ocr_v5_mobile_detection:latest \
  python web_detection.py --model_path model/paddle_ocr_v5_mobile_detection.hef --video_path video/test.mp4

Modelldetails

reComputer R Series (CM5 + Hailo-10H)

PaddleOCR v5 Mobile Detection auf reComputer R Series (CM5 + Hailo-10H)

PaddleOCR v5 Mobile Detection lokalisiert Textbereiche in Dokumentbildern auf Hailo-10H über HailoRT 5.1.1. Der DB-Head (Differentiable Binarization) gibt eine Wahrscheinlichkeitskarte aus; die Bereiche werden auf der CPU als viereckige Polygone extrahiert.

Diese Seite zielt auf die reComputer R Series (CM5 + Hailo-10H) mit einem PCIe-Hailo-10H-Beschleuniger ab.

Modellinformationen

EigenschaftWert
ArchitekturPaddleOCR v5 Mobile (DB-Textdetektor)
AufgabeTexterkennung (Detection)
Eingabe544×960×3
AusgabeWahrscheinlichkeitskarte der Textbereiche → Vierecke
Parameter1,2M
Operationen6,5G
GenauigkeitDB-Metrik 4,60 (Model-Zoo-Referenz)
HEFHailo Model Zoo v5.4.0, Hailo-10H

Hardware- und Host-Einrichtung

ElementWert
BoardreComputer R Series mit Raspberry Pi CM5
BeschleunigerHailo-10H über PCIe, verfügbar als /dev/hailo0
Host-Treiberapt-Paket hailo-h10-all
RuntimeHailoRT 5.1.1 (Host/Container müssen major.minor teilen)
Python im Container3.13, aarch64
bash
sudo apt update
sudo apt install hailo-h10-all -y
sudo reboot

# Nach dem Neustart
hailortcli fw-control identify
ls /dev/hailo0

# Docker
curl -fsSL https://get.docker.com -o get-docker.sh
sudo sh get-docker.sh --mirror Aliyun
sudo systemctl enable docker
sudo systemctl start docker

Mit Demovideo ausführen

bash
sudo docker run --rm \
  --name hailo10h-paddle-ocr-detection \
  --privileged \
  --net=host \
  -e PYTHONUNBUFFERED=1 \
  --device /dev/hailo0:/dev/hailo0 \
  -v /usr/lib/libhailort.so.5.1.1:/usr/lib/libhailort.so.5.1.1:ro \
  -v /usr/lib/libhailort.so:/usr/lib/libhailort.so:ro \
  ghcr.io/seeed-projects/recomputer-hailo10h-cv/paddle_ocr_v5_mobile_detection:latest \
  python web_detection.py --model_path model/paddle_ocr_v5_mobile_detection.hef --video_path video/test.mp4

Öffnen Sie http://<Board_IP>:8000, um die Webvorschau anzusehen (grüne Vierecke um die erkannten Textbereiche).

USB-Kameramodus

bash
sudo docker run --rm \
  --name hailo10h-paddle-ocr-detection \
  --privileged \
  --net=host \
  -e PYTHONUNBUFFERED=1 \
  --device /dev/hailo0:/dev/hailo0 \
  --device /dev/video0:/dev/video0 \
  -v /usr/lib/libhailort.so.5.1.1:/usr/lib/libhailort.so.5.1.1:ro \
  -v /usr/lib/libhailort.so:/usr/lib/libhailort.so:ro \
  ghcr.io/seeed-projects/recomputer-hailo10h-cv/paddle_ocr_v5_mobile_detection:latest \
  python web_detection.py --model_path model/paddle_ocr_v5_mobile_detection.hef --camera_id 0

REST-API

bash
curl -X POST "http://<Board_IP>:8000/api/models/paddle_ocr_v5_mobile_detection/predict" \
  -F "file=@test.png"
EndpunktMethodeZweck
/api/models/paddle_ocr_v5_mobile_detection/predictPOSTPolygone der Textbereiche (JSON)
/api/video_feedGETMJPEG-Vorschaustream
/api/configGET / POSTBox-Score-/Binarisierungsschwellen

Entwicklungsnotizen

  • Quellmodul: src/hailo10h_paddle_ocr_v5_mobile_detection/
  • Dockerfile: docker/hailo10h/paddle_ocr_v5_mobile_detection.dockerfile
  • Container: ghcr.io/seeed-projects/recomputer-hailo10h-cv/paddle_ocr_v5_mobile_detection:latest
  • DB-Nachbearbeitung unverändert; der Executor verwendet die create_infer_model-API von HailoRT 5.1.1.
  • Kombinieren Sie es auf Anwendungsebene mit dem Erkennungsmodul, um den Text in jedem erkannten Bereich zu lesen.

Eingaben und Ausgaben

Input: full document image or demo video frame. Output: detected text-region polygons, bounding boxes, and annotated MJPEG preview.