PaddleOCR

Awesome multilingual OCR toolkits based on PaddlePaddle (practical ultra lightweight OCR system, support 80+ languages recognition, provide data annotation and synthesis tools, support training and deployment among server, mobile, embedded and IoT devices)

chineseocr crnn db ocr ocrlite

Go to file

Wang Xin 5b54ac4606 update kie doc (#13799 )		2024-09-02 19:28:02 +08:00
.github	cache Python dependencies and PaddleOCR files (#13682 )	2024-08-16 16:08:32 +08:00
applications	docs: Remove old applications docs (#13551 )	2024-07-31 10:58:36 +08:00
benchmark	update common pre-commit configs and commit the results of running pre-commit run -a (#12516 )	2024-05-29 15:26:09 +08:00
configs	Add Syriac script support (#13800 )	2024-09-01 20:10:42 +08:00
deploy	rename MKLDNN to OneDNN (#13757 )	2024-08-27 12:52:00 +08:00
doc	docs: Remove doc/datasets directory and fix docs/datasets documents (#13700 )	2024-08-19 22:00:22 +08:00
docs	update kie doc (#13799 )	2024-09-02 19:28:02 +08:00
overrides/partials	docs: Add a new document site (#13375 )	2024-07-24 20:00:15 +08:00
ppocr	Add Syriac script support (#13800 )	2024-09-01 20:10:42 +08:00
ppstructure	update kie doc (#13799 )	2024-09-02 19:28:02 +08:00
test_tipc	update common pre-commit configs and commit the results of running pre-commit run -a (#12516 )	2024-05-29 15:26:09 +08:00
tests	add test for cls_postprocess (#12534 )	2024-05-31 11:22:13 +08:00
tools	Repair the bug in the inference script for LaTeX OCR (#13750 )	2024-08-26 14:21:41 +08:00
.clang_format.hook	…
.gitignore	update common pre-commit configs and commit the results of running pre-commit run -a (#12516 )	2024-05-29 15:26:09 +08:00
.pre-commit-config.yaml	docs: Add a new document site (#13375 )	2024-07-24 20:00:15 +08:00
.style.yapf	…
LICENSE	…
MANIFEST.in	use setuptools-scm extracts PaddleOCR versions (#13716 )	2024-08-23 17:27:31 +08:00
README.md	Remove channel links from documentation (#13674 )	2024-08-19 14:41:18 +08:00
README_en.md	docs: Update docs and remove out-of-date event (#13660 )	2024-08-15 08:23:10 +08:00
__init__.py	fix layout recovery import error (#13434 )	2024-07-20 21:19:09 +08:00
mkdocs.yml	docs: Update docs and fix markdown render error (#13678 )	2024-08-16 08:21:40 +08:00
paddleocr.py	remove unused enumerate (#13760 )	2024-08-28 09:10:33 +08:00
pyproject.toml	use setuptools-scm extracts PaddleOCR versions (#13716 )	2024-08-23 17:27:31 +08:00
requirements.txt	remove some of the less common dependencies (#13461 )	2024-07-24 19:29:58 +08:00
setup.py	…
train.sh	update common pre-commit configs and commit the results of running pre-commit run -a (#12516 )	2024-05-29 15:26:09 +08:00

README_en.md

English | 简体中文

Introduction

PaddleOCR aims to create multilingual, awesome, leading, and practical OCR tools that help users train better models and apply them into practice.

🚀 Community

PaddleOCR is being oversight by a PMC. Issues and PRs will be reviewed on a best-effort basis. For a complete overview of PaddlePaddle community, please visit community.

⚠️ Note: The Issues module is only for reporting program 🐞 bugs, for the rest of the questions, please move to the Discussions. Please note that if the Issue mentioned is not a bug, it will be moved to the Discussions module.

📣 Recent updates (more)

🔥2024.7 Added PaddleOCR Algorithm Model Challenge Champion Solutions:
- Challenge One, OCR End-to-End Recognition Task Champion Solution: Scene Text Recognition Algorithm-SVTRv2;
- Challenge Two, General Table Recognition Task Champion Solution: Table Recognition Algorithm-SLANet-LCNetV2.

📚 Documentation

Full documentation can be found on docs.

🌟 Features

PaddleOCR support a variety of cutting-edge algorithms related to OCR, and developed industrial featured models/solution PP-OCR、 PP-Structure and PP-ChatOCR on this basis, and get through the whole process of data production, model training, compression, inference and deployment.

It is recommended to start with the “quick experience” in the document tutorial

⚡ Quick Start

📖 Technical exchange and cooperation

PaddleX provides a one-stop full-process high-efficiency development platform for flying paddle ecological model training, pressure, and push. Its mission is to help AI technology quickly land, and its vision is to make everyone an AI Developer!

PaddleX currently covers areas such as image classification, object detection, image segmentation, 3D, OCR, and time series prediction, and has built-in 36 basic single models, such as RP-DETR, PP-YOLOE, PP-HGNet, PP-LCNet, PP- LiteSeg, etc.; integrated 12 practical industrial solutions, such as PP-OCRv4, PP-ChatOCR, PP-ShiTu, PP-TS, vehicle-mounted road waste detection, identification of prohibited wildlife products, etc.
PaddleX provides two AI development modes: "Toolbox" and "Developer". The toolbox mode can tune key hyperparameters without code, and the developer mode can perform single-model training, push and multi-model serial inference with low code, and supports both cloud and local terminals.
PaddleX also supports joint innovation and development, profit sharing! At present, PaddleX is rapidly iterating, and welcomes the participation of individual developers and enterprise developers to create a prosperous AI technology ecosystem!

📚 E-book: Dive Into OCR

Dive Into OCR

🎖 Contributors

⭐️ Star

🇺🇳 Guideline for New Language Requests

If you want to request a new language support, a PR with 1 following files are needed：

In folder ppocr/utils/dict, it is necessary to submit the dict text to this path and name it with {language}_dict.txt that contains a list of all characters. Please see the format example from other files in that folder.

If your language has unique elements, please tell me in advance within any way, such as useful links, wikipedia and so on.

More details, please refer to Multilingual OCR Development Plan.

📄 License

This project is released under Apache License Version 2.0.

README_en.md Unescape Escape