PaddleOCR

Awesome multilingual OCR toolkits based on PaddlePaddle (practical ultra lightweight OCR system, support 80+ languages recognition, provide data annotation and synthesis tools, support training and deployment among server, mobile, embedded and IoT devices)

chineseocr crnn db ocr ocrlite

Go to file

zhangyubo0722 5e22578a96 update paddlex of readme for 2.7.1 (#11425 )		2023-12-28 14:24:28 +08:00
.github/ISSUE_TEMPLATE	Update newfeature.md	2023-10-16 17:55:59 +08:00
PPOCRLabel	refactored splitTrainVal and added multiOS path support (#11069 )	2023-10-13 10:27:26 +08:00
StyleText	upgrade pillow to 10.0.0 (#10405 )	2023-07-17 14:42:30 +08:00
applications	compat_pillow (#10596 )	2023-08-10 16:41:35 +08:00
benchmark	benchmark: unset python version (#10300 )	2023-07-05 17:30:45 +08:00
configs	add cppd u14m train model and doc (#11052 )	2023-10-11 17:15:01 +08:00
deploy	修复测试服务中图片转Base64的引用地址错误。 (#8334 )	2023-10-16 17:44:41 +08:00
doc	Update FAQ.md (#10349 )	2023-10-16 17:51:43 +08:00
ppocr	bugfix	2023-10-17 14:59:45 +08:00
ppstructure	修复代码错误 (#11107 )	2023-10-19 14:16:11 +08:00
test_tipc	Add new recognition method "ParseQ" (#10836 )	2023-09-07 16:36:47 +08:00
tools	[MLU] add mlu device for infer (#10249 )	2023-10-16 17:55:42 +08:00
.clang_format.hook	keep same version of clang-format with paddle (#9286 )	2023-03-03 14:51:38 +08:00
.gitignore	…
.pre-commit-config.yaml	修改数据增强导致的DSR报错 (#10662 )	2023-08-17 15:32:56 +08:00
.style.yapf	…
LICENSE	…
MANIFEST.in	…
README.md	update paddlex of readme for 2.7.1 (#11425 )	2023-12-28 14:24:28 +08:00
README_ch.md	revert README_ch.md update	2023-10-16 19:12:43 +08:00
README_en.md	update readme_en,fix_documents (#10592 )	2023-08-09 22:36:43 +08:00
__init__.py	…
paddleocr.py	Add preprocessing common to OCR tasks (#10217 )	2023-10-16 17:55:27 +08:00
requirements.txt	compat_pillow (#10596 )	2023-08-10 16:41:35 +08:00
setup.py	改进文档`deploy/hubserving/readme.md`和`doc/doc_ch/models_list.md` (#9110 )	2023-10-16 17:44:34 +08:00
train.sh	…

README_en.md

Introduction

PaddleOCR aims to create multilingual, awesome, leading, and practical OCR tools that help users train better models and apply them into practice.

📣 Recent updates

🔥2023.8.7 Release PaddleOCRrelease/2.7
- Release PP-OCRv4, support mobile version and server version
  - PP-OCRv4-mobile：When the speed is comparable, the effect of the Chinese scene is improved by 4.5% compared with PP-OCRv3, the English scene is improved by 10%, and the average recognition accuracy of the 80-language multilingual model is increased by more than 8%.
  - PP-OCRv4-server：Release the OCR model with the highest accuracy at present, the detection model accuracy increased by 4.9% in the Chinese and English scenes, and the recognition model accuracy increased by 2% refer quickstart quick use by one line command, At the same time, the whole process of model training, reasoning, and high-performance deployment can also be completed with few code in the General OCR Industry Solution in PaddleX.
- ReleasePP-ChatOCR, a new scheme for extracting key information of general scenes using PP-OCR model and ERNIE LLM.
🔨2022.11 Add implementation of 4 cutting-edge algorithms：Text Detection DRRG, Text Recognition RFL, Image Super-Resolution Text Telescope，Handwritten Mathematical Expression Recognition CAN
2022.10 release optimized JS version PP-OCRv3 model with 4.3M model size, 8x faster inference time, and a ready-to-use web demo
💥 Live Playback: Introduction to PP-StructureV2 optimization strategy. Scan the QR code below using WeChat, follow the PaddlePaddle official account and fill out the questionnaire to join the WeChat group, get the live link and 20G OCR learning materials (including PDF2Word application, 10 models in vertical scenarios, etc.)
🔥2022.8.24 Release PaddleOCR release/2.6
- Release PP-StructureV2，with functions and performance fully upgraded, adapted to Chinese scenes, and new support for Layout Recovery and one line command to convert PDF to Word;
- Layout Analysis optimization: model storage reduced by 95%, while speed increased by 11 times, and the average CPU time-cost is only 41ms;
- Table Recognition optimization: 3 optimization strategies are designed, and the model accuracy is improved by 6% under comparable time consumption;
- Key Information Extraction optimization：a visual-independent model structure is designed, the accuracy of semantic entity recognition is increased by 2.8%, and the accuracy of relation extraction is increased by 9.1%.
🔥2022.8 Release OCR scene application collection
- Release 9 vertical models such as digital tube, LCD screen, license plate, handwriting recognition model, high-precision SVTR model, etc, covering the main OCR vertical applications in general, manufacturing, finance, and transportation industries.
2022.8 Add implementation of 8 cutting-edge algorithms
- Text Detection: FCENet, DB++
- Text Recognition: ViTSTR, ABINet, VisionLAN, SPIN, RobustScanner
- Table Recognition: TableMaster
2022.5.9 Release PaddleOCR release/2.5
- Release PP-OCRv3: With comparable speed, the effect of Chinese scene is further improved by 5% compared with PP-OCRv2, the effect of English scene is improved by 11%, and the average recognition accuracy of 80 language multilingual models is improved by more than 5%.
- Release PPOCRLabelv2: Add the annotation function for table recognition task, key information extraction task and irregular text image.
- Release interactive e-book "Dive into OCR", covers the cutting-edge theory and code practice of OCR full stack technology.
more

🌟 Features

PaddleOCR support a variety of cutting-edge algorithms related to OCR, and developed industrial featured models/solution PP-OCR、 PP-Structure and PP-ChatOCR on this basis, and get through the whole process of data production, model training, compression, inference and deployment.

It is recommended to start with the “quick experience” in the document tutorial

⚡ Quick Experience

Web online experience
- PP-OCRv4 online experience：https://aistudio.baidu.com/aistudio/projectdetail/6611435
- PP-ChatOCR online experience：https://aistudio.baidu.com/aistudio/projectdetail/6488689
One line of code quick use: Quick Start（Chinese/English/Multilingual/Document Analysis
Full-process experience of training, inference, and high-performance deployment in the Paddle AI suite (PaddleX)：
- PP-OCRv4：https://aistudio.baidu.com/aistudio/modelsdetail?modelId=286
- PP-ChatOCR：https://aistudio.baidu.com/aistudio/modelsdetail?modelId=332
Mobile demo experience：Installation DEMO(Based on EasyEdge and Paddle-Lite, support iOS and Android systems)

📖 Technical exchange and cooperation

(PaddleX)provides a one-stop full-process high-efficiency development platform for flying paddle ecological model training, pressure, and push. Its mission is to help AI technology quickly land, and its vision is to make everyone an AI Developer!
- PaddleX currently covers areas such as image classification, object detection, image segmentation, 3D, OCR, and time series prediction, and has built-in 36 basic single models, such as RP-DETR, PP-YOLOE, PP-HGNet, PP-LCNet, PP- LiteSeg, etc.; integrated 12 practical industrial solutions, such as PP-OCRv4, PP-ChatOCR, PP-ShiTu, PP-TS, vehicle-mounted road waste detection, identification of prohibited wildlife products, etc.
- PaddleX provides two AI development modes: "Toolbox" and "Developer". The toolbox mode can tune key hyperparameters without code, and the developer mode can perform single-model training, push and multi-model serial inference with low code, and supports both cloud and local terminals.
- PaddleX also supports joint innovation and development, profit sharing! At present, PaddleX is rapidly iterating, and welcomes the participation of individual developers and enterprise developers to create a prosperous AI technology ecosystem!

Scan the QR code below on WeChat to add operation students, and reply [paddlex], operation students will invite you to join the official communication group for more efficient questions and answers.

[PaddleX] technology exchange group QR code

📚 E-book: Dive Into OCR

Dive Into OCR

👫 Community

For international developers, we regard PaddleOCR Discussions as our international community platform. All ideas and questions can be discussed here in English.
For Chinese develops, Scan the QR code below with your Wechat, you can join the official technical discussion group. For richer community content, please refer to 中文README, looking forward to your participation.

🛠️ PP-OCR Series Model List（Update on September 8th）

Model introduction	Model name	Recommended scene	Detection model	Direction classifier	Recognition model
Chinese and English ultra-lightweight PP-OCRv4 model（16.2M）	ch_PP-OCRv4_xx	Mobile & Server	inference model / trained model	inference model / trained model	inference model / trained model
Chinese and English ultra-lightweight PP-OCRv3 model（16.2M）	ch_PP-OCRv3_xx	Mobile & Server	inference model / trained model	inference model / trained model	inference model / trained model
English ultra-lightweight PP-OCRv3 model（13.4M）	en_PP-OCRv3_xx	Mobile & Server	inference model / trained model	inference model / trained model	inference model / trained model

For more model downloads (including multiple languages), please refer to PP-OCR series model downloads.
For a new language request, please refer to Guideline for new language_requests.
For structural document analysis models, please refer to PP-Structure models.

📖 Tutorials

👀 Visualization more

PP-OCRv3 Chinese model

PP-OCRv3 English model

PP-OCRv3 Multilingual model

PP-StructureV2

layout analysis + table recognition

SER (Semantic entity recognition)

RE (Relation Extraction)

🇺🇳 Guideline for New Language Requests

If you want to request a new language support, a PR with 1 following files are needed：

In folder ppocr/utils/dict, it is necessary to submit the dict text to this path and name it with {language}_dict.txt that contains a list of all characters. Please see the format example from other files in that folder.

If your language has unique elements, please tell me in advance within any way, such as useful links, wikipedia and so on.

More details, please refer to Multilingual OCR Development Plan.

📄 License

This project is released under Apache 2.0 license

README_en.md Unescape Escape