PaddleOCR/ppstructure/recovery/README.md

English | [简体中文](README_ch.md)

- [Getting Started](#getting-started)
  - [1.  Introduction](#1)
  - [2. Install](#2)
    - [2.1 Installation dependencies](#2.1)
    - [2.2 Install PaddleOCR](#2.2)
  - [3. Quick Start](#3)

<a name="1"></a>

## 1. Introduction

Layout recovery means that after OCR recognition, the content is still arranged like the original document pictures, and the paragraphs are output to word document in the same order.

Layout recovery combines [layout analysis](../layout/README.md)、[table recognition](../table/README.md) to better recover images, tables, titles, etc.
The following figure shows the result：

<div align="center">
<img src="../docs/table/recovery.jpg"  width = "700" />
</div>
<a name="2"></a>

## 2. Install

<a name="2.1"></a>

### 2.1 Install dependencies

- **(1) Install PaddlePaddle**

```bash
python3 -m pip install --upgrade pip

# GPU installation
python3 -m pip install "paddlepaddle-gpu" -i https://mirror.baidu.com/pypi/simple

# CPU installation
python3 -m pip install "paddlepaddle" -i https://mirror.baidu.com/pypi/simple

````

For more requirements, please refer to the instructions in [Installation Documentation](https://www.paddlepaddle.org.cn/en/install/quick?docurl=/documentation/docs/en/install/pip/macos-pip_en.html).

<a name="2.2"></a>

### 2.2 Install PaddleOCR

- **(1) Download source code**

```bash
[Recommended] git clone https://github.com/PaddlePaddle/PaddleOCR

# If the pull cannot be successful due to network problems, you can also choose to use the hosting on the code cloud:
git clone https://gitee.com/paddlepaddle/PaddleOCR

# Note: Code cloud hosting code may not be able to synchronize the update of this github project in real time, there is a delay of 3 to 5 days, please use the recommended method first.
````

- **(2) Install recovery's `requirements`**

```bash
python3 -m pip install -r ppstructure/recovery/requirements.txt
````

<a name="3"></a>

## 3. Quick Start

<a name="3.1"></a>
### 3.1 下载模型

If input is English document, download English models:

```python
cd PaddleOCR/ppstructure

# download model
mkdir inference && cd inference
# Download the detection model of the ultra-lightweight English PP-OCRv3 model and unzip it
https://paddleocr.bj.bcebos.com/PP-OCRv3/english/en_PP-OCRv3_det_infer.tar && tar xf en_PP-OCRv3_det_infer.tar
# Download the recognition model of the ultra-lightweight English PP-OCRv3 model and unzip it
wget https://paddleocr.bj.bcebos.com/PP-OCRv3/english/en_PP-OCRv3_rec_infer.tar && tar xf en_PP-OCRv3_rec_infer.tar
# Download the ultra-lightweight English table inch model and unzip it
wget https://paddleocr.bj.bcebos.com/ppstructure/models/slanet/en_ppstructure_mobile_v2.0_SLANet_infer.tar && tar xf en_ppstructure_mobile_v2.0_SLANet_infer.tar
# Download the layout model of publaynet dataset and unzip it
wget https://paddleocr.bj.bcebos.com/ppstructure/models/layout/picodet_lcnet_x1_0_fgd_layout_infer.tar && tar xf picodet_lcnet_x1_0_fgd_layout_infer.tar
cd ..
```
If input is Chinese document，download Chinese models:
[Chinese and English ultra-lightweight PP-OCRv3 model](https://github.com/PaddlePaddle/PaddleOCR/blob/dygraph/README.md#pp-ocr-series-model-listupdate-on-september-8th)、[表格识别模型](https://github.com/PaddlePaddle/PaddleOCR/blob/dygraph/ppstructure/docs/models_list.md#22-表格识别模型)、[版面分析模型](https://github.com/PaddlePaddle/PaddleOCR/blob/dygraph/ppstructure/docs/models_list.md#1-版面分析模型)

<a name="3.2"></a>
### 3.2 版面恢复


```bash
python3 predict_system.py \
    --image_dir=./docs/table/1.png \
    --det_model_dir=inference/en_PP-OCRv3_det_infer \
    --rec_model_dir=inference/en_PP-OCRv3_rec_infer \
    --rec_char_dict_path=../ppocr/utils/en_dict.txt \
    --table_model_dir=inference/en_ppstructure_mobile_v2.0_SLANet_infer \
    --table_char_dict_path=../ppocr/utils/dict/table_structure_dict.txt \
    --layout_model_dir=inference/picodet_lcnet_x1_0_fgd_layout_infer \
    --layout_dict_path=../ppocr/utils/dict/layout_dict/layout_publaynet_dict.txt \
    --vis_font_path=../doc/fonts/simfang.ttf \
    --recovery=True \
    --save_pdf=False \
    --output=../output/
```

After running, the docx of each picture will be saved in the directory specified by the output field

Field：

- image_dir：test file测试文件， can be picture, picture directory, pdf file, pdf file directory
- det_model_dir：OCR detection model path
- rec_model_dir：OCR recognition model path
- rec_char_dict_path：OCR recognition dict path. If the Chinese model is used, change to "../ppocr/utils/ppocr_keys_v1.txt". And if you trained the model on your own dataset, change to the trained dictionary
- table_model_dir：tabel recognition model path
- table_char_dict_path：tabel recognition dict path. If the Chinese model is used, no need to change
- layout_model_dir：layout analysis model path
- layout_dict_path：layout analysis dict path. If the Chinese model is used, change to "../ppocr/utils/dict/layout_dict/layout_cdla_dict.txt"
- recovery：whether to enable layout of recovery, default False
- save_pdf：when recovery file, whether to save pdf file, default False
- output：save the recovery result path
-												add recovery

											
										
										
											2022-05-07 16:55:20 +08:00
+								English | [简体中文](README_ch.md)
 								- [Getting Started](#getting-started)
 								  - [1.  Introduction](#1)
-												modify recovery

											
										
										
											2022-05-09 16:17:40 +08:00
+								  - [2. Install](#2)
 								    - [2.1 Installation dependencies](#2.1)
 								    - [2.2 Install PaddleOCR](#2.2)
 								  - [3. Quick Start](#3)
-												add recovery

											
										
										
											2022-05-07 16:55:20 +08:00
 								<a name="1"></a>
-												update doc

											
										
										
											2022-08-22 11:48:18 +08:00
+								## 1. Introduction
-												add recovery

											
										
										
											2022-05-07 16:55:20 +08:00
 								Layout recovery means that after OCR recognition, the content is still arranged like the original document pictures, and the paragraphs are output to word document in the same order.
 								Layout recovery combines [layout analysis](../layout/README.md)、[table recognition](../table/README.md) to better recover images, tables, titles, etc.
 								The following figure shows the result：
 								<div align="center">
 								<img src="../docs/table/recovery.jpg"  width = "700" />
 								</div>
 								<a name="2"></a>
-												modify recovery

											
										
										
											2022-05-09 16:17:40 +08:00
+								## 2. Install
 								<a name="2.1"></a>
 								### 2.1 Install dependencies
 								- **(1) Install PaddlePaddle**
 								```bash
 								python3 -m pip install --upgrade pip
 								# GPU installation
-												update doc

											
										
										
											2022-08-22 11:48:18 +08:00
+								python3 -m pip install "paddlepaddle-gpu" -i https://mirror.baidu.com/pypi/simple
-												modify recovery

											
										
										
											2022-05-09 16:17:40 +08:00
 								# CPU installation
-												update doc

											
										
										
											2022-08-22 11:48:18 +08:00
+								python3 -m pip install "paddlepaddle" -i https://mirror.baidu.com/pypi/simple
-												modify recovery

											
										
										
											2022-05-09 16:17:40 +08:00
 								````
-												update doc

											
										
										
											2022-08-22 11:48:18 +08:00
+								For more requirements, please refer to the instructions in [Installation Documentation](https://www.paddlepaddle.org.cn/en/install/quick?docurl=/documentation/docs/en/install/pip/macos-pip_en.html).
-												modify recovery

											
										
										
											2022-05-09 16:17:40 +08:00
 								<a name="2.2"></a>
 								### 2.2 Install PaddleOCR
 								- **(1) Download source code**
 								```bash
 								[Recommended] git clone https://github.com/PaddlePaddle/PaddleOCR
 								# If the pull cannot be successful due to network problems, you can also choose to use the hosting on the code cloud:
 								git clone https://gitee.com/paddlepaddle/PaddleOCR
 								# Note: Code cloud hosting code may not be able to synchronize the update of this github project in real time, there is a delay of 3 to 5 days, please use the recommended method first.
 								````
 								- **(2) Install recovery's `requirements`**
 								```bash
 								python3 -m pip install -r ppstructure/recovery/requirements.txt
 								````
 								<a name="3"></a>
 								## 3. Quick Start
-												add recovery

											
										
										
											2022-05-07 16:55:20 +08:00
-												update doc

											
										
										
											2022-08-22 11:48:18 +08:00
+								<a name="3.1"></a>
 								### 3.1 下载模型
 								If input is English document, download English models:
-												add recovery

											
										
										
											2022-05-07 16:55:20 +08:00
+								```python
 								cd PaddleOCR/ppstructure
 								# download model
 								mkdir inference && cd inference
 								# Download the detection model of the ultra-lightweight English PP-OCRv3 model and unzip it
-												update doc

											
										
										
											2022-08-22 11:48:18 +08:00
+								https://paddleocr.bj.bcebos.com/PP-OCRv3/english/en_PP-OCRv3_det_infer.tar && tar xf en_PP-OCRv3_det_infer.tar
-												add recovery

											
										
										
											2022-05-07 16:55:20 +08:00
+								# Download the recognition model of the ultra-lightweight English PP-OCRv3 model and unzip it
-												update doc

											
										
										
											2022-08-22 11:48:18 +08:00
+								wget https://paddleocr.bj.bcebos.com/PP-OCRv3/english/en_PP-OCRv3_rec_infer.tar && tar xf en_PP-OCRv3_rec_infer.tar
-												add recovery

											
										
										
											2022-05-07 16:55:20 +08:00
+								# Download the ultra-lightweight English table inch model and unzip it
-												update doc

											
										
										
											2022-08-22 11:48:18 +08:00
+								wget https://paddleocr.bj.bcebos.com/ppstructure/models/slanet/en_ppstructure_mobile_v2.0_SLANet_infer.tar && tar xf en_ppstructure_mobile_v2.0_SLANet_infer.tar
-												update recovery (#7259)

* update recovery

* update recovery

* update recovery

* update recovery

* update recovery
											
										
										
											2022-08-19 20:15:37 +08:00
+								# Download the layout model of publaynet dataset and unzip it
-												update doc

											
										
										
											2022-08-22 11:48:18 +08:00
+								wget https://paddleocr.bj.bcebos.com/ppstructure/models/layout/picodet_lcnet_x1_0_fgd_layout_infer.tar && tar xf picodet_lcnet_x1_0_fgd_layout_infer.tar
-												add recovery

											
										
										
											2022-05-07 16:55:20 +08:00
+								cd ..
-												update doc

											
										
										
											2022-08-22 11:48:18 +08:00
+								```
 								If input is Chinese document，download Chinese models:
 								[Chinese and English ultra-lightweight PP-OCRv3 model](https://github.com/PaddlePaddle/PaddleOCR/blob/dygraph/README.md#pp-ocr-series-model-listupdate-on-september-8th)、[表格识别模型](https://github.com/PaddlePaddle/PaddleOCR/blob/dygraph/ppstructure/docs/models_list.md#22-表格识别模型)、[版面分析模型](https://github.com/PaddlePaddle/PaddleOCR/blob/dygraph/ppstructure/docs/models_list.md#1-版面分析模型)
 								<a name="3.2"></a>
 								### 3.2 版面恢复
 								```bash
-												update recovery (#7259)

* update recovery

* update recovery

* update recovery

* update recovery

* update recovery
											
										
										
											2022-08-19 20:15:37 +08:00
+								python3 predict_system.py \
 								    --image_dir=./docs/table/1.png \
 								    --det_model_dir=inference/en_PP-OCRv3_det_infer \
-												update doc

											
										
										
											2022-08-22 11:48:18 +08:00
+								    --rec_model_dir=inference/en_PP-OCRv3_rec_infer \
-												update recovery (#7259)

* update recovery

* update recovery

* update recovery

* update recovery

* update recovery
											
										
										
											2022-08-19 20:15:37 +08:00
+								    --rec_char_dict_path=../ppocr/utils/en_dict.txt \
-												update doc

											
										
										
											2022-08-22 11:48:18 +08:00
+								    --table_model_dir=inference/en_ppstructure_mobile_v2.0_SLANet_infer \
-												update recovery (#7259)

* update recovery

* update recovery

* update recovery

* update recovery

* update recovery
											
										
										
											2022-08-19 20:15:37 +08:00
+								    --table_char_dict_path=../ppocr/utils/dict/table_structure_dict.txt \
-												update doc

											
										
										
											2022-08-22 11:48:18 +08:00
+								    --layout_model_dir=inference/picodet_lcnet_x1_0_fgd_layout_infer \
-												update recovery (#7259)

* update recovery

* update recovery

* update recovery

* update recovery

* update recovery
											
										
										
											2022-08-19 20:15:37 +08:00
+								    --layout_dict_path=../ppocr/utils/dict/layout_dict/layout_publaynet_dict.txt \
 								    --vis_font_path=../doc/fonts/simfang.ttf \
 								    --recovery=True \
-												update doc

											
										
										
											2022-08-22 11:48:18 +08:00
+								    --save_pdf=False \
 								    --output=../output/
-												add recovery

											
										
										
											2022-05-07 16:55:20 +08:00
+								```
-												update doc

											
										
										
											2022-08-22 11:48:18 +08:00
+								After running, the docx of each picture will be saved in the directory specified by the output field
 								Field：
 								- image_dir：test file测试文件， can be picture, picture directory, pdf file, pdf file directory
 								- det_model_dir：OCR detection model path
 								- rec_model_dir：OCR recognition model path
 								- rec_char_dict_path：OCR recognition dict path. If the Chinese model is used, change to "../ppocr/utils/ppocr_keys_v1.txt". And if you trained the model on your own dataset, change to the trained dictionary
 								- table_model_dir：tabel recognition model path
 								- table_char_dict_path：tabel recognition dict path. If the Chinese model is used, no need to change
 								- layout_model_dir：layout analysis model path
 								- layout_dict_path：layout analysis dict path. If the Chinese model is used, change to "../ppocr/utils/dict/layout_dict/layout_cdla_dict.txt"
 								- recovery：whether to enable layout of recovery, default False
 								- save_pdf：when recovery file, whether to save pdf file, default False
 								- output：save the recovery result path