• Stars
    star
    856
  • Rank 53,268 (Top 2 %)
  • Language
    Python
  • Created over 4 years ago
  • Updated 12 months ago

Reviews

There are no reviews yet. Be the first to send feedback to the community and the maintainers!

Repository Details

收集并整理有关OCR的数据集并统一标注格式,以便实验需要

Todo

  • 提供数据集百度云链接
  • 数据集转换为统一格式(检测和识别)
    • icdar2015
    • MLT2019
    • COCO-Text_v2
    • ReCTS
    • SROIE
    • ArT
    • LSVT
    • Synth800k
    • icdar2017rctw
    • MTWI 2018
    • 百度中文场景文字识别
    • mjsynth
    • Synthetic Chinese String Dataset(360万中文数据集)
    • 英文识别数据大礼包
  • 提供读取脚本

下载

下载数据集之后,记得修改标注文件里对应的路径为自己的路径 百度云 提取码:9s4x

数据集

数据集 主页 适用情况 数据情况 标注形式 说明
ICDAR2015 https://rrc.cvc.uab.es/?ch=4 检测&识别 语言: 英文 train:1,000 test:500 x1, y1, x2, y2, x3, y3, x4, y4, transcription 坐标: x1, y1, x2, y2, x3, y3, x4, y4 transcription : 框内的文字信息
MLT2019 https://rrc.cvc.uab.es/?ch=15 检测&识别 语言: 混合 train:10,000 test:10,000 x1,y1,x2,y2,x3,y3,x4,y4,script,transcription 坐标: x1, y1, x2, y2, x3, y3, x4, y4 script: 文字所属语言 transcription : 框内的文字信息
COCO-Text_v2 https://bgshih.github.io/cocotext/ 检测&识别 语言: 混合 train:43,686 validation:10,000 test:10,000 json
ReCTS https://rrc.cvc.uab.es/?ch=12&com=introduction 检测&识别 语言: 混合 train:20,000 test:5,000 { “chars”: [ {“points”: [x1,y1,x2,y2,x3,y3,x4,y4], “transcription” : “trans1”, "ignore":0 }, {“points”: [x1,y1,x2,y2,x3,y3,x4,y4], “transcription” : “trans2”, " ignore ":0 }], “lines”: [ {“points”: [x1,y1,x2,y2,x3,y3,x4,y4] , “transcription” : “trans3”, "ignore ":0 }], } points: x1,y1,x2,y2,x3,y3,x4,y4 chars: 字符级别的标注 lines: 行级别的标注. transcription : 框内的文字信息 ignore: 0:不忽略,1:忽略
SROIE https://rrc.cvc.uab.es/?ch=13 检测&识别 语言: 英文 train:699 test:400 x1, y1, x2, y2, x3, y3, x4, y4, transcription 坐标: x1, y1, x2, y2, x3, y3, x4, y4 transcription : 框内的文字信息
ArT(已包含Total-Text和SCUT-CTW1500) https://rrc.cvc.uab.es/?ch=14 检测&识别 语言: 混合 train: 5,603 test: 4,563 { “gt_1”: [ {“points”: [[x1, y1], [x2, y2], …, [xn, yn]], “transcription” : “trans1”, “language” : “Latin”, "illegibility": false }, {“points”: [[x1, y1], [x2, y2], …, [xn, yn]], “transcription” : “trans2”, “language” : “Chinese”, "illegibility": false }], } points: x1,y1,x2,y2,x3,y3,x4,y4…xn,yn transcription : 框内的文字信息 language: 语言信息 illegibility: 是否模糊
LSVT https://rrc.cvc.uab.es/?ch=16 检测&识别 语言: 混合 全标注 train: 30,000 test: 20,000 只标注文本 400,000 { “gt_1”: [ {“points”: [[x1, y1], [x2, y2], …, [xn, yn]], “transcription” : “trans1”, "illegibility": false }, {“points”: [[x1, y1], [x2, y2], …, [xn, yn]], “transcription” : “trans2”, "illegibility": false }], } points: x1,y1,x2,y2,x3,y3,x4,y4…xn,yn transcription : 框内的文字信息 illegibility: 是否模糊
Synth800k http://www.robots.ox.ac.uk/~vgg/data/scenetext/ 检测&识别 语言: 英文 800,000 imnames: wordBB: charBB: txt: imnames: 文件名称 wordBB: 24n,每张图像内的文本框 charBB: 24n,每张图像内的字符框 txt: 每张图形内的字符串
icdar2017rctw https://blog.csdn.net/wl1710582732/article/details/89761818 检测&识别 语言: 混合 train:8,034 test:4,229 x1,y1,x2,y2,x3,y3,x4,y4,<识别难易程度>,transcription 坐标: x1, y1, x2, y2, x3, y3, x4, y4 transcription : 框内的文字信息
MTWI 2018 识别: https://tianchi.aliyun.com/competition/entrance/231684/introduction 检测: https://tianchi.aliyun.com/competition/entrance/231685/introduction 检测&识别 语言: 混合 train:10,000 test:10,000 x1, y1, x2, y2, x3, y3, x4, y4, transcription 坐标: x1, y1, x2, y2, x3, y3, x4, y4 transcription : 框内的文字信息
百度中文场景文字识别 https://aistudio.baidu.com/aistudio/competition/detail/20 识别 语言: 混合 train:未统计 test:未统计 h,w,name,value h: 图片高度 w: 图片宽度 name: 图片名 value: 图片上文字
mjsynth http://www.robots.ox.ac.uk/~vgg/data/text/ 识别 语言: 英文 9,000,000 - -
Synthetic Chinese String Dataset(360万中文数据集) 链接:https://pan.baidu.com/s/1jefn4Jh4jHjQdiWoanjKpQ 提取码:spyi 识别 语言: 混合 300k - -
英文识别数据大礼包(https://github.com/clovaai/deep-text-recognition-benchmark) 训练:MJSynth和SynthText 验证:IIIT, SVT, IC03, IC13, IC15, SVTP, CUTE 链接:https://pan.baidu.com/s/1KSNLv4EY3zFWHpBYlpFCBQ 提取码:rryk 识别 语言: 英文 - -

数据生成工具

https://github.com/TianzhongSong/awesome-SynthText

数据集读取脚本

More Repositories

1

PytorchOCR

基于Pytorch的OCR工具库,支持常用的文字检测和识别算法
Python
1,345
star
2

DBNet.pytorch

A pytorch re-implementation of Real-time Scene Text Detection with Differentiable Binarization
Python
939
star
3

PSENet.pytorch

A pytorch re-implementation of PSENet: Shape Robust Text Detection with Progressive Scale Expansion Network
C++
462
star
4

PAN.pytorch

A unofficial pytorch implementation of PAN(PSENet2): Efficient and Accurate Arbitrary-Shaped Text Detection with Pixel Aggregation Network
C++
413
star
5

TableGeneration

通过浏览器渲染生成表格图像
Python
185
star
6

flask_pytorch

using flask to run pytorch model
Python
48
star
7

crnn.gluon

A gluon re-implementation of Convolutional recurrent network in gluon
Python
21
star
8

reprod_log

Python
16
star
9

Segmentation-Free_OCR

recognize chinese and english without segmentation
Python
11
star
10

Torch_Quant_Demo

一个使用torch进行量化训练的demo
Python
9
star
11

ctpn.pytorch

Python
9
star
12

crypto

Python
7
star
13

dl_docker

用于深度学习的docker环境,cuda支持cuda10.1和cuda10.2,框架支持各种框架
Dockerfile
6
star
14

IcdarToCOCO

Python
5
star
15

gluon_mnist

learning gluon with mnist dataset
Python
5
star
16

mxnet_cifar10

Python
4
star
17

crnn.paddle

Python
4
star
18

leetcode

learning data struct with python
Jupyter Notebook
4
star
19

UCDIR.paddle

Python
4
star
20

TableMASTER_mmocr

Python
3
star
21

rust_python

use rust speed up python
Rust
3
star
22

pytorch_mnist

learning pytorch with mnist dataset
Python
3
star
23

WenmuZhou.github.io

个人博客
HTML
2
star
24

keras_mnist

learning keras with mnist
Python
2
star
25

gitment-comments

2
star
26

DABNet_Paddle

a paddle reproduce of DABNet
Python
1
star
27

simple_nlp

Python
1
star