Repository files navigation

MRN: Multiplexed Routing Network
for Incremental Multilingual Text Recognition

ICCV 2023ArXiv preprintBlogLICENSE

Method |IMLTR Dataset | Getting Started | Citation

It started as code for the paper:

MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition (Accepted by ICCV 2023)

This project is a toolkit for the novel scenario of Incremental Multilingual Text Recognition (IMLTR), the project supports many incremental learning methods and proposes a more applicable method for IMLTR: Multiplexed Routing Network (MRN) and the corresponding dataset. The project provides an efficient framework to assist in developing new methods and analyzing existing ones under the IMLTR task, and we hope it will advance the IMLTR community.

image

Methods

Incremental Learning Methods

  • Base: Baseline method which simply updates parameters on new tasks.
  • Joint: Bound method: data for all tasks are trained at once, an upper bound for the method
    (Joint_mix means all tasks data mixed in batch, Joint_loader means the consistent proportion of data from each task in a batch)
  • EWC[PNAS2017]: Overcoming catastrophic forgetting in neural networks
  • LwF[ECCV2016]: Learning without Forgetting
  • WA[CVPR2020]: Maintaining Discrimination and Fairness in Class Incremental Learning
  • DER[CVPR2021]: DER: Dynamically Expandable Representation for Class Incremental Learning
  • MRN[ICCV2023]: MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition

you can change config config/crnn_mrn.py for different il methods or setting.

common=dict(
il="mrn", # joint_mix | joint_loader | base | lwf | wa | ewc | der | mrn
memory="random", # None | random
memory_num=2000,
start_task = 0 # checkpoint start
)

Text Recognition Methods

  • CRNN[TPAMI2017]: An End-to-End Trainable Neural Network for Image-Based Sequence Recognition and Its Application to Scene Text Recognition
  • TRBA[ICCV2019]: What Is Wrong With Scene Text Recognition Model Comparisons? Dataset and Model Analysis
  • SVTR[IJCAI2022]: SVTR: Scene Text Recognition with a Single Visual Model

you can change config config/crnn_mrn.py for different text recognition modules or setting.

""" Model Architecture """
common=dict(
batch_max_length = 25,
imgH = 32,
imgW = 256,
)
model=dict(
model_name="TRBA",
Transformation = "TPS", #None TPS
FeatureExtraction = "ResNet", #VGG ResNet SVTR
SequenceModeling = "BiLSTM", #None BiLSTM
Prediction = "Attn", #CTC Attn
num_fiducial=20,
input_channel=4,
output_channel=512,
hidden_size=256,
)

IMLTR Dataset

The Dataset can be downloaded from BaiduNetdisk(passwd:c07h).

dataset
├── MLT17_IL
│ ├── test_2017
│ ├── train_2017
├── MLT19_IL
│ ├── test_2019
│ ├── train_2019

Incremental MLT17: MLT17 has 68,613 training instances and 16,255 validation instances, which are from 6 scripts and 9 languages: Chinese, Japanese, Korean, Bangla, Arabic, Italian, English, French, and German. The last four use Latin script. Incremental MLT17 use the validation set for test due to the unavailability of test data. Tasks are split by scripts and modeled sequentially. Special symbols are discarded at the preprocessing step as with no linguistic meaning.

Incremental MLT19: MLT19 has 89,177 text instances coming from 7 scripts. Since the inaccessibility of test set, Incremental MLT19 randomly split the training instances to 9:1 script-by-script, for model training and test. To be consistent with Incremental MLT17 dataset, we discard the Hindi script and also special symbols. Statistics of the two datasets are shown in the following.

DatasetCategories
Task1Task2Task3Task4Task5Task6
ChineseLatinJapaneseKoreanArabicBangla
MLT171Train Instance2687474114609563137113237
Test Instance5291107313501230983713
Train Class18953251620112473112
MLT192Train Instance2897529215324610742303542
Test Instance3225882590679470393
Train Class20862201728116073102

Getting Started

Dependency

  • This work was tested with PyTorch 1.6.0, CUDA 10.1 and python 3.6.
conda create -n mrn python=3.7 -y
conda activate mrn
conda install pytorch==1.9.1 torchvision==0.10.1 torchaudio==0.9.1 cudatoolkit=11.3 -c pytorch -c conda-forge
pip install torch==1.9.1+cu111 torchvision==0.10.1+cu111 torchaudio==0.9.1 -f https://download.pytorch.org/whl/torch_stable.html
  • requirements :
pip3 install lmdb pillow torchvision nltk natsort fire tensorboard tqdm opencv-python einops timm mmcv shapely scipy
pip3 install mmcv-full -f https://download.openmmlab.com/mmcv/dist/cu111/torch1.9.1/index.html

Training

python3 tiny_train.py --config=config/crnn_mrn.py --exp_name CRNN_real

Arguments

tiny_train.py (as a default, evaluate trained model on IMLTR datasets at the end of training.

  • --select_data: folder path to training lmdb datasets.
    [" ../dataset/MLT17_IL/train_2017", "../dataset/MLT19_IL/train_2019"]
  • --valid_datas: folder path to testing lmdb dataset.
    [" ../dataset/MLT17_IL/test_2017", "../dataset/MLT19_IL/test_2019"]
  • --batch_ratio: assign ratio for each selected data in the batch. default is '1 / number of datasets'.
  • --Aug: whether to use augmentation |None|Blur|Crop|Rot|

Config Detail

For detailed configuration modifications please use the config file config/crnn_mrn.py

common=dict(
exp_name="TRBA_MRN", # Where to store logs and models
il="mrn", # joint_mix | joint_loader | base | lwf | wa | ewc | der | mrn
memory="random", # None | random
memory_num=2000,
batch_max_length = 25,
imgH = 32,
imgW = 256,
manual_seed=111,
start_task = 0
)
""" Model Architecture """
model=dict(
model_name="TRBA",
Transformation = "TPS", #None TPS
FeatureExtraction = "ResNet", #VGG ResNet
SequenceModeling = "BiLSTM", #None BiLSTM
Prediction = "Attn", #CTC Attn
num_fiducial=20,
input_channel=4,
output_channel=512,
hidden_size=256,
)
""" Optimizer """
optimizer=dict(
schedule="super", #default is super for super convergence, 1 for None, [0.6, 0.8] for the same setting with ASTER
optimizer="adam",
lr=0.0005,
sgd_momentum=0.9,
sgd_weight_decay=0.000001,
milestones=[2000,4000],
lrate_decay=0.1,
rho=0.95,
eps=1e-8,
lr_drop_rate=0.1
)
""" Data processing """
train = dict(
saved_model="", # "path to model to continue training"
Aug="None", # |None|Blur|Crop|Rot|ABINet
workers=4,
lan_list=["Chinese","Latin","Japanese", "Korean", "Arabic", "Bangla"],
valid_datas=[
"../dataset/MLT17_IL/test_2017",
"../dataset/MLT19_IL/test_2019"
],
select_data=[
"../dataset/MLT17_IL/train_2017",
"../dataset/MLT19_IL/train_2019"
],
batch_ratio="0.5-0.5",
total_data_usage_ratio="1.0",
NED=True,
batch_size=256,
num_iter=10000,
val_interval=5000,
log_multiple_test=None,
grad_clip=5,
)

Data Analysis

The experimental results of each task are recorded in data_any.txt and can be used for analysis of the data.

Acknowledgements

This implementation has been based on these repositories:

Citation

Please consider citing this work in your publications if it helps your research.

@article{zheng2023mrn,
title={MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition},
author={Zheng, Tianlun and Chen, Zhineng and Huang, BingChen and Zhang, Wei and Jiang, Yu-Gang},
journal={Proceedings of the IEEE/CVF International Conference on Computer Vision},
year={2023}
}

License

This project is released under the Apache 2.0 license.

Footnotes

  1. Nayef, N., et al. (2017). MLT 2017.

  2. Nayef, N., et al. (2019). MLT 2019.

About

Official Pytorch implementations of MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition (ICCV 2023)

Topics

Resources

Stars

46 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

MRN: Multiplexed Routing Network
for Incremental Multilingual Text Recognition

ICCV 2023ArXiv preprintBlogLICENSE

Method |IMLTR Dataset | Getting Started | Citation

It started as code for the paper:

MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition (Accepted by ICCV 2023)

This project is a toolkit for the novel scenario of Incremental Multilingual Text Recognition (IMLTR), the project supports many incremental learning methods and proposes a more applicable method for IMLTR: Multiplexed Routing Network (MRN) and the corresponding dataset. The project provides an efficient framework to assist in developing new methods and analyzing existing ones under the IMLTR task, and we hope it will advance the IMLTR community.

image

Methods

Incremental Learning Methods

  • Base: Baseline method which simply updates parameters on new tasks.
  • Joint: Bound method: data for all tasks are trained at once, an upper bound for the method
    (Joint_mix means all tasks data mixed in batch, Joint_loader means the consistent proportion of data from each task in a batch)
  • EWC[PNAS2017]: Overcoming catastrophic forgetting in neural networks
  • LwF[ECCV2016]: Learning without Forgetting
  • WA[CVPR2020]: Maintaining Discrimination and Fairness in Class Incremental Learning
  • DER[CVPR2021]: DER: Dynamically Expandable Representation for Class Incremental Learning
  • MRN[ICCV2023]: MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition

you can change config config/crnn_mrn.py for different il methods or setting.

common=dict(
il="mrn", # joint_mix | joint_loader | base | lwf | wa | ewc | der | mrn
memory="random", # None | random
memory_num=2000,
start_task = 0 # checkpoint start
)

Text Recognition Methods

  • CRNN[TPAMI2017]: An End-to-End Trainable Neural Network for Image-Based Sequence Recognition and Its Application to Scene Text Recognition
  • TRBA[ICCV2019]: What Is Wrong With Scene Text Recognition Model Comparisons? Dataset and Model Analysis
  • SVTR[IJCAI2022]: SVTR: Scene Text Recognition with a Single Visual Model

you can change config config/crnn_mrn.py for different text recognition modules or setting.

""" Model Architecture """
common=dict(
batch_max_length = 25,
imgH = 32,
imgW = 256,
)
model=dict(
model_name="TRBA",
Transformation = "TPS", #None TPS
FeatureExtraction = "ResNet", #VGG ResNet SVTR
SequenceModeling = "BiLSTM", #None BiLSTM
Prediction = "Attn", #CTC Attn
num_fiducial=20,
input_channel=4,
output_channel=512,
hidden_size=256,
)

IMLTR Dataset

The Dataset can be downloaded from BaiduNetdisk(passwd:c07h).

dataset
├── MLT17_IL
│ ├── test_2017
│ ├── train_2017
├── MLT19_IL
│ ├── test_2019
│ ├── train_2019

Incremental MLT17: MLT17 has 68,613 training instances and 16,255 validation instances, which are from 6 scripts and 9 languages: Chinese, Japanese, Korean, Bangla, Arabic, Italian, English, French, and German. The last four use Latin script. Incremental MLT17 use the validation set for test due to the unavailability of test data. Tasks are split by scripts and modeled sequentially. Special symbols are discarded at the preprocessing step as with no linguistic meaning.

Incremental MLT19: MLT19 has 89,177 text instances coming from 7 scripts. Since the inaccessibility of test set, Incremental MLT19 randomly split the training instances to 9:1 script-by-script, for model training and test. To be consistent with Incremental MLT17 dataset, we discard the Hindi script and also special symbols. Statistics of the two datasets are shown in the following.

DatasetCategories
Task1Task2Task3Task4Task5Task6
ChineseLatinJapaneseKoreanArabicBangla
MLT171Train Instance2687474114609563137113237
Test Instance5291107313501230983713
Train Class18953251620112473112
MLT192Train Instance2897529215324610742303542
Test Instance3225882590679470393
Train Class20862201728116073102

Getting Started

Dependency

  • This work was tested with PyTorch 1.6.0, CUDA 10.1 and python 3.6.
conda create -n mrn python=3.7 -y
conda activate mrn
conda install pytorch==1.9.1 torchvision==0.10.1 torchaudio==0.9.1 cudatoolkit=11.3 -c pytorch -c conda-forge
pip install torch==1.9.1+cu111 torchvision==0.10.1+cu111 torchaudio==0.9.1 -f https://download.pytorch.org/whl/torch_stable.html
  • requirements :
pip3 install lmdb pillow torchvision nltk natsort fire tensorboard tqdm opencv-python einops timm mmcv shapely scipy
pip3 install mmcv-full -f https://download.openmmlab.com/mmcv/dist/cu111/torch1.9.1/index.html

Training

python3 tiny_train.py --config=config/crnn_mrn.py --exp_name CRNN_real

Arguments

tiny_train.py (as a default, evaluate trained model on IMLTR datasets at the end of training.

  • --select_data: folder path to training lmdb datasets.
    [" ../dataset/MLT17_IL/train_2017", "../dataset/MLT19_IL/train_2019"]
  • --valid_datas: folder path to testing lmdb dataset.
    [" ../dataset/MLT17_IL/test_2017", "../dataset/MLT19_IL/test_2019"]
  • --batch_ratio: assign ratio for each selected data in the batch. default is '1 / number of datasets'.
  • --Aug: whether to use augmentation |None|Blur|Crop|Rot|

Config Detail

For detailed configuration modifications please use the config file config/crnn_mrn.py

common=dict(
exp_name="TRBA_MRN", # Where to store logs and models
il="mrn", # joint_mix | joint_loader | base | lwf | wa | ewc | der | mrn
memory="random", # None | random
memory_num=2000,
batch_max_length = 25,
imgH = 32,
imgW = 256,
manual_seed=111,
start_task = 0
)
""" Model Architecture """
model=dict(
model_name="TRBA",
Transformation = "TPS", #None TPS
FeatureExtraction = "ResNet", #VGG ResNet
SequenceModeling = "BiLSTM", #None BiLSTM
Prediction = "Attn", #CTC Attn
num_fiducial=20,
input_channel=4,
output_channel=512,
hidden_size=256,
)
""" Optimizer """
optimizer=dict(
schedule="super", #default is super for super convergence, 1 for None, [0.6, 0.8] for the same setting with ASTER
optimizer="adam",
lr=0.0005,
sgd_momentum=0.9,
sgd_weight_decay=0.000001,
milestones=[2000,4000],
lrate_decay=0.1,
rho=0.95,
eps=1e-8,
lr_drop_rate=0.1
)
""" Data processing """
train = dict(
saved_model="", # "path to model to continue training"
Aug="None", # |None|Blur|Crop|Rot|ABINet
workers=4,
lan_list=["Chinese","Latin","Japanese", "Korean", "Arabic", "Bangla"],
valid_datas=[
"../dataset/MLT17_IL/test_2017",
"../dataset/MLT19_IL/test_2019"
],
select_data=[
"../dataset/MLT17_IL/train_2017",
"../dataset/MLT19_IL/train_2019"
],
batch_ratio="0.5-0.5",
total_data_usage_ratio="1.0",
NED=True,
batch_size=256,
num_iter=10000,
val_interval=5000,
log_multiple_test=None,
grad_clip=5,
)

Data Analysis

The experimental results of each task are recorded in data_any.txt and can be used for analysis of the data.

Acknowledgements

This implementation has been based on these repositories:

Citation

Please consider citing this work in your publications if it helps your research.

@article{zheng2023mrn,
title={MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition},
author={Zheng, Tianlun and Chen, Zhineng and Huang, BingChen and Zhang, Wei and Jiang, Yu-Gang},
journal={Proceedings of the IEEE/CVF International Conference on Computer Vision},
year={2023}
}

License

This project is released under the Apache 2.0 license.

Footnotes

  1. Nayef, N., et al. (2017). MLT 2017.

  2. Nayef, N., et al. (2019). MLT 2019.

About

Official Pytorch implementations of MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition (ICCV 2023)

Topics

Resources

Stars

46 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

MRN: Multiplexed Routing Network
for Incremental Multilingual Text Recognition

ICCV 2023ArXiv preprintBlogLICENSE

Method |IMLTR Dataset | Getting Started | Citation

It started as code for the paper:

MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition (Accepted by ICCV 2023)

This project is a toolkit for the novel scenario of Incremental Multilingual Text Recognition (IMLTR), the project supports many incremental learning methods and proposes a more applicable method for IMLTR: Multiplexed Routing Network (MRN) and the corresponding dataset. The project provides an efficient framework to assist in developing new methods and analyzing existing ones under the IMLTR task, and we hope it will advance the IMLTR community.

image

Methods

Incremental Learning Methods

  • Base: Baseline method which simply updates parameters on new tasks.
  • Joint: Bound method: data for all tasks are trained at once, an upper bound for the method
    (Joint_mix means all tasks data mixed in batch, Joint_loader means the consistent proportion of data from each task in a batch)
  • EWC[PNAS2017]: Overcoming catastrophic forgetting in neural networks
  • LwF[ECCV2016]: Learning without Forgetting
  • WA[CVPR2020]: Maintaining Discrimination and Fairness in Class Incremental Learning
  • DER[CVPR2021]: DER: Dynamically Expandable Representation for Class Incremental Learning
  • MRN[ICCV2023]: MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition

you can change config config/crnn_mrn.py for different il methods or setting.

common=dict(
il="mrn", # joint_mix | joint_loader | base | lwf | wa | ewc | der | mrn
memory="random", # None | random
memory_num=2000,
start_task = 0 # checkpoint start
)

Text Recognition Methods

  • CRNN[TPAMI2017]: An End-to-End Trainable Neural Network for Image-Based Sequence Recognition and Its Application to Scene Text Recognition
  • TRBA[ICCV2019]: What Is Wrong With Scene Text Recognition Model Comparisons? Dataset and Model Analysis
  • SVTR[IJCAI2022]: SVTR: Scene Text Recognition with a Single Visual Model

you can change config config/crnn_mrn.py for different text recognition modules or setting.

""" Model Architecture """
common=dict(
batch_max_length = 25,
imgH = 32,
imgW = 256,
)
model=dict(
model_name="TRBA",
Transformation = "TPS", #None TPS
FeatureExtraction = "ResNet", #VGG ResNet SVTR
SequenceModeling = "BiLSTM", #None BiLSTM
Prediction = "Attn", #CTC Attn
num_fiducial=20,
input_channel=4,
output_channel=512,
hidden_size=256,
)

IMLTR Dataset

The Dataset can be downloaded from BaiduNetdisk(passwd:c07h).

dataset
├── MLT17_IL
│ ├── test_2017
│ ├── train_2017
├── MLT19_IL
│ ├── test_2019
│ ├── train_2019

Incremental MLT17: MLT17 has 68,613 training instances and 16,255 validation instances, which are from 6 scripts and 9 languages: Chinese, Japanese, Korean, Bangla, Arabic, Italian, English, French, and German. The last four use Latin script. Incremental MLT17 use the validation set for test due to the unavailability of test data. Tasks are split by scripts and modeled sequentially. Special symbols are discarded at the preprocessing step as with no linguistic meaning.

Incremental MLT19: MLT19 has 89,177 text instances coming from 7 scripts. Since the inaccessibility of test set, Incremental MLT19 randomly split the training instances to 9:1 script-by-script, for model training and test. To be consistent with Incremental MLT17 dataset, we discard the Hindi script and also special symbols. Statistics of the two datasets are shown in the following.

DatasetCategories
Task1Task2Task3Task4Task5Task6
ChineseLatinJapaneseKoreanArabicBangla
MLT171Train Instance2687474114609563137113237
Test Instance5291107313501230983713
Train Class18953251620112473112
MLT192Train Instance2897529215324610742303542
Test Instance3225882590679470393
Train Class20862201728116073102

Getting Started

Dependency

  • This work was tested with PyTorch 1.6.0, CUDA 10.1 and python 3.6.
conda create -n mrn python=3.7 -y
conda activate mrn
conda install pytorch==1.9.1 torchvision==0.10.1 torchaudio==0.9.1 cudatoolkit=11.3 -c pytorch -c conda-forge
pip install torch==1.9.1+cu111 torchvision==0.10.1+cu111 torchaudio==0.9.1 -f https://download.pytorch.org/whl/torch_stable.html
  • requirements :
pip3 install lmdb pillow torchvision nltk natsort fire tensorboard tqdm opencv-python einops timm mmcv shapely scipy
pip3 install mmcv-full -f https://download.openmmlab.com/mmcv/dist/cu111/torch1.9.1/index.html

Training

python3 tiny_train.py --config=config/crnn_mrn.py --exp_name CRNN_real

Arguments

tiny_train.py (as a default, evaluate trained model on IMLTR datasets at the end of training.

  • --select_data: folder path to training lmdb datasets.
    [" ../dataset/MLT17_IL/train_2017", "../dataset/MLT19_IL/train_2019"]
  • --valid_datas: folder path to testing lmdb dataset.
    [" ../dataset/MLT17_IL/test_2017", "../dataset/MLT19_IL/test_2019"]
  • --batch_ratio: assign ratio for each selected data in the batch. default is '1 / number of datasets'.
  • --Aug: whether to use augmentation |None|Blur|Crop|Rot|

Config Detail

For detailed configuration modifications please use the config file config/crnn_mrn.py

common=dict(
exp_name="TRBA_MRN", # Where to store logs and models
il="mrn", # joint_mix | joint_loader | base | lwf | wa | ewc | der | mrn
memory="random", # None | random
memory_num=2000,
batch_max_length = 25,
imgH = 32,
imgW = 256,
manual_seed=111,
start_task = 0
)
""" Model Architecture """
model=dict(
model_name="TRBA",
Transformation = "TPS", #None TPS
FeatureExtraction = "ResNet", #VGG ResNet
SequenceModeling = "BiLSTM", #None BiLSTM
Prediction = "Attn", #CTC Attn
num_fiducial=20,
input_channel=4,
output_channel=512,
hidden_size=256,
)
""" Optimizer """
optimizer=dict(
schedule="super", #default is super for super convergence, 1 for None, [0.6, 0.8] for the same setting with ASTER
optimizer="adam",
lr=0.0005,
sgd_momentum=0.9,
sgd_weight_decay=0.000001,
milestones=[2000,4000],
lrate_decay=0.1,
rho=0.95,
eps=1e-8,
lr_drop_rate=0.1
)
""" Data processing """
train = dict(
saved_model="", # "path to model to continue training"
Aug="None", # |None|Blur|Crop|Rot|ABINet
workers=4,
lan_list=["Chinese","Latin","Japanese", "Korean", "Arabic", "Bangla"],
valid_datas=[
"../dataset/MLT17_IL/test_2017",
"../dataset/MLT19_IL/test_2019"
],
select_data=[
"../dataset/MLT17_IL/train_2017",
"../dataset/MLT19_IL/train_2019"
],
batch_ratio="0.5-0.5",
total_data_usage_ratio="1.0",
NED=True,
batch_size=256,
num_iter=10000,
val_interval=5000,
log_multiple_test=None,
grad_clip=5,
)

Data Analysis

The experimental results of each task are recorded in data_any.txt and can be used for analysis of the data.

Acknowledgements

This implementation has been based on these repositories:

Citation

Please consider citing this work in your publications if it helps your research.

@article{zheng2023mrn,
title={MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition},
author={Zheng, Tianlun and Chen, Zhineng and Huang, BingChen and Zhang, Wei and Jiang, Yu-Gang},
journal={Proceedings of the IEEE/CVF International Conference on Computer Vision},
year={2023}
}

License

This project is released under the Apache 2.0 license.

Footnotes

  1. Nayef, N., et al. (2017). MLT 2017.

  2. Nayef, N., et al. (2019). MLT 2019.

About

Official Pytorch implementations of MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition (ICCV 2023)

Topics

Resources

Stars

46 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

MRN: Multiplexed Routing Network
for Incremental Multilingual Text Recognition

ICCV 2023ArXiv preprintBlogLICENSE

Method |IMLTR Dataset | Getting Started | Citation

It started as code for the paper:

MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition (Accepted by ICCV 2023)

This project is a toolkit for the novel scenario of Incremental Multilingual Text Recognition (IMLTR), the project supports many incremental learning methods and proposes a more applicable method for IMLTR: Multiplexed Routing Network (MRN) and the corresponding dataset. The project provides an efficient framework to assist in developing new methods and analyzing existing ones under the IMLTR task, and we hope it will advance the IMLTR community.

image

Methods

Incremental Learning Methods

  • Base: Baseline method which simply updates parameters on new tasks.
  • Joint: Bound method: data for all tasks are trained at once, an upper bound for the method
    (Joint_mix means all tasks data mixed in batch, Joint_loader means the consistent proportion of data from each task in a batch)
  • EWC[PNAS2017]: Overcoming catastrophic forgetting in neural networks
  • LwF[ECCV2016]: Learning without Forgetting
  • WA[CVPR2020]: Maintaining Discrimination and Fairness in Class Incremental Learning
  • DER[CVPR2021]: DER: Dynamically Expandable Representation for Class Incremental Learning
  • MRN[ICCV2023]: MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition

you can change config config/crnn_mrn.py for different il methods or setting.

common=dict(
il="mrn", # joint_mix | joint_loader | base | lwf | wa | ewc | der | mrn
memory="random", # None | random
memory_num=2000,
start_task = 0 # checkpoint start
)

Text Recognition Methods

  • CRNN[TPAMI2017]: An End-to-End Trainable Neural Network for Image-Based Sequence Recognition and Its Application to Scene Text Recognition
  • TRBA[ICCV2019]: What Is Wrong With Scene Text Recognition Model Comparisons? Dataset and Model Analysis
  • SVTR[IJCAI2022]: SVTR: Scene Text Recognition with a Single Visual Model

you can change config config/crnn_mrn.py for different text recognition modules or setting.

""" Model Architecture """
common=dict(
batch_max_length = 25,
imgH = 32,
imgW = 256,
)
model=dict(
model_name="TRBA",
Transformation = "TPS", #None TPS
FeatureExtraction = "ResNet", #VGG ResNet SVTR
SequenceModeling = "BiLSTM", #None BiLSTM
Prediction = "Attn", #CTC Attn
num_fiducial=20,
input_channel=4,
output_channel=512,
hidden_size=256,
)

IMLTR Dataset

The Dataset can be downloaded from BaiduNetdisk(passwd:c07h).

dataset
├── MLT17_IL
│ ├── test_2017
│ ├── train_2017
├── MLT19_IL
│ ├── test_2019
│ ├── train_2019

Incremental MLT17: MLT17 has 68,613 training instances and 16,255 validation instances, which are from 6 scripts and 9 languages: Chinese, Japanese, Korean, Bangla, Arabic, Italian, English, French, and German. The last four use Latin script. Incremental MLT17 use the validation set for test due to the unavailability of test data. Tasks are split by scripts and modeled sequentially. Special symbols are discarded at the preprocessing step as with no linguistic meaning.

Incremental MLT19: MLT19 has 89,177 text instances coming from 7 scripts. Since the inaccessibility of test set, Incremental MLT19 randomly split the training instances to 9:1 script-by-script, for model training and test. To be consistent with Incremental MLT17 dataset, we discard the Hindi script and also special symbols. Statistics of the two datasets are shown in the following.

DatasetCategories
Task1Task2Task3Task4Task5Task6
ChineseLatinJapaneseKoreanArabicBangla
MLT171Train Instance2687474114609563137113237
Test Instance5291107313501230983713
Train Class18953251620112473112
MLT192Train Instance2897529215324610742303542
Test Instance3225882590679470393
Train Class20862201728116073102

Getting Started

Dependency

  • This work was tested with PyTorch 1.6.0, CUDA 10.1 and python 3.6.
conda create -n mrn python=3.7 -y
conda activate mrn
conda install pytorch==1.9.1 torchvision==0.10.1 torchaudio==0.9.1 cudatoolkit=11.3 -c pytorch -c conda-forge
pip install torch==1.9.1+cu111 torchvision==0.10.1+cu111 torchaudio==0.9.1 -f https://download.pytorch.org/whl/torch_stable.html
  • requirements :
pip3 install lmdb pillow torchvision nltk natsort fire tensorboard tqdm opencv-python einops timm mmcv shapely scipy
pip3 install mmcv-full -f https://download.openmmlab.com/mmcv/dist/cu111/torch1.9.1/index.html

Training

python3 tiny_train.py --config=config/crnn_mrn.py --exp_name CRNN_real

Arguments

tiny_train.py (as a default, evaluate trained model on IMLTR datasets at the end of training.

  • --select_data: folder path to training lmdb datasets.
    [" ../dataset/MLT17_IL/train_2017", "../dataset/MLT19_IL/train_2019"]
  • --valid_datas: folder path to testing lmdb dataset.
    [" ../dataset/MLT17_IL/test_2017", "../dataset/MLT19_IL/test_2019"]
  • --batch_ratio: assign ratio for each selected data in the batch. default is '1 / number of datasets'.
  • --Aug: whether to use augmentation |None|Blur|Crop|Rot|

Config Detail

For detailed configuration modifications please use the config file config/crnn_mrn.py

common=dict(
exp_name="TRBA_MRN", # Where to store logs and models
il="mrn", # joint_mix | joint_loader | base | lwf | wa | ewc | der | mrn
memory="random", # None | random
memory_num=2000,
batch_max_length = 25,
imgH = 32,
imgW = 256,
manual_seed=111,
start_task = 0
)
""" Model Architecture """
model=dict(
model_name="TRBA",
Transformation = "TPS", #None TPS
FeatureExtraction = "ResNet", #VGG ResNet
SequenceModeling = "BiLSTM", #None BiLSTM
Prediction = "Attn", #CTC Attn
num_fiducial=20,
input_channel=4,
output_channel=512,
hidden_size=256,
)
""" Optimizer """
optimizer=dict(
schedule="super", #default is super for super convergence, 1 for None, [0.6, 0.8] for the same setting with ASTER
optimizer="adam",
lr=0.0005,
sgd_momentum=0.9,
sgd_weight_decay=0.000001,
milestones=[2000,4000],
lrate_decay=0.1,
rho=0.95,
eps=1e-8,
lr_drop_rate=0.1
)
""" Data processing """
train = dict(
saved_model="", # "path to model to continue training"
Aug="None", # |None|Blur|Crop|Rot|ABINet
workers=4,
lan_list=["Chinese","Latin","Japanese", "Korean", "Arabic", "Bangla"],
valid_datas=[
"../dataset/MLT17_IL/test_2017",
"../dataset/MLT19_IL/test_2019"
],
select_data=[
"../dataset/MLT17_IL/train_2017",
"../dataset/MLT19_IL/train_2019"
],
batch_ratio="0.5-0.5",
total_data_usage_ratio="1.0",
NED=True,
batch_size=256,
num_iter=10000,
val_interval=5000,
log_multiple_test=None,
grad_clip=5,
)

Data Analysis

The experimental results of each task are recorded in data_any.txt and can be used for analysis of the data.

Acknowledgements

This implementation has been based on these repositories:

Citation

Please consider citing this work in your publications if it helps your research.

@article{zheng2023mrn,
title={MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition},
author={Zheng, Tianlun and Chen, Zhineng and Huang, BingChen and Zhang, Wei and Jiang, Yu-Gang},
journal={Proceedings of the IEEE/CVF International Conference on Computer Vision},
year={2023}
}

License

This project is released under the Apache 2.0 license.

Footnotes

  1. Nayef, N., et al. (2017). MLT 2017.

  2. Nayef, N., et al. (2019). MLT 2019.

About

Official Pytorch implementations of MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition (ICCV 2023)

Topics

Resources

Stars

46 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

MRN: Multiplexed Routing Network
for Incremental Multilingual Text Recognition

ICCV 2023ArXiv preprintBlogLICENSE

Method |IMLTR Dataset | Getting Started | Citation

It started as code for the paper:

MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition (Accepted by ICCV 2023)

This project is a toolkit for the novel scenario of Incremental Multilingual Text Recognition (IMLTR), the project supports many incremental learning methods and proposes a more applicable method for IMLTR: Multiplexed Routing Network (MRN) and the corresponding dataset. The project provides an efficient framework to assist in developing new methods and analyzing existing ones under the IMLTR task, and we hope it will advance the IMLTR community.

image

Methods

Incremental Learning Methods

  • Base: Baseline method which simply updates parameters on new tasks.
  • Joint: Bound method: data for all tasks are trained at once, an upper bound for the method
    (Joint_mix means all tasks data mixed in batch, Joint_loader means the consistent proportion of data from each task in a batch)
  • EWC[PNAS2017]: Overcoming catastrophic forgetting in neural networks
  • LwF[ECCV2016]: Learning without Forgetting
  • WA[CVPR2020]: Maintaining Discrimination and Fairness in Class Incremental Learning
  • DER[CVPR2021]: DER: Dynamically Expandable Representation for Class Incremental Learning
  • MRN[ICCV2023]: MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition

you can change config config/crnn_mrn.py for different il methods or setting.

common=dict(
il="mrn", # joint_mix | joint_loader | base | lwf | wa | ewc | der | mrn
memory="random", # None | random
memory_num=2000,
start_task = 0 # checkpoint start
)

Text Recognition Methods

  • CRNN[TPAMI2017]: An End-to-End Trainable Neural Network for Image-Based Sequence Recognition and Its Application to Scene Text Recognition
  • TRBA[ICCV2019]: What Is Wrong With Scene Text Recognition Model Comparisons? Dataset and Model Analysis
  • SVTR[IJCAI2022]: SVTR: Scene Text Recognition with a Single Visual Model

you can change config config/crnn_mrn.py for different text recognition modules or setting.

""" Model Architecture """
common=dict(
batch_max_length = 25,
imgH = 32,
imgW = 256,
)
model=dict(
model_name="TRBA",
Transformation = "TPS", #None TPS
FeatureExtraction = "ResNet", #VGG ResNet SVTR
SequenceModeling = "BiLSTM", #None BiLSTM
Prediction = "Attn", #CTC Attn
num_fiducial=20,
input_channel=4,
output_channel=512,
hidden_size=256,
)

IMLTR Dataset

The Dataset can be downloaded from BaiduNetdisk(passwd:c07h).

dataset
├── MLT17_IL
│ ├── test_2017
│ ├── train_2017
├── MLT19_IL
│ ├── test_2019
│ ├── train_2019

Incremental MLT17: MLT17 has 68,613 training instances and 16,255 validation instances, which are from 6 scripts and 9 languages: Chinese, Japanese, Korean, Bangla, Arabic, Italian, English, French, and German. The last four use Latin script. Incremental MLT17 use the validation set for test due to the unavailability of test data. Tasks are split by scripts and modeled sequentially. Special symbols are discarded at the preprocessing step as with no linguistic meaning.

Incremental MLT19: MLT19 has 89,177 text instances coming from 7 scripts. Since the inaccessibility of test set, Incremental MLT19 randomly split the training instances to 9:1 script-by-script, for model training and test. To be consistent with Incremental MLT17 dataset, we discard the Hindi script and also special symbols. Statistics of the two datasets are shown in the following.

DatasetCategories
Task1Task2Task3Task4Task5Task6
ChineseLatinJapaneseKoreanArabicBangla
MLT171Train Instance2687474114609563137113237
Test Instance5291107313501230983713
Train Class18953251620112473112
MLT192Train Instance2897529215324610742303542
Test Instance3225882590679470393
Train Class20862201728116073102

Getting Started

Dependency

  • This work was tested with PyTorch 1.6.0, CUDA 10.1 and python 3.6.
conda create -n mrn python=3.7 -y
conda activate mrn
conda install pytorch==1.9.1 torchvision==0.10.1 torchaudio==0.9.1 cudatoolkit=11.3 -c pytorch -c conda-forge
pip install torch==1.9.1+cu111 torchvision==0.10.1+cu111 torchaudio==0.9.1 -f https://download.pytorch.org/whl/torch_stable.html
  • requirements :
pip3 install lmdb pillow torchvision nltk natsort fire tensorboard tqdm opencv-python einops timm mmcv shapely scipy
pip3 install mmcv-full -f https://download.openmmlab.com/mmcv/dist/cu111/torch1.9.1/index.html

Training

python3 tiny_train.py --config=config/crnn_mrn.py --exp_name CRNN_real

Arguments

tiny_train.py (as a default, evaluate trained model on IMLTR datasets at the end of training.

  • --select_data: folder path to training lmdb datasets.
    [" ../dataset/MLT17_IL/train_2017", "../dataset/MLT19_IL/train_2019"]
  • --valid_datas: folder path to testing lmdb dataset.
    [" ../dataset/MLT17_IL/test_2017", "../dataset/MLT19_IL/test_2019"]
  • --batch_ratio: assign ratio for each selected data in the batch. default is '1 / number of datasets'.
  • --Aug: whether to use augmentation |None|Blur|Crop|Rot|

Config Detail

For detailed configuration modifications please use the config file config/crnn_mrn.py

common=dict(
exp_name="TRBA_MRN", # Where to store logs and models
il="mrn", # joint_mix | joint_loader | base | lwf | wa | ewc | der | mrn
memory="random", # None | random
memory_num=2000,
batch_max_length = 25,
imgH = 32,
imgW = 256,
manual_seed=111,
start_task = 0
)
""" Model Architecture """
model=dict(
model_name="TRBA",
Transformation = "TPS", #None TPS
FeatureExtraction = "ResNet", #VGG ResNet
SequenceModeling = "BiLSTM", #None BiLSTM
Prediction = "Attn", #CTC Attn
num_fiducial=20,
input_channel=4,
output_channel=512,
hidden_size=256,
)
""" Optimizer """
optimizer=dict(
schedule="super", #default is super for super convergence, 1 for None, [0.6, 0.8] for the same setting with ASTER
optimizer="adam",
lr=0.0005,
sgd_momentum=0.9,
sgd_weight_decay=0.000001,
milestones=[2000,4000],
lrate_decay=0.1,
rho=0.95,
eps=1e-8,
lr_drop_rate=0.1
)
""" Data processing """
train = dict(
saved_model="", # "path to model to continue training"
Aug="None", # |None|Blur|Crop|Rot|ABINet
workers=4,
lan_list=["Chinese","Latin","Japanese", "Korean", "Arabic", "Bangla"],
valid_datas=[
"../dataset/MLT17_IL/test_2017",
"../dataset/MLT19_IL/test_2019"
],
select_data=[
"../dataset/MLT17_IL/train_2017",
"../dataset/MLT19_IL/train_2019"
],
batch_ratio="0.5-0.5",
total_data_usage_ratio="1.0",
NED=True,
batch_size=256,
num_iter=10000,
val_interval=5000,
log_multiple_test=None,
grad_clip=5,
)

Data Analysis

The experimental results of each task are recorded in data_any.txt and can be used for analysis of the data.

Acknowledgements

This implementation has been based on these repositories:

Citation

Please consider citing this work in your publications if it helps your research.

@article{zheng2023mrn,
title={MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition},
author={Zheng, Tianlun and Chen, Zhineng and Huang, BingChen and Zhang, Wei and Jiang, Yu-Gang},
journal={Proceedings of the IEEE/CVF International Conference on Computer Vision},
year={2023}
}

License

This project is released under the Apache 2.0 license.

Footnotes

  1. Nayef, N., et al. (2017). MLT 2017.

  2. Nayef, N., et al. (2019). MLT 2019.

About

Official Pytorch implementations of MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition (ICCV 2023)

Topics

Resources

Stars

46 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

MRN: Multiplexed Routing Network
for Incremental Multilingual Text Recognition

ICCV 2023ArXiv preprintBlogLICENSE

Method |IMLTR Dataset | Getting Started | Citation

It started as code for the paper:

MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition (Accepted by ICCV 2023)

This project is a toolkit for the novel scenario of Incremental Multilingual Text Recognition (IMLTR), the project supports many incremental learning methods and proposes a more applicable method for IMLTR: Multiplexed Routing Network (MRN) and the corresponding dataset. The project provides an efficient framework to assist in developing new methods and analyzing existing ones under the IMLTR task, and we hope it will advance the IMLTR community.

image

Methods

Incremental Learning Methods

  • Base: Baseline method which simply updates parameters on new tasks.
  • Joint: Bound method: data for all tasks are trained at once, an upper bound for the method
    (Joint_mix means all tasks data mixed in batch, Joint_loader means the consistent proportion of data from each task in a batch)
  • EWC[PNAS2017]: Overcoming catastrophic forgetting in neural networks
  • LwF[ECCV2016]: Learning without Forgetting
  • WA[CVPR2020]: Maintaining Discrimination and Fairness in Class Incremental Learning
  • DER[CVPR2021]: DER: Dynamically Expandable Representation for Class Incremental Learning
  • MRN[ICCV2023]: MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition

you can change config config/crnn_mrn.py for different il methods or setting.

common=dict(
il="mrn", # joint_mix | joint_loader | base | lwf | wa | ewc | der | mrn
memory="random", # None | random
memory_num=2000,
start_task = 0 # checkpoint start
)

Text Recognition Methods

  • CRNN[TPAMI2017]: An End-to-End Trainable Neural Network for Image-Based Sequence Recognition and Its Application to Scene Text Recognition
  • TRBA[ICCV2019]: What Is Wrong With Scene Text Recognition Model Comparisons? Dataset and Model Analysis
  • SVTR[IJCAI2022]: SVTR: Scene Text Recognition with a Single Visual Model

you can change config config/crnn_mrn.py for different text recognition modules or setting.

""" Model Architecture """
common=dict(
batch_max_length = 25,
imgH = 32,
imgW = 256,
)
model=dict(
model_name="TRBA",
Transformation = "TPS", #None TPS
FeatureExtraction = "ResNet", #VGG ResNet SVTR
SequenceModeling = "BiLSTM", #None BiLSTM
Prediction = "Attn", #CTC Attn
num_fiducial=20,
input_channel=4,
output_channel=512,
hidden_size=256,
)

IMLTR Dataset

The Dataset can be downloaded from BaiduNetdisk(passwd:c07h).

dataset
├── MLT17_IL
│ ├── test_2017
│ ├── train_2017
├── MLT19_IL
│ ├── test_2019
│ ├── train_2019

Incremental MLT17: MLT17 has 68,613 training instances and 16,255 validation instances, which are from 6 scripts and 9 languages: Chinese, Japanese, Korean, Bangla, Arabic, Italian, English, French, and German. The last four use Latin script. Incremental MLT17 use the validation set for test due to the unavailability of test data. Tasks are split by scripts and modeled sequentially. Special symbols are discarded at the preprocessing step as with no linguistic meaning.

Incremental MLT19: MLT19 has 89,177 text instances coming from 7 scripts. Since the inaccessibility of test set, Incremental MLT19 randomly split the training instances to 9:1 script-by-script, for model training and test. To be consistent with Incremental MLT17 dataset, we discard the Hindi script and also special symbols. Statistics of the two datasets are shown in the following.

DatasetCategories
Task1Task2Task3Task4Task5Task6
ChineseLatinJapaneseKoreanArabicBangla
MLT171Train Instance2687474114609563137113237
Test Instance5291107313501230983713
Train Class18953251620112473112
MLT192Train Instance2897529215324610742303542
Test Instance3225882590679470393
Train Class20862201728116073102

Getting Started

Dependency

  • This work was tested with PyTorch 1.6.0, CUDA 10.1 and python 3.6.
conda create -n mrn python=3.7 -y
conda activate mrn
conda install pytorch==1.9.1 torchvision==0.10.1 torchaudio==0.9.1 cudatoolkit=11.3 -c pytorch -c conda-forge
pip install torch==1.9.1+cu111 torchvision==0.10.1+cu111 torchaudio==0.9.1 -f https://download.pytorch.org/whl/torch_stable.html
  • requirements :
pip3 install lmdb pillow torchvision nltk natsort fire tensorboard tqdm opencv-python einops timm mmcv shapely scipy
pip3 install mmcv-full -f https://download.openmmlab.com/mmcv/dist/cu111/torch1.9.1/index.html

Training

python3 tiny_train.py --config=config/crnn_mrn.py --exp_name CRNN_real

Arguments

tiny_train.py (as a default, evaluate trained model on IMLTR datasets at the end of training.

  • --select_data: folder path to training lmdb datasets.
    [" ../dataset/MLT17_IL/train_2017", "../dataset/MLT19_IL/train_2019"]
  • --valid_datas: folder path to testing lmdb dataset.
    [" ../dataset/MLT17_IL/test_2017", "../dataset/MLT19_IL/test_2019"]
  • --batch_ratio: assign ratio for each selected data in the batch. default is '1 / number of datasets'.
  • --Aug: whether to use augmentation |None|Blur|Crop|Rot|

Config Detail

For detailed configuration modifications please use the config file config/crnn_mrn.py

common=dict(
exp_name="TRBA_MRN", # Where to store logs and models
il="mrn", # joint_mix | joint_loader | base | lwf | wa | ewc | der | mrn
memory="random", # None | random
memory_num=2000,
batch_max_length = 25,
imgH = 32,
imgW = 256,
manual_seed=111,
start_task = 0
)
""" Model Architecture """
model=dict(
model_name="TRBA",
Transformation = "TPS", #None TPS
FeatureExtraction = "ResNet", #VGG ResNet
SequenceModeling = "BiLSTM", #None BiLSTM
Prediction = "Attn", #CTC Attn
num_fiducial=20,
input_channel=4,
output_channel=512,
hidden_size=256,
)
""" Optimizer """
optimizer=dict(
schedule="super", #default is super for super convergence, 1 for None, [0.6, 0.8] for the same setting with ASTER
optimizer="adam",
lr=0.0005,
sgd_momentum=0.9,
sgd_weight_decay=0.000001,
milestones=[2000,4000],
lrate_decay=0.1,
rho=0.95,
eps=1e-8,
lr_drop_rate=0.1
)
""" Data processing """
train = dict(
saved_model="", # "path to model to continue training"
Aug="None", # |None|Blur|Crop|Rot|ABINet
workers=4,
lan_list=["Chinese","Latin","Japanese", "Korean", "Arabic", "Bangla"],
valid_datas=[
"../dataset/MLT17_IL/test_2017",
"../dataset/MLT19_IL/test_2019"
],
select_data=[
"../dataset/MLT17_IL/train_2017",
"../dataset/MLT19_IL/train_2019"
],
batch_ratio="0.5-0.5",
total_data_usage_ratio="1.0",
NED=True,
batch_size=256,
num_iter=10000,
val_interval=5000,
log_multiple_test=None,
grad_clip=5,
)

Data Analysis

The experimental results of each task are recorded in data_any.txt and can be used for analysis of the data.

Acknowledgements

This implementation has been based on these repositories:

Citation

Please consider citing this work in your publications if it helps your research.

@article{zheng2023mrn,
title={MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition},
author={Zheng, Tianlun and Chen, Zhineng and Huang, BingChen and Zhang, Wei and Jiang, Yu-Gang},
journal={Proceedings of the IEEE/CVF International Conference on Computer Vision},
year={2023}
}

License

This project is released under the Apache 2.0 license.

Footnotes

  1. Nayef, N., et al. (2017). MLT 2017.

  2. Nayef, N., et al. (2019). MLT 2019.

About

Official Pytorch implementations of MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition (ICCV 2023)

Topics

Resources

Stars

46 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

MRN: Multiplexed Routing Network
for Incremental Multilingual Text Recognition

ICCV 2023ArXiv preprintBlogLICENSE

Method |IMLTR Dataset | Getting Started | Citation

It started as code for the paper:

MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition (Accepted by ICCV 2023)

This project is a toolkit for the novel scenario of Incremental Multilingual Text Recognition (IMLTR), the project supports many incremental learning methods and proposes a more applicable method for IMLTR: Multiplexed Routing Network (MRN) and the corresponding dataset. The project provides an efficient framework to assist in developing new methods and analyzing existing ones under the IMLTR task, and we hope it will advance the IMLTR community.

image

Methods

Incremental Learning Methods

  • Base: Baseline method which simply updates parameters on new tasks.
  • Joint: Bound method: data for all tasks are trained at once, an upper bound for the method
    (Joint_mix means all tasks data mixed in batch, Joint_loader means the consistent proportion of data from each task in a batch)
  • EWC[PNAS2017]: Overcoming catastrophic forgetting in neural networks
  • LwF[ECCV2016]: Learning without Forgetting
  • WA[CVPR2020]: Maintaining Discrimination and Fairness in Class Incremental Learning
  • DER[CVPR2021]: DER: Dynamically Expandable Representation for Class Incremental Learning
  • MRN[ICCV2023]: MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition

you can change config config/crnn_mrn.py for different il methods or setting.

common=dict(
il="mrn", # joint_mix | joint_loader | base | lwf | wa | ewc | der | mrn
memory="random", # None | random
memory_num=2000,
start_task = 0 # checkpoint start
)

Text Recognition Methods

  • CRNN[TPAMI2017]: An End-to-End Trainable Neural Network for Image-Based Sequence Recognition and Its Application to Scene Text Recognition
  • TRBA[ICCV2019]: What Is Wrong With Scene Text Recognition Model Comparisons? Dataset and Model Analysis
  • SVTR[IJCAI2022]: SVTR: Scene Text Recognition with a Single Visual Model

you can change config config/crnn_mrn.py for different text recognition modules or setting.

""" Model Architecture """
common=dict(
batch_max_length = 25,
imgH = 32,
imgW = 256,
)
model=dict(
model_name="TRBA",
Transformation = "TPS", #None TPS
FeatureExtraction = "ResNet", #VGG ResNet SVTR
SequenceModeling = "BiLSTM", #None BiLSTM
Prediction = "Attn", #CTC Attn
num_fiducial=20,
input_channel=4,
output_channel=512,
hidden_size=256,
)

IMLTR Dataset

The Dataset can be downloaded from BaiduNetdisk(passwd:c07h).

dataset
├── MLT17_IL
│ ├── test_2017
│ ├── train_2017
├── MLT19_IL
│ ├── test_2019
│ ├── train_2019

Incremental MLT17: MLT17 has 68,613 training instances and 16,255 validation instances, which are from 6 scripts and 9 languages: Chinese, Japanese, Korean, Bangla, Arabic, Italian, English, French, and German. The last four use Latin script. Incremental MLT17 use the validation set for test due to the unavailability of test data. Tasks are split by scripts and modeled sequentially. Special symbols are discarded at the preprocessing step as with no linguistic meaning.

Incremental MLT19: MLT19 has 89,177 text instances coming from 7 scripts. Since the inaccessibility of test set, Incremental MLT19 randomly split the training instances to 9:1 script-by-script, for model training and test. To be consistent with Incremental MLT17 dataset, we discard the Hindi script and also special symbols. Statistics of the two datasets are shown in the following.

DatasetCategories
Task1Task2Task3Task4Task5Task6
ChineseLatinJapaneseKoreanArabicBangla
MLT171Train Instance2687474114609563137113237
Test Instance5291107313501230983713
Train Class18953251620112473112
MLT192Train Instance2897529215324610742303542
Test Instance3225882590679470393
Train Class20862201728116073102

Getting Started

Dependency

  • This work was tested with PyTorch 1.6.0, CUDA 10.1 and python 3.6.
conda create -n mrn python=3.7 -y
conda activate mrn
conda install pytorch==1.9.1 torchvision==0.10.1 torchaudio==0.9.1 cudatoolkit=11.3 -c pytorch -c conda-forge
pip install torch==1.9.1+cu111 torchvision==0.10.1+cu111 torchaudio==0.9.1 -f https://download.pytorch.org/whl/torch_stable.html
  • requirements :
pip3 install lmdb pillow torchvision nltk natsort fire tensorboard tqdm opencv-python einops timm mmcv shapely scipy
pip3 install mmcv-full -f https://download.openmmlab.com/mmcv/dist/cu111/torch1.9.1/index.html

Training

python3 tiny_train.py --config=config/crnn_mrn.py --exp_name CRNN_real

Arguments

tiny_train.py (as a default, evaluate trained model on IMLTR datasets at the end of training.

  • --select_data: folder path to training lmdb datasets.
    [" ../dataset/MLT17_IL/train_2017", "../dataset/MLT19_IL/train_2019"]
  • --valid_datas: folder path to testing lmdb dataset.
    [" ../dataset/MLT17_IL/test_2017", "../dataset/MLT19_IL/test_2019"]
  • --batch_ratio: assign ratio for each selected data in the batch. default is '1 / number of datasets'.
  • --Aug: whether to use augmentation |None|Blur|Crop|Rot|

Config Detail

For detailed configuration modifications please use the config file config/crnn_mrn.py

common=dict(
exp_name="TRBA_MRN", # Where to store logs and models
il="mrn", # joint_mix | joint_loader | base | lwf | wa | ewc | der | mrn
memory="random", # None | random
memory_num=2000,
batch_max_length = 25,
imgH = 32,
imgW = 256,
manual_seed=111,
start_task = 0
)
""" Model Architecture """
model=dict(
model_name="TRBA",
Transformation = "TPS", #None TPS
FeatureExtraction = "ResNet", #VGG ResNet
SequenceModeling = "BiLSTM", #None BiLSTM
Prediction = "Attn", #CTC Attn
num_fiducial=20,
input_channel=4,
output_channel=512,
hidden_size=256,
)
""" Optimizer """
optimizer=dict(
schedule="super", #default is super for super convergence, 1 for None, [0.6, 0.8] for the same setting with ASTER
optimizer="adam",
lr=0.0005,
sgd_momentum=0.9,
sgd_weight_decay=0.000001,
milestones=[2000,4000],
lrate_decay=0.1,
rho=0.95,
eps=1e-8,
lr_drop_rate=0.1
)
""" Data processing """
train = dict(
saved_model="", # "path to model to continue training"
Aug="None", # |None|Blur|Crop|Rot|ABINet
workers=4,
lan_list=["Chinese","Latin","Japanese", "Korean", "Arabic", "Bangla"],
valid_datas=[
"../dataset/MLT17_IL/test_2017",
"../dataset/MLT19_IL/test_2019"
],
select_data=[
"../dataset/MLT17_IL/train_2017",
"../dataset/MLT19_IL/train_2019"
],
batch_ratio="0.5-0.5",
total_data_usage_ratio="1.0",
NED=True,
batch_size=256,
num_iter=10000,
val_interval=5000,
log_multiple_test=None,
grad_clip=5,
)

Data Analysis

The experimental results of each task are recorded in data_any.txt and can be used for analysis of the data.

Acknowledgements

This implementation has been based on these repositories:

Citation

Please consider citing this work in your publications if it helps your research.

@article{zheng2023mrn,
title={MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition},
author={Zheng, Tianlun and Chen, Zhineng and Huang, BingChen and Zhang, Wei and Jiang, Yu-Gang},
journal={Proceedings of the IEEE/CVF International Conference on Computer Vision},
year={2023}
}

License

This project is released under the Apache 2.0 license.

Footnotes

  1. Nayef, N., et al. (2017). MLT 2017.

  2. Nayef, N., et al. (2019). MLT 2019.

About

Official Pytorch implementations of MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition (ICCV 2023)

Topics

Resources

Stars

46 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

MRN: Multiplexed Routing Network
for Incremental Multilingual Text Recognition

ICCV 2023ArXiv preprintBlogLICENSE

Method |IMLTR Dataset | Getting Started | Citation

It started as code for the paper:

MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition (Accepted by ICCV 2023)

This project is a toolkit for the novel scenario of Incremental Multilingual Text Recognition (IMLTR), the project supports many incremental learning methods and proposes a more applicable method for IMLTR: Multiplexed Routing Network (MRN) and the corresponding dataset. The project provides an efficient framework to assist in developing new methods and analyzing existing ones under the IMLTR task, and we hope it will advance the IMLTR community.

image

Methods

Incremental Learning Methods

  • Base: Baseline method which simply updates parameters on new tasks.
  • Joint: Bound method: data for all tasks are trained at once, an upper bound for the method
    (Joint_mix means all tasks data mixed in batch, Joint_loader means the consistent proportion of data from each task in a batch)
  • EWC[PNAS2017]: Overcoming catastrophic forgetting in neural networks
  • LwF[ECCV2016]: Learning without Forgetting
  • WA[CVPR2020]: Maintaining Discrimination and Fairness in Class Incremental Learning
  • DER[CVPR2021]: DER: Dynamically Expandable Representation for Class Incremental Learning
  • MRN[ICCV2023]: MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition

you can change config config/crnn_mrn.py for different il methods or setting.

common=dict(
il="mrn", # joint_mix | joint_loader | base | lwf | wa | ewc | der | mrn
memory="random", # None | random
memory_num=2000,
start_task = 0 # checkpoint start
)

Text Recognition Methods

  • CRNN[TPAMI2017]: An End-to-End Trainable Neural Network for Image-Based Sequence Recognition and Its Application to Scene Text Recognition
  • TRBA[ICCV2019]: What Is Wrong With Scene Text Recognition Model Comparisons? Dataset and Model Analysis
  • SVTR[IJCAI2022]: SVTR: Scene Text Recognition with a Single Visual Model

you can change config config/crnn_mrn.py for different text recognition modules or setting.

""" Model Architecture """
common=dict(
batch_max_length = 25,
imgH = 32,
imgW = 256,
)
model=dict(
model_name="TRBA",
Transformation = "TPS", #None TPS
FeatureExtraction = "ResNet", #VGG ResNet SVTR
SequenceModeling = "BiLSTM", #None BiLSTM
Prediction = "Attn", #CTC Attn
num_fiducial=20,
input_channel=4,
output_channel=512,
hidden_size=256,
)

IMLTR Dataset

The Dataset can be downloaded from BaiduNetdisk(passwd:c07h).

dataset
├── MLT17_IL
│ ├── test_2017
│ ├── train_2017
├── MLT19_IL
│ ├── test_2019
│ ├── train_2019

Incremental MLT17: MLT17 has 68,613 training instances and 16,255 validation instances, which are from 6 scripts and 9 languages: Chinese, Japanese, Korean, Bangla, Arabic, Italian, English, French, and German. The last four use Latin script. Incremental MLT17 use the validation set for test due to the unavailability of test data. Tasks are split by scripts and modeled sequentially. Special symbols are discarded at the preprocessing step as with no linguistic meaning.

Incremental MLT19: MLT19 has 89,177 text instances coming from 7 scripts. Since the inaccessibility of test set, Incremental MLT19 randomly split the training instances to 9:1 script-by-script, for model training and test. To be consistent with Incremental MLT17 dataset, we discard the Hindi script and also special symbols. Statistics of the two datasets are shown in the following.

DatasetCategories
Task1Task2Task3Task4Task5Task6
ChineseLatinJapaneseKoreanArabicBangla
MLT171Train Instance2687474114609563137113237
Test Instance5291107313501230983713
Train Class18953251620112473112
MLT192Train Instance2897529215324610742303542
Test Instance3225882590679470393
Train Class20862201728116073102

Getting Started

Dependency

  • This work was tested with PyTorch 1.6.0, CUDA 10.1 and python 3.6.
conda create -n mrn python=3.7 -y
conda activate mrn
conda install pytorch==1.9.1 torchvision==0.10.1 torchaudio==0.9.1 cudatoolkit=11.3 -c pytorch -c conda-forge
pip install torch==1.9.1+cu111 torchvision==0.10.1+cu111 torchaudio==0.9.1 -f https://download.pytorch.org/whl/torch_stable.html
  • requirements :
pip3 install lmdb pillow torchvision nltk natsort fire tensorboard tqdm opencv-python einops timm mmcv shapely scipy
pip3 install mmcv-full -f https://download.openmmlab.com/mmcv/dist/cu111/torch1.9.1/index.html

Training

python3 tiny_train.py --config=config/crnn_mrn.py --exp_name CRNN_real

Arguments

tiny_train.py (as a default, evaluate trained model on IMLTR datasets at the end of training.

  • --select_data: folder path to training lmdb datasets.
    [" ../dataset/MLT17_IL/train_2017", "../dataset/MLT19_IL/train_2019"]
  • --valid_datas: folder path to testing lmdb dataset.
    [" ../dataset/MLT17_IL/test_2017", "../dataset/MLT19_IL/test_2019"]
  • --batch_ratio: assign ratio for each selected data in the batch. default is '1 / number of datasets'.
  • --Aug: whether to use augmentation |None|Blur|Crop|Rot|

Config Detail

For detailed configuration modifications please use the config file config/crnn_mrn.py

common=dict(
exp_name="TRBA_MRN", # Where to store logs and models
il="mrn", # joint_mix | joint_loader | base | lwf | wa | ewc | der | mrn
memory="random", # None | random
memory_num=2000,
batch_max_length = 25,
imgH = 32,
imgW = 256,
manual_seed=111,
start_task = 0
)
""" Model Architecture """
model=dict(
model_name="TRBA",
Transformation = "TPS", #None TPS
FeatureExtraction = "ResNet", #VGG ResNet
SequenceModeling = "BiLSTM", #None BiLSTM
Prediction = "Attn", #CTC Attn
num_fiducial=20,
input_channel=4,
output_channel=512,
hidden_size=256,
)
""" Optimizer """
optimizer=dict(
schedule="super", #default is super for super convergence, 1 for None, [0.6, 0.8] for the same setting with ASTER
optimizer="adam",
lr=0.0005,
sgd_momentum=0.9,
sgd_weight_decay=0.000001,
milestones=[2000,4000],
lrate_decay=0.1,
rho=0.95,
eps=1e-8,
lr_drop_rate=0.1
)
""" Data processing """
train = dict(
saved_model="", # "path to model to continue training"
Aug="None", # |None|Blur|Crop|Rot|ABINet
workers=4,
lan_list=["Chinese","Latin","Japanese", "Korean", "Arabic", "Bangla"],
valid_datas=[
"../dataset/MLT17_IL/test_2017",
"../dataset/MLT19_IL/test_2019"
],
select_data=[
"../dataset/MLT17_IL/train_2017",
"../dataset/MLT19_IL/train_2019"
],
batch_ratio="0.5-0.5",
total_data_usage_ratio="1.0",
NED=True,
batch_size=256,
num_iter=10000,
val_interval=5000,
log_multiple_test=None,
grad_clip=5,
)

Data Analysis

The experimental results of each task are recorded in data_any.txt and can be used for analysis of the data.

Acknowledgements

This implementation has been based on these repositories:

Citation

Please consider citing this work in your publications if it helps your research.

@article{zheng2023mrn,
title={MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition},
author={Zheng, Tianlun and Chen, Zhineng and Huang, BingChen and Zhang, Wei and Jiang, Yu-Gang},
journal={Proceedings of the IEEE/CVF International Conference on Computer Vision},
year={2023}
}

License

This project is released under the Apache 2.0 license.

Footnotes

  1. Nayef, N., et al. (2017). MLT 2017.

  2. Nayef, N., et al. (2019). MLT 2019.

About

Official Pytorch implementations of MRN: Multiplexed Routing Network for Incremental Multilingual Text Recognition (ICCV 2023)

Topics

Resources

Stars

46 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages