Repository files navigation

Superpixel-based Refinement for Object Proposal Generation

Superpixel-based Refinement for Object Proposal Generation (ICPR 2020)

Precise segmentation of objects is an important problem in tasks like class-agnostic object proposal generation or instance segmentation. Deep learning-based systems usually generate segmentations of objects based on coarse feature maps, due to the inherent downsampling in CNNs. This leads to segmentation boundaries not adhering well to the object boundaries in the image. To tackle this problem, we introduce a new superpixel-based refinement approach on top of the state-of-the-art object proposal system AttentionMask. The refinement utilizes superpixel pooling for feature extraction and a novel superpixel classifier to determine if a high precision superpixel belongs to an object or not. Our experiments show an improvement of up to 26.0% in terms of average recall compared to original AttentionMask. Furthermore, qualitative and quantitative analyses of the segmentations reveal significant improvements in terms of boundary adherence for the proposed refinement compared to various deep learning-based state-of-the-art object proposal generation systems.

Example

The system is based on AttentionMask and FastMask.

If you find this software useful in your research, please cite our paper.

@inproceedings{WilmsFrintropICPR2020,
title = {{Superpixel-based Refinement for Object Proposal Generation}, author = {Christian Wilms and Simone Frintrop},
booktitle = {International Conference on Pattern Recognition (ICPR)},
year = {2020}
}

Requirements

  • Ubuntu 18.04
  • Cuda 10.0
  • Python 2.7
  • OpenCV-Python
  • Python packages: scipy, numpy, python-cjson, setproctitle, scikit-image
  • COCOApi
  • Caffe (already part of this git)
  • Alchemy (already part of this git)

Hardware specifications

For the results in the paper we used the following hardware:

  • Intel i7-5930K 6 core CPU
  • 64 GB RAM
  • GTX Titan X GPU with 12 GB RAM

Installation

Follow the installation instructions in the AttentionMask git

Usage

Our superpixel-based refinement system can be used without any re-training, utilizing our provided weights and segmentation. Just download the weights,the LVIS dataset and the segmentations.

Download dataset

Download the train2014 splits from COCO dataset for training and the validation split form the LVIS dataset. After downloading, extract the data in the following structure:

spxattmask
|
---- data
|
---- coco
|
---- annotations
| |
| ---- instances_train2014.json
| |
| ---- instances_val2017LVIS.json
|
---- train2014
| |
| ---- COCO_train2014_000000000009.jpg
| |
| ---- ...
|
---- val2017LVIS
|
---- COCO_val2017LVIS_000000000139.jpg
|
---- ...

Download weights

Download our weights for the superpixel-based refinement system: Link to caffemodel.

For training the system on your own dataset, download the initial ImageNet weights for the ResNet-34.

All weight files (.caffemodel) should be moved into the params subdirectory.

Creating segmentations

Essential to our superpixel-based refinement system are the superpixel segmentations. We generated and optimized all superpixel segmentations using the framework by Stutz et al. For training and testing eight segmentations needs to be generated per image, one segmentation per AttentionMask scale.

Due to the size we will not provide the segmentations. However, the segmentations used in the paper can be reproduced using the framework by Stutz et al. Follow the instalation insturctions in that repo. Note that only the segmentation algorithm by Felzenszwalb and Huttenlocher has to be build. The following table provides the parameters (scale (-t), minimum-size (-m), sigma (-g) in the framework by Stutz et al.) for generating the segmentations for each of the eight scales in AttentionMask.

ScaleParameter scale (-t)Parameter minimum-size (-m)Parameter sigma (-g)
810101
1660150
2460300
32120300
4810601
6430901
96601201
128101800

The segmentation size, i.e., the image size during segmentation, as well information about flipping the image and the segmentation (training only) can be found in the following json files for training data and test data. The json files contain a mapping from the image id to the height and width of the segmentation as well as a flag for denoting a left-right-flip (training only).

Segmentations for training

For training, the segmentations have to be provided in two different ways. First, all segmentations are expected as compressed csv-file (csv.gz) in the subdirectory segmentations/train2014/ with an indidividual folder per scale. Additionally, from those segmentations the superpixelized ground truth needs to be generated as json-file with scipt generateSpxJson.py followed by the script splitJson.py. Generate the segmentations using the framework by Stutz et al. and the parameters discussed above and paste the results into the directory structure shown below.

Segmentations for testing

During testing, only the superpixel segmentations are necessary. The segmentations are expeted as csv-file in the subdirectory segmentations/val2017LVIS with an individual folder per scale. Generate the segmentations using the framework by Stutz et al. and the parameters discussed above and paste the results into the directory structure shown below.

spxattmask
|
---- spxGT_train2014_FH_128.json
|
---- spxGT_train2014_FH_16.json
|
---- spxGT_train2014_FH_24.json
|
---- spxGT_train2014_FH_32.json
|
---- spxGT_train2014_FH_48.json
|
---- spxGT_train2014_FH_64.json
|
---- spxGT_train2014_FH_8.json
|
---- spxGT_train2014_FH_96.json
|
---- segmentations
|
---- train2014
| |
| ---- fh-8-8000
| | |
| | ---- 132574.csv.gz
| | |
| | ---- ...
| |
| ---- ...
|
---- val2017LVIS
|
---- fh-8-8000
| |
| ---- 1000.csv
| |
| ---- ...
|
---- ...

Inference

For inference on the LVIS dataset, first use the script generateIntermediateResults.py that runs the images thorugh the CNN and generates intermediate reuslts. Those results are stored in the folder intermediateResults, which has to be created first. Call the script with the gpu id, the model name, the weights and the dataset you want to test on (e.g., val2017LVIS):

$ python generateIntermediateResults.py 0 spxRefinedAttMask --init_weights spxrefinedattmask-final.caffemodel --dataset val2017LVIS --end 5000

Second, to apply the post-processing to the results and to stitch the proposals back into the image, call generateFinalResults.py with the model name and the dataset:

$ python generateFinalResults.py spxRefinedAttMask --dataset val2017LVIS --end 5000

You can find an example for both calls as well as the evaluation (see below) in the script test.sh.

Evaluation

Use evalCOCONMS.py to evaluate on the LVIS dataset with the model name and the dataset used. --useSegm is a flag for using segmentation masks instead of bounding boxes.

$ python evalCOCONMS.py spxRefinedAttMask --dataset val2017LVIS --useSegm True --end 5000

Training

To train our superpixel-based refinement system on the COCO dataset, you can use the train.sh script. The training runs for 13 epochs to generate the final model weights.

$ export EPOCH=1
$ ./train.sh

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Superpixel-based Refinement for Object Proposal Generation

Superpixel-based Refinement for Object Proposal Generation (ICPR 2020)

Precise segmentation of objects is an important problem in tasks like class-agnostic object proposal generation or instance segmentation. Deep learning-based systems usually generate segmentations of objects based on coarse feature maps, due to the inherent downsampling in CNNs. This leads to segmentation boundaries not adhering well to the object boundaries in the image. To tackle this problem, we introduce a new superpixel-based refinement approach on top of the state-of-the-art object proposal system AttentionMask. The refinement utilizes superpixel pooling for feature extraction and a novel superpixel classifier to determine if a high precision superpixel belongs to an object or not. Our experiments show an improvement of up to 26.0% in terms of average recall compared to original AttentionMask. Furthermore, qualitative and quantitative analyses of the segmentations reveal significant improvements in terms of boundary adherence for the proposed refinement compared to various deep learning-based state-of-the-art object proposal generation systems.

Example

The system is based on AttentionMask and FastMask.

If you find this software useful in your research, please cite our paper.

@inproceedings{WilmsFrintropICPR2020,
title = {{Superpixel-based Refinement for Object Proposal Generation}, author = {Christian Wilms and Simone Frintrop},
booktitle = {International Conference on Pattern Recognition (ICPR)},
year = {2020}
}

Requirements

  • Ubuntu 18.04
  • Cuda 10.0
  • Python 2.7
  • OpenCV-Python
  • Python packages: scipy, numpy, python-cjson, setproctitle, scikit-image
  • COCOApi
  • Caffe (already part of this git)
  • Alchemy (already part of this git)

Hardware specifications

For the results in the paper we used the following hardware:

  • Intel i7-5930K 6 core CPU
  • 64 GB RAM
  • GTX Titan X GPU with 12 GB RAM

Installation

Follow the installation instructions in the AttentionMask git

Usage

Our superpixel-based refinement system can be used without any re-training, utilizing our provided weights and segmentation. Just download the weights,the LVIS dataset and the segmentations.

Download dataset

Download the train2014 splits from COCO dataset for training and the validation split form the LVIS dataset. After downloading, extract the data in the following structure:

spxattmask
|
---- data
|
---- coco
|
---- annotations
| |
| ---- instances_train2014.json
| |
| ---- instances_val2017LVIS.json
|
---- train2014
| |
| ---- COCO_train2014_000000000009.jpg
| |
| ---- ...
|
---- val2017LVIS
|
---- COCO_val2017LVIS_000000000139.jpg
|
---- ...

Download weights

Download our weights for the superpixel-based refinement system: Link to caffemodel.

For training the system on your own dataset, download the initial ImageNet weights for the ResNet-34.

All weight files (.caffemodel) should be moved into the params subdirectory.

Creating segmentations

Essential to our superpixel-based refinement system are the superpixel segmentations. We generated and optimized all superpixel segmentations using the framework by Stutz et al. For training and testing eight segmentations needs to be generated per image, one segmentation per AttentionMask scale.

Due to the size we will not provide the segmentations. However, the segmentations used in the paper can be reproduced using the framework by Stutz et al. Follow the instalation insturctions in that repo. Note that only the segmentation algorithm by Felzenszwalb and Huttenlocher has to be build. The following table provides the parameters (scale (-t), minimum-size (-m), sigma (-g) in the framework by Stutz et al.) for generating the segmentations for each of the eight scales in AttentionMask.

ScaleParameter scale (-t)Parameter minimum-size (-m)Parameter sigma (-g)
810101
1660150
2460300
32120300
4810601
6430901
96601201
128101800

The segmentation size, i.e., the image size during segmentation, as well information about flipping the image and the segmentation (training only) can be found in the following json files for training data and test data. The json files contain a mapping from the image id to the height and width of the segmentation as well as a flag for denoting a left-right-flip (training only).

Segmentations for training

For training, the segmentations have to be provided in two different ways. First, all segmentations are expected as compressed csv-file (csv.gz) in the subdirectory segmentations/train2014/ with an indidividual folder per scale. Additionally, from those segmentations the superpixelized ground truth needs to be generated as json-file with scipt generateSpxJson.py followed by the script splitJson.py. Generate the segmentations using the framework by Stutz et al. and the parameters discussed above and paste the results into the directory structure shown below.

Segmentations for testing

During testing, only the superpixel segmentations are necessary. The segmentations are expeted as csv-file in the subdirectory segmentations/val2017LVIS with an individual folder per scale. Generate the segmentations using the framework by Stutz et al. and the parameters discussed above and paste the results into the directory structure shown below.

spxattmask
|
---- spxGT_train2014_FH_128.json
|
---- spxGT_train2014_FH_16.json
|
---- spxGT_train2014_FH_24.json
|
---- spxGT_train2014_FH_32.json
|
---- spxGT_train2014_FH_48.json
|
---- spxGT_train2014_FH_64.json
|
---- spxGT_train2014_FH_8.json
|
---- spxGT_train2014_FH_96.json
|
---- segmentations
|
---- train2014
| |
| ---- fh-8-8000
| | |
| | ---- 132574.csv.gz
| | |
| | ---- ...
| |
| ---- ...
|
---- val2017LVIS
|
---- fh-8-8000
| |
| ---- 1000.csv
| |
| ---- ...
|
---- ...

Inference

For inference on the LVIS dataset, first use the script generateIntermediateResults.py that runs the images thorugh the CNN and generates intermediate reuslts. Those results are stored in the folder intermediateResults, which has to be created first. Call the script with the gpu id, the model name, the weights and the dataset you want to test on (e.g., val2017LVIS):

$ python generateIntermediateResults.py 0 spxRefinedAttMask --init_weights spxrefinedattmask-final.caffemodel --dataset val2017LVIS --end 5000

Second, to apply the post-processing to the results and to stitch the proposals back into the image, call generateFinalResults.py with the model name and the dataset:

$ python generateFinalResults.py spxRefinedAttMask --dataset val2017LVIS --end 5000

You can find an example for both calls as well as the evaluation (see below) in the script test.sh.

Evaluation

Use evalCOCONMS.py to evaluate on the LVIS dataset with the model name and the dataset used. --useSegm is a flag for using segmentation masks instead of bounding boxes.

$ python evalCOCONMS.py spxRefinedAttMask --dataset val2017LVIS --useSegm True --end 5000

Training

To train our superpixel-based refinement system on the COCO dataset, you can use the train.sh script. The training runs for 13 epochs to generate the final model weights.

$ export EPOCH=1
$ ./train.sh

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Superpixel-based Refinement for Object Proposal Generation

Superpixel-based Refinement for Object Proposal Generation (ICPR 2020)

Precise segmentation of objects is an important problem in tasks like class-agnostic object proposal generation or instance segmentation. Deep learning-based systems usually generate segmentations of objects based on coarse feature maps, due to the inherent downsampling in CNNs. This leads to segmentation boundaries not adhering well to the object boundaries in the image. To tackle this problem, we introduce a new superpixel-based refinement approach on top of the state-of-the-art object proposal system AttentionMask. The refinement utilizes superpixel pooling for feature extraction and a novel superpixel classifier to determine if a high precision superpixel belongs to an object or not. Our experiments show an improvement of up to 26.0% in terms of average recall compared to original AttentionMask. Furthermore, qualitative and quantitative analyses of the segmentations reveal significant improvements in terms of boundary adherence for the proposed refinement compared to various deep learning-based state-of-the-art object proposal generation systems.

Example

The system is based on AttentionMask and FastMask.

If you find this software useful in your research, please cite our paper.

@inproceedings{WilmsFrintropICPR2020,
title = {{Superpixel-based Refinement for Object Proposal Generation}, author = {Christian Wilms and Simone Frintrop},
booktitle = {International Conference on Pattern Recognition (ICPR)},
year = {2020}
}

Requirements

  • Ubuntu 18.04
  • Cuda 10.0
  • Python 2.7
  • OpenCV-Python
  • Python packages: scipy, numpy, python-cjson, setproctitle, scikit-image
  • COCOApi
  • Caffe (already part of this git)
  • Alchemy (already part of this git)

Hardware specifications

For the results in the paper we used the following hardware:

  • Intel i7-5930K 6 core CPU
  • 64 GB RAM
  • GTX Titan X GPU with 12 GB RAM

Installation

Follow the installation instructions in the AttentionMask git

Usage

Our superpixel-based refinement system can be used without any re-training, utilizing our provided weights and segmentation. Just download the weights,the LVIS dataset and the segmentations.

Download dataset

Download the train2014 splits from COCO dataset for training and the validation split form the LVIS dataset. After downloading, extract the data in the following structure:

spxattmask
|
---- data
|
---- coco
|
---- annotations
| |
| ---- instances_train2014.json
| |
| ---- instances_val2017LVIS.json
|
---- train2014
| |
| ---- COCO_train2014_000000000009.jpg
| |
| ---- ...
|
---- val2017LVIS
|
---- COCO_val2017LVIS_000000000139.jpg
|
---- ...

Download weights

Download our weights for the superpixel-based refinement system: Link to caffemodel.

For training the system on your own dataset, download the initial ImageNet weights for the ResNet-34.

All weight files (.caffemodel) should be moved into the params subdirectory.

Creating segmentations

Essential to our superpixel-based refinement system are the superpixel segmentations. We generated and optimized all superpixel segmentations using the framework by Stutz et al. For training and testing eight segmentations needs to be generated per image, one segmentation per AttentionMask scale.

Due to the size we will not provide the segmentations. However, the segmentations used in the paper can be reproduced using the framework by Stutz et al. Follow the instalation insturctions in that repo. Note that only the segmentation algorithm by Felzenszwalb and Huttenlocher has to be build. The following table provides the parameters (scale (-t), minimum-size (-m), sigma (-g) in the framework by Stutz et al.) for generating the segmentations for each of the eight scales in AttentionMask.

ScaleParameter scale (-t)Parameter minimum-size (-m)Parameter sigma (-g)
810101
1660150
2460300
32120300
4810601
6430901
96601201
128101800

The segmentation size, i.e., the image size during segmentation, as well information about flipping the image and the segmentation (training only) can be found in the following json files for training data and test data. The json files contain a mapping from the image id to the height and width of the segmentation as well as a flag for denoting a left-right-flip (training only).

Segmentations for training

For training, the segmentations have to be provided in two different ways. First, all segmentations are expected as compressed csv-file (csv.gz) in the subdirectory segmentations/train2014/ with an indidividual folder per scale. Additionally, from those segmentations the superpixelized ground truth needs to be generated as json-file with scipt generateSpxJson.py followed by the script splitJson.py. Generate the segmentations using the framework by Stutz et al. and the parameters discussed above and paste the results into the directory structure shown below.

Segmentations for testing

During testing, only the superpixel segmentations are necessary. The segmentations are expeted as csv-file in the subdirectory segmentations/val2017LVIS with an individual folder per scale. Generate the segmentations using the framework by Stutz et al. and the parameters discussed above and paste the results into the directory structure shown below.

spxattmask
|
---- spxGT_train2014_FH_128.json
|
---- spxGT_train2014_FH_16.json
|
---- spxGT_train2014_FH_24.json
|
---- spxGT_train2014_FH_32.json
|
---- spxGT_train2014_FH_48.json
|
---- spxGT_train2014_FH_64.json
|
---- spxGT_train2014_FH_8.json
|
---- spxGT_train2014_FH_96.json
|
---- segmentations
|
---- train2014
| |
| ---- fh-8-8000
| | |
| | ---- 132574.csv.gz
| | |
| | ---- ...
| |
| ---- ...
|
---- val2017LVIS
|
---- fh-8-8000
| |
| ---- 1000.csv
| |
| ---- ...
|
---- ...

Inference

For inference on the LVIS dataset, first use the script generateIntermediateResults.py that runs the images thorugh the CNN and generates intermediate reuslts. Those results are stored in the folder intermediateResults, which has to be created first. Call the script with the gpu id, the model name, the weights and the dataset you want to test on (e.g., val2017LVIS):

$ python generateIntermediateResults.py 0 spxRefinedAttMask --init_weights spxrefinedattmask-final.caffemodel --dataset val2017LVIS --end 5000

Second, to apply the post-processing to the results and to stitch the proposals back into the image, call generateFinalResults.py with the model name and the dataset:

$ python generateFinalResults.py spxRefinedAttMask --dataset val2017LVIS --end 5000

You can find an example for both calls as well as the evaluation (see below) in the script test.sh.

Evaluation

Use evalCOCONMS.py to evaluate on the LVIS dataset with the model name and the dataset used. --useSegm is a flag for using segmentation masks instead of bounding boxes.

$ python evalCOCONMS.py spxRefinedAttMask --dataset val2017LVIS --useSegm True --end 5000

Training

To train our superpixel-based refinement system on the COCO dataset, you can use the train.sh script. The training runs for 13 epochs to generate the final model weights.

$ export EPOCH=1
$ ./train.sh

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Superpixel-based Refinement for Object Proposal Generation

Superpixel-based Refinement for Object Proposal Generation (ICPR 2020)

Precise segmentation of objects is an important problem in tasks like class-agnostic object proposal generation or instance segmentation. Deep learning-based systems usually generate segmentations of objects based on coarse feature maps, due to the inherent downsampling in CNNs. This leads to segmentation boundaries not adhering well to the object boundaries in the image. To tackle this problem, we introduce a new superpixel-based refinement approach on top of the state-of-the-art object proposal system AttentionMask. The refinement utilizes superpixel pooling for feature extraction and a novel superpixel classifier to determine if a high precision superpixel belongs to an object or not. Our experiments show an improvement of up to 26.0% in terms of average recall compared to original AttentionMask. Furthermore, qualitative and quantitative analyses of the segmentations reveal significant improvements in terms of boundary adherence for the proposed refinement compared to various deep learning-based state-of-the-art object proposal generation systems.

Example

The system is based on AttentionMask and FastMask.

If you find this software useful in your research, please cite our paper.

@inproceedings{WilmsFrintropICPR2020,
title = {{Superpixel-based Refinement for Object Proposal Generation}, author = {Christian Wilms and Simone Frintrop},
booktitle = {International Conference on Pattern Recognition (ICPR)},
year = {2020}
}

Requirements

  • Ubuntu 18.04
  • Cuda 10.0
  • Python 2.7
  • OpenCV-Python
  • Python packages: scipy, numpy, python-cjson, setproctitle, scikit-image
  • COCOApi
  • Caffe (already part of this git)
  • Alchemy (already part of this git)

Hardware specifications

For the results in the paper we used the following hardware:

  • Intel i7-5930K 6 core CPU
  • 64 GB RAM
  • GTX Titan X GPU with 12 GB RAM

Installation

Follow the installation instructions in the AttentionMask git

Usage

Our superpixel-based refinement system can be used without any re-training, utilizing our provided weights and segmentation. Just download the weights,the LVIS dataset and the segmentations.

Download dataset

Download the train2014 splits from COCO dataset for training and the validation split form the LVIS dataset. After downloading, extract the data in the following structure:

spxattmask
|
---- data
|
---- coco
|
---- annotations
| |
| ---- instances_train2014.json
| |
| ---- instances_val2017LVIS.json
|
---- train2014
| |
| ---- COCO_train2014_000000000009.jpg
| |
| ---- ...
|
---- val2017LVIS
|
---- COCO_val2017LVIS_000000000139.jpg
|
---- ...

Download weights

Download our weights for the superpixel-based refinement system: Link to caffemodel.

For training the system on your own dataset, download the initial ImageNet weights for the ResNet-34.

All weight files (.caffemodel) should be moved into the params subdirectory.

Creating segmentations

Essential to our superpixel-based refinement system are the superpixel segmentations. We generated and optimized all superpixel segmentations using the framework by Stutz et al. For training and testing eight segmentations needs to be generated per image, one segmentation per AttentionMask scale.

Due to the size we will not provide the segmentations. However, the segmentations used in the paper can be reproduced using the framework by Stutz et al. Follow the instalation insturctions in that repo. Note that only the segmentation algorithm by Felzenszwalb and Huttenlocher has to be build. The following table provides the parameters (scale (-t), minimum-size (-m), sigma (-g) in the framework by Stutz et al.) for generating the segmentations for each of the eight scales in AttentionMask.

ScaleParameter scale (-t)Parameter minimum-size (-m)Parameter sigma (-g)
810101
1660150
2460300
32120300
4810601
6430901
96601201
128101800

The segmentation size, i.e., the image size during segmentation, as well information about flipping the image and the segmentation (training only) can be found in the following json files for training data and test data. The json files contain a mapping from the image id to the height and width of the segmentation as well as a flag for denoting a left-right-flip (training only).

Segmentations for training

For training, the segmentations have to be provided in two different ways. First, all segmentations are expected as compressed csv-file (csv.gz) in the subdirectory segmentations/train2014/ with an indidividual folder per scale. Additionally, from those segmentations the superpixelized ground truth needs to be generated as json-file with scipt generateSpxJson.py followed by the script splitJson.py. Generate the segmentations using the framework by Stutz et al. and the parameters discussed above and paste the results into the directory structure shown below.

Segmentations for testing

During testing, only the superpixel segmentations are necessary. The segmentations are expeted as csv-file in the subdirectory segmentations/val2017LVIS with an individual folder per scale. Generate the segmentations using the framework by Stutz et al. and the parameters discussed above and paste the results into the directory structure shown below.

spxattmask
|
---- spxGT_train2014_FH_128.json
|
---- spxGT_train2014_FH_16.json
|
---- spxGT_train2014_FH_24.json
|
---- spxGT_train2014_FH_32.json
|
---- spxGT_train2014_FH_48.json
|
---- spxGT_train2014_FH_64.json
|
---- spxGT_train2014_FH_8.json
|
---- spxGT_train2014_FH_96.json
|
---- segmentations
|
---- train2014
| |
| ---- fh-8-8000
| | |
| | ---- 132574.csv.gz
| | |
| | ---- ...
| |
| ---- ...
|
---- val2017LVIS
|
---- fh-8-8000
| |
| ---- 1000.csv
| |
| ---- ...
|
---- ...

Inference

For inference on the LVIS dataset, first use the script generateIntermediateResults.py that runs the images thorugh the CNN and generates intermediate reuslts. Those results are stored in the folder intermediateResults, which has to be created first. Call the script with the gpu id, the model name, the weights and the dataset you want to test on (e.g., val2017LVIS):

$ python generateIntermediateResults.py 0 spxRefinedAttMask --init_weights spxrefinedattmask-final.caffemodel --dataset val2017LVIS --end 5000

Second, to apply the post-processing to the results and to stitch the proposals back into the image, call generateFinalResults.py with the model name and the dataset:

$ python generateFinalResults.py spxRefinedAttMask --dataset val2017LVIS --end 5000

You can find an example for both calls as well as the evaluation (see below) in the script test.sh.

Evaluation

Use evalCOCONMS.py to evaluate on the LVIS dataset with the model name and the dataset used. --useSegm is a flag for using segmentation masks instead of bounding boxes.

$ python evalCOCONMS.py spxRefinedAttMask --dataset val2017LVIS --useSegm True --end 5000

Training

To train our superpixel-based refinement system on the COCO dataset, you can use the train.sh script. The training runs for 13 epochs to generate the final model weights.

$ export EPOCH=1
$ ./train.sh

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Superpixel-based Refinement for Object Proposal Generation

Superpixel-based Refinement for Object Proposal Generation (ICPR 2020)

Precise segmentation of objects is an important problem in tasks like class-agnostic object proposal generation or instance segmentation. Deep learning-based systems usually generate segmentations of objects based on coarse feature maps, due to the inherent downsampling in CNNs. This leads to segmentation boundaries not adhering well to the object boundaries in the image. To tackle this problem, we introduce a new superpixel-based refinement approach on top of the state-of-the-art object proposal system AttentionMask. The refinement utilizes superpixel pooling for feature extraction and a novel superpixel classifier to determine if a high precision superpixel belongs to an object or not. Our experiments show an improvement of up to 26.0% in terms of average recall compared to original AttentionMask. Furthermore, qualitative and quantitative analyses of the segmentations reveal significant improvements in terms of boundary adherence for the proposed refinement compared to various deep learning-based state-of-the-art object proposal generation systems.

Example

The system is based on AttentionMask and FastMask.

If you find this software useful in your research, please cite our paper.

@inproceedings{WilmsFrintropICPR2020,
title = {{Superpixel-based Refinement for Object Proposal Generation}, author = {Christian Wilms and Simone Frintrop},
booktitle = {International Conference on Pattern Recognition (ICPR)},
year = {2020}
}

Requirements

  • Ubuntu 18.04
  • Cuda 10.0
  • Python 2.7
  • OpenCV-Python
  • Python packages: scipy, numpy, python-cjson, setproctitle, scikit-image
  • COCOApi
  • Caffe (already part of this git)
  • Alchemy (already part of this git)

Hardware specifications

For the results in the paper we used the following hardware:

  • Intel i7-5930K 6 core CPU
  • 64 GB RAM
  • GTX Titan X GPU with 12 GB RAM

Installation

Follow the installation instructions in the AttentionMask git

Usage

Our superpixel-based refinement system can be used without any re-training, utilizing our provided weights and segmentation. Just download the weights,the LVIS dataset and the segmentations.

Download dataset

Download the train2014 splits from COCO dataset for training and the validation split form the LVIS dataset. After downloading, extract the data in the following structure:

spxattmask
|
---- data
|
---- coco
|
---- annotations
| |
| ---- instances_train2014.json
| |
| ---- instances_val2017LVIS.json
|
---- train2014
| |
| ---- COCO_train2014_000000000009.jpg
| |
| ---- ...
|
---- val2017LVIS
|
---- COCO_val2017LVIS_000000000139.jpg
|
---- ...

Download weights

Download our weights for the superpixel-based refinement system: Link to caffemodel.

For training the system on your own dataset, download the initial ImageNet weights for the ResNet-34.

All weight files (.caffemodel) should be moved into the params subdirectory.

Creating segmentations

Essential to our superpixel-based refinement system are the superpixel segmentations. We generated and optimized all superpixel segmentations using the framework by Stutz et al. For training and testing eight segmentations needs to be generated per image, one segmentation per AttentionMask scale.

Due to the size we will not provide the segmentations. However, the segmentations used in the paper can be reproduced using the framework by Stutz et al. Follow the instalation insturctions in that repo. Note that only the segmentation algorithm by Felzenszwalb and Huttenlocher has to be build. The following table provides the parameters (scale (-t), minimum-size (-m), sigma (-g) in the framework by Stutz et al.) for generating the segmentations for each of the eight scales in AttentionMask.

ScaleParameter scale (-t)Parameter minimum-size (-m)Parameter sigma (-g)
810101
1660150
2460300
32120300
4810601
6430901
96601201
128101800

The segmentation size, i.e., the image size during segmentation, as well information about flipping the image and the segmentation (training only) can be found in the following json files for training data and test data. The json files contain a mapping from the image id to the height and width of the segmentation as well as a flag for denoting a left-right-flip (training only).

Segmentations for training

For training, the segmentations have to be provided in two different ways. First, all segmentations are expected as compressed csv-file (csv.gz) in the subdirectory segmentations/train2014/ with an indidividual folder per scale. Additionally, from those segmentations the superpixelized ground truth needs to be generated as json-file with scipt generateSpxJson.py followed by the script splitJson.py. Generate the segmentations using the framework by Stutz et al. and the parameters discussed above and paste the results into the directory structure shown below.

Segmentations for testing

During testing, only the superpixel segmentations are necessary. The segmentations are expeted as csv-file in the subdirectory segmentations/val2017LVIS with an individual folder per scale. Generate the segmentations using the framework by Stutz et al. and the parameters discussed above and paste the results into the directory structure shown below.

spxattmask
|
---- spxGT_train2014_FH_128.json
|
---- spxGT_train2014_FH_16.json
|
---- spxGT_train2014_FH_24.json
|
---- spxGT_train2014_FH_32.json
|
---- spxGT_train2014_FH_48.json
|
---- spxGT_train2014_FH_64.json
|
---- spxGT_train2014_FH_8.json
|
---- spxGT_train2014_FH_96.json
|
---- segmentations
|
---- train2014
| |
| ---- fh-8-8000
| | |
| | ---- 132574.csv.gz
| | |
| | ---- ...
| |
| ---- ...
|
---- val2017LVIS
|
---- fh-8-8000
| |
| ---- 1000.csv
| |
| ---- ...
|
---- ...

Inference

For inference on the LVIS dataset, first use the script generateIntermediateResults.py that runs the images thorugh the CNN and generates intermediate reuslts. Those results are stored in the folder intermediateResults, which has to be created first. Call the script with the gpu id, the model name, the weights and the dataset you want to test on (e.g., val2017LVIS):

$ python generateIntermediateResults.py 0 spxRefinedAttMask --init_weights spxrefinedattmask-final.caffemodel --dataset val2017LVIS --end 5000

Second, to apply the post-processing to the results and to stitch the proposals back into the image, call generateFinalResults.py with the model name and the dataset:

$ python generateFinalResults.py spxRefinedAttMask --dataset val2017LVIS --end 5000

You can find an example for both calls as well as the evaluation (see below) in the script test.sh.

Evaluation

Use evalCOCONMS.py to evaluate on the LVIS dataset with the model name and the dataset used. --useSegm is a flag for using segmentation masks instead of bounding boxes.

$ python evalCOCONMS.py spxRefinedAttMask --dataset val2017LVIS --useSegm True --end 5000

Training

To train our superpixel-based refinement system on the COCO dataset, you can use the train.sh script. The training runs for 13 epochs to generate the final model weights.

$ export EPOCH=1
$ ./train.sh

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Superpixel-based Refinement for Object Proposal Generation

Superpixel-based Refinement for Object Proposal Generation (ICPR 2020)

Precise segmentation of objects is an important problem in tasks like class-agnostic object proposal generation or instance segmentation. Deep learning-based systems usually generate segmentations of objects based on coarse feature maps, due to the inherent downsampling in CNNs. This leads to segmentation boundaries not adhering well to the object boundaries in the image. To tackle this problem, we introduce a new superpixel-based refinement approach on top of the state-of-the-art object proposal system AttentionMask. The refinement utilizes superpixel pooling for feature extraction and a novel superpixel classifier to determine if a high precision superpixel belongs to an object or not. Our experiments show an improvement of up to 26.0% in terms of average recall compared to original AttentionMask. Furthermore, qualitative and quantitative analyses of the segmentations reveal significant improvements in terms of boundary adherence for the proposed refinement compared to various deep learning-based state-of-the-art object proposal generation systems.

Example

The system is based on AttentionMask and FastMask.

If you find this software useful in your research, please cite our paper.

@inproceedings{WilmsFrintropICPR2020,
title = {{Superpixel-based Refinement for Object Proposal Generation}, author = {Christian Wilms and Simone Frintrop},
booktitle = {International Conference on Pattern Recognition (ICPR)},
year = {2020}
}

Requirements

  • Ubuntu 18.04
  • Cuda 10.0
  • Python 2.7
  • OpenCV-Python
  • Python packages: scipy, numpy, python-cjson, setproctitle, scikit-image
  • COCOApi
  • Caffe (already part of this git)
  • Alchemy (already part of this git)

Hardware specifications

For the results in the paper we used the following hardware:

  • Intel i7-5930K 6 core CPU
  • 64 GB RAM
  • GTX Titan X GPU with 12 GB RAM

Installation

Follow the installation instructions in the AttentionMask git

Usage

Our superpixel-based refinement system can be used without any re-training, utilizing our provided weights and segmentation. Just download the weights,the LVIS dataset and the segmentations.

Download dataset

Download the train2014 splits from COCO dataset for training and the validation split form the LVIS dataset. After downloading, extract the data in the following structure:

spxattmask
|
---- data
|
---- coco
|
---- annotations
| |
| ---- instances_train2014.json
| |
| ---- instances_val2017LVIS.json
|
---- train2014
| |
| ---- COCO_train2014_000000000009.jpg
| |
| ---- ...
|
---- val2017LVIS
|
---- COCO_val2017LVIS_000000000139.jpg
|
---- ...

Download weights

Download our weights for the superpixel-based refinement system: Link to caffemodel.

For training the system on your own dataset, download the initial ImageNet weights for the ResNet-34.

All weight files (.caffemodel) should be moved into the params subdirectory.

Creating segmentations

Essential to our superpixel-based refinement system are the superpixel segmentations. We generated and optimized all superpixel segmentations using the framework by Stutz et al. For training and testing eight segmentations needs to be generated per image, one segmentation per AttentionMask scale.

Due to the size we will not provide the segmentations. However, the segmentations used in the paper can be reproduced using the framework by Stutz et al. Follow the instalation insturctions in that repo. Note that only the segmentation algorithm by Felzenszwalb and Huttenlocher has to be build. The following table provides the parameters (scale (-t), minimum-size (-m), sigma (-g) in the framework by Stutz et al.) for generating the segmentations for each of the eight scales in AttentionMask.

ScaleParameter scale (-t)Parameter minimum-size (-m)Parameter sigma (-g)
810101
1660150
2460300
32120300
4810601
6430901
96601201
128101800

The segmentation size, i.e., the image size during segmentation, as well information about flipping the image and the segmentation (training only) can be found in the following json files for training data and test data. The json files contain a mapping from the image id to the height and width of the segmentation as well as a flag for denoting a left-right-flip (training only).

Segmentations for training

For training, the segmentations have to be provided in two different ways. First, all segmentations are expected as compressed csv-file (csv.gz) in the subdirectory segmentations/train2014/ with an indidividual folder per scale. Additionally, from those segmentations the superpixelized ground truth needs to be generated as json-file with scipt generateSpxJson.py followed by the script splitJson.py. Generate the segmentations using the framework by Stutz et al. and the parameters discussed above and paste the results into the directory structure shown below.

Segmentations for testing

During testing, only the superpixel segmentations are necessary. The segmentations are expeted as csv-file in the subdirectory segmentations/val2017LVIS with an individual folder per scale. Generate the segmentations using the framework by Stutz et al. and the parameters discussed above and paste the results into the directory structure shown below.

spxattmask
|
---- spxGT_train2014_FH_128.json
|
---- spxGT_train2014_FH_16.json
|
---- spxGT_train2014_FH_24.json
|
---- spxGT_train2014_FH_32.json
|
---- spxGT_train2014_FH_48.json
|
---- spxGT_train2014_FH_64.json
|
---- spxGT_train2014_FH_8.json
|
---- spxGT_train2014_FH_96.json
|
---- segmentations
|
---- train2014
| |
| ---- fh-8-8000
| | |
| | ---- 132574.csv.gz
| | |
| | ---- ...
| |
| ---- ...
|
---- val2017LVIS
|
---- fh-8-8000
| |
| ---- 1000.csv
| |
| ---- ...
|
---- ...

Inference

For inference on the LVIS dataset, first use the script generateIntermediateResults.py that runs the images thorugh the CNN and generates intermediate reuslts. Those results are stored in the folder intermediateResults, which has to be created first. Call the script with the gpu id, the model name, the weights and the dataset you want to test on (e.g., val2017LVIS):

$ python generateIntermediateResults.py 0 spxRefinedAttMask --init_weights spxrefinedattmask-final.caffemodel --dataset val2017LVIS --end 5000

Second, to apply the post-processing to the results and to stitch the proposals back into the image, call generateFinalResults.py with the model name and the dataset:

$ python generateFinalResults.py spxRefinedAttMask --dataset val2017LVIS --end 5000

You can find an example for both calls as well as the evaluation (see below) in the script test.sh.

Evaluation

Use evalCOCONMS.py to evaluate on the LVIS dataset with the model name and the dataset used. --useSegm is a flag for using segmentation masks instead of bounding boxes.

$ python evalCOCONMS.py spxRefinedAttMask --dataset val2017LVIS --useSegm True --end 5000

Training

To train our superpixel-based refinement system on the COCO dataset, you can use the train.sh script. The training runs for 13 epochs to generate the final model weights.

$ export EPOCH=1
$ ./train.sh

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Superpixel-based Refinement for Object Proposal Generation

Superpixel-based Refinement for Object Proposal Generation (ICPR 2020)

Precise segmentation of objects is an important problem in tasks like class-agnostic object proposal generation or instance segmentation. Deep learning-based systems usually generate segmentations of objects based on coarse feature maps, due to the inherent downsampling in CNNs. This leads to segmentation boundaries not adhering well to the object boundaries in the image. To tackle this problem, we introduce a new superpixel-based refinement approach on top of the state-of-the-art object proposal system AttentionMask. The refinement utilizes superpixel pooling for feature extraction and a novel superpixel classifier to determine if a high precision superpixel belongs to an object or not. Our experiments show an improvement of up to 26.0% in terms of average recall compared to original AttentionMask. Furthermore, qualitative and quantitative analyses of the segmentations reveal significant improvements in terms of boundary adherence for the proposed refinement compared to various deep learning-based state-of-the-art object proposal generation systems.

Example

The system is based on AttentionMask and FastMask.

If you find this software useful in your research, please cite our paper.

@inproceedings{WilmsFrintropICPR2020,
title = {{Superpixel-based Refinement for Object Proposal Generation}, author = {Christian Wilms and Simone Frintrop},
booktitle = {International Conference on Pattern Recognition (ICPR)},
year = {2020}
}

Requirements

  • Ubuntu 18.04
  • Cuda 10.0
  • Python 2.7
  • OpenCV-Python
  • Python packages: scipy, numpy, python-cjson, setproctitle, scikit-image
  • COCOApi
  • Caffe (already part of this git)
  • Alchemy (already part of this git)

Hardware specifications

For the results in the paper we used the following hardware:

  • Intel i7-5930K 6 core CPU
  • 64 GB RAM
  • GTX Titan X GPU with 12 GB RAM

Installation

Follow the installation instructions in the AttentionMask git

Usage

Our superpixel-based refinement system can be used without any re-training, utilizing our provided weights and segmentation. Just download the weights,the LVIS dataset and the segmentations.

Download dataset

Download the train2014 splits from COCO dataset for training and the validation split form the LVIS dataset. After downloading, extract the data in the following structure:

spxattmask
|
---- data
|
---- coco
|
---- annotations
| |
| ---- instances_train2014.json
| |
| ---- instances_val2017LVIS.json
|
---- train2014
| |
| ---- COCO_train2014_000000000009.jpg
| |
| ---- ...
|
---- val2017LVIS
|
---- COCO_val2017LVIS_000000000139.jpg
|
---- ...

Download weights

Download our weights for the superpixel-based refinement system: Link to caffemodel.

For training the system on your own dataset, download the initial ImageNet weights for the ResNet-34.

All weight files (.caffemodel) should be moved into the params subdirectory.

Creating segmentations

Essential to our superpixel-based refinement system are the superpixel segmentations. We generated and optimized all superpixel segmentations using the framework by Stutz et al. For training and testing eight segmentations needs to be generated per image, one segmentation per AttentionMask scale.

Due to the size we will not provide the segmentations. However, the segmentations used in the paper can be reproduced using the framework by Stutz et al. Follow the instalation insturctions in that repo. Note that only the segmentation algorithm by Felzenszwalb and Huttenlocher has to be build. The following table provides the parameters (scale (-t), minimum-size (-m), sigma (-g) in the framework by Stutz et al.) for generating the segmentations for each of the eight scales in AttentionMask.

ScaleParameter scale (-t)Parameter minimum-size (-m)Parameter sigma (-g)
810101
1660150
2460300
32120300
4810601
6430901
96601201
128101800

The segmentation size, i.e., the image size during segmentation, as well information about flipping the image and the segmentation (training only) can be found in the following json files for training data and test data. The json files contain a mapping from the image id to the height and width of the segmentation as well as a flag for denoting a left-right-flip (training only).

Segmentations for training

For training, the segmentations have to be provided in two different ways. First, all segmentations are expected as compressed csv-file (csv.gz) in the subdirectory segmentations/train2014/ with an indidividual folder per scale. Additionally, from those segmentations the superpixelized ground truth needs to be generated as json-file with scipt generateSpxJson.py followed by the script splitJson.py. Generate the segmentations using the framework by Stutz et al. and the parameters discussed above and paste the results into the directory structure shown below.

Segmentations for testing

During testing, only the superpixel segmentations are necessary. The segmentations are expeted as csv-file in the subdirectory segmentations/val2017LVIS with an individual folder per scale. Generate the segmentations using the framework by Stutz et al. and the parameters discussed above and paste the results into the directory structure shown below.

spxattmask
|
---- spxGT_train2014_FH_128.json
|
---- spxGT_train2014_FH_16.json
|
---- spxGT_train2014_FH_24.json
|
---- spxGT_train2014_FH_32.json
|
---- spxGT_train2014_FH_48.json
|
---- spxGT_train2014_FH_64.json
|
---- spxGT_train2014_FH_8.json
|
---- spxGT_train2014_FH_96.json
|
---- segmentations
|
---- train2014
| |
| ---- fh-8-8000
| | |
| | ---- 132574.csv.gz
| | |
| | ---- ...
| |
| ---- ...
|
---- val2017LVIS
|
---- fh-8-8000
| |
| ---- 1000.csv
| |
| ---- ...
|
---- ...

Inference

For inference on the LVIS dataset, first use the script generateIntermediateResults.py that runs the images thorugh the CNN and generates intermediate reuslts. Those results are stored in the folder intermediateResults, which has to be created first. Call the script with the gpu id, the model name, the weights and the dataset you want to test on (e.g., val2017LVIS):

$ python generateIntermediateResults.py 0 spxRefinedAttMask --init_weights spxrefinedattmask-final.caffemodel --dataset val2017LVIS --end 5000

Second, to apply the post-processing to the results and to stitch the proposals back into the image, call generateFinalResults.py with the model name and the dataset:

$ python generateFinalResults.py spxRefinedAttMask --dataset val2017LVIS --end 5000

You can find an example for both calls as well as the evaluation (see below) in the script test.sh.

Evaluation

Use evalCOCONMS.py to evaluate on the LVIS dataset with the model name and the dataset used. --useSegm is a flag for using segmentation masks instead of bounding boxes.

$ python evalCOCONMS.py spxRefinedAttMask --dataset val2017LVIS --useSegm True --end 5000

Training

To train our superpixel-based refinement system on the COCO dataset, you can use the train.sh script. The training runs for 13 epochs to generate the final model weights.

$ export EPOCH=1
$ ./train.sh

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Superpixel-based Refinement for Object Proposal Generation

Superpixel-based Refinement for Object Proposal Generation (ICPR 2020)

Precise segmentation of objects is an important problem in tasks like class-agnostic object proposal generation or instance segmentation. Deep learning-based systems usually generate segmentations of objects based on coarse feature maps, due to the inherent downsampling in CNNs. This leads to segmentation boundaries not adhering well to the object boundaries in the image. To tackle this problem, we introduce a new superpixel-based refinement approach on top of the state-of-the-art object proposal system AttentionMask. The refinement utilizes superpixel pooling for feature extraction and a novel superpixel classifier to determine if a high precision superpixel belongs to an object or not. Our experiments show an improvement of up to 26.0% in terms of average recall compared to original AttentionMask. Furthermore, qualitative and quantitative analyses of the segmentations reveal significant improvements in terms of boundary adherence for the proposed refinement compared to various deep learning-based state-of-the-art object proposal generation systems.

Example

The system is based on AttentionMask and FastMask.

If you find this software useful in your research, please cite our paper.

@inproceedings{WilmsFrintropICPR2020,
title = {{Superpixel-based Refinement for Object Proposal Generation}, author = {Christian Wilms and Simone Frintrop},
booktitle = {International Conference on Pattern Recognition (ICPR)},
year = {2020}
}

Requirements

  • Ubuntu 18.04
  • Cuda 10.0
  • Python 2.7
  • OpenCV-Python
  • Python packages: scipy, numpy, python-cjson, setproctitle, scikit-image
  • COCOApi
  • Caffe (already part of this git)
  • Alchemy (already part of this git)

Hardware specifications

For the results in the paper we used the following hardware:

  • Intel i7-5930K 6 core CPU
  • 64 GB RAM
  • GTX Titan X GPU with 12 GB RAM

Installation

Follow the installation instructions in the AttentionMask git

Usage

Our superpixel-based refinement system can be used without any re-training, utilizing our provided weights and segmentation. Just download the weights,the LVIS dataset and the segmentations.

Download dataset

Download the train2014 splits from COCO dataset for training and the validation split form the LVIS dataset. After downloading, extract the data in the following structure:

spxattmask
|
---- data
|
---- coco
|
---- annotations
| |
| ---- instances_train2014.json
| |
| ---- instances_val2017LVIS.json
|
---- train2014
| |
| ---- COCO_train2014_000000000009.jpg
| |
| ---- ...
|
---- val2017LVIS
|
---- COCO_val2017LVIS_000000000139.jpg
|
---- ...

Download weights

Download our weights for the superpixel-based refinement system: Link to caffemodel.

For training the system on your own dataset, download the initial ImageNet weights for the ResNet-34.

All weight files (.caffemodel) should be moved into the params subdirectory.

Creating segmentations

Essential to our superpixel-based refinement system are the superpixel segmentations. We generated and optimized all superpixel segmentations using the framework by Stutz et al. For training and testing eight segmentations needs to be generated per image, one segmentation per AttentionMask scale.

Due to the size we will not provide the segmentations. However, the segmentations used in the paper can be reproduced using the framework by Stutz et al. Follow the instalation insturctions in that repo. Note that only the segmentation algorithm by Felzenszwalb and Huttenlocher has to be build. The following table provides the parameters (scale (-t), minimum-size (-m), sigma (-g) in the framework by Stutz et al.) for generating the segmentations for each of the eight scales in AttentionMask.

ScaleParameter scale (-t)Parameter minimum-size (-m)Parameter sigma (-g)
810101
1660150
2460300
32120300
4810601
6430901
96601201
128101800

The segmentation size, i.e., the image size during segmentation, as well information about flipping the image and the segmentation (training only) can be found in the following json files for training data and test data. The json files contain a mapping from the image id to the height and width of the segmentation as well as a flag for denoting a left-right-flip (training only).

Segmentations for training

For training, the segmentations have to be provided in two different ways. First, all segmentations are expected as compressed csv-file (csv.gz) in the subdirectory segmentations/train2014/ with an indidividual folder per scale. Additionally, from those segmentations the superpixelized ground truth needs to be generated as json-file with scipt generateSpxJson.py followed by the script splitJson.py. Generate the segmentations using the framework by Stutz et al. and the parameters discussed above and paste the results into the directory structure shown below.

Segmentations for testing

During testing, only the superpixel segmentations are necessary. The segmentations are expeted as csv-file in the subdirectory segmentations/val2017LVIS with an individual folder per scale. Generate the segmentations using the framework by Stutz et al. and the parameters discussed above and paste the results into the directory structure shown below.

spxattmask
|
---- spxGT_train2014_FH_128.json
|
---- spxGT_train2014_FH_16.json
|
---- spxGT_train2014_FH_24.json
|
---- spxGT_train2014_FH_32.json
|
---- spxGT_train2014_FH_48.json
|
---- spxGT_train2014_FH_64.json
|
---- spxGT_train2014_FH_8.json
|
---- spxGT_train2014_FH_96.json
|
---- segmentations
|
---- train2014
| |
| ---- fh-8-8000
| | |
| | ---- 132574.csv.gz
| | |
| | ---- ...
| |
| ---- ...
|
---- val2017LVIS
|
---- fh-8-8000
| |
| ---- 1000.csv
| |
| ---- ...
|
---- ...

Inference

For inference on the LVIS dataset, first use the script generateIntermediateResults.py that runs the images thorugh the CNN and generates intermediate reuslts. Those results are stored in the folder intermediateResults, which has to be created first. Call the script with the gpu id, the model name, the weights and the dataset you want to test on (e.g., val2017LVIS):

$ python generateIntermediateResults.py 0 spxRefinedAttMask --init_weights spxrefinedattmask-final.caffemodel --dataset val2017LVIS --end 5000

Second, to apply the post-processing to the results and to stitch the proposals back into the image, call generateFinalResults.py with the model name and the dataset:

$ python generateFinalResults.py spxRefinedAttMask --dataset val2017LVIS --end 5000

You can find an example for both calls as well as the evaluation (see below) in the script test.sh.

Evaluation

Use evalCOCONMS.py to evaluate on the LVIS dataset with the model name and the dataset used. --useSegm is a flag for using segmentation masks instead of bounding boxes.

$ python evalCOCONMS.py spxRefinedAttMask --dataset val2017LVIS --useSegm True --end 5000

Training

To train our superpixel-based refinement system on the COCO dataset, you can use the train.sh script. The training runs for 13 epochs to generate the final model weights.

$ export EPOCH=1
$ ./train.sh

Releases

Packages

Used by

Contributors

Languages