Skip to content

Repository files navigation

Deep Learning Model that Colors Grayscale Images

Paper by Federico Baldassarre, Diego González Morín, Lucas Rodés-Guirao: arXiv:1712.03400 [cs.CV] Deep Koalarization

Some Predicted Results

Test Images

Test Image 1Test Image 2Test Image 3

Generated Images

Predicted Image 1Predicted Image 2Predicted Image 3

Network Architecture

Network Architecture

Encoder Network Architecture

LayerFiltersKernel SizeStridesPaddingActivation
Conv2D_E164(3 × 3)(2 × 2)sameReLU
Conv2D_E2128(3 × 3)(1 × 1)sameReLU
Conv2D_E3128(3 × 3)(2 × 2)sameReLU
Conv2D_E4256(3 × 3)(1 × 1)sameReLU
Conv2D_E5256(3 × 3)(2 × 2)sameReLU
Conv2D_E6512(3 × 3)(1 × 1)sameReLU
Conv2D_E7512(3 × 3)(1 × 1)sameReLU
Conv2D_E8256(3 × 3)(1 × 1)sameReLU

Fusion Network Architecture

LayerFiltersKernel SizeStridesPaddingActivation
Conv2D_F1256(1 × 1)(1 × 1)sameReLU

Decoder Network Architecture

LayerFiltersKernel SizeStridesPaddingActivation
Conv2D_D1128(3 × 3)(1 × 1)sameReLU
UpSamp2D_D1-----
Conv2D_D264(3 × 3)(1 × 1)sameReLU
Conv2D_D364(3 × 3)(1 × 1)sameReLU
UpSamp2D_D2-----
Conv2D_D432(3 × 3)(1 × 1)sameReLU
Conv2D_D52(3 × 3)(1 × 1)sametanh
UpSamp2D_D2-----

Fusion Layer Architecture

Fusion Layer

High-Level Feature Extraction through Inception Resnet v2

The Inception Resnet v2 Model extracts the high-level features of the input grayscale image. The last layer before the softmax activation outputs a vector of size 1000 or dimension (1000 × 1 × 1) (feature-vector). This vector is repeated 28 × 28 times and then reshaped into a volume of (28 × 28 × 1000). This volume is then concatenated depth-wise to the Conv2D_E8 layer. This whole block of size (28 * 28 * 1256) is then passed through Conv2D_F1.

Releases

Packages

Contributors

Languages