ESPNetv2: A Light-Weight, Power Efficient, and General Purpose Convolutional Neural Network

Mehta, Sachin; Rastegari, Mohammad; Shapiro, Linda G.; Hajishirzi, Hannaneh

doi:10.1109/cvpr.2019.00941

Cited by 485 publications

(354 citation statements)

References 45 publications

Supporting

Mentioning

304

Contrasting

Unclassified

Order By: Relevance

“…Standard [21] n 2 cĉ n × n Group [15] n 2 cĉ/g n × n 1D-factorized [13] 2ncĉ n × n DW [11,12] n 2 c + cĉ n × n DDW [14] n 2 c + cĉ n r × n r our FDDWC 2nc + cĉ n r × n r…”

Section: Convolutional Type Parameters Size Of Receptive Fieldmentioning

confidence: 99%

“…Towards this end, this section introduces EERM, the core unit of FDDWNet, to approach the representational power of larger and denser layers, but at a considerably lower computational budgets. EERM unit, which leverages the residual connections and FDDWC, combines the strength of 1D-factorized convolution [13] and dilated depth-wise separable convolution [14]. More specifically, as shown in Fig.…”

Section: Eerm Unitmentioning

confidence: 99%

“…ERFNet [13] decomposes a 2D convolution (e.g., 3 × 3) into two 1D-factorized convolution (e.g., 3 × 1 and 1 × 3). An alternative efficient approach to lighten CNNs depends on group convolution [14,15], where input channels and filter kernels are accordingly factored into a set of groups and each group is convolved independently. In spite of achieving impressive results, these lightweight networks prefer to adopt shallow network architectures to reduce model complexity, which may weaken the representation ability of visual data, leading to the degradation of performance.…”

Section: Introductionmentioning

confidence: 99%

“…In this paper, we introduce a novel lightweight network, called FDDWNet, for the task of real-time semantic segmentation. Different from previous methods [11,14,16,17,18], our FDDWNet makes an effort to design more deeper network architecture to develop the ability of feature representation, while maintains very fewer model parameters to accelerate inference speed. In addition, FDDWNet has multiple branches of skipped connections from intermediate convolution layers, in spite of adding a bit of computational burden, but helping to gather more context.…”

Section: Introductionmentioning

confidence: 99%

See 3 more Smart Citations

FDDWNet: A Lightweight Convolutional Neural Network for Real-Time Semantic Segmentation

Liu

Zhou

Qiang

et al. 2020

ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

View full text Add to dashboard Cite

This paper introduces a lightweight convolutional neural network, called FDDWNet, for real-time accurate semantic segmentation. In contrast to recent advances of lightweight networks that prefer to utilize shallow structure, FDDWNet makes an effort to design more deeper network architecture, while maintains faster inference speed and higher segmentation accuracy. Our network uses factorized dilated depth-wise separable convolutions (FDDWC) to learn feature representations from different scale receptive fields with fewer model parameters. Additionally, FDDWNet has multiple branches of skipped connections to gather context cues from intermediate convolution layers. The experiments show that FDDWNet only has 0.8M model size, while achieves 60 FPS running speed on a single RTX 2080Ti GPU with a 1024 × 512 input image.The comprehensive experiments demonstrate that our model achieves state-of-the-art results in terms of available speed and accuracy trade-off on CityScapes and CamVid datasets.

show abstract

“…Standard [21] n 2 cĉ n × n Group [15] n 2 cĉ/g n × n 1D-factorized [13] 2ncĉ n × n DW [11,12] n 2 c + cĉ n × n DDW [14] n 2 c + cĉ n r × n r our FDDWC 2nc + cĉ n r × n r…”

Section: Convolutional Type Parameters Size Of Receptive Fieldmentioning

confidence: 99%

Section: Eerm Unitmentioning

confidence: 99%

Section: Introductionmentioning

confidence: 99%

Section: Introductionmentioning

confidence: 99%

See 2 more Smart Citations

FDDWNet: A Lightweight Convolutional Neural Network for Real-Time Semantic Segmentation

Liu

Zhou

Qiang

et al. 2020

ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

View full text Add to dashboard Cite

show abstract

“…Recently, there has been a growing interest into developing space-time efficient neural networks for real time and resource restricted applications [10,6,7,9,8,11,12]. Depthwise Separable Convolutions.…”

Section: Related Workmentioning

confidence: 99%

Depthwise-STFT Based Separable Convolutional Neural Networks

Kumawat

Raman

2020

ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

View full text Add to dashboard Cite

In this paper, we propose a new convolutional layer called Depthwise-STFT Separable layer that can serve as an alternative to the standard depthwise separable convolutional layer. The construction of the proposed layer is inspired by the fact that the Fourier coefficients can accurately represent important features such as edges in an image. It utilizes the Fourier coefficients computed (channelwise) in the 2D local neighborhood (e.g., 3 × 3) of each position of the input map to obtain the feature maps. The Fourier coefficients are computed using 2D Short Term Fourier Transform (STFT) at multiple fixed low frequency points in the 2D local neighborhood at each position. These feature maps at different frequency points are then linearly combined using trainable pointwise (1 × 1) convolutions. We show that the proposed layer outperforms the standard depthwise separable layer based models on the CIFAR-10 and CIFAR-100 image classification datasets with reduced space-time complexity.

show abstract

ASFNet: Adaptive multiscale segmentation fusion network for real‐time semantic segmentation

Zha

Liu

Yang

et al. 2021

Computer Animation & Virtual

View full text Add to dashboard Cite

Recently, the development of deep learning has facilitated continuous progress in the field of computer vision. Pixel-level semantic segmentation serves as a fundamental task in computer vision. It achieves significant results by connecting wider and deeper backbone networks and building fine-grained segmentation heads. However, applications such as self-driving cars are more critical to the computational speed of the algorithms. The trade-off between accuracy and real-time performance of existing algorithms is still a challenging task. To address this challenge, this article proposes an adaptive multiscale segmentation fusion network to fuse multiscale contextual, which designs an adaptive multiscale segmentation fusion module based on an attention mechanism. Using segmentation fusion instead of feature fusion, the multiscale segmentation results are aggregated to obtain more precise segmentation results. The final results achieved 70.9% mIoU of accuracy in the Cityspace test set, processing images at 61 FPS when the input is 1024 × 2048. In addition, when adjusting the input size to 512 × 1024, the images are processed at 185 FPS.

show abstract

ESPNetv2: A Light-Weight, Power Efficient, and General Purpose Convolutional Neural Network

Cited by 485 publications

References 45 publications

FDDWNet: A Lightweight Convolutional Neural Network for Real-Time Semantic Segmentation

FDDWNet: A Lightweight Convolutional Neural Network for Real-Time Semantic Segmentation

Depthwise-STFT Based Separable Convolutional Neural Networks

ASFNet: Adaptive multiscale segmentation fusion network for real‐time semantic segmentation

Contact Info

Product

Resources

About