Weeping and Gnashing of Teeth: Teaching Deep Learning in Image and Video Processing Classes

Bovik, A.C.

doi:10.1109/ssiai49293.2020.9094606

Cited by 4 publications

(2 citation statements)

References 22 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…Researches showed that transferring existing classification networks to regression tasks performed not well [18,20,22]. As noted by A. C. Bovik [23], "Unlike human participation in crowdsourced picture labeling experiments like ImageNet, where each human label might need only 0.5-1.0 seconds to apply, human quality judgments on pictures generally required 10-20x that amount to time for a subject to feel comfortable in making their assessments on a Likert scale [24]." In general, a clip of video has a long duration including hundreds of images, and its perceptual quality difference from other videos with different content is extremely subtle.…”

Section: Introductionmentioning

confidence: 99%

Starvqa: Space-Time Attention for Video Quality Assessment

Xing

Wang

et al. 2022

2022 IEEE International Conference on Image Processing (ICIP)

View full text Add to dashboard Cite

Transformer based on self-attention mechanism is blooming in computer vision nowadays. However, its application to video quality assessment (VQA) has not been reported. Evaluating the quality of in-the-wild videos is challenging due to the unknown of pristine reference and shooting distortion. This paper presents a novel space-time attention network for the VQA problem, named StarVQA. StarVQA builds a Transformer by alternately concatenating the divided space-time attention. To adapt the Transformer architecture for training, StarVQA designs a vectorized regression loss by encoding the mean opinion score (MOS) to the probability vector and embedding a special vectorized label token as the learnable variable. To capture the long-range spatiotemporal dependencies of a video sequence, StarVQA encodes the space-time position information of each patch to the input of the Transformer. Various experiments are conducted on the de-facto in-the-wild video datasets, including LIVE-VQC, KoNViD-1k, LSVQ, and LSVQ-1080p. Experimental results demonstrate the superiority of StarVQA over the state-of-theart. The source code is available at https://github.com/GZHU-DVL/StarVQA.

show abstract

Section: Introductionmentioning

confidence: 99%

Starvqa: Space-Time Attention for Video Quality Assessment

Xing

Wang

et al. 2022

2022 IEEE International Conference on Image Processing (ICIP)

View full text Add to dashboard Cite

show abstract

“…Researches showed that transferring existing classification networks to regression tasks performed not well [8], [3], [32]. As noted by A. C. Bovik [4], "Unlike human participation in crowdsourced picture labeling experiments like ImageNet, where each human label might need only 0.5-1.0 seconds to apply, human quality judgments on pictures generally required 10-20x that amount to time for a subject to feel comfortable in making their assessments on a Likert scale [10]." In general, a clip of video has a long duration including hundreds of images, and its perceptual quality difference from other videos with different content is extremely subtle.…”

Section: Introductionmentioning

confidence: 99%

StarVQA: Space-Time Attention for Video Quality Assessment

Xing

Wang

et al. 2021

Preprint

View full text Add to dashboard Cite

The attention mechanism is blooming in computer vision nowadays. However, its application to video quality assessment (VQA) has not been reported. Evaluating the quality of inthe-wild videos is challenging due to the unknown of pristine reference and shooting distortion. This paper presents a novel spacetime attention network for the VQA problem, named StarVQA. StarVQA builds a Transformer by alternately concatenating the divided space-time attention. To adapt the Transformer architecture for training, StarVQA designs a vectorized regression loss by encoding the mean opinion score (MOS) to the probability vector and embedding a special vectorized label token as the learnable variable. To capture the long-range spatiotemporal dependencies of a video sequence, StarVQA encodes the space-time position information of each patch to the input of the Transformer. Various experiments are conducted on the de-facto in-the-wild video datasets, including LIVE-VQC, KoNViD-1k, LSVQ, and LSVQ-1080p. Experimental results demonstrate the superiority of the proposed StarVQA over the state-of-the-art. Code and model will be available at: https://github.com/DVL/StarVQA.

show abstract