Location Dependency in Video Prediction

Azizi, Niloofar; Farazi, Hafez; Behnke, Sven

doi:10.1007/978-3-030-01424-7_62

Cited by 8 publications

(5 citation statements)

References 6 publications

(7 reference statements)

Supporting

Mentioning

Contrasting

Order By: Relevance

“…We evaluated our LFDTN model on this dataset. We compared our model against many well-known models like Conv-PGP [8], VLN-ResNet [23], VLN-LDC [21], HPNetT [6], PredRNN [24], † http://ais.uni-bonn.de/~hfarazi/LFDTN/ Fig. 7: Internal states' formation for two samples for color a) "Moving MNIST on STL", and b) "NGSIM" datasets.…”

Section: Resultsmentioning

confidence: 99%

“…To get a better result and help the model decide how reliable each local velocity is, we also include each direction's variance [𝜎 2 𝑥 , 𝜎 2 𝑦 ] and concatenate it channel-wise. Since a convolutional model cannot learn location-dependent features, similar to the positional encoding proposed by Azizi et.al [21], we add two additional channels to the input. To grant the model the ability account for former velocities, we channelwise concatenate the same saved features from previous time steps.…”

Section: A Transform Modelmentioning

confidence: 99%

See 1 more Smart Citation

Local Frequency Domain Transformer Networks for Video Prediction

Farazi

Nogga

Behnke

2021

2021 International Joint Conference on Neural Networks (IJCNN)

Self Cite

View full text Add to dashboard Cite

Video prediction is commonly referred to as forecasting future frames of a video sequence provided several past frames thereof. It remains a challenging domain as visual scenes evolve according to complex underlying dynamics, such as the camera's egocentric motion or the distinct motility per individual object viewed. These are mostly hidden from the observer and manifest as often highly non-linear transformations between consecutive video frames. Therefore, video prediction is of interest not only in anticipating visual changes in the real world but has, above all, emerged as an unsupervised learning rule targeting the formation and dynamics of the observed environment. Many of the deep learning-based state-of-the-art models for video prediction utilize some form of recurrent layers like Long Short-Term Memory (LSTMs) or Gated Recurrent Units (GRUs) at the core of their models. Although these models can predict the future frames, they rely entirely on these recurrent structures to simultaneously perform three distinct tasks: extracting transformations, projecting them into the future, and transforming the current frame. In order to completely interpret the formed internal representations, it is crucial to disentangle these tasks. This paper proposes a fully differentiable building block that can perform all of those tasks separately while maintaining interpretability. We derive the relevant theoretical foundations and showcase results on synthetic as well as real data. We demonstrate that our method is readily extended to perform motion segmentation and account for the scene's composition, and learns to produce reliable predictions in an entirely interpretable manner by only observing unlabeled video data.

show abstract

Section: Resultsmentioning

confidence: 99%

Section: A Transform Modelmentioning

confidence: 99%

Local Frequency Domain Transformer Networks for Video Prediction

Farazi

Nogga

Behnke

2021

2021 International Joint Conference on Neural Networks (IJCNN)

Self Cite

View full text Add to dashboard Cite

show abstract

“…While requiring a large number of parameters, they lack interpretability. Two successful examples are Location-Dependent Video Ladder Networks [5] and PredRNN++ [6], which employs a stack of LSTM modules to predict plausible future frames.…”

Section: Related Workmentioning

confidence: 99%

Semantic Prediction: Which One Should Come First, Recognition or Prediction?

Farazi¹,

Nogga²,

Behnke³

2021

Preprint

Self Cite

View full text Add to dashboard Cite

The ultimate goal of video prediction is not forecasting future pixel-values given some previous frames. Rather, the end goal of video prediction is to discover valuable internal representations from the vast amount of available unlabeled video data in a self-supervised fashion for downstream tasks. One of the primary downstream tasks is interpreting the scene's semantic composition and using it for decision-making. For example, by predicting human movements, an observer can anticipate human activities and collaborate in a shared workspace. There are two main ways to achieve the same outcome, given a pre-trained video prediction and pre-trained semantic extraction model; one can first apply predictions and then extract semantics or first extract semantics and then predict. We investigate these configurations using the Local Frequency Domain Transformer Network (LFDTN) as the video prediction model and U-Net as the semantic extraction model on synthetic and real datasets.

show abstract

“…Transpose-convolutional layers are used for up-sampling the representations. To use location-dependent features, we used newly proposed location-dependent convolutional layer [15]. In order to limit the number of parameters used, a shared learnable bias between both output heads is implemented.…”

Section: Deep Learning Visual Perceptionmentioning

confidence: 99%

RoboCup 2019 AdultSize Winner NimbRo: Deep Learning Perception, In-Walk Kick, Push Recovery, and Team Play Capabilities

Rodríguez¹,

Farazi²,

Ficht³

et al. 2019

RoboCup 2019: Robot World Cup XXIII

Self Cite

View full text Add to dashboard Cite

Individual and team capabilities are challenged every year by rule changes and the increasing performance of the soccer teams at RoboCup Humanoid League. For RoboCup 2019 in the AdultSize class, the number of players (2 vs. 2 games) and the field dimensions were increased, which demanded for team coordination and robust visual perception and localization modules. In this paper, we present the latest developments that lead team NimbRo to win the soccer tournament, drop-in games, technical challenges and the Best Humanoid Award of the RoboCup Humanoid League 2019 in Sydney. These developments include a deep learning vision system, in-walk kicks, step-based pushrecovery, and team play strategies.

show abstract

Location Dependency in Video Prediction

Cited by 8 publications

References 6 publications

Local Frequency Domain Transformer Networks for Video Prediction

Local Frequency Domain Transformer Networks for Video Prediction

Semantic Prediction: Which One Should Come First, Recognition or Prediction?

RoboCup 2019 AdultSize Winner NimbRo: Deep Learning Perception, In-Walk Kick, Push Recovery, and Team Play Capabilities

Contact Info

Product

Resources

About