Jan Nogga scite author profile

Jan Nogga

4Publications

3Citation Statements Received

15Citation Statements Given

How they've been cited

How they cite others

Affiliations

University of Bonn

Publications

Order By: Most citations

Local Frequency Domain Transformer Networks for Video Prediction

Farazi

Nogga

Behnke

2021

View full text Add to dashboard Cite

Video prediction is commonly referred to as forecasting future frames of a video sequence provided several past frames thereof. It remains a challenging domain as visual scenes evolve according to complex underlying dynamics, such as the camera's egocentric motion or the distinct motility per individual object viewed. These are mostly hidden from the observer and manifest as often highly non-linear transformations between consecutive video frames. Therefore, video prediction is of interest not only in anticipating visual changes in the real world but has, above all, emerged as an unsupervised learning rule targeting the formation and dynamics of the observed environment. Many of the deep learning-based state-of-the-art models for video prediction utilize some form of recurrent layers like Long Short-Term Memory (LSTMs) or Gated Recurrent Units (GRUs) at the core of their models. Although these models can predict the future frames, they rely entirely on these recurrent structures to simultaneously perform three distinct tasks: extracting transformations, projecting them into the future, and transforming the current frame. In order to completely interpret the formed internal representations, it is crucial to disentangle these tasks. This paper proposes a fully differentiable building block that can perform all of those tasks separately while maintaining interpretability. We derive the relevant theoretical foundations and showcase results on synthetic as well as real data. We demonstrate that our method is readily extended to perform motion segmentation and account for the scene's composition, and learns to produce reliable predictions in an entirely interpretable manner by only observing unlabeled video data.

show abstract

Semantic Prediction: Which One Should Come First, Recognition or Prediction?

Farazi¹,

Nogga²,

Behnke³

2021

Preprint

View full text Add to dashboard Cite

The ultimate goal of video prediction is not forecasting future pixel-values given some previous frames. Rather, the end goal of video prediction is to discover valuable internal representations from the vast amount of available unlabeled video data in a self-supervised fashion for downstream tasks. One of the primary downstream tasks is interpreting the scene's semantic composition and using it for decision-making. For example, by predicting human movements, an observer can anticipate human activities and collaborate in a shared workspace. There are two main ways to achieve the same outcome, given a pre-trained video prediction and pre-trained semantic extraction model; one can first apply predictions and then extract semantics or first extract semantics and then predict. We investigate these configurations using the Local Frequency Domain Transformer Network (LFDTN) as the video prediction model and U-Net as the semantic extraction model on synthetic and real datasets.

show abstract

Semantic Prediction: Which One Should Come First, Recognition or Prediction?

Farazi¹,

Nogga²,

Behnke³

2021

View full text Add to dashboard Cite

show abstract

Local Frequency Domain Transformer Networks for Video Prediction

Farazi¹,

Nogga²,

Behnke³

2021

Preprint

View full text Add to dashboard Cite

scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.

Contact Info

hi@scite.ai

10624 S. Eastern Ave., Ste. A-614

Henderson, NV 89052, USA

Blog Terms and Conditions API Terms Privacy Policy Contact Cookie Preferences Do Not Sell or Share My Personal Information

Made with 💙 for researchers

Part of the Research Solutions Family.

Jan Nogga

Local Frequency Domain Transformer Networks for Video Prediction

Semantic Prediction: Which One Should Come First, Recognition or Prediction?

Semantic Prediction: Which One Should Come First, Recognition or Prediction?

Local Frequency Domain Transformer Networks for Video Prediction

Contact Info

Product

Resources

About