Self-supervised Dance Video Synthesis Conditioned on Music

Ren, Xuanchi; Li, Haoran; Huang, Zijian; Chen, Qifeng

doi:10.1145/3394171.3413932

Cited by 48 publications

(29 citation statements)

References 22 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…Fig. 5 shows a group of translation results by using our method and previous state of the art methods on the music "Sorry" (also used in the previous work (Ren et al 2020)). It can be seen that the music-dance video generated by our method not only accurately capture the rhythm in the song, but also contain rich musical feelings and movement strength.…”

Section: Music-to-dance Translation Resultsmentioning

confidence: 99%

“…Lee et al propose a decomposition-tocomposition framework for music-to-dance generation (Lee et al 2019), where they use a VAE to model dance units and use a Generative Adversarial Network (GAN) to organize the dance units based on input music. Ren et al integrate the local temporal discriminator and the global content discriminator for helping generate coherent dance sequences based on the noisy dataset, and then use pose-toappearance mapping to generate human dance videos (Ren et al 2020). However, all the above methods directly generate the dance movements from music, which inevitably leads to a problem of motion degradation and is not yet able to meet the requirements of expert-level music-to-dance translation.…”

Section: Music-to-dance Translationmentioning

confidence: 99%

“…To better evaluate the quality of the generated dance phrases, subjective evaluations are further conducted. In this experiment, we first collect three groups of music: 1) music used in the previous method (Ren et al 2020), 2) music from our unlabeled test set, 3) unseen style music outside of our dataset. Note that all these musics are not shown in our training dataset.…”

Section: Subjective Evaluationmentioning

confidence: 99%

“…Recently, music-to-dance translation has drawn increasing research attention due to its wide applications in the game industry and virtual reality. Deep learning based methods have shown great potential in this task (Alemi, Franc ¸oise, and Pasquier 2017;Ren et al 2020). However, these methods are difficult to apply to in-game expert-level music-to-dance applications.…”

Section: Introductionmentioning

confidence: 99%

See 3 more Smart Citations

Semi-Supervised Learning for In-Game Expert-Level Music-to-Dance Translation

Duan,

Shi,

Zou

et al. 2020

Preprint

View full text Add to dashboard Cite

Music-to-dance translation is a brand-new and powerful feature in recent role-playing games. Players can now let their characters dance along with specified music clips and even generate fan-made dance videos. Previous works of this topic consider music-to-dance as a supervised motion generation problem based on time-series data. However, these methods suffer from limited training data pairs and the degradation of movements. This paper provides a new perspective for this task where we re-formulate the translation problem as a piece-wise dance phrase retrieval problem based on the choreography theory. With such a design, players are allowed to further edit the dance movements on top of our generation while other regression based methods ignore such user interactivity. Considering that the dance motion capture is an expensive and time-consuming procedure which requires the assistance of professional dancers, we train our method under a semi-supervised learning framework with a large unlabeled dataset (20x than labeled data) collected. A coascent mechanism is introduced to improve the robustness of our network. Using this unlabeled dataset, we also introduce self-supervised pre-training so that the translator can understand the melody, rhythm, and other components of music phrases. We show that the pre-training significantly improves the translation accuracy than that of training from scratch. Experimental results suggest that our method not only generalizes well over various styles of music but also succeeds in expert-level choreography for game players.

show abstract

Section: Music-to-dance Translation Resultsmentioning

confidence: 99%

Section: Music-to-dance Translationmentioning

confidence: 99%

Section: Subjective Evaluationmentioning

confidence: 99%

Section: Introductionmentioning

confidence: 99%

See 2 more Smart Citations

Semi-Supervised Learning for In-Game Expert-Level Music-to-Dance Translation

Duan,

Shi,

Zou

et al. 2020

Preprint

View full text Add to dashboard Cite

show abstract

“…For the reconstruction term in Eq. 3, we adopt a VGG perceptual loss (Simonyan & Zisserman, 2015;Ren et al, 2020), which is widely used in unsupervised disentanglement methods (Wu et al, 2020;2019c). For the Ψ-constraint, i.e.…”

Section: Proposed C-s Disentanglement Modulementioning

confidence: 99%

Rethinking Content and Style: Exploring Bias for Unsupervised Disentanglement

Ren¹,

Yang²,

Wang³

et al. 2021

Preprint

Self Cite

View full text Add to dashboard Cite

Content and style (C-S) disentanglement intends to decompose the underlying explanatory factors of objects into two independent subspaces. From the unsupervised disentanglement perspective, we rethink content and style and propose a formulation for unsupervised C-S disentanglement based on our assumption that different factors are of different importance and popularity for image reconstruction, which serves as a data bias. The corresponding model inductive bias is introduced by our proposed C-S disentanglement Module (C-S DisMo), which assigns different and independent roles to content and style when approximating the real data distributions. Specifically, each content embedding from the dataset, which encodes the most dominant factors for image reconstruction, is assumed to be sampled from a shared distribution across the dataset. The style embedding for a particular image, encoding the remaining factors, is used to customize the shared distribution through an affine transformation. The experiments on several popular datasets demonstrate that our method achieves the state-of-theart unsupervised C-S disentanglement, which is comparable or even better than supervised methods. We verify the effectiveness of our method by downstream tasks: domain translation and singleview 3D reconstruction. Project page at https: //github.com/xrenaa/CS-DisMo.

show abstract