DSE-GAN: Dynamic Semantic Evolution Generative Adversarial Network for Text-to-Image Generation

Huang, Mengqi; Mao, Zhendong; Wang, Peng-Hui; Wang, Quan; Zhang, Yongdong

doi:10.1145/3503161.3547881

Cited by 10 publications

(2 citation statements)

References 43 publications

(97 reference statements)

Supporting

Mentioning

Contrasting

Order By: Relevance

“…Early works mainly adopted GAN (Goodfellow et al 2014) as the foundational generative model for this task. Various works have been proposed (Zhang et al 2021;Zhu et al 2019;Xu et al 2018;Zhang et al 2017;Zhang, Xie, and Yang 2018;Liang, Pei, and Lu 2020;Cheng et al 2020;Ruan et al 2021;Tao et al 2020;Li et al 2019;Huang et al 2022) with well-designed textual representations, elegant text-image interactions, and effective loss functions. However, GAN-based models often suffer from training instability and model collapse, making it hard to be trained on largescale datasets (Brock, Donahue, and Simonyan 2018;Kang et al 2023;Schuhmann et al 2021).…”

Section: Related Work Text-to-image Generationmentioning

confidence: 99%

DreamIdentity: Enhanced Editability for Efficient Face-Identity Preserved Image Generation

Chen,

Fang,

Liu

et al. 2024

AAAI

View full text Add to dashboard Cite

While large-scale pre-trained text-to-image models can synthesize diverse and high-quality human-centric images, an intractable problem is how to preserve the face identity and follow the text prompts simultaneously for conditioned input face images and texts. Despite existing encoder-based methods achieving high efficiency and decent face similarity, the generated image often fails to follow the textual prompts. To ease this editability issue, we present DreamIdentity, to learn edit-friendly and accurate face-identity representations in the word embedding space. Specifically, we propose self-augmented editability learning to enhance the editability for projected embedding, which is achieved by constructing paired generated celebrity's face and edited celebrity images for training, aiming at transferring mature editability of off-the-shelf text-to-image models in celebrity to unseen identities. Furthermore, we design a novel dedicated face-identity encoder to learn an accurate representation of human faces, which applies multi-scale ID-aware features followed by a multi-embedding projector to generate the pseudo words in the text embedding space directly. Extensive experiments show that our method can generate more text-coherent and ID-preserved images with negligible time overhead compared to the standard text-to-image generation process.

show abstract

Section: Related Work Text-to-image Generationmentioning

confidence: 99%

DreamIdentity: Enhanced Editability for Efficient Face-Identity Preserved Image Generation

Chen,

Fang,

Liu

et al. 2024

AAAI

View full text Add to dashboard Cite

show abstract

“…Then, the top-𝑘 measures, such as P@𝑘, mAP@𝑘, and NDCG@𝑘, will not drop significantly while the specified label has been ranked far behind the top-𝑘 position. In the extreme case, the classifier can never detect certain important categories and produce degraded classification performance, which may result in serious consequences for multimedia applications, such as multimodal emotion recognition [35] and text-to-image generation [20].…”

Section: Introductionmentioning

confidence: 99%

When Measures are Unreliable: Imperceptible Adversarial Perturbations toward Top-k Multi-Label Learning

Sun,

Xu,

Wang

et al. 2023

Proceedings of the 31st ACM International Conference on Multimedia

View full text Add to dashboard Cite

With the great success of deep neural networks, adversarial learning has received widespread attention in various studies, ranging from multi-class learning to multi-label learning. However, existing adversarial attacks toward multi-label learning only pursue the traditional visual imperceptibility but ignore the new perceptible problem coming from measures such as Precision@𝑘 and mAP@𝑘. Specifically, when a well-trained multi-label classifier performs far below the expectation on some samples, the victim can easily realize that this performance degeneration stems from attack, rather than the model itself. Therefore, an ideal multi-labeling adversarial attack should manage to not only deceive visual perception but also evade monitoring of measures. To this end, this paper first proposes the concept of measure imperceptibility. Then, a novel loss function is devised to generate such adversarial perturbations that could achieve both visual and measure imperceptibility. Furthermore, an efficient algorithm, which enjoys a convex objective, is established to optimize this objective. Finally, extensive experiments on large-scale benchmark datasets, such as PASCAL VOC 2012, MS COCO, and NUS WIDE, demonstrate the superiority of our proposed method in attacking the top-𝑘 multi-label systems. Our code is available at: https://github.com/Yuchen-Sunflower/TKMIA. CCS CONCEPTS• Computing methodologies → Ranking; Classification and regression trees.

show abstract

GH-DDM: the generalized hybrid denoising diffusion model for medical image generation

Zhang

Liu

et al. 2023

Multimedia Systems

Self Cite

View full text Add to dashboard Cite

show abstract

DSE-GAN: Dynamic Semantic Evolution Generative Adversarial Network for Text-to-Image Generation

Cited by 10 publications

References 43 publications

DreamIdentity: Enhanced Editability for Efficient Face-Identity Preserved Image Generation

DreamIdentity: Enhanced Editability for Efficient Face-Identity Preserved Image Generation

When Measures are Unreliable: Imperceptible Adversarial Perturbations toward Top-k Multi-Label Learning

GH-DDM: the generalized hybrid denoising diffusion model for medical image generation

Contact Info

Product

Resources

About