Xinchen Yan scite author profile

Abstract. This paper investigates a novel problem of generating images from visual attributes. We model the image as a composite of foreground and background and develop a layered generative model with disentangled latent variables that can be learned end-to-end using a variational auto-encoder. We experiment with natural images of faces and birds and demonstrate that the proposed models are capable of generating realistic and diverse samples with disentangled latent representations. We use a general energy minimization algorithm for posterior inference of latent variables given novel images. Therefore, the learned generative models show excellent quantitative and visual results in the tasks of attributeconditioned image reconstruction and completion.

show abstract

MT-VAE: Learning Motion Transformations to Generate Multimodal Human Dynamics

Yan

Rastogi

Villegas

et al. 2018

111

View full text Add to dashboard Cite

Long-term human motion can be represented as a series of motion modes-motion sequences that capture short-term temporal dynamics-with transitions between them. We leverage this structure and present a novel Motion Transformation Variational Auto-Encoders (MT-VAE) for learning motion sequence generation. Our model jointly learns a feature embedding for motion modes (that the motion sequence can be reconstructed from) and a feature transformation that represents the transition of one motion mode to the next motion mode. Our model is able to generate multiple diverse and plausible motion sequences in the future from the same input. We apply our approach to both facial and full body motion, and demonstrate applications like analogy-based motion transfer and video synthesis. * Work partially done during internship with Adobe Research.

show abstract

Block-NeRF: Scalable Large Scene Neural View Synthesis

Tancik

Casser²,

Yan³

et al. 2022

297

View full text Add to dashboard Cite

Learning 6-DOF Grasping Interaction via Deep Geometry-Aware 3D Representations

Yan¹,

Hsu

Khansari³

et al. 2018

104

View full text Add to dashboard Cite

This paper focuses on the problem of learning 6-DOF grasping with a parallel jaw gripper in simulation. Our key idea is constraining and regularizing grasping interaction learning through 3D geometry prediction. We introduce a deep geometry-aware grasping network (DGGN) that decomposes the learning into two steps. First, we learn to build mental geometry-aware representation by reconstructing the scene (i.e., 3D occupancy grid) from RGBD input via generative 3D shape modeling. Second, we learn to predict grasping outcome with its internal geometry-aware representation. The learned outcome prediction model is used to sequentially propose grasping solutions via analysis-by-synthesis optimization. Our contributions are fourfold: (1) To best of our knowledge, we are presenting for the first time a method to learn a 6-DOF grasping net from RGBD input; (2) We build a grasping dataset from demonstrations in virtual reality with rich sensory and interaction annotations. This dataset includes 101 everyday objects spread across 7 categories, additionally, we propose a data augmentation strategy for effective learning; (3) We demonstrate that the learned geometry-aware representation leads to about 10% relative performance improvement over the baseline CNN on grasping objects from our dataset. (4) We further demonstrate that the model generalizes to novel viewpoints and object instances.

show abstract

SemanticAdv: Generating Adversarial Examples via Attribute-Conditioned Image Editing

et al. 2020

View full text Add to dashboard Cite

scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.

Contact Info

hi@scite.ai

10624 S. Eastern Ave., Ste. A-614

Henderson, NV 89052, USA

Blog Terms and Conditions API Terms Privacy Policy Contact Cookie Preferences Do Not Sell or Share My Personal Information

Made with 💙 for researchers

Part of the Research Solutions Family.

Xinchen Yan

Attribute2Image: Conditional Image Generation from Visual Attributes

MT-VAE: Learning Motion Transformations to Generate Multimodal Human Dynamics

Block-NeRF: Scalable Large Scene Neural View Synthesis

Learning 6-DOF Grasping Interaction via Deep Geometry-Aware 3D Representations

SemanticAdv: Generating Adversarial Examples via Attribute-Conditioned Image Editing

Contact Info

Product

Resources

About