End-to-End Face Parsing via Interlinked Convolutional Neural Networks

Yin, Zi; Yiu, Valentin; Hu, Xiaolin; Tang, Liang

doi:10.48550/arxiv.2002.04831

Cited by 2 publications

(5 citation statements)

References 33 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…Region-based parsing takes the scale discrepancy into account and predicts each component respectively, which is advantageous in capturing elaborate details [12], [13], [14], [15], [16]. Zhou et al present an interlinked CNN that takes multi-scale images as input and allows bidirectional information passing [15].…”

Section: Related Work a Face Parsingmentioning

confidence: 99%

“…This method demonstrates high performance especially for hair segmentation [13]. Yin et al introduce the Spatial Transformer Network and build a training connection between traditional interlinked CNNs, which makes the end-to-end joint training process possible [14]. Nevertheless, this class of methods often neglect the correlation among components to characterize long range dependencies.…”

Section: Related Work a Face Parsingmentioning

confidence: 99%

“…This however may neglect the scale discrepancy in different facial components, sometimes resulting in the lack of details. 2) Region-based parsing, which addresses the above problem by first predicting the bounding box and then acquiring the parsing map of each facial component individually [12], [13], [14], [15], [16]. However, there still exist limitations and challenges.…”

Section: Introductionmentioning

confidence: 99%

See 2 more Smart Citations

AGRNet: Adaptive Graph Representation Learning and Reasoning for Face Parsing

Te,

Hu,

Liu

et al. 2021

Preprint

View full text Add to dashboard Cite

Face parsing infers a pixel-wise label to each facial component, which has drawn much attention recently. Previous methods have shown their success in face parsing, which however overlook the correlation among facial components. As a matter of fact, the component-wise relationship is a critical clue in discriminating ambiguous pixels in facial area. To address this issue, we propose adaptive graph representation learning and reasoning over facial components, aiming to learn representative vertices that describe each component, exploit the componentwise relationship and thereby produce accurate parsing results against ambiguity. In particular, we devise an adaptive and differentiable graph abstraction method to represent the components on a graph via pixel-to-vertex projection under the initial condition of a predicted parsing map, where pixel features within a certain facial region are aggregated onto a vertex. Further, we explicitly incorporate the image edge as a prior in the model, which helps to discriminate edge and non-edge pixels during the projection, thus leading to refined parsing results along the edges. Then, our model learns and reasons over the relations among components by propagating information across vertices on the graph. Finally, the refined vertex features are projected back to pixel grids for the prediction of the final parsing map. To train our model, we propose a discriminative loss to penalize small distances between vertices in the feature space, which leads to distinct vertices with strong semantics. Experimental results show the superior performance of the proposed model on multiple face parsing datasets, along with the validation on the human parsing task to demonstrate the generalizability of our model.

show abstract

Section: Related Work a Face Parsingmentioning

confidence: 99%

Section: Related Work a Face Parsingmentioning

confidence: 99%

Section: Introductionmentioning

confidence: 99%

See 1 more Smart Citation

AGRNet: Adaptive Graph Representation Learning and Reasoning for Face Parsing

Te,

Hu,

Liu

et al. 2021

Preprint

View full text Add to dashboard Cite

show abstract

“…We conduct experiments on the broadly acknowledged Helen dataset to demonstrate the superiority of the proposed model. To keep consistent with the previous works [6,4,38,5,33], we employ the overall F1 score to measure the performance, which is computed by combining the merged eyes, brows, nose and mouth categories. As Table 3 shows, Our model surpasses state-of-the-art methods and achieves 93.2% on this dataset.…”

Section: Comparison With the State-of-the-artmentioning

confidence: 99%

“…The region-based methods have been recently proposed to model the facial components separately [4,5,6], and achieved state-of-the-art performance on the current benchmarks. However, these methods are based on the individual information within each region, and the correlation among regions is not exploited yet to capture long range dependencies.…”

Section: Introductionmentioning

confidence: 99%

Edge-aware Graph Representation Learning and Reasoning for Face Parsing

Liu

et al. 2020

Preprint

View full text Add to dashboard Cite

Face parsing infers a pixel-wise label to each facial component, which has drawn much attention recently. Previous methods have shown their efficiency in face parsing, which however overlook the correlation among different face regions. The correlation is a critical clue about the facial appearance, pose, expression, etc., and should be taken into account for face parsing. To this end, we propose to model and reason the region-wise relations by learning graph representations, and leverage the edge information between regions for optimized abstraction. Specifically, we encode a facial image onto a global graph representation where a collection of pixels ("regions") with similar features are projected to each vertex. Our model learns and reasons over relations between the regions by propagating information across vertices on the graph. Furthermore, we incorporate the edge information to aggregate the pixel-wise features onto vertices, which emphasizes on the features around edges for fine segmentation along edges. The finally learned graph representation is projected back to pixel grids for parsing. Experiments demonstrate that our model outperforms state-of-the-art methods on the widely used Helen dataset, and also exhibits the superior performance on the large-scale CelebAMask-HQ and LaPa dataset. The code is available at https://github.com/tegusi/EAGRNet.

show abstract

End-to-End Face Parsing via Interlinked Convolutional Neural Networks

Cited by 2 publications

References 33 publications

AGRNet: Adaptive Graph Representation Learning and Reasoning for Face Parsing

AGRNet: Adaptive Graph Representation Learning and Reasoning for Face Parsing

Edge-aware Graph Representation Learning and Reasoning for Face Parsing

Contact Info

Product

Resources

About