From Parts to Whole: A Unified Reference Framework for Controllable Human Image Generation

Huang, Zehuan; Fan, Hongxing; Wang, Lipeng; Sheng, Lu

Computer Science > Computer Vision and Pattern Recognition

arXiv:2404.15267v1 (cs)

[Submitted on 23 Apr 2024]

Title:From Parts to Whole: A Unified Reference Framework for Controllable Human Image Generation

Authors:Zehuan Huang, Hongxing Fan, Lipeng Wang, Lu Sheng

View PDF HTML (experimental)

Abstract:Recent advancements in controllable human image generation have led to zero-shot generation using structural signals (e.g., pose, depth) or facial appearance. Yet, generating human images conditioned on multiple parts of human appearance remains challenging. Addressing this, we introduce Parts2Whole, a novel framework designed for generating customized portraits from multiple reference images, including pose images and various aspects of human appearance. To achieve this, we first develop a semantic-aware appearance encoder to retain details of different human parts, which processes each image based on its textual label to a series of multi-scale feature maps rather than one image token, preserving the image dimension. Second, our framework supports multi-image conditioned generation through a shared self-attention mechanism that operates across reference and target features during the diffusion process. We enhance the vanilla attention mechanism by incorporating mask information from the reference human images, allowing for the precise selection of any part. Extensive experiments demonstrate the superiority of our approach over existing alternatives, offering advanced capabilities for multi-part controllable human image customization. See our project page at this https URL.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2404.15267 [cs.CV]
	(or arXiv:2404.15267v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2404.15267

Submission history

From: Hongxing Fan [view email]
[v1] Tue, 23 Apr 2024 17:56:08 UTC (10,997 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:From Parts to Whole: A Unified Reference Framework for Controllable Human Image Generation

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:From Parts to Whole: A Unified Reference Framework for Controllable Human Image Generation

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators