Return
CPoser: An Optimization-after-Parsing Approach for Text-to-PoseGeneration Using Large Language Models
DOI:10.1145/3687932.png)
Abstract
En 中文
Text-to-pose generation is challenging due to the complexity of naturallanguage and human posture semantics. Utilizing large language models(LLMs) for text-to-pose generation is appealing due to their strong capabili-ties in text understanding and reasoning. However, as LLMs are designed forgeneral-purpose language processing and not specifically trained for posegeneration, it remains nontrivial to generate precise articulation targets for the full body using LLMs directly. To this end, we propose CPoser, a novelapproach to harness the power of LLMs for text-to-pose generation, featur-ing a prompt parsing stage and a pose optimization stage. The parsing stageutilizes LLMs to turn text prompts into pose intermediate representations(Pose-IRs) through a set of predefined structured queries. These Pose-IRsexplicitly describe specific pose conditions, such as squatting depth and kneebending angle, naturally forming an objective function that a target poseshould satisfy. The optimization stage solves for expressive poses and handgestures based on the Pose-IR objective function via robust optimizationin a quantized pose prior space. The results are further refined to enhancenaturalness and incorporate facial expressions. Experiments show that ourapproach effectively understands diverse text prompts for pose generation,surpassing existing text-to-pose methods
Keywords:
Human posture
text-to-pose generation
zero-shot learning
pose priors
large language models
Journal
IF:
9.5
Papers:
4.7K
Citations:
3.6W

