Guided Multimodal Virtual Try-On
Abstract
The technology is generally directed to using a machine learning model to generate a visualization of an intended wearer wearing one or more source garments by preserving the appearance and shape of the intended wearer and garments, even when the pose and/or environmental details are adjusted based on natural language inputs. The model may include a plurality of layers, each conditioned to warp a different type of garment, adjust the pose of the intended wearer, adjust the environmental details in the output image, etc. The plurality of layers may allow for the garments to be warped simultaneously, as compared to sequentially. When executing the model, the model may receive input images and/or natural language inputs corresponding to the intended wearer, garments, pose, environmental details, or the like. The output of the model may be a realistic visualization of the intended wearer wearing the input garments in an intended pose.
Claims
exact text as granted — not AI-modified1 . A method, comprising:
receiving, by one or more processors from a user device, inputs including natural language inputs corresponding to a description of a wearer pose and at least one intended garment; executing, by the one or more processors based on the received inputs, a machine learning model, wherein the executing comprises:
adjusting a pose of a wearer based on the natural language inputs corresponding to the description of the wearer pose; and
conditioning, using a respective machine learning layer of the machine learning model, the at least one intended garment to the adjusted pose of the wearer; and
providing for output on the user device, by the one or more processors, the at least one intended garment on the wearer in the wearer pose.
2 . The method of claim 1 , wherein when the at least one intended garment includes two or more intended garments, each of the two or more intended garments are conditioned simultaneously to the adjusted pose of the wearer.
3 . The method of claim 2 , wherein:
executing the machine learning model further comprises simultaneously conditioning, using respective machine learning layers of the machine learning model, the two or more intended garments to the adjusted pose of the wearer, the respective machine learning layers being different machine learning layers.
4 . The method of claim 1 , wherein the natural language inputs further include environmental details, and wherein the output further includes an environment for the wearer based on the environmental details.
5 . The method of claim 4 , wherein the environmental details correspond to at least one of lighting, texture, background, color, angle, or image filter.
6 . The method of claim 4 , wherein executing the machine learning model further comprises adjusting an output environment based on the natural language inputs corresponding to the environmental details.
7 . The method of claim 1 , wherein:
the machine learning model comprises one or more machine learning layers, each machine learning layer corresponds to a type of garment, and each machine learning layer, when executed, provides as output a corresponding garment layer.
8 . The method of claim 7 , wherein the type of garment includes at least one of a lower body garment, an upper body garment, an accessory, or shoes.
9 . The method of claim 7 , wherein executing the machine learning model further comprises ordering the garment layers based on the received input.
10 . The method of claim 1 , further comprising
receiving, by the one or more processors, second input; and executing, by the one or more processors based on the received second input, the machine learning model, wherein executing the machine learning model further comprises:
adjusting a depiction of a second garment in a second garment layer based on the second input.
11 . The method of claim 10 , wherein adjusting the second garment in the second garment layer does not change the depiction of the at least one intended garment in a first garment layer.
12 . The method of claim 10 , wherein the second input includes at least one of: an addition of a second intended garment, removal of the intended garment, or an adjustment of an order of the layers of the at least one intended garment.
13 . A system, comprising:
one or more processors, the one or more processors configured to:
receive, from a user device, inputs including natural language inputs corresponding to a description of a wearer pose and at least one intended garment;
execute, based on the received inputs, a machine learning model, wherein the executing comprises:
adjusting a pose of a wearer based on the natural language inputs corresponding to the description of the wearer pose; and
conditioning, using a respective machine learning layer of the machine learning model, the at least one intended garment to the adjusted pose of the wearer; and
provide for output on the user device, by the one or more processors, the at least one intended garment on the wearer in the wearer pose.
14 . The system of claim 13 , wherein when the at least one intended garment includes two or more intended garments, each of the two or more intended garments are conditioned simultaneously to the adjusted pose of the wearer.
15 . The system of claim 14 , wherein the one or more processors, when executing the machine learning model, are further configured to simultaneously condition, using respective machine learning layers of the machine learning model, the two or more intended garments to the adjusted pose of the wearer, the respective machine learning layers being different machine learning layers.
16 . The system of claim 13 , wherein the natural language inputs further include environmental details, and wherein the output further includes an environment for the wearer based on the environmental details.
17 . The system of claim 16 , wherein the environmental details correspond to at least one of lighting, texture, background, color, angle, or image filter.
18 . The system of claim 16 , wherein executing the machine learning model further comprises adjusting an output environment based on the natural language inputs corresponding to the environmental details.
19 . The system of claim 13 , wherein:
the machine learning model comprises one or more machine learning layers, each machine learning layer corresponds to a type of garment, and each machine learning layer, when executed, provides as output a corresponding garment layer.
20 . A non-transitory computer-readable medium storing instructions, which when executed by one or more processors, cause the one or more processors to:
receive, from a user device, inputs including natural language inputs corresponding to a description of a wearer pose and at least one intended garment; execute, based on the received inputs, a machine learning model, wherein the executing comprises:
adjusting a pose of a wearer based on the natural language inputs corresponding to the description of the wearer pose; and
conditioning, using a respective machine learning layer of the machine learning model, the at least one intended garment to the adjusted pose of the wearer; and
provide for output on the user device, by the one or more processors, the at least one intended garment on the wearer in the wearer pose.Join the waitlist — get patent alerts
Track US2025037343A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.