Extraction of standardized images from a single-view or multi-view capture
Abstract
According to various embodiments, component information may be identified for each input image of an object. The component information may indicate a portion of the input image in which a particular component of the object is depicted. A viewpoint may be determined for each input image that indicates a camera pose for the input image relative to the object. A three-dimensional skeleton of the object may be determined based on the viewpoints and the component information. A multi-view panel corresponding to the designated component of the object that is navigable in three dimensions and that the portions of the input images in which the designated component of the object is depicted may be stored on a storage device.
Claims
exact text as granted — not AI-modified1 . A method comprising:
receiving a request to generate a multi-view panel of an object, wherein the object is a vehicle, wherein the multi-view panel is a top-down multi-view panel showing a plurality of designated object components including a vehicle roof, a plurality of vehicle doors, and a vehicle windshield, the request identifying a plurality of designated viewpoints of the designated object component, each of the designated viewpoints specifying a respective camera pose with respect to the designated object component; identifying via a processor respective component information for each of a plurality of input images of the object, the plurality of input images selected from a multi-view capture of the object, the multi-view capture used to generate a multi-view interactive digital media representation navigable in multiple dimensions by a viewer, the respective component information indicating a respective portion of the respective input image in which the designated component of the object is depicted; determining a respective viewpoint for each input image via the processor, the respective viewpoint indicating a respective camera pose for the respective input image relative to the object; mapping each of the input images to a top-down view associated with the top-down multi-view panel of the object; evaluating via the processor the plurality of input images based on the mapping of the input images to the top-down view to select a subset of the images that includes the designated object component and that is associated with a respective viewpoint that matches one or more of the designated viewpoints; and storing on a storage device the multi-view panel.
2 . The method of claim 1 , wherein the processor generates the three-dimensional skeleton of the object based on the respective viewpoints, a top-down view of the object, and the respective component information, wherein the three-dimensional skeleton is determined by putting each of the plurality of input images in a neural network, detecting a two-dimensional skeleton in each of the plurality of input images using the neural network, and combing the two-dimensional skeletons to form the three-dimensional skeleton.
3 . The method of claim 2 , wherein the processor creates a 3D mesh from the three-dimensional skeleton.
4 . The method of claim 3 , wherein mapping each of the input images to the top-down view includes by projecting a plurality of pixels in the input images onto the three-dimensional skeleton of the object.
5 . The method of claim 4 , wherein mapping each of the input images to the top-down view includes selecting a pixel in a perspective frame and projecting the pixel onto a 3D mesh by simulating a camera ray passing by the pixel's position into the 3D mesh.
6 . The method of claim 5 , wherein mapping each of the input images to a top-down view of the object includes extracting barycentric coordinates of an intersection point with respect to vertices of an intersection face.
7 . The method of claim 1 , wherein the multi-view panel is generated based on target viewpoint information defined in the top-down view of the object.
8 . The method of claim 1 , wherein the plurality of images form a multi-view capture of the object navigable in three dimensions, the multi-view capture constructed based in part on inertial measurement unit (IMU) data collected from an IMU in a mobile phone.
9 . The method of claim 5 , wherein the object is a vehicle, and wherein the three-dimensional skeleton includes a door and a windshield.
10 . The method of claim 1 , wherein the respective viewpoint further includes a respective distance of the camera from the object.
11 . The method of claim 1 , wherein the respective camera pose includes a respective vertical angle identifying a respective angular height of the viewpoint relative to a 2D plane parallel to a surface on which the object is situated.
12 . The method of claim 1 , wherein the respective camera pose includes a respective rotational angle identifying a respective degree of rotation of the viewpoint relative to a designated fixed position of an object.
13 . The method of claim 1 , wherein the respective camera pose includes a respective position identifying a respective position of the viewpoint relative to a designated fixed position of the object.
14 . A system comprising:
an interface configured to receive a request to generate a multi-view panel of an object, wherein the object is a vehicle, wherein the multi-view panel is a top-down multi-view panel showing a plurality of designated object components including a vehicle roof, a plurality of vehicle doors, and a vehicle windshield, the request identifying a plurality of designated viewpoints of the designated object component, each of the designated viewpoints specifying a respective camera pose with respect to the designated object component; a processor configured to identify via a processor respective component information for each of a plurality of input images of the object, the plurality of input images selected from a multi-view capture of the object, the multi-view capture used to generate a multi-view interactive digital media representation navigable in multiple dimensions by a viewer, the respective component information indicating a respective portion of the respective input image in which the designated component of the object is depicted; wherein the processor is further configured to determine a respective viewpoint for each input image via the processor, the respective viewpoint indicating a respective camera pose for the respective input image relative to the object, map each of the input images to a top-down view associated with the top-down multi-view panel of the object, and evaluate the plurality of input images based on the mapping of the input images to the top-down view to select a subset of the images that includes the designated object component and that is associated with a respective viewpoint that matches one or more of the designated viewpoints; and memory configured to store the multi-view panel.
15 . The system of claim 14 , wherein the processor generates the three-dimensional skeleton of the object based on the respective viewpoints, a top-down view of the object, and the respective component information, wherein the three-dimensional skeleton is determined by putting each of the plurality of input images in a neural network, detecting a two-dimensional skeleton in each of the plurality of input images using the neural network, and combing the two-dimensional skeletons to form the three-dimensional skeleton.
16 . The system of claim 15 , wherein the processor creates a 3D mesh from the three-dimensional skeleton.
17 . The system of claim 16 , wherein mapping each of the input images to the top-down view includes by projecting a plurality of pixels in the input images onto the three-dimensional skeleton of the object.
18 . The system of claim 17 , wherein mapping each of the input images to the top-down view includes selecting a pixel in a perspective frame and projecting the pixel onto a 3D mesh by simulating a camera ray passing by the pixel's position into the 3D mesh.
19 . The system of claim 18 , wherein mapping each of the input images to a top-down view of the object includes extracting barycentric coordinates of an intersection point with respect to vertices of an intersection face.
20 . The system of claim 14 , wherein the multi-view panel is generated based on target viewpoint information defined in the top-down view of the object.Join the waitlist — get patent alerts
Track US2023419438A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.