US2023410276A1PendingUtilityA1

Systems and methods for object dimensioning based on partial visual information

Assignee: PACKSIZE LLCPriority: Dec 20, 2018Filed: Sep 6, 2023Published: Dec 21, 2023
Est. expiryDec 20, 2038(~12.4 yrs left)· nominal 20-yr term from priority
G06T 7/0002G06T 7/10H04N 13/20G06N 3/084G06T 3/4046G06T 2210/12G06T 2207/10012G06T 2207/10024G06T 2207/10028G06T 2207/20081G06T 2207/20084G06T 17/00G06T 15/00
76
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method for estimating tightly enclosing bounding boxes by a computing system includes: controlling a scanning system including one or more depth cameras to capture visual information of the scene including one or more objects; detecting the one or more objects of the scene based on the visual information; singulating each the one or more objects from the frame of the scene to generate one or more 3D models corresponding to the one or more objects, the one or more 3D models including a partial 3D model of a corresponding one of the one or more objects; extrapolating a more complete 3D model of the corresponding one of the one or more objects based on the partial 3D model; and estimating a tightly enclosing bounding box of the corresponding one of the one or more objects based on the more complete 3D model.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A computing system for estimating tightly enclosing bounding boxes comprising:
 one or more processors; and   one or more computer-readable media having stored thereon executable instructions that when executed by the one or more processors configure the computing system to:
 control, by a computing system, a scanning system comprising one or more depth cameras to capture visual information of a scene comprising one or more objects; 
 detect, by the computing system, the one or more objects of the scene based on the visual information; 
 singulate, by the computing system, each of the one or more objects from a frame of the scene to generate one or more 3D models corresponding to the one or more objects, the one or more 3D models comprising a partial 3D model of a corresponding one of the one or more objects; 
 extrapolate, by the computing system, a more complete 3D model of the corresponding one of the one or more objects based on the partial 3D model, wherein:
 the extrapolating the more complete 3D model comprises supplying the partial 3D model to a generative model trained to predict a generated 3D model based on an input partial 3D model, the more complete 3D model comprising the generated 3D model; and 
 
 estimate, by the computing system, a tightly enclosing bounding box of the corresponding one of the one or more objects based on the more complete 3D model. 
   
     
     
         2 . The computing system of  claim 1 , wherein the scanning system further comprises one or more color cameras separate from the one or more depth cameras. 
     
     
         3 . The computing system of  claim 1 , wherein the one or more depth cameras comprises:
 a time-of-flight depth camera;   a structured light depth camera;   a stereo depth camera comprising at least two color cameras;   a stereo depth camera comprising:
 at least two color cameras, and 
 a color projector; 
   a stereo depth camera comprising at least two infrared cameras; or   a stereo depth camera comprising:
 a color camera, 
 a plurality of infrared cameras, and 
 an infrared projector configured to emit light in a wavelength interval that is detectable by the plurality of infrared cameras. 
   
     
     
         4 . The computing system of  claim 1 , wherein the executable instructions for detecting the one or more objects in the scene comprise one or more instructions that when executed separate the one or more objects from depictions of background and ground plane in the visual information. 
     
     
         5 . The computing system of  claim 1 , wherein the generative model comprises a conditional generative adversarial network. 
     
     
         6 . The computing system of  claim 1 , wherein the executable instructions for extrapolating the more complete 3D model comprise one or more instructions that when executed:
 classify the partial 3D model to compute a matching classification;   load one or more heuristic rules for generating more complete 3D models for the matching classification; and   generate the more complete 3D model from the partial 3D model in accordance with the one or more heuristic rules.   
     
     
         7 . The computing system of  claim 6 , wherein the one or more heuristic rules comprise one or more assumed axes of symmetry of the more complete 3D model based on the matching classification, or a canonical general shape of the more complete 3D model based on the matching classification. 
     
     
         8 . The computing system of  claim 1 , wherein the one or more objects comprise a plurality of objects, and
 wherein the singulating each the one or more objects from the frame of the scene comprises singulating the plurality of objects by applying an appearance-based segmentation to the visual information.   
     
     
         9 . The computing system of  claim 1 , wherein the one or more objects comprise a plurality of objects, and
 wherein the singulating each the one or more objects from the frame of the scene comprises singulating the plurality of objects by applying semantic segmentation to the visual information.   
     
     
         10 . The computing system of  claim 9 , wherein the executable instructions for applying semantic segmentation comprise one or more instructions that when executed supply the visual information to a trained fully convolutional neural network to compute a segmentation map, and
 wherein each partial 3D model corresponds to one segment of the segmentation map.   
     
     
         11 . The computing system of  claim 1 , further comprising associating the tightly enclosing bounding box with an item descriptor. 
     
     
         12 . A computer-implemented method comprising:
 controlling, by a computing system, a scanning system comprising one or more depth cameras to capture visual information of a scene comprising one or more objects;   detecting, by the computing system, the one or more objects of the scene based on the visual information;   singulating, by the computing system, each of the one or more objects from a frame of the scene to generate one or more 3D models corresponding to the one or more objects, the one or more 3D models comprising a partial 3D model of a corresponding one of the one or more objects;   extrapolating, by the computing system, a more complete 3D model of the corresponding one of the one or more objects based on the partial 3D model, wherein:
 the extrapolating the more complete 3D model comprises supplying the partial 3D model to a generative model trained to predict a generated 3D model based on an input partial 3D model, the more complete 3D model comprising the generated 3D model; and 
   estimating, by the computing system, a tightly enclosing bounding box of the corresponding one of the one or more objects based on the more complete 3D model.   
     
     
         13 . The computer-implemented method of  claim 12 , wherein the scanning system further comprises one or more color cameras separate from the one or more depth cameras. 
     
     
         14 . The computer-implemented method of  claim 12 , wherein the one or more depth cameras comprises:
 a time-of-flight depth camera;   a structured light depth camera;   a stereo depth camera comprising at least two color cameras;   a stereo depth camera comprising:
 at least two color cameras, and 
 a color projector; 
   a stereo depth camera comprising at least two infrared cameras; or   a stereo depth camera comprising:
 a color camera, 
 a plurality of infrared cameras, and 
 an infrared projector configured to emit light in a wavelength interval that is detectable by the plurality of infrared cameras. 
   
     
     
         15 . The computer-implemented method of  claim 12 , wherein detecting the one or more objects in the scene comprises separating the one or more objects from depictions of background and ground plane in the visual information. 
     
     
         16 . The computer-implemented method of  claim 12 , wherein the generative model comprises a conditional generative adversarial network. 
     
     
         17 . The computer-implemented method of  claim 12 , wherein extrapolating the more complete 3D model comprises:
 classify the partial 3D model to compute a matching classification;   load one or more heuristic rules for generating more complete 3D models for the matching classification; and   generate the more complete 3D model from the partial 3D model in accordance with the one or more heuristic rules.   
     
     
         18 . The computer-implemented method of  claim 17 , wherein the one or more heuristic rules comprise one or more assumed axes of symmetry of the more complete 3D model based on the matching classification, or a canonical general shape of the more complete 3D model based on the matching classification. 
     
     
         19 . The computer-implemented method of  claim 12 , wherein the one or more objects comprise a plurality of objects, and
 wherein the singulating each the one or more objects from the frame of the scene comprises singulating the plurality of objects by applying an appearance-based segmentation to the visual information.   
     
     
         20 . The computer-implemented method of  claim 12 , wherein the one or more objects comprise a plurality of objects, and
 wherein the singulating each the one or more objects from the frame of the scene comprises singulating the plurality of objects by applying semantic segmentation to the visual information.

Join the waitlist — get patent alerts

Track US2023410276A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.