US2023316458A1PendingUtilityA1

Image stitching with dynamic seam placement based on object saliency for surround view visualization

Assignee: NVIDIA CORPPriority: Apr 1, 2022Filed: Feb 23, 2023Published: Oct 5, 2023
Est. expiryApr 1, 2042(~15.7 yrs left)· nominal 20-yr term from priority
G06V 10/16B60W 60/001G06T 19/20G06T 17/20B60W 30/06G06V 20/58H04N 5/2624G06T 3/4038G06T 7/74G06V 20/56G06T 7/70G06T 15/20G06T 2219/2004B60W 2510/0638B60W 2420/403B60W 2420/408B60W 2520/10
75
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

In various examples, dynamic seam placement is used to position seams in regions of overlapping image data to avoid crossing salient objects or regions. Objects may be detected from image frames representing overlapping views of an environment surrounding an ego-object such as a vehicle. The images may be aligned to create an aligned composite image or surface (e.g., a panorama, a 360° image, bowl shaped surface) with regions of overlapping image data, and a representation of the detected objects and/or salient regions (e.g., a saliency mask) may be generated and projected onto the aligned composite image or surface. Seams may be positioned in the overlapping regions to avoid or minimize crossing salient pixels represented in the projected masks, and the image data may be blended at the seams to create a stitched image or surface (e.g., a stitched panorama, stitched 360° image, stitched textured surface).

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method comprising:
 generating, using sensor data captured during a first time slice, two or more image frames representative of two or more at least partially overlapping viewpoints around an ego-object in an environment;   updating a candidate position for a seam to an updated position based at least on an intersection of the seam at the candidate position with one or more pixels of one or more detected objects in at least one projected mask of the two or more projected masks; and   generating a composite image frame based at least on stitching the two or more image frames using the updated position of the seam.   
     
     
         2 . The method of  claim 1 , wherein the updating of the candidate position for the seam to the updated position comprises:
 generating two or more projected masks comprising at least partially overlapping representations of one or more detected objects depicted in the two or more image frames;   determining a candidate position for a seam in an overlapping region of the two or more image frames; and   updating the candidate position for the seam to an updated position based at least on an intersection of the seam at the candidate position with one or more pixels of the one or more detected objects in at least one projected mask of the two or more projected masks.   
     
     
         3 . The method of  claim 1 , wherein the updating of the candidate position for the seam to the updated position causes the updated position to be at least partially horizontal or an at least partially vertical. 
     
     
         4 . The method of  claim 1 , further comprising generating the two or more projected masks based at least on projecting two or more corresponding binary object masks onto an overlapping portion of the two or more image frames, the two or more binary object masks being indicative of whether or not each pixel in a corresponding one of the two or more image frames corresponds to one of the one or more detected objects. 
     
     
         5 . The method of  claim 1 , further comprising generating the two or more projected masks based at least on projecting two or more corresponding weighted saliency masks onto an overlapping portion of the two or more image frames, the two or more weighted saliency masks representing a measure of saliency of each pixel in a corresponding one of the two or more image frames. 
     
     
         6 . The method of  claim 1 , further comprising generating the two or more projected masks based at least on omitting a subset of an initial set of detected objects determined to be beyond a threshold distance from the ego-object. 
     
     
         7 . The method of  claim 1 , further comprising generating the two or more projected masks based at least on projecting two or more corresponding weighted saliency masks, prioritizing one or more classes of the one or more detected objects, onto an overlapping portion of the two or more image frames. 
     
     
         8 . The method of  claim 1 , further comprising generating the two or more projected masks based at least on projecting two or more corresponding weighted saliency masks, weighting corresponding binary object masks based at least on proximity to the one or more detected objects, onto an overlapping portion of the two or more image frames. 
     
     
         9 . The method of  claim 1 , wherein the ego-object is a medical probe, and the two or more viewpoints comprise at least one of a multi-view or a proximity view corresponding to the medical probe. 
     
     
         10 . The method of  claim 1 , wherein the method is performed by at least one of:
 a control system for an autonomous or semi-autonomous machine;   a perception system for an autonomous or semi-autonomous machine;   a system for performing simulation operations;   a system for performing digital twin operations;   a system for performing light transport simulation;   a system for performing collaborative content creation for 3D assets;   a system for performing deep learning operations;   a system implemented using an edge device;   a system implemented using a robot;   a system for performing conversational AI operations;   a system for generating synthetic data;   a system incorporating one or more virtual machines (VMs);   a system implemented at least partially in a data center; or   a system implemented at least partially using cloud computing resources.   
     
     
         11 . A processor comprising:
 one or more circuits to:
 obtain image data of two or more image frames corresponding to two or more separate viewpoints that share at least a portion of an area in an environment; 
 detect one or more objects depicted in the two or more image frames; 
 determine a candidate position for a seam in an aligned representation of the two or more image frames; 
 update the candidate position of the seam to an updated position based at least on an intersection of the seam at the candidate position with one or more pixels in the aligned representation that at least partially depict the one or more objects; and 
 generate a composite image based at least on stitching the two or more image frames using the updated position of the seam. 
   
     
     
         12 . The processor of  claim 11 , the one or more circuits further to update the candidate position for the seam to the updated position based at least on reducing a number of the one or more pixels in the aligned representation of the two or more image frames that at least partially depict the one or more objects. 
     
     
         13 . The processor of  claim 11 , the one or more circuits further to update the candidate position for the seam to the updated position corresponding to an at least partially horizontal position or an at least partially vertical position. 
     
     
         14 . The processor of  claim 11 , the one or more circuits further to determine whether the seam at the candidate position intersects the one or more pixels in the aligned representation that at least partially depict the one or more objects based at least on comparing the candidate position of the seam with corresponding pixels in two or more projected masks representing positions of the one or more detected objects in the aligned representation. 
     
     
         15 . The processor of  claim 11 , wherein the processor is comprised in at least one of:
 a control system for an autonomous or semi-autonomous machine;   a perception system for an autonomous or semi-autonomous machine;   a system for performing simulation operations;   a system for performing digital twin operations;   a system for performing light transport simulation;   a system for performing collaborative content creation for 3D assets;   a system for performing deep learning operations;   a system implemented using an edge device;   a system implemented using a robot;   a system for performing conversational AI operations;   a system for generating synthetic data;   a system incorporating one or more virtual machines (VMs);   a system implemented at least partially in a data center; or   a system implemented at least partially using cloud computing resources.   
     
     
         16 . A system comprising:
 one or more processing units to update a candidate position for a seam in an overlapping region of two or more image frames based at least on an intersection with an initial candidate position of the seam and one or more pixels corresponding to one or more detected objects of a designated object class that are at least partially depicted in the overlapping region, and to generate a composite image based at least on stitching the two or more image frames using a seam at the updated candidate position.   
     
     
         17 . The system of  claim 16 , wherein the one or more processing units update the candidate position for the seam by generating two or more projected masks corresponding to the one or more detected objects based at least on projecting two or more corresponding weighted saliency masks onto an overlapping portion of the two or more image frames. 
     
     
         18 . The system of  claim 17 , wherein the two or more weighted saliency masks represent a measure of saliency of at least one pixel in a corresponding one of the two or more image frames. 
     
     
         19 . The system of  claim 17 , wherein projecting the two or more corresponding weighted saliency masks onto an overlapping portion of the two or more image frames comprises prioritizing one or more object classes of the one or more detected objects. 
     
     
         20 . The system of  claim 16 , wherein the system is comprised in at least one of:
 a control system for an autonomous or semi-autonomous machine;   a perception system for an autonomous or semi-autonomous machine;   a system for performing simulation operations;   a system for performing digital twin operations;   a system for performing deep learning operations;   a system implemented using an edge device;   a system implemented using a robot;   a system incorporating one or more virtual machines (VMs);   a system implemented at least partially in a data center;   a system for performing light transport simulation;   a system for performing collaborative content creation for 3D assets;   a system for generating synthetic data; or   a system implemented at least partially using cloud computing resources.

Join the waitlist — get patent alerts

Track US2023316458A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.