US2023037731A1PendingUtilityA1

Systems and methods for self-supervised depth estimation

Assignee: TOYOTA RES INST INCPriority: Sep 15, 2020Filed: Oct 13, 2022Published: Feb 9, 2023
Est. expirySep 15, 2040(~14.1 yrs left)· nominal 20-yr term from priority
G06T 7/55G06T 2207/30252G06T 2207/20084G06T 2207/20081G06T 3/0093G06T 3/18
69
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Systems and methods for self-supervised depth estimation using image frames captured from cameras, may include: receiving a first image captured by a first camera while the camera is mounted at a first location, the first image comprising pixels representing a first scene of an environment of a vehicle; receiving a reference image captured by a second camera while the second camera is mounted at a second location, the reference image comprising pixels representing a second scene of the environment; warping the first image to a perspective of the second camera at the second location on the vehicle to arrive at a warped first image; projecting the warped first image onto the reference image; determining a loss based on the projection; and updating predicted depth values for the first image.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method of self-supervised depth estimation using image frames captured from cameras, comprising:
 receiving a first image captured by a first camera mounted at a first camera mounting location, the first image comprising pixels representing a first scene of an environment of a vehicle;   receiving a reference image from a second camera mounted at a second camera mounting location, the reference image comprising pixels representing a second scene of the environment of the vehicle;   predicting a depth map for the first image, the depth map comprising predicted depth values for pixels of the first image;   warping the first image to a perspective of the second camera at the second camera mounting location to arrive at a warped first image;   projecting the warped first image onto the reference image; and   determining a loss in the predicted depth values based on the projection.   
     
     
         2 . The method of  claim 1 , wherein the first camera mounting location is a first location on the vehicle and the second camera mounting location is a second location on the vehicle. 
     
     
         3 . The method of  claim 1 , further comprising:
 updating the predicted depth values for the first image based on the loss; and   reiterating the operations of warping the first image, projecting the warped first image and determining a loss using the updated predicted depth values for the first image.   
     
     
         4 . The method of  claim 1 , wherein projecting is performed using a neural camera model to model intrinsic parameters of the first camera. 
     
     
         5 . The method of  claim 1 , further comprising predicting a transformation from the first camera mounting location to the second camera mounting location based on loss calculations between the warped first image and the reference image. 
     
     
         6 . The method of  claim 1 , wherein projecting the warped first image onto the reference image comprises lifting 2D points of the warped first image to 3D points, determining a transformation between the first camera mounting location and the second camera mounting location, and using the transformation to project the 3D points onto the reference image in 2D. 
     
     
         7 . The method of  claim 6 , wherein the transformation comprises a distance in three dimensions between image sensors of the first and second cameras. 
     
     
         8 . A system for self-supervised learning depth estimation using image frames captured by cameras, the system comprising:
 a non-transitory memory configured to store instructions;   a processor configured to execute the instructions to perform the operations of:
 receiving a first image captured by a first camera mounted at a first camera mounting location, the first image comprising pixels representing a first scene of an environment of a vehicle; 
 receiving a reference image captured by a second camera mounted at a second camera mounting location, the reference image comprising pixels representing a second scene of the environment of the vehicle; 
 predicting a depth map for the first image, the depth map comprising predicted depth values for pixels of the first image; 
 warping the first image to a perspective of the second camera at the second camera mounting location to arrive at a warped first image; 
 projecting the warped first image onto the reference image; and 
 determining a loss in the predicted depth values based on the projection. 
   
     
     
         9 . The system of  claim 8 , wherein the first camera mounting location is a first location on the vehicle and the second camera mounting location is a second location on the vehicle. 
     
     
         10 . The system of  claim 8 , wherein the operations further comprise:
 updating the predicted depth values for the first image based on the loss; and   reiterating the operations of warping the first image, projecting the warped first image and determining a loss using the updated predicted depth values for the first image.   
     
     
         11 . The system of  claim 8 , wherein projecting is performed using a neural camera model to model intrinsic parameters of the first camera. 
     
     
         12 . The system of  claim 8 , wherein the operations further comprise predicting a transformation from the first camera mounting location to the second camera mounting location based on loss calculations between the warped first image and the reference image. 
     
     
         13 . The system of  claim 8 , wherein projecting the warped first image onto the reference image comprises lifting 2D points of the warped first image to 3D points, determining a transformation between the first camera mounting location and the second camera mounting location and using the transformation to project the 3D points onto the reference image in 2D. 
     
     
         14 . The system of  claim 13 , wherein the transformation comprises a distance in three dimensions between image sensors of the first and second cameras. 
     
     
         15 . A system for self-supervised learning depth estimation, the system comprising:
 a first camera mounted at a first camera mounting location, wherein the first camera captures a first image, the first image comprising pixels representing a first scene of an environment of a vehicle;   a second camera mounted at a second camera mounting location, wherein the second camera captures a reference image, the reference image comprising pixels representing a second scene of the environment of the vehicle;   an ECU including machine executable instructions in non-transitory memory to perform a method comprising:
 predicting a depth map for the first image, the depth map comprising predicted depth values for pixels of the first image; 
 warping the first image to a perspective of the second camera at the second camera mounting location to arrive at a warped first image; 
 projecting the warped first image onto the reference image; and 
 determining a loss in the predicted depth values based on the projection. 
   
     
     
         16 . The system of  claim 15 , wherein the first camera mounting location is a first location on the vehicle and the second camera mounting location is a second location on the vehicle. 
     
     
         17 . The system of  claim 15 , further comprising a neural camera model configured to model intrinsic parameters of the first camera. 
     
     
         18 . The system of  claim 15 , wherein the operations further comprise predicting a transformation from the first camera mounting location to the second camera mounting location based on loss calculations between the warped first image and the reference image. 
     
     
         19 . The system of  claim 15 , wherein projecting the warped first image onto the reference image comprises lifting 2D points of the warped first image to 3D points, determining a transformation between the first camera mounting location and the second camera mounting location and using the transformation to project the 3D points onto the reference image in 2D. 
     
     
         20 . The system of  claim 15 , wherein the transformation comprises a distance in three dimensions between image sensors of the first and second cameras.

Join the waitlist — get patent alerts

Track US2023037731A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.