US2025182419A1PendingUtilityA1

Spatial video capture and replay

Assignee: APPLE INCPriority: May 13, 2020Filed: Feb 12, 2025Published: Jun 5, 2025
Est. expiryMay 13, 2040(~13.8 yrs left)· nominal 20-yr term from priority
G06T 2200/08G06T 7/579G06T 17/00G06T 2210/56G06T 2207/20021G06T 2207/10028G06T 2207/10024G06T 2207/10016G06T 2200/04G06T 7/70G06T 7/30H04N 23/683H04N 23/6812H04N 23/6811G06T 7/38G06T 19/006G06T 15/20
68
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Various implementations disclosed herein include devices, systems, and methods that create a 3D video that includes determining first adjustments (e.g., first transforms) to video frames (e.g., one or more RGB images and depth images per frame) to align content in a coordinate system to remove the effects of capturing camera motion. Various implementations disclosed herein include devices, systems, and methods that playback a 3D video and includes determining second adjustments (e.g., second transforms) to remove the effects of movement of a viewing electronic device relative to a viewing environment during playback of the 3D video. Some implementations distinguish static content and moving content of the video frames to playback only moving objects or facilitate concurrent playback of multiple spatially related 3D videos. The 3D video may include images, audio, or 3D video of a video-capture-device user.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method comprising:
 at an electronic device having a processor:
 obtaining a 3D video comprising images and depth data; 
 obtaining first adjustments to align content represented in the images and depth data, the first adjustments accounting for movement of a device that captured the images and depth data; 
 determining second adjustments to align the content represented in the images and depth data in an environment presented by the electronic device, the second adjustments determined based on movement of the electronic device during presentation of the environment; and 
 presenting the 3D video in the environment based on the first adjustments and the second adjustments. 
   
     
     
         2 . The method of  claim 1 , wherein the environment is a CGR environment. 
     
     
         3 . The method of  claim 1 , wherein the 3D video file comprises RGB images, depth maps, confidences, segmentations, point cloud files for static reconstruction, camera metadata, or spatial audio. 
     
     
         4 . The method of  claim 1 , wherein the first adjustments comprise one or more transforms based on motion sensor data corresponding to movement of sensors that generated the images and the depth data of the 3D video. 
     
     
         5 . The method of  claim 1 , wherein the first adjustments comprise one or more transforms based on at least one static object identified in the images and depth data of the 3D video. 
     
     
         6 . The method of  claim 1 , wherein the second adjustments comprise one or more transforms that remove the movement of the electronic device during presentation of the 3D video in the environment. 
     
     
         7 . The method of  claim 1 , wherein the 3D video comprises one or more files including segmentations, wherein the segmentations include static objects and moving objects, and wherein presenting the 3D video in the environment removes the static objects from the presentation. 
     
     
         8 . The method of  claim 1 , wherein presenting the 3D video in the environment comprises adjusting the presentation of the environment based on lighting or shadowing of the 3D video or adjusting the presentation of the 3D video based on lighting or the shadowing of the environment. 
     
     
         9 . The method of  claim 1 , wherein the 3D video identifies a ground plane in the 3D video, and wherein presenting the 3D video in the environment comprises aligning the ground plane of the 3D video to a ground plane of the environment. 
     
     
         10 . The method of  claim 1 , wherein the 3D video identifies a single coordinate system for the 3D video, and wherein presenting the 3D video in the environment matches spatialized audio data to the single coordinate system. 
     
     
         11 . The method of  claim 1 , wherein presenting the 3D video in the environment comprises providing a visual buffer around the 3D video in the environment. 
     
     
         12 . The method of  claim 1 , wherein presenting the 3D video in the environment comprises determining a starting position for the 3D video in the environment, and wherein the method further comprises:
 re-mapping the 3D video back to the starting position when the 3D video moves beyond a preset spatial threshold distance from the starting position.   
     
     
         13 . The method of  claim 1 , further comprising:
 obtaining a second 3D video comprising second images and second depth data;   obtaining capture adjustments to align second content represented in the second images and the second depth data, the capture adjustments accounting for movement of a second device that captured the second images and the second depth data;   determining playback adjustments to align the second content represented in the second images and the second depth data in the environment presented by the electronic device, the playback adjustments determined based on movement of the electronic device during presentation of the environment; and   presenting the second 3D video in the environment based on the capture adjustments and the playback adjustments,   wherein the 3D video comprises static reconstructions representing static objects in a physical environment, wherein the second 3D video comprises second static reconstructions representing static objects in a second physical environment, and wherein a spatial relationship between the 3D video and the second 3D video in the environment is based on the static reconstructions and the second static reconstructions.   
     
     
         14 . The method of  claim 13 , wherein the second physical environment is the physical environment. 
     
     
         15 . The method of  claim 1 , wherein the 3D video comprises a sequence of images and depth data of a user of the device that captured the images and depth data or a sequence of audio inputs and orientation data of the user of the device that captured the images and depth data. 
     
     
         16 . A system comprising:
 a non-transitory computer-readable storage medium; and   one or more processors coupled to the non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium comprises program instructions that, when executed on the one or more processors, cause the system to perform operations comprising:   obtaining a 3D video comprising images and depth data;   obtaining first adjustments to align content represented in the images and depth data, the first adjustments accounting for movement of a device that captured the images and depth data;   determining second adjustments to align the content represented in the images and depth data in an environment presented by the electronic device, the second adjustments determined based on movement of the electronic device during presentation of the environment; and   presenting the 3D video in the environment based on the first adjustments and the second adjustments.   
     
     
         17 . The system of  claim 16 , wherein the environment is a CGR environment. 
     
     
         18 . The system of  claim 16 , wherein the 3D video file comprises RGB images, depth maps, confidences, segmentations, point cloud files for static reconstruction, camera metadata, or spatial audio. 
     
     
         19 . The system of  claim 16 , wherein the first adjustments comprise one or more transforms based on motion sensor data corresponding to movement of sensors that generated the images and the depth data of the 3D video. 
     
     
         20 . A non-transitory computer-readable storage medium, storing program instructions computer-executable on a computer to perform operations comprising:
 obtaining a 3D video comprising images and depth data;   obtaining first adjustments to align content represented in the images and depth data, the first adjustments accounting for movement of a device that captured the images and depth data;   determining second adjustments to align the content represented in the images and depth data in an environment presented by the electronic device, the second adjustments determined based on movement of the electronic device during presentation of the environment; and   presenting the 3D video in the environment based on the first adjustments and the second adjustments.

Join the waitlist — get patent alerts

Track US2025182419A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.