Spatial video capture and replay
Abstract
Various implementations disclosed herein include devices, systems, and methods that create a 3D video that includes determining first adjustments (e.g., first transforms) to video frames (e.g., one or more RGB images and depth images per frame) to align content in a coordinate system to remove the effects of capturing camera motion. Various implementations disclosed herein include devices, systems, and methods that playback a 3D video and includes determining second adjustments (e.g., second transforms) to remove the effects of movement of a viewing electronic device relative to a viewing environment during playback of the 3D video. Some implementations distinguish static content and moving content of the video frames to playback only moving objects or facilitate concurrent playback of multiple spatially related 3D videos. The 3D video may include images, audio, or 3D video of a video-capture-device user.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method comprising:
at an electronic device having a processor:
obtaining a 3D video comprising images and depth data;
obtaining first adjustments to align content represented in the images and depth data, the first adjustments accounting for movement of a device that captured the images and depth data;
determining second adjustments to align the content represented in the images and depth data in an environment presented by the electronic device, the second adjustments determined based on movement of the electronic device during presentation of the environment; and
presenting the 3D video in the environment based on the first adjustments and the second adjustments.
2 . The method of claim 1 , wherein the environment is a CGR environment.
3 . The method of claim 1 , wherein the 3D video file comprises RGB images, depth maps, confidences, segmentations, point cloud files for static reconstruction, camera metadata, or spatial audio.
4 . The method of claim 1 , wherein the first adjustments comprise one or more transforms based on motion sensor data corresponding to movement of sensors that generated the images and the depth data of the 3D video.
5 . The method of claim 1 , wherein the first adjustments comprise one or more transforms based on at least one static object identified in the images and depth data of the 3D video.
6 . The method of claim 1 , wherein the second adjustments comprise one or more transforms that remove the movement of the electronic device during presentation of the 3D video in the environment.
7 . The method of claim 1 , wherein the 3D video comprises one or more files including segmentations, wherein the segmentations include static objects and moving objects, and wherein presenting the 3D video in the environment removes the static objects from the presentation.
8 . The method of claim 1 , wherein presenting the 3D video in the environment comprises adjusting the presentation of the environment based on lighting or shadowing of the 3D video or adjusting the presentation of the 3D video based on lighting or the shadowing of the environment.
9 . The method of claim 1 , wherein the 3D video identifies a ground plane in the 3D video, and wherein presenting the 3D video in the environment comprises aligning the ground plane of the 3D video to a ground plane of the environment.
10 . The method of claim 1 , wherein the 3D video identifies a single coordinate system for the 3D video, and wherein presenting the 3D video in the environment matches spatialized audio data to the single coordinate system.
11 . The method of claim 1 , wherein presenting the 3D video in the environment comprises providing a visual buffer around the 3D video in the environment.
12 . The method of claim 1 , wherein presenting the 3D video in the environment comprises determining a starting position for the 3D video in the environment, and wherein the method further comprises:
re-mapping the 3D video back to the starting position when the 3D video moves beyond a preset spatial threshold distance from the starting position.
13 . The method of claim 1 , further comprising:
obtaining a second 3D video comprising second images and second depth data; obtaining capture adjustments to align second content represented in the second images and the second depth data, the capture adjustments accounting for movement of a second device that captured the second images and the second depth data; determining playback adjustments to align the second content represented in the second images and the second depth data in the environment presented by the electronic device, the playback adjustments determined based on movement of the electronic device during presentation of the environment; and presenting the second 3D video in the environment based on the capture adjustments and the playback adjustments, wherein the 3D video comprises static reconstructions representing static objects in a physical environment, wherein the second 3D video comprises second static reconstructions representing static objects in a second physical environment, and wherein a spatial relationship between the 3D video and the second 3D video in the environment is based on the static reconstructions and the second static reconstructions.
14 . The method of claim 13 , wherein the second physical environment is the physical environment.
15 . The method of claim 1 , wherein the 3D video comprises a sequence of images and depth data of a user of the device that captured the images and depth data or a sequence of audio inputs and orientation data of the user of the device that captured the images and depth data.
16 . A system comprising:
a non-transitory computer-readable storage medium; and one or more processors coupled to the non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium comprises program instructions that, when executed on the one or more processors, cause the system to perform operations comprising: obtaining a 3D video comprising images and depth data; obtaining first adjustments to align content represented in the images and depth data, the first adjustments accounting for movement of a device that captured the images and depth data; determining second adjustments to align the content represented in the images and depth data in an environment presented by the electronic device, the second adjustments determined based on movement of the electronic device during presentation of the environment; and presenting the 3D video in the environment based on the first adjustments and the second adjustments.
17 . The system of claim 16 , wherein the environment is a CGR environment.
18 . The system of claim 16 , wherein the 3D video file comprises RGB images, depth maps, confidences, segmentations, point cloud files for static reconstruction, camera metadata, or spatial audio.
19 . The system of claim 16 , wherein the first adjustments comprise one or more transforms based on motion sensor data corresponding to movement of sensors that generated the images and the depth data of the 3D video.
20 . A non-transitory computer-readable storage medium, storing program instructions computer-executable on a computer to perform operations comprising:
obtaining a 3D video comprising images and depth data; obtaining first adjustments to align content represented in the images and depth data, the first adjustments accounting for movement of a device that captured the images and depth data; determining second adjustments to align the content represented in the images and depth data in an environment presented by the electronic device, the second adjustments determined based on movement of the electronic device during presentation of the environment; and presenting the 3D video in the environment based on the first adjustments and the second adjustments.Join the waitlist — get patent alerts
Track US2025182419A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.