Automatic Generation of Video Using Location-Based Metadata Generated from Wireless Beacons
A spherical content capture system captures spherical video content. A spherical video sharing platform enables users to share the captured spherical content and enables users to access spherical content shared by other users. In one embodiment, captured metadata provides proximity information indicating which cameras were in proximity to a target device during a particular time frame. The platform can then generate an output video from spherical video captured from those cameras. The output video may include a non-spherical reduced field of view such as those commonly associated with conventional camera systems. Particularly, relevant sub-frames having a reduced field of view may be extracted from frames of one or more spherical videos to generate an output video that tracks a particular individual or object of interest.
1 . A method for generating an output video, the method comprising:
receiving captured first proximity data corresponding to a target device, the first proximity data indicating for a first time frame, first camera identifiers associated with a first plurality of cameras from which respective first beacon signals were detected by of the target device during the first time frame and the first proximity data indicating first signal strengths associated with the first beacon signals received from each of the first plurality of cameras;
determining based on the first signal strengths, a first selected camera having a highest signal strength for the first time frame;
querying a video database for video captured during the first time frame by the first selected camera to obtain a first spherical video captured during the first time frame by the first selected camera;
processing each frame of the first spherical video corresponding to the first time frame to identify a first sequence of sub-frames corresponding to a first location of the target device relative to the first selected camera during the first time frame, the first sequence of sub-frames having a reduced field of view relative to a field of view of the first spherical video; and
generating a first portion of the output video comprising the first sequence of sub-frames.
2 . The method of claim 1 , further comprising:
receiving captured second proximity data corresponding to the target device, the second proximity data indicating for a second time frame, second camera identifiers associated with a second plurality of cameras that were detected to be within the threshold proximity of the target device during the second time frame and the second proximity data indicating second signal strengths associated with respective second beacon signals received from each of the second plurality of cameras;
determining based on the second signal strengths, a second selected camera having a highest signal strength for the second time frame;
querying the video database for video captured during the second time frame by the second selected camera to obtain a second spherical video captured during the second time frame by the second selected camera;
processing each frame of the second spherical video corresponding to the second time frame to identify a second sequence of sub-frames corresponding to a second location of the target device relative to the second selected camera during the second time frame, the second sequence of sub-frames having the reduced field of view relative to the field of view of the second spherical video; and
generating a second portion of the output video comprising the second sequence of sub-frames.
3 . The method of claim 1 , wherein processing each frame of the first spherical video corresponding to the first time frame to identify the first sequence of sub-frames corresponding to the first location of the target device relative to the first selected camera, comprises:
performing a facial recognition algorithm on the first spherical video to identify a face depicted in the first spherical video; and
determining the first location of the target device based on a location of the face.
4 . The method of claim 1 , wherein processing each frame of the first spherical video corresponding to the first time frame to identify the first sequence of sub-frames corresponding to the first location of the target device relative to the first selected camera, comprises:
performing an object recognition algorithm on first spherical video to recognize the target device depicted in the first spherical video.
5 . The method of claim 1 , wherein processing each frame of the first spherical video corresponding to the first time frame to identify the first sequence of sub-frames corresponding to the first location of the target device relative to the first selected camera, comprises:
performing an audio analysis of an audio track of the first spherical video to recognize an audio signal from the target device; and
determining the first location of the target device based on a direction of the audio signal.
6 . The method of claim 1 , wherein processing each frame of the first spherical video corresponding to the first time frame to identify the first sequence of sub-frames corresponding to the first location of the target device relative to the first selected camera, comprises
determining GPS positions of the target device and the first selected camera; and
determining the first location of the target device based on the GPS positions.
7 . The method of claim 1 , wherein the target device comprises a camera that captures video, wherein processing each frame of the first spherical video corresponding to the first time frame to identify the first sequence of sub-frames corresponding to the first location of the target device relative to the first selected camera, comprises
determining one or more matching objects recognized in the first spherical video from the selected camera and the video captured by the target device; and
determining the first location of the target device based on relative directions of the one or more matching objects in the first spherical video and the video captured by the target device.
8 . The method of claim 1 , wherein processing each frame of the first spherical video corresponding to the first time frame to identify the first sequence of sub-frames corresponding to the first location of the target device relative to the first selected camera, comprises:
determining estimated distances between the target device and one or more additional cameras and between the first selected camera and the one or more additional cameras, the estimated distances determined based on signal strengths of the beacon signals received by the one or more additional cameras; and
determining the first location of the target device based on the estimated distances.
9 . A non-transitory computer-readable storage medium storing instructions for generating an output video, the instructions to cause the one or more processors to perform steps including:
indicating for a first time frame, first camera identifiers associated with a first plurality of cameras from which respective first beacon signals were detected by of the target device during the first time frame and the first proximity data indicating first signal strengths associated with the first beacon signals received from each of the first plurality of cameras;
determining based on the first signal strengths, a first selected camera having a highest signal strength for the first time frame;
querying a video database for video captured during the first time frame by the first selected camera to obtain a first spherical video captured during the first time frame by the first selected camera;
processing each frame of the first spherical video corresponding to the first time frame to identify a first sequence of sub-frames corresponding to a first location of the target device relative to the first selected camera during the first time frame, the first sequence of sub-frames having a reduced field of view relative to a field of view of the first spherical video; and
generating a first portion of the output video comprising the first sequence of sub-frames.
10 . The non-transitory computer-readable storage medium of claim 9 , wherein the instructions when executed further cause the one or more processors to perform steps including:
receiving captured second proximity data corresponding to the target device, the second proximity data indicating for a second time frame, second camera identifiers associated with a second plurality of cameras that were detected to be within the threshold proximity of the target device during the second time frame and the second proximity data indicating second signal strengths associated with respective second beacon signals received from each of the second plurality of cameras;
determining based on the second signal strengths, a second selected camera having a highest signal strength for the second time frame;
querying the video database for video captured during the second time frame by the second selected camera to obtain a second spherical video captured during the second time frame by the second selected camera;
processing each frame of the second spherical video corresponding to the second time frame to identify a second sequence of sub-frames corresponding to a second location of the target device relative to the second selected camera during the second time frame, the second sequence of sub-frames having the reduced field of view relative to the field of view of the second spherical video; and
generating a second portion of the output video comprising the second sequence of sub-frames.
11 . The non-transitory computer-readable storage medium of claim 9 , wherein processing each frame of the first spherical video corresponding to the first time frame to identify the first sequence of sub-frames corresponding to the first location of the target device relative to the first selected camera, comprises:
performing a facial recognition algorithm on the first spherical video to identify a face depicted in the first spherical video; and
determining the first location of the target device based on a location of the face.
12 . The non-transitory computer-readable storage medium of claim 9 , wherein processing each frame of the first spherical video corresponding to the first time frame to identify the first sequence of sub-frames corresponding to the first location of the target device relative to the first selected camera, comprises:
performing an object recognition algorithm on first spherical video to recognize the target device depicted in the first spherical video.
13 . The non-transitory computer-readable storage medium of claim 9 , wherein processing each frame of the first spherical video corresponding to the first time frame to identify the first sequence of sub-frames corresponding to the first location of the target device relative to the first selected camera, comprises:
performing an audio analysis of an audio track of the first spherical video to recognize an audio signal from the target device; and
determining the first location of the target device based on a direction of the audio signal.
14 . The non-transitory computer-readable storage medium of claim 9 , wherein processing each frame of the first spherical video corresponding to the first time frame to identify the first sequence of sub-frames corresponding to the first location of the target device relative to the first selected camera, comprises
determining GPS positions of the target device and the first selected camera; and
determining the first location of the target device based on the GPS positions.
15 . The non-transitory computer-readable storage medium of claim 9 , wherein the target device comprises a camera that captures video, wherein processing each frame of the first spherical video corresponding to the first time frame to identify the first sequence of sub-frames corresponding to the first location of the target device relative to the first selected camera, comprises
determining one or more matching objects recognized in the first spherical video from the selected camera and the video captured by the target device; and
determining the first location of the target device based on relative directions of the one or more matching objects in the first spherical video and the video captured by the target device.
16 . The non-transitory computer-readable storage medium of claim 9 , wherein processing each frame of the first spherical video corresponding to the first time frame to identify the first sequence of sub-frames corresponding to the first location of the target device relative to the first selected camera, comprises:
determining estimated distances between the target device and one or more additional cameras and between the first selected camera and the one or more additional cameras, the estimated distances determined based on signal strengths of the beacon signals received by the one or more additional cameras; and
determining the first location of the target device based on the estimated distances.
17 . A video server for generating an output video, the video server comprising:
one or more processors; and
a non-transitory computer-readable storage medium storing instructions for generating an output video from spherical video content, the instructions to cause the one or more processors to perform steps including:
indicating for a first time frame, first camera identifiers associated with a first plurality of cameras from which respective first beacon signals were detected by of the target device during the first time frame and the first proximity data indicating first signal strengths associated with the first beacon signals received from each of the first plurality of cameras;
determining based on the first signal strengths, a first selected camera having a highest signal strength for the first time frame;
querying a video database for video captured during the first time frame by the first selected camera to obtain a first spherical video captured during the first time frame by the first selected camera;
processing each frame of the first spherical video corresponding to the first time frame to identify a first sequence of sub-frames corresponding to a first location of the target device relative to the first selected camera during the first time frame, the first sequence of sub-frames having a reduced field of view relative to a field of view of the first spherical video; and
generating a first portion of the output video comprising the first sequence of sub-frames.
18 . The video server of claim 17 , wherein the instructions when executed further cause the one or more processors to perform steps including:
receiving captured second proximity data corresponding to the target device, the second proximity data indicating for a second time frame, second camera identifiers associated with a second plurality of cameras that were detected to be within the threshold proximity of the target device during the second time frame and the second proximity data indicating second signal strengths associated with respective second beacon signals received from each of the second plurality of cameras;
determining based on the second signal strengths, a second selected camera having a highest signal strength for the second time frame;
querying the video database for video captured during the second time frame by the second selected camera to obtain a second spherical video captured during the second time frame by the second selected camera;
processing each frame of the second spherical video corresponding to the second time frame to identify a second sequence of sub-frames corresponding to a second location of the target device relative to the second selected camera during the second time frame, the second sequence of sub-frames having the reduced field of view relative to the field of view of the second spherical video; and
generating a second portion of the output video comprising the second sequence of sub-frames.
19 . The video server of claim 17 , wherein processing each frame of the first spherical video corresponding to the first time frame to identify the first sequence of sub-frames corresponding to the first location of the target device relative to the first selected camera, comprises:
performing a facial recognition algorithm on the first spherical video to identify a face depicted in the first spherical video; and
determining the first location of the target device based on a location of the face.
20 . The video server of claim 17 , wherein processing each frame of the first spherical video corresponding to the first time frame to identify the first sequence of sub-frames corresponding to the first location of the target device relative to the first selected camera, comprises:
performing an object recognition algorithm on first spherical video to recognize the target device depicted in the first spherical video.