IP Library Granted Patent US 10,056,115
Granted Patent B2
US 10,056,115 · App. 15/396,385 · Granted Aug 21, 2018

Automatic generation of video and directional audio from spherical content

Inventors: Scott Patrick Campbell (Belmont, CA); Zhinian Jing (Belmont, CA); Timothy Macmillan (La Honda, CA); David A. Newman (San Diego, CA); Balineedu Chowdary Adsumilli (San Mateo, CA)
Assignee: GoPro, Inc.
G11B27/3081H04N5/23238H04N9/806H04N9/8211
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,056,115
App. No.
15/396,385
Granted
Aug 21, 2018
Kind
B2
Abstract

A spherical content capture system captures spherical video and audio content. In one embodiment, captured metadata or video/audio processing is used to identify content relevant to a particular user based on time and location information. The platform can then generate an output video from one or more shared spherical content files relevant to the user. The output video may include a non-spherical reduced field of view such as those commonly associated with conventional camera systems. Particularly, relevant sub-frames having a reduced field of view may be extracted from each frame of spherical video to generate an output video that tracks a particular individual or object of interest. For each sub-frame, a corresponding portion of an audio track is generated that includes a directional audio signal having a directionality based on the selected sub-frame.

Claims (61)

1. A method for generating a video with corresponding audio, the method performing by a computing system including one or more processors, the method comprising:

receiving a video, by the computing system, the video comprising frames including a target, the video having a field of view;

receiving, by the computing system, audio channels representing audio captured concurrently with the video, individual audio channels comprising directional audio corresponding to respective different directions;

defining, by the computing system, respective spatial sub-regions to subdivide the frames of the video;

mapping, by the computing system, the audio channels to the spatial sub-regions of the video such that each sub-region is associated with at least one audio channel;

determining, by the computing system, a time-varying path of the target within the video based on an analysis of content of the video and/or information associated with the video;

extracting, by the computing system, sub-frames from the frames based on the time-varying path of the target, the sub-frames having a reduced field of view relative to the field of view of the video, the sub-frames including the target;

for individual sub-frames, by the computing system:

determining one or more of the spatial sub-regions overlapping a given sub-frame;

determining a composite audio channel, the composite audio channel including one or more of the audio channels mapped to the one or more of the spatial sub-regions overlapping the given sub-frame;

generating a portion of an audio stream from the composite audio channel; and

outputting, by the computing system, the sub-frames and the audio stream.

2. The method of claim 1 , wherein determining the composite audio channel further comprises:

weighting the audio channels based on an amount of overlap between the spatial sub-regions and the sub-frames; and

combining the weighted audio channels.

3. The method of claim 1 , wherein the individual audio channels corresponds to a direction perpendicular to faces of at least one member of a group consisting of: a tetrahedron, a cube, or a pyramid.

4. The method of claim 1 , wherein determining the time-varying path of the target within the video based on the analysis of the content of the video and/or the information associated with the video includes determining the time-varying path of the target within the video based on analysis of location information generated by a tracking device carried by the target.

5. The method of claim 1 , wherein determining the time-varying path of the target within the video based on the analysis of the content of the video and/or the information associated with the video includes determining the time-varying path of the target within the video based on visual recognition of the target within the video.

6. The method of claim 5 , wherein the visual recognition of the target includes facial recognition, object recognition, motion recognition, or gesture recognition.

7. The method of claim 1 , wherein determining the time-varying path of the target within the video based on the analysis of the content of the video and/or the information associated with the video includes determining the time-varying path of the target within the video based directionality of an audio source of the audio captured concurrently with the video.

8. The method of claim 1 , wherein mapping the audio channels to the spatial sub-regions is dependent on an orientation of a camera that was used to capture the video.

9. A non-transitory computer-readable storage medium storing instructions for generating a video with corresponding audio, the instructions when executed by one or more processors causing the one or more processors to perform steps including:

receiving a video, the video comprising frames including a target, the video having a field of view;

receiving audio channels representing audio captured concurrently with the video, individual audio channels comprising directional audio corresponding to respective different directions;

defining respective spatial sub-regions to subdivide the frames of the video;

mapping the audio channels to the spatial sub-regions of the video such that each sub-region is associated with at least one audio channel;

determining a time-varying path of the target within the video based on an analysis of content of the video and/or information associated with the video;

extracting sub-frames from the frames based on the time-varying path of the target, the sub-frames having a reduced field of view relative to the field of view of the video, the sub-frames including the target;

for individual sub-frames:

determining one or more spatial sub-regions overlapping a given sub-frame;

determining a composite audio channel, the composite audio channel including one or more of the audio channels mapped to the one or more of the spatial sub-regions overlapping the given sub-frame;

generating a portion of an audio stream from the composite audio channel; and

outputting the sub-frames and the audio stream.

10. The non-transitory computer-readable storage medium of claim 9 , wherein determining the composite audio channel comprises:

weighting the audio channels based on an amount of overlap between the spatial sub-regions and the sub-frames; and

combining the weighted audio channels.

11. The non-transitory computer-readable storage medium of claim 9 , wherein the individual audio channels corresponds to a direction perpendicular to the faces of at least one member of a group consisting of: a tetrahedron, a cube, or a pyramid.

12. The non-transitory computer-readable storage medium of claim 9 , wherein mapping the audio channels to the spatial sub-regions is dependent on an orientation of a camera that was used to capture the video.

13. The non-transitory computer-readable storage medium of claim 9 , wherein determining the time-varying path of the target within the video based on the analysis of the content of the video and/or the information associated with the video includes determining the time-varying path of the target within the video based on analysis of location information generated by a tracking device carried by the target.

14. The non-transitory computer-readable storage medium of claim 9 , wherein determining the time-varying path of the target within the video based on the analysis of the content of the video and/or the information associated with the video includes determining the time-varying path of the target within the video based on visual recognition of the target within the video.

15. A system for generating a video with corresponding audio, the system comprising:

one or more processors; and

a non-transitory computer-readable storage medium storing instructions that when executed by the one or more processors causes the one or more processors to perform steps including:

receiving a video, the video comprising frames including a target, the video having a field of view;

receiving audio channels representing audio captured concurrently with the video, individual audio channels comprising directional audio corresponding to respective different directions;

defining respective spatial sub-regions to subdivide the frames of the video, the spatial sub-regions corresponding to the respective different directions associated with the audio channels;

mapping the audio channels to the spatial sub-regions of the video such that each sub-region is associated with at least one audio channel;

determining a time-varying path of the target within the video based on an analysis of content of the video and/or information associated with the video;

extracting sub-frames from the frames based on the time-varying path of the target, the sub-frames having a reduced field of view relative to the field of view of the video, the sub-frames including the target;

for individual sub-frames:

determining one or more spatial sub-regions overlapping a given sub-frame;

determining a composite audio channel, the composite audio channel including one or more of the audio channels mapped to the one or more spatial sub-regions overlapping the given sub-frame;

generating a portion of an audio stream from the composite audio channel; and

outputting the sub-frames and the audio stream.

16. The system of claim 15 , wherein determining the composite audio channel further comprises:

weighting the audio channels based on an amount of overlap between the spatial sub-regions and the sub-frames; and

combining the weighted audio channels.

17. The system of claim 15 , wherein the individual audio channels corresponds to a direction perpendicular to faces of at least one member of a group consisting of: a tetrahedron, a cube, or a pyramid.

18. The system of claim 15 , wherein determining the time-varying path of the target within the video based on the analysis of the content of the video and/or the information associated with the video includes determining the time-varying path of the target within the video based on analysis of location information generated by a tracking device carried by the target.

19. The system of claim 15 , wherein determining the time-varying path of the target within the video based on the analysis of the content of the video and/or the information associated with the video includes determining the time-varying path of the target within the video based on visual recognition of the target within the video.

20. The system of claim 15 , wherein determining the time-varying path of the target within the video based on the analysis of the content of the video and/or the information associated with the video includes determining the time-varying path of the target within the video based directionality of an audio source of the audio captured concurrently with the video.

Assignments (7)
SECURITY INTEREST Recorded Aug 4, 2025
From: GOPRO, INC.
To: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
Reel/Frame 072358/0001 →
SECURITY INTEREST Recorded Aug 4, 2025
From: GOPRO, INC.
To: FARALLON CAPITAL MANAGEMENT, L.L.C., AS AGENT
Reel/Frame 072340/0676 →
RELEASE OF PATENT SECURITY INTEREST Recorded Jan 25, 2021
From: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
To: GOPRO, INC.
Reel/Frame 055106/0434 →
CORRECTIVE ASSIGNMENT TO CORRECT THE SCHEDULE TO REMOVE APPLICATION 15387383 AND REPLACE WITH 15385383 PREVIOUSLY RECORDED ON REEL 042665 FRAME 0065. ASSIGNOR(S) HEREBY CONFIRMS THE SECURITY INTEREST. Recorded Oct 23, 2019
From: GOPRO, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 050808/0824 →
SECURITY INTEREST Recorded Dec 3, 2018
From: GOPRO, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 047713/0309 →
SECURITY INTEREST Recorded Jun 1, 2017
From: GOPRO, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 042665/0065 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 3, 2017
From: CAMPBELL, SCOTT PATRICK; JING, ZHINIAN; MACMILLAN, TIMOTHY; NEWMAN, DAVID A.; ADSUMILLI, BALINEEDU CHOWDARY
To: GOPRO, INC.
Reel/Frame 040827/0581 →
Continuity (3)
Continuation 14789706 · Jul 1, 2015
Provisional Application 62020867 · Jul 3, 2014
Related Publication 20170110155A1 · Apr 20, 2017