IP Library Granted Patent US 10,410,680
Granted Patent B2
US 10,410,680 · App. 16/105,304 · Granted Sep 10, 2019

Automatic generation of video and directional audio from spherical content

Inventors: Scott Patrick Campbell (Belmont, CA); Zhinian Jing (Belmont, CA); Timothy Macmillan (La Honda, CA); David A. Newman (San Diego, CA); Balineedu Chowdary Adsumilli (San Mateo, CA)
Assignee: GoPro, Inc.
G11B27/3081H04N5/23238H04N5/77H04N9/806H04N9/8205H04N9/8211
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,410,680
App. No.
16/105,304
Granted
Sep 10, 2019
Kind
B2
Abstract

A spherical content capture system captures spherical video and audio content. In one embodiment, captured metadata or video/audio processing is used to identify content relevant to a particular user based on time and location information. The platform can then generate an output video from one or more shared spherical content files relevant to the user. The output video may include a non-spherical reduced field of view such as those commonly associated with conventional camera systems. Particularly, relevant sub-frames having a reduced field of view may be extracted from each frame of spherical video to generate an output video that tracks a particular individual or object of interest. For each sub-frame, a corresponding portion of an audio track is generated that includes a directional audio signal having a directionality based on the selected sub-frame.

Claims (39)

1. A method for generating a video with corresponding audio, the method performing by a computing system including one or more processors, the method comprising:

receiving, by the computing system, a video, the video comprising frames including a target, the video having a field of view;

receiving, by the computing system, directional audio signals captured concurrently with the video;

determining, by the computing system, a time-varying path of the target within the video based on an analysis of content of the video or information associated with the video;

identifying, by the computing system, sub-frames from the frames based on the time-varying path of the target, the sub-frames having a reduced field of view relative to the field of view of the video, the sub-frames including the target;

generating, by the computing system, an audio stream from the directional audio signals based on the time-varying path of the target, the audio stream including portions of one or more of the directional audio signals corresponding to a direction of the target; and

outputting, by the computing system, the sub-frames and the audio stream.

2. The method of claim 1 , wherein determining the audio stream further comprises selecting the portions based on matching directionalities of the directional audio signals with the direction of the target.

3. The method of claim 1 , wherein the directional audio signals correspond to directions perpendicular to faces of at least one member of a group consisting of: a tetrahedron, a cube, or a pyramid.

4. The method of claim 1 , wherein the information associated with the video includes location information generated by a tracking device carried by the target.

5. The method of claim 1 , wherein the analysis of the content of the video includes visual recognition of the target within the video.

6. The method of claim 5 , wherein the visual recognition of the target includes facial recognition, object recognition, motion recognition, or gesture recognition.

7. The method of claim 1 , wherein the information associated with the video includes a directionality of an audio source of one or more of the directional audio signals.

8. The method of claim 1 , wherein a scene motion analysis is performed based on the directional audio signals.

9. A non-transitory computer-readable storage medium storing instructions for generating a video with corresponding audio, the instructions when executed by one or more processors causing the one or more processors to perform steps including:

receiving a video, the video comprising frames including a target, the video having a field of view;

receiving directional audio signals captured concurrently with the video;

determining a time-varying path of the target within the video based on an analysis of content of the video or information associated with the video;

identifying sub-frames from the frames based on the time-varying path of the target, the sub-frames having a reduced field of view relative to the field of view of the video, the sub-frames including the target;

generating an audio stream from the directional audio signals based on the time-varying path of the target, the audio stream including portions of one or more of the directional audio signals corresponding to a direction of the target; and outputting the sub-frames and the audio stream.

10. The non-transitory computer-readable storage medium of claim 9 , wherein determining the audio stream further comprises selecting the portions based on matching directionalities of the directional audio signals with the direction of the target.

11. The non-transitory computer-readable storage medium of claim 9 , wherein the directional audio signals correspond to directions perpendicular to the faces of at least one member of a group consisting of: a tetrahedron, a cube, or a pyramid.

12. The non-transitory computer-readable storage medium of claim 9 , wherein the information associated with the video includes location information generated by a tracking device carried by the target.

13. The non-transitory computer-readable storage medium of claim 9 , wherein the analysis of the content of the video includes visual recognition of the target within the video.

14. The non-transitory computer-readable storage medium of claim 13 , wherein the visual recognition of the target includes facial recognition, object recognition, motion recognition, or gesture recognition.

15. A system for generating a video with corresponding audio, the system comprising:

one or more processors; and

a non-transitory computer-readable storage medium storing instruction that when executed by the one or more processors causes the one or more processors to perform steps including:

receiving a video, the video comprising frames including a target, the video having a field of view;

receiving directional audio signals captured concurrently with the video;

determining a time-varying path of the target within the video based on an analysis of content of the video or information associated with the video;

identifying sub-frames from the frames based on the time-varying path of the target, the sub-frames having a reduced field of view relative to the field of view of the video, the sub-frames including the target;

generating an audio stream from the directional audio signals based on the time-varying path of the target, the audio stream including portions of one or more of the directional audio signals corresponding to a direction of the target; and

outputting the sub-frames and the audio stream.

16. The system of claim 15 , wherein determining the audio stream further comprises selecting the portions based on matching directionalities of the directional audio signals with the direction of the target.

17. The system of claim 15 , wherein the directional audio signals correspond to directions perpendicular to faces of at least one member of a group consisting of: a tetrahedron, a cube, or a pyramid.

18. The system of claim 15 , wherein the information associated with the video includes location information generated by a tracking device carried by the target.

19. The system of claim 15 , wherein the analysis of the content of the video includes visual recognition of the target within the video.

20. The system of claim 15 , wherein the information associated with the video includes a directionality of an audio source of one or more of the directional audio signals.

Assignments (6)
SECURITY INTEREST Recorded Aug 4, 2025
From: GOPRO, INC.
To: FARALLON CAPITAL MANAGEMENT, L.L.C., AS AGENT
Reel/Frame 072340/0676 →
SECURITY INTEREST Recorded Aug 4, 2025
From: GOPRO, INC.
To: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
Reel/Frame 072358/0001 →
RELEASE OF PATENT SECURITY INTEREST Recorded Jan 25, 2021
From: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
To: GOPRO, INC.
Reel/Frame 055106/0434 →
SECURITY INTEREST Recorded Oct 19, 2020
From: GOPRO, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 054113/0594 →
SECURITY INTEREST Recorded Dec 3, 2018
From: GOPRO, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 047713/0309 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 24, 2018
From: CAMPBELL, SCOTT PATRICK; JING, ZHINIAN; MACMILLAN, TIMOTHY; NEWMAN, DAVID A.; ADSUMILLI, BALINEEDU CHOWDARY
To: GOPRO, INC.
Reel/Frame 046699/0948 →
Continuity (4)
Continuation 15396385 · Dec 30, 2016
Continuation 14789706 · Jul 1, 2015
Provisional Application 62020867 · Jul 3, 2014
Related Publication 20190005987A1 · Jan 3, 2019