IP Library Granted Patent US 9,756,281
Granted Patent B2
US 9,756,281 · App. 15/017,541 · Granted Sep 5, 2017

Apparatus and method for audio based video synchronization

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,756,281
App. No.
15/017,541
Granted
Sep 5, 2017
Kind
B2
Abstract

Multiple video recordings may be synchronized using audio features of the recordings. A synchronization process may compare energy tracks of each recording within a multi-resolution framework to correlate audio features of one recording to another.

Claims (50)

1. A method for synchronizing audio tracks, the method comprising:

(a) obtaining a first feature track for the first audio track including information representing sound captured from a live occurrence, the first feature track representing individual feature samples associated with the sound captured on the first audio track from the live occurrence, the first feature track conveying feature magnitude for individual sampling periods of a first sampling period length;

(b) obtaining a second feature track for the first audio track, the second feature track representing individual feature samples associated with the sound captured on the first audio track from the live occurrence, the second feature track conveying feature magnitude for individual sampling periods of a second sampling period length, the second sampling period length being different than the first sampling period length;

(c) obtaining a third feature track for the second audio track including information representing sound captured from the live occurrence, the third feature track representing individual feature samples associated with the sound captured on the second audio track from the live occurrence, the third feature track conveying feature magnitude for individual sampling periods of the first sampling period length;

(d) obtaining a fourth feature track for the second audio track, the fourth feature track representing individual feature samples associated with the sound captured on the second audio track from the live occurrence, the fourth feature track conveying feature magnitude for individual sampling periods of the second sampling period length;

(e) comparing at least a portion of the first feature track against at least a portion of the third feature track to correlate one or more features in the first feature track with one or more features in the third feature track, the correlated features being identified as potentially representing features in the same sounds from the live occurrence;

(f) comparing a portion of the second feature track against a portion of the fourth feature track to correlate one or more features in the second feature track with one or more features in the fourth feature track, the portion of the second feature track and the portion of the fourth feature track being selected for the comparison based on the correlation of features between the first feature track and the third feature track;

(g) determining, from the correlations of features in the feature tracks, a temporal alignment estimate between the first audio track and the second audio track, the temporal alignment estimate reflecting an offset in time between commencement of sound capture for the first audio track and commencement of sound capture for the second audio track;

(h) obtaining a temporal alignment threshold;

(i) determining whether to continue comparing feature tracks associated with the first audio track and the second audio track based on the comparison of the temporal alignment estimate and the temporal alignment threshold, including determining to not continue comparing feature tracks associated with the first audio track and the second track in response to the temporal alignment estimate being smaller than the alignment threshold;

(j) responsive to the determining to not continue comparing feature tracks associated with the first audio track and the second audio track, using the temporal alignment estimate to synchronize the first audio track with the second audio track;

(k) responsive to the determining to continue comparing feature tracks associated with the first audio track and the second audio track, iterating after operations (a) through (k).

2. The method of claim 1 , wherein the first feature track represents individual energy samples associated with the sound captured on the first audio track from the live occurrence.

3. The method of claim 1 , wherein the temporal alignment estimate reflects corresponding feature samples between the first feature track and the third feature track.

4. The method of claim 1 , further comprising:

selecting a comparison window to at least one portion of the first feature track and to at least one portion the third feature track, the comparison window having a start position and an end position, such that the start position corresponding with the individual feature sample having been selected at random, the end position corresponding with the individual feature sample having a predetermined value.

5. The method of claim 4 , wherein the comparison window is selected to at least one portion of the second feature track and to at least one portion of the fourth feature track based on the start position and the end position first feature track and the third feature track.

6. The method of claim 1 , further comprising:

determining whether to continue comparing feature tracks associated with the first audio track and the second audio track by assessing whether a stopping criteria has been satisfied, such determination being based on the temporal alignment estimate and the stopping criteria.

7. The method of claim 6 , wherein the stopping criteria is satisfied by multiple, consecutive determinations of the temporal alignment estimate falling within a specific range or ranges.

8. The method of claim 7 , wherein the specific range or ranges are bounded by a temporal alignment threshold or thresholds.

9. The method of claim 1 , wherein the first audio track is generated from a first media file captured from the live occurrence by a first sound capture device, the first media file including audio and video information.

10. The method of claim 1 , wherein the second audio track is generated from a second media file captured from the live occurrence by a second sound capture device, the second media file including audio and video information.

11. The method of claim 1 , further comprising:

applying at least one frequency band to the first audio track and applying at least one frequency band to the second audio track.

12. A system for synchronizing audio tracks, the system comprising:

one or more physical computer processors is further configured by computer readable instructions to:

(a) obtain a first feature track for the first audio track including information representing sound captured from a live occurrence, the first feature track representing individual feature samples associated with the sound captured on the first audio track from the live occurrence, the first feature track conveying feature magnitude for individual sampling periods of a first sampling period length;

(b) obtain a second feature track for the first audio track, the second feature track representing individual feature samples associated with the sound captured on the first audio track from the live occurrence, the second feature track conveying feature magnitude for individual sampling periods of a second sampling period length, the second sampling period length being different than the first sampling period length;

(c) obtain a third feature track for the second audio track, the third feature track representing individual feature samples associated with the sound captured on the second audio track from the live occurrence, the third feature track conveying feature magnitude for individual sampling periods of the first sampling period length;

(d) obtain a fourth feature track for the second audio track, the fourth feature track representing individual feature samples associated with the sound captured on the second audio track from the live occurrence, the fourth feature track conveying feature magnitude for individual sampling periods of the second sampling period length;

(e) compare at least a portion of the first feature track against at least a portion of the third feature track to correlate one or more features in the first feature track with one or more features in the third feature track, the correlated features being identified as potentially representing features in the same sounds from the live occurrence;

(f) compare a portion of the second feature track against a portion of the fourth feature track to correlate one or more features in the second feature track with one or more features in the fourth feature track, the portion of the second feature track and the portion of the fourth feature track being selected for the comparison based on the correlation of features between the first feature track and the third feature track;

(g) determine, from the correlations of features in the feature tracks, a temporal alignment estimate between the first audio track and the second audio track, the temporal alignment estimate reflecting an offset in time between commencement of sound capture for the first audio track and commencement of sound capture for the second audio track;

(h) obtain a temporal alignment threshold;

(i) determine whether to continue comparing feature tracks associated with the first audio track and the second audio track based on the comparison of the temporal alignment estimate and the temporal alignment threshold, including determining to not continue comparing feature tracks associated with the first audio track and the second track in response to the temporal alignment estimate being smaller than the alignment threshold;

(j) responsive to the determining to not continue comparing feature tracks associated with the first audio track and the second audio track, using the temporal alignment estimate to synchronize the first audio track with the second audio track;

(k) responsive to the determining to continue comparing feature tracks associated with the first audio track and the second audio track, iterating after operations (a) through (k).

13. The system of claim 12 , wherein the temporal alignment estimate reflects corresponding feature samples between the first feature track and the third feature track.

14. The system of claim 12 , further comprising:

selecting a comparison window to at least one portion of the first feature track and to at least one portion the third feature track, the comparison window having a start position and an end position, such that the start position corresponding with the individual feature sample having been selected at random, the end position corresponding with the individual feature sample having a predetermined value.

15. The system of claim 14 , wherein the comparison window is selected to at least one portion of the second feature track and to at least one portion of the fourth feature track based on the start position and the end position first feature track and the third feature track.

16. The system of claim 12 , wherein the one or more physical computer processors is further configured by computer readable instructions to:

determine whether to continue comparing feature tracks associated with the first audio track and the second audio track by assessing whether a stopping criteria has been satisfied, such determination being based on the temporal alignment estimate and the stopping criteria.

17. The system of claim 16 , wherein the stopping criteria is satisfied by multiple, consecutive determinations of the temporal alignment estimate falling within a specific range or ranges.

18. The system of claim 17 , wherein the specific range or ranges are bounded by a temporal alignment threshold or thresholds.

19. The system of claim 12 , wherein the first audio track is generated from a first media file captured from the live occurrence by a first sound capture device, the first media file including audio and video information.

20. The system of claim 12 , wherein the second audio track is generated from a second media file captured from the live occurrence by a second sound capture device, the second media file including audio and video information.

21. The system of claim 12 , wherein the one or more physical computer processors is further configured by computer readable instructions to:

apply at least one frequency band to the first audio track and applying at least one frequency band to the second audio track.

Assignments (5)
SECURITY INTEREST Recorded Aug 4, 2025
From: GOPRO, INC.
To: FARALLON CAPITAL MANAGEMENT, L.L.C., AS AGENT
Reel/Frame 072340/0676 →
SECURITY INTEREST Recorded Aug 4, 2025
From: GOPRO, INC.
To: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
Reel/Frame 072358/0001 →
RELEASE OF PATENT SECURITY INTEREST Recorded Jan 25, 2021
From: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
To: GOPRO, INC.
Reel/Frame 055106/0434 →
SECURITY INTEREST Recorded Feb 22, 2017
From: GOPRO, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 041777/0440 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 23, 2016
From: TCHENG, DAVID K.
To: GOPRO, INC.
Reel/Frame 039504/0498 →