IP Library › Granted Patent US 12,430,914
Granted Patent B1
US 12,430,914 · App. 18/186,533 · Granted Sep 30, 2025

Generating summaries of events based on sound intensities

Inventors: Eli Alshan (Kfar-Saba, IL); Gilad Cohen (Raanana, IL); Ido Yerushalmy (Tel-Aviv, IL)
Assignee: Amazon Technologies, Inc.
G06V20/47G06F16/7834G06F16/7867G06F16/787G06V20/42
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,430,914
App. No.
18/186,533
Granted
Sep 30, 2025
Kind
B1
Abstract

Summaries of broadcasts of live events are generated based on intensities of audio signals captured during the live events. Streams of multimedia including video signals and audio signals are captured by one or more cameras. The video signals are processed to identify specific activities of interest (e.g., plays of a sporting event) during the live event. Audio signals captured concurrently with the video signals are processed to determine their respective intensities or other acoustic characteristics. The activities of interest are ranked based on intensities of the audio signals. A multimedia stream representing a summary of a media program and includes the highest-ranking video signals and corresponding audio signals is generated and transmitted to one or more devices of viewers.

Claims (59)

1. A computer-implemented method comprising:

identifying a plurality of multimedia streams captured by a plurality of cameras during an event, wherein each one of the multimedia streams comprises a set of video signals and a set of audio signals captured simultaneously by one of the plurality of cameras;

determining that each one of a first plurality of sets of video signals depicts at least one of a plurality of activities of the event, wherein each one of the first plurality of the sets of video signals is included in one of the plurality of multimedia streams;

identifying a first plurality of sets of audio signals, wherein each one of the first plurality of sets of audio signals is included in one of the plurality of multimedia streams with one of the first plurality of sets of video signals;

determining at least one acoustic characteristic of each one of the first plurality of sets of audio signals;

selecting a first activity based at least in part on a first acoustic characteristic of a first set of audio signals, wherein the first set of audio signals is included in one of the plurality of multimedia streams with a first set of video signals, and wherein the first set of video signals depicts the first activity;

generating a first multimedia stream representing at least the first activity, wherein the first multimedia stream comprises the first set of video signals and the first set of audio signals concurrent with the first set of video signals; and

storing the first multimedia stream in association with the event.

2. The computer-implemented method of claim 1 , further comprising:

identifying metadata associated with the event,

wherein the metadata comprises identifiers of at least one of:

a start time of at least one of the plurality of activities;

an end time of at least one of the plurality of activities;

a location of at least one of the plurality of activities;

a participant in at least one of the plurality of activities;

a time at which a change in a status of the at least one of the plurality of activities occurred;

the change in the status of the at least one of the plurality of activities; or

a description of the at least one of the plurality of activities, and

wherein that each of the first plurality of sets of video signals depicts one of the plurality of activities is determined based at least in part on the metadata.

3. The computer-implemented method of claim 1 , wherein the event is a live sporting event, and

wherein each of the plurality of activities is a play at the live sporting event.

4. The computer-implemented method of claim 1 , further comprising:

determining that each one of a second plurality of sets of video signals depicts one of the plurality of activities of the event, wherein each one of the second plurality of the sets of video signals is included in one of the plurality of multimedia streams;

identifying a second plurality of sets of audio signals, wherein each one of the second plurality of sets of audio signals is included in one of the plurality of multimedia streams with one of the second plurality of sets of video;

determining at least one acoustic characteristic of each one of the second plurality of sets of audio signals; and

determining a ranking of the plurality of activities based at least in part on the at least one acoustic characteristic of each one of the first plurality of sets of audio signals and the at least one acoustic characteristic of each one of the second plurality of sets of audio signals,

wherein the first activity is selected based at least in part on the ranking.

5. The computer-implemented method of claim 1 , further comprising:

selecting a second activity based at least in part on a second acoustic characteristic of a second set of audio signals, wherein the second set of audio signals is included in one of the plurality of multimedia streams with a second set of video signals, and wherein the second set of video signals depicts the second activity,

wherein the first multimedia stream comprises the first set of video signals depicting at least the first activity in series with the second set of video signals depicting at least the second activity, and

wherein the first multimedia stream further comprises the first set of audio signals concurrent with the first set of video signals and the second set of audio signals concurrent with the second set of video signals.

6. The computer-implemented method of claim 5 , wherein an order of the first set of video signals and the second set of video signals in the first multimedia stream is selected based at least in part on the first acoustic characteristic and the second acoustic characteristic.

7. The computer-implemented method of claim 5 , wherein an order of the first set of video signals and the second set of video signals in the first multimedia stream is selected based at least in part on a first time at which the first activity occurred during the event and a second time at which the second activity occurred during the event.

8. The computer-implemented method of claim 5 , wherein the first multimedia stream further comprises at least one of:

a video transition between the first set of video signals and the second set of video signals; or

an audio transition between the first set of audio signals and the second set of audio signals.

9. The computer-implemented method of claim 5 , wherein the first set of video signals was captured by a first camera, and

wherein the second set of video signals was captured by a second camera.

10. The computer-implemented method of claim 1 , wherein each one of the plurality of cameras is in communication with a media distribution system, and

wherein the computer-implemented method further comprises:

broadcasting a second multimedia stream live to at least one personal device over at least one network by the media distribution system,

wherein the second multimedia stream comprises at least some of the first plurality of sets of video signals and at least some of the first plurality of sets of audio signals.

11. The computer-implemented method of claim 10 , further comprising:

broadcasting at least the first multimedia stream to the at least one personal device over the at least one network by the media distribution system,

wherein at least the first multimedia stream is broadcast to the at least one personal device after the second multimedia stream.

12. The computer-implemented method of claim 1 , wherein each one of the plurality of cameras is mounted within a venue, and

wherein the venue is one of an amphitheater, an arena, an auditorium, a ballpark, a convention center, a resort, a restaurant, a stadium or a theater.

13. The computer-implemented method of claim 1 , wherein the event is one of:

a baseball game;

a basketball game;

a concert;

a football game;

a golf match;

a hockey game;

a parade;

a public meeting;

a soccer game;

a social gathering; or

a tennis match.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 20, 2023
From: ALSHAN, ELI; COHEN, GILAD; YERUSHALMY, IDO
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 063035/0552 →
References Cited (36)
US 5351075A · Herz et al. · 1994 [cited by applicant]
US 6018768A · Ullman et al. · 2000 [cited by applicant]
US 6457010B1 · Eldering et al. · 2002 [cited by applicant]
US 6760916B2 · Holtz et al. · 2004 [cited by applicant]
US 7493636B2 · Kitsukawa et al. · 2009 [cited by applicant]
US 7752642B2 · Lemmons · 2010 [cited by applicant]
US 8627379B2 · Kokenos et al. · 2014 [cited by applicant]
US 8630844B1 · Nichols · 2014 [cited by examiner]
US 8949890B2 · Evans et al. · 2015 [cited by applicant]
US 9697178B1 · Nichols · 2017 [cited by examiner]
US 10363488B1 · Willette et al. · 2019 [cited by applicant]
US 20030028873A1 · Lemmons · 2003 [cited by applicant]
US 20070029112A1 · Li et al. · 2007 [cited by applicant]
US 20080037951A1 · Cho et al. · 2008 [cited by applicant]
US 20100050202A1 · Kandekar et al. · 2010 [cited by applicant]
US 20110217019A1 · Kamezawa · 2011 [cited by examiner]
US 20150248917A1 · Chang · 2015 [cited by examiner]
US 20160065884A1 · Di Censo · 2016 [cited by examiner]
US 20160117928A1 · Hodges · 2016 [cited by examiner]
US 20160117940A1 · Gomory · 2016 [cited by examiner]
US 20160225410A1 · Lee · 2016 [cited by examiner]
US 20160321506A1 · Fridental · 2016 [cited by examiner]
US 20170099526A1 · Hua et al. · 2017 [cited by applicant]
US 20170201779A1 · Publicover · 2017 [cited by examiner]
US 20170259115A1 · Hall · 2017 [cited by examiner]
US 20170266491A1 · Rissanen · 2017 [cited by examiner]
US 20180084291A1 · Wei et al. · 2018 [cited by applicant]
US 20190356948A1 · Stojancic · 2019 [cited by examiner]
US 20200097502A1 · Trim · 2020 [cited by examiner]
US 20220067384A1 · Kaushik · 2022 [cited by examiner]
Dosovitskiy, Alexey, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani et al. “An image is worth 16×16 words: Transformers for image recognition at scale.” arXiv pre… [cited by applicant]
Frome, Andrea, Greg S. Corrado, Jon Shlens, Samy Bengio, Jeff Dean, Marc'Aurelio Ranzato, and Tomas Mikolov. “DeViSE: A deep visual-semantic embedding model.” Advances in Neural Information Processing Systems 26 (2013).… [cited by applicant]
Li, Ang, Allan Jabri, Armand Joulin, and Laurens Van Der Maaten. “Learning visual n-grams from web data.” In Proceedings of the IEEE International Conference on Computer Vision, pp. 4183-4192. 2017. URL: https://openacc… [cited by applicant]
Radford, Alec, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry et al. “Learning transferable visual models from natural language supervision.” In International Conference on Mac… [cited by applicant]
Socher, Richard, Milind Ganjoo, Christopher D. Manning, and Andrew Ng. “Zero-shot learning through cross-modal transfer.” Advances in Neural Information Processing Systems 26 (2013). 10 pages. URL: https://proceedings.n… [cited by applicant]
Vaswani, A. et al., 2017, Attention is All you Need. In Annual Conference on Neural Information Processing Systems 2017, Dec. 4-9, 2017, Long Beach, CA, USA, pp. 6000-6010, Retrieved: https://arxiv.org/pdf/1706.03762.pd… [cited by applicant]