IP Library Granted Patent US 12,219,215
Granted Patent B2
US 12,219,215 · App. 18/521,657 · Granted Feb 4, 2025

Systems and methods for displaying subjects of a video portion of content

Inventors: Gabriel C Dalbec (Morgan Hill, CA); Nicholas Lovell (Santa Clara, CA); Lance G. O'Connor (Sunnyvale, CA)
Assignee: Adeia Guides Inc.
H04N21/47217G06T7/246G06V20/40G06V40/164H04N21/4316H04N21/8456
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,219,215
App. No.
18/521,657
Granted
Feb 4, 2025
Kind
B2
Abstract

Systems and methods are described herein for displaying subjects of a portion of content. Media data of content is analyzed during playback, and a number of action signatures are identified. Each action signature is associated with a particular subject within the content. The action signature is stored, along with a timestamp corresponding to a playback position at which the action signature begins, in association with an identifier of the particular subject. Upon receiving a command, icons representing each of a number of action signatures at or near the current playback position are displayed. Upon receiving user selection of an icon corresponding to a particular signature, a portion of the content corresponding to the action signature is played back.

Claims (42)

1. A method for displaying subjects of a portion of audio of content, the method comprising:

receiving at least one video frame from a series of video frames of a content item;

identifying, using audio processing, an audio signature in the content item, wherein the audio signature is associated with the at least one video frame;

determining, using video processing, whether one or more subjects are displayed in the at least one video frame;

based at least in part on the determining whether one or more subjects are displayed in the at least one video frame, associating a subject with the audio signature based at least in part on determining that a mouth of the subject in the at least one video frame is moving;

receiving an input command; and

based on the input command, generating for display an icon of the associated subject representing the audio signature.

2. The method of claim 1 wherein video processing is one of edge detection, facial recognition, or object recognition.

3. The method of claim 1 wherein associating a subject with the audio signature is based at least in part on audio characteristics of the audio signature.

4. The method of claim 3 wherein the audio characteristics of the audio signature include at least either audio frequency or speech pattern.

5. The method of claim 1 wherein:

determining, using video processing, whether one or more subjects are displayed in the at least one video frame indicates no subjects are displayed in the at least one video frame; and

wherein the audio signature is associated with a new subject.

6. The method of claim 1 wherein determining, using video processing, whether one or more subjects are displayed in the at least one video frame indicates no subjects are displayed in the at least one video frame, the method further comprising:

comparing the audio signature to a database of known subjects.

7. The method of claim 1 wherein the at least one video frame and audio of the audio signature occur simultaneously in the content item.

8. A system comprising:

control circuitry configured to:

receive at least one video frame from a series of video frames of a content item;

identify, using audio processing, an audio signature in the content item, wherein the audio signature is associated with the at least one video frame;

determine, using video processing, whether one or more subjects are displayed in the at least one video frame;

based at least in part on the determining whether one or more subjects are displayed in the at least one video frame, associate a subject with the audio signature based on at least in part on determining that a mouth of the subject in the at least one video frame is moving;

receive an input command; and

based on the input command, generate for display an icon of the associated subject representing the audio signature.

9. The system of claim 8 wherein video processing is one of edge detection, facial recognition, or object recognition.

10. The system of claim 8 wherein to associate a subject with the audio signature is based at least in part on audio characteristics of the audio signature.

11. The system of claim 10 wherein the audio characteristics of the audio signature include at least either audio frequency or speech pattern.

12. The system of claim 8 wherein:

to determine, using video processing, whether one or more subjects are displayed in the at least one video frame indicates no subjects are displayed in the at least one video frame; and

wherein the audio signature is associated with a new subject.

13. The system of claim 8 wherein to determine, using video processing, whether one or more subjects are displayed in the at least one video frame indicates no subjects are displayed in the at least one video frame, the control circuitry is further configured to:

compare the audio signature to a database of known subjects.

14. The system of claim 8 wherein the at least one video frame and audio of the audio signature occur simultaneously in the content item.

15. A non-transitory computer-readable medium having instructions encoded thereon that when executed by control circuitry causes the control circuitry to:

receive at least one video frame from a series of video frames of a content item;

identify, using audio processing, an audio signature in the content item, wherein the audio signature is associated with the at least one video frame;

determine, using video processing, whether one or more subjects are displayed in the at least one video frame;

based at least in part on the determining whether one or more subjects are displayed in the at least one video frame, associate a subject with the audio signature based at least in part on determining that a mouth of the subject in the at least one video frame is moving;

receive an input command; and

based on the input command, generate for display an icon of the associated subject representing the audio signature.

16. The non-transitory computer-readable medium of claim 15 wherein video processing is one of edge detection, facial recognition, or object recognition.

17. The non-transitory computer-readable medium of claim 15 wherein to associate a subject with the audio signature is based at least in part on audio characteristics of the audio signature.

Assignments (2)
CHANGE OF NAME Recorded Oct 3, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069106/0171 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 29, 2023
From: DALBEC, GABRIEL C.; LOVELL, NICHOLAS; O'CONNOR, LANCE G.
To: ROVI GUIDES, INC.
Reel/Frame 065710/0053 →
Continuity (4)
Continuation 17965140 · Oct 13, 2022
Continuation 17491153 · Sep 30, 2021
Continuation 16226916 · Dec 20, 2018
Related Publication 20240214640A1 · Jun 27, 2024
References Cited (59)
US 6567775B1 · Maali et al. · 2003 [cited by applicant]
US 6748356B1 · Beigi et al. · 2004 [cited by applicant]
US 6922478B1 · Konen et al. · 2005 [cited by applicant]
US 9258604B1 · Bilobrov et al. · 2016 [cited by applicant]
US 9454993B1 · Lawson et al. · 2016 [cited by applicant]
US 9558749B1 · Secker-Walker et al. · 2017 [cited by applicant]
US 9578377B1 · Malik et al. · 2017 [cited by applicant]
US 20030149881A1 · Patel et al. · 2003 [cited by applicant]
US 20050254685A1 · Miyamori · 2005 [cited by applicant]
US 20070279494A1 · Aman et al. · 2007 [cited by applicant]
US 20080235724A1 · Sassenscheidt et al. · 2008 [cited by applicant]
US 20090122198A1 · Thorn · 2009 [cited by applicant]
US 20090278937A1 · Botchen et al. · 2009 [cited by applicant]
US 20100115542A1 · Lee · 2010 [cited by applicant]
US 20110022589A1 · Bauer et al. · 2011 [cited by applicant]
US 20110035221A1 · Zhang · 2011 [cited by examiner]
US 20110211802A1 · Kamada et al. · 2011 [cited by applicant]
US 20120114233A1 · Gunatilake · 2012 [cited by applicant]
US 20120163677A1 · Thorn · 2012 [cited by examiner]
US 20120203757A1 · Ravindran · 2012 [cited by applicant]
US 20130011121A1 · Forsyth et al. · 2013 [cited by applicant]
US 20130152139A1 · Davis et al. · 2013 [cited by applicant]
US 20130160038A1 · Slaney et al. · 2013 [cited by applicant]
US 20140037264A1 · Jackson et al. · 2014 [cited by applicant]
US 20140114656A1 · Cheung · 2014 [cited by applicant]
US 20140245339A1 · Zhang et al. · 2014 [cited by applicant]
US 20150356332A1 · Turner et al. · 2015 [cited by applicant]
US 20160275588A1 · Ye et al. · 2016 [cited by applicant]
US 20160295273A1 · Ehlers et al. · 2016 [cited by applicant]
US 20160301972A1 · Liu et al. · 2016 [cited by applicant]
US 20160337701A1 · Khare et al. · 2016 [cited by applicant]
US 20170041684A1 · Krishnamurthy et al. · 2017 [cited by applicant]
US 20170199934A1 · Nongpiur et al. · 2017 [cited by applicant]
US 20170264970A1 · Mitra et al. · 2017 [cited by applicant]
US 20170339446A1 · Arms · 2017 [cited by applicant]
US 20170352380A1 · Doumbouya et al. · 2017 [cited by applicant]
US 20180070008A1 · Tyagi · 2018 [cited by examiner]
US 20180192101A1 · Bilobrov · 2018 [cited by applicant]
US 20180220195A1 · Panchaksharaiah et al. · 2018 [cited by applicant]
US 20190013047A1 · Wait et al. · 2019 [cited by applicant]
US 20190180149A1 · Knittel · 2019 [cited by applicant]
US 20190289359A1 · Sekar et al. · 2019 [cited by applicant]
US 20190341011A1 · Neuhauser et al. · 2019 [cited by applicant]
US 20220021942A1 · Dalbec et al. · 2022 [cited by applicant]
US 20230188794A1 · Dalbec et al. · 2023 [cited by applicant]
JP 2005175525A · 2005 [cited by applicant]
JP 2005210573A · 2005 [cited by applicant]
JP 2005286524A · 2005 [cited by applicant]
WO 2017157428A1 · 2017 [cited by applicant]
International Search Report and Written Opinion of PCT/US2019/067498 dated Jul. 3, 2020. [cited by applicant]
Partial International Search Report of PCT/US2019/067498 dated Apr. 28, 2020. [cited by applicant]
Friedland , et al., “Using artistic markers and speaker identification for narrative-theme navigation of Seinfeld episodes,” 2009 11th IEEE International symposium on Multimedia, IEEE Computer Society 511-516. [cited by applicant]
Miro , et al., “Speaker diarization: A review of recent research,” IEEE Transactions on Audio Speech, and Language Processing 20: (2) 356-370 (2012). [cited by applicant]
Shih H, “A Survey of Content-Aware Video Analysis for Sports”, IEEE Transactions on Circuits and Systems for Video Technology, vol. 28. No. 5, pp. 1212-1231, May 2018. [cited by applicant]
Tejero-De-Pablos Antonio, et al., “Summarization of User-Generated Sports Video by Using Deep Action Recognition Features”, IEEE Transactions on Multimedia, vol. 20, No. 8, pp. 2000-2011 (2018). [cited by applicant]
Tim Stefan Chan Wai, “Rejection-Based classification for action recognition using a spatio-temporal dictionary”, Stefan Chan Wai Tim et al., “Rejection-Based classification for action recognition using a spatio-temporal… [cited by applicant]
Donald G Kimber et al: “Speaker segmentation for browsing recorded audio”, Human Factors in Computing Systems / CHI '95, Conference on Human Factors in Computing Systems, May 7-11, 1995. [cited by applicant]
Lie Lu et al: “Speaker change detection and tracking in real-time news broadcasting analysis”, Proceedings ACM Multimedia 2002. 10th. International Conference on Multimedia, Dec. 1, 2002, pp. 602-610. [cited by applicant]
Shamsoddini A et al: “A System for Speech Separation”, 4th European Conference on Speech Communication and Technology. Eurospeech '95. Madrid, Spain, Sep. 18-21, 1995. [cited by applicant]