IP Library › Granted Patent US 12,190,867
Granted Patent B2
US 12,190,867 · App. 17/804,603 · Granted Jan 7, 2025

Keyword detection for audio content

Inventor: Zvi Figov (Modiin, IL)
Assignee: Microsoft Technology Licensing, LLC
G10L15/08G06F40/279G06F40/40G10L15/04G10L15/22G10L25/57G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,190,867
App. No.
17/804,603
Granted
Jan 7, 2025
Kind
B2
Abstract

Examples of the present disclosure describe improved systems and methods for detecting keywords in audio content. In one example implementation, audio content is segmented into one or more audio segments. One or more text segments is generated, each text segment corresponding to each of the audio segments. For each text segment, one or more phrase candidate values is generated using a textual analysis, and one or more sentence embedding values is generated using a sentence embedding analysis. Next, an average sentence embedding value is calculated using the one or more sentence embedding values. Each of the one or more phrase candidate values is compared to the average sentence embedding value. Each phrase candidate value having a comparison value above a threshold value is labeled as representing a keyword.

Claims (42)

1. A method for detecting keywords for audio content, the method comprising:

segmenting the audio content into a plurality of audio segments;

generating a plurality of text segments corresponding to the plurality of audio segments;

generating a plurality of phrase candidate values using a textual analysis of the plurality of text segments;

generating a plurality of sentence embedding values using a sentence embedding analysis of the plurality of text segments;

calculating an average sentence embedding value using the plurality of sentence embedding values;

comparing each phrase candidate value of the plurality of phrase candidate values to the average sentence embedding value;

labeling each phrase candidate value having a comparison value above a threshold value as a keyword; and

presenting the keyword as a stream of the audio content progresses.

2. The method of claim 1 , wherein the detecting keywords for audio content occurs during playback of the audio content.

3. The method of claim 1 , wherein the average sentence embedding value approaches the threshold value as playback of the audio content progresses.

4. The method of claim 1 , wherein each text segment of the plurality of text segments corresponds to an audio segment in the plurality of text segments.

5. The method of claim 1 , wherein video content comprises the audio content.

6. The method of claim 1 , wherein the textual analysis comprises a stopword analysis.

7. The method of claim 1 , wherein the plurality of sentence embedding values comprise a plurality of vectors, respectively.

8. The method of claim 1 , wherein the threshold value is a predetermined value based on relevance.

9. A system for detecting keywords for audio content, the system comprising:

a language conversion module stored in memory and executable to segment audio content into a plurality of audio segments and to generate a plurality of text segments, each text segment of the plurality of text segments corresponding to a particular audio segment of the plurality of audio segments;

a phrase generator module stored in the memory and executable to: generate a plurality of phrase candidate values using a textual analysis of the plurality of text segments, generate a plurality of sentence embedding values using a sentence embedding analysis of the plurality of text segments, and calculate an average sentence embedding value using the plurality of sentence embedding values; and

a candidate extraction module stored in the memory and executable to: compare each of the plurality of phrase candidate values to the average sentence embedding value, and to label each phrase candidate value having a comparison value above a threshold value as representing a keyword, and present the keyword as a stream of the audio content progresses.

10. The system of claim 9 , wherein the average sentence embedding value approaches the threshold value as playback of the audio content progresses.

11. The system of claim 9 , wherein the language conversion module, the phrase generator module, and the phrase candidate extraction module are stored in the memory on a personal computing device.

12. The system of claim 9 , wherein the audio content comprises a livestream.

13. The system of claim 9 , wherein each text segment of the plurality of text segments exists in a one-to-one relationship with an audio segment in the plurality of audio segments.

14. A device comprising:

a processor; and

memory coupled to the processor, the memory comprising computer executable instructions that, when executed by the processor, perform operations comprising:

segmenting audio content into one or more audio segments;

generating one or more text segments, each text segment corresponding to each of the audio segments;

for each text segment:

generating one or more phrase candidate values using a textual analysis; and

generating one or more sentence embedding values using a sentence embedding analysis;

calculating an average sentence embedding value using the one or more sentence embedding values;

comparing each of the one or more phrase candidate values to the average sentence embedding value;

labeling each phrase candidate value having a comparison value above a threshold value as representing a keyword; and

presenting the keyword as a stream of the audio content progresses.

15. The device of claim 14 , wherein the audio content comprises a livestream.

16. The device of claim 14 , wherein the operations are performed during playback of the audio content, and wherein the average sentence embedding value approaches the threshold value as playback of the audio content progresses.

17. The device of claim 14 , wherein the textual analysis comprises a stopword analysis.

18. The device of claim 14 , wherein the one or more sentence embedding values comprise one or more vectors.

19. The device of claim 14 , wherein the threshold value is a predetermined value based on relevance.

20. The device of claim 14 , wherein the processor and the memory are comprised on a personal computing device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2022
From: FIGOV, ZVI
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 060364/0494 →
Continuity (2)
Provisional Application 63363283 · Apr 20, 2022
Related Publication 20230343329A1 · Oct 26, 2023
References Cited (8)
US 20200226154A1 · Muffat · 2020 [cited by examiner]
US 20200401910A1 · Hassanzadeh · 2020 [cited by examiner]
US 20210327413A1 · Suwandy · 2021 [cited by examiner]
Grootendorst, et al., “KeyBERT: Minimal Keyword Extraction with BERT”, Retrieved From: https://github.com/MaartenGr/KeyBERT, Oct. 27, 2020, 7 Pages. [cited by applicant]
Sharma, et al., “Self-Supervised Contextual Keyword and Keyphrase Retrieval with Self-Labelling”, In Publication of Preprints, Aug. 6, 2019, 6 Pages. [cited by applicant]
Papagiannopoulou, et al., “A Review of Keyphrase Extraction”, In repository of arXiv:1905.05044v2, Jul. 30, 2019, 59 Pages. [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US23/013442”, Mailed Date: May 4, 2023, 9 Pages. [cited by applicant]
Smires, et al., “Simple Unsupervised Keyphrase Extraction using Sentence Embeddings”, In repository of arXiv:1801.04470v3, Sep. 5, 2018, 9 Pages. [cited by applicant]