IP Library Granted Patent US 12,525,232
Granted Patent B2
US 12,525,232 · App. 18/298,672 · Granted Jan 13, 2026

System and method for assessing and correcting potential underserved content in natural language understanding applications

Inventors: Aaron Springer (Santa Cruz, CA); Henriette Cramer (San Francisco, CA); Sravana Reddy (Cambridge, MA)
Assignee: Spotify AB
G10L15/187G06F40/295G10L15/1815G10L15/22G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,525,232
App. No.
18/298,672
Granted
Jan 13, 2026
Kind
B2
Abstract

Methods, systems, and related products that provide detection of media content items that are under-locatable by machine voice-driven retrieval of uttered requests for retrieval of the media items. For a given media item, a resolvability value and/or an utterance resolve frequency is calculated by a number of playbacks of the media item by a speech retrieval modality to a total number of playbacks of the media item regardless of retrieval modality. In some examples, the methods, systems and related products also provide for improvement in the locatability of an under-locatable media item by collecting and/or generating one or more pronunciation aliases for the under-locatable item.

Claims (48)

1 . A method comprising:

receiving a playback request containing text associated with a transcribed utterance;

looking up the text in a table that associates respective media content items each with at least one alias;

identifying, in the table, an alias that matches the text, wherein the alias is associated with a media content item, wherein the alias was generated by an alias generation rule, the alias generation rule using at least one of a classification of a second media content item or a second alias of the second media content item; and

playing back the media content item corresponding to the alias.

2 . The method of claim 1 , wherein the media content item is determined to be under-locatable, and further wherein the media content item corresponds to a content item identifier.

3 . The method of claim 2 , wherein the media content item is determined to be under-locatable by:

retrieving, from a database, a name entity associated with the media content item, the name entity having a name entity text; and

determining, based on a classification tag associated with the name entity, that the name entity is under-locatable by a machine voice-driven retrieval of a playback command utterance commanding playback of the media content item, the classification tag being based on the name entity text.

4 . The method of claim 2 , wherein the media content item is determined to be under-locatable by:

ascertaining a first number of playbacks of the media content item;

ascertaining a second number of playbacks of the media content item, where each of the second number of playbacks corresponds to one of the first number of playbacks triggered by a machine voice-driven retrieval of a playback command utterance; and

comparing the first number of playbacks to the second number of playbacks to: 1) generate a resolvability value for the media content item and compare the resolvability value to a predefined threshold resolvability value and, if the resolvability value is less than the predefined threshold resolvability value, determine that the media content item is under-locatable; or 2) generate an utterance resolve frequency for the media content item and compare the utterance resolve frequency to a predefined threshold frequency and, if the utterance resolve frequency is less than the predefined threshold frequency, determine that the media content item is under-locatable.

5 . A system, comprising:

one or more processors adapted to:

receive a playback request containing text associated with a transcribed utterance;

look up the text in a table that associates respective media content items each with at least one alias;

identify, in the table, an alias that matches the text, wherein the alias is associated with a media content item, wherein the alias was generated by an alias generation rule, the alias generation rule using at least one of a classification of a second media content item or a second alias of the second media content item; and

play back the media content item corresponding to the alias.

6 . The system of claim 5 , wherein the media content item is determined to be under-locatable, and further wherein the media content item corresponds to a content item identifier.

7 . The system of claim 6 , wherein the one or more processers are further adapted to:

determine that the media content item is under-locatable, wherein to determine that the media content item is under-locatable includes to:

retrieve, from a database, a name entity associated with the media content item, the name entity having a name entity text; and

determine, based on a classification tag associated with the name entity, that the name entity is under-locatable by a machine voice-driven retrieval of a playback command utterance commanding playback of the media content item, the classification tag being based on the name entity text.

8 . The system of claim 6 , wherein the one or more processers are further adapted to:

determine that the media content item is under-locatable, wherein to determine that the media content item is under-locatable includes to:

ascertain a first number of playbacks of the media content item;

ascertain a second number of playbacks of the media content item, where each of the second number of playbacks corresponds to one of the first number of playbacks triggered by a machine voice-driven retrieval of a playback command utterance; and

compare the first number of playbacks to the second number of playbacks to: 1) generate a resolvability value for the media content item and compare the resolvability value to a predefined threshold resolvability value and, if the resolvability value is less than the predefined threshold resolvability value, determine that the media content item is under-locatable; or 2) generate an utterance resolve frequency for the media content item and compare the utterance resolve frequency to a predefined threshold frequency and, if the utterance resolve frequency is less than the predefined threshold frequency, determine that the media content item is under-locatable.

9 . A non-transitory computer-readable medium, comprising:

one or more sequences of instructions that, when executed by one or more processors, causes the one or more processors to make media content more locatable by:

receiving a playback request containing text associated with a transcribed utterance;

looking up the text in a table that associates respective media content items each with at least one alias, wherein the at least one alias was generated by an alias generation rule, the alias generation rule using at least one of a classification of a second media content item or a second alias of the second media content item;

identifying, in the table, an alias that matches the text wherein the alias is associated with a media content item; and

playing back the media content item corresponding to the alias.

10 . The non-transitory computer-readable medium of claim 9 , wherein the media content item is determined to be under-locatable, and further wherein the media content item corresponds to a content item identifier.

11 . The non-transitory computer-readable medium of claim 10 , wherein the one or more sequences of instructions, when executed by the one or more processors, further causes the one or more processors to make the media content item more locatable by:

determining that the media content item is under-locatable, wherein to determine that the media content item is under-locatable includes:

retrieving, from a database, a name entity associated with the media content item, the name entity having a name entity text; and

determining, based on a classification tag associated with the name entity, that the name entity is under-locatable by a machine voice-driven retrieval of a playback command utterance commanding playback of the media content item, the classification tag being based on the name entity text.

12 . The non-transitory computer-readable medium of claim 10 , wherein the one or more sequences of instructions, when executed by the one or more processors, further causes the one or more processors to make a media content item more locatable by:

determining that the media content item is under-locatable, wherein to determine that the media content item is under-locatable includes:

ascertaining a first number of playbacks of the media content item;

ascertaining a second number of playbacks of the media content item, where each of the second number of playbacks corresponds to one of the first number of playbacks triggered by a machine voice-driven retrieval of a playback command utterance; and

comparing the first number of playbacks to the second number of playbacks to: 1) generate a resolvability value for the media content item and compare the resolvability value to a predefined threshold resolvability value and, if the resolvability value is less than the predefined threshold resolvability value, determine that the media content item is under-locatable; or 2) generate an utterance resolve frequency for the media content item and compare the utterance resolve frequency to a predefined threshold frequency and, if the utterance resolve frequency is less than the predefined threshold frequency, determine that the media content item is under-locatable.

13 . The method of claim 1 , wherein the classification of the second media content item is based on one or more tags applied to a name entity of the second media content item, wherein the one or more tags correspond to a class of name entities that are under-locatable.

14 . The system of claim 5 , wherein the classification of the second media content item is based on one or more tags applied to a name entity of the second media content item, wherein the one or more tags correspond to a class of name entities that are under-locatable.

15 . The non-transitory computer-readable medium of claim 9 , wherein the classification of the second media content item is based on one or more tags applied to a name entity of the second media content item, wherein the one or more tags correspond to a class of name entities that are under-locatable.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 11, 2023
From: SPRINGER, AARON; CRAMER, HENRIETTE; REDDY, SRAVANA
To: SPOTIFY AB
Reel/Frame 063290/0473 →
Continuity (5)
Continuation 17136836 · Dec 29, 2020
Continuation 16124697 · Sep 7, 2018
Provisional Application 62567582 · Oct 3, 2017
Provisional Application 62557265 · Sep 12, 2017
Related Publication 20230274735A1 · Aug 31, 2023
References Cited (47)
US 7028252B1 · Baru · 2006 [cited by examiner]
US 8959020B1 · Strope · 2015 [cited by examiner]
US 9413891B2 · Dwyer · 2016 [cited by examiner]
US 10313520B2 · Dwyer · 2019 [cited by examiner]
US 10582056B2 · Dwyer · 2020 [cited by examiner]
US 10601992B2 · Dwyer · 2020 [cited by examiner]
US 10645224B2 · Dwyer · 2020 [cited by examiner]
US 10902847B2 · Springer · 2021 [cited by examiner]
US 10963497B1 · Tablan · 2021 [cited by examiner]
US 10992807B2 · Dwyer · 2021 [cited by examiner]
US 11277516B2 · Dwyer · 2022 [cited by examiner]
US 11501764B2 · Bromand · 2022 [cited by examiner]
US 11657809B2 · Springer · 2023 [cited by examiner]
US 12137186B2 · Dwyer · 2024 [cited by examiner]
US 12219093B2 · Dwyer · 2025 [cited by examiner]
US 20030182111A1 · Handal · 2003 [cited by examiner]
US 20040172258A1 · Dominach · 2004 [cited by examiner]
US 20060206339A1 · Silvera · 2006 [cited by applicant]
US 20110276335A1 · Silvera · 2011 [cited by examiner]
US 20120227115A1 · Kidron · 2012 [cited by examiner]
US 20130179170A1 · Cath · 2013 [cited by examiner]
US 20130275899A1 · Schubert · 2013 [cited by examiner]
US 20140207468A1 · Bartnik · 2014 [cited by applicant]
US 20140222415A1 · Legat · 2014 [cited by examiner]
US 20150185964A1 · Stout · 2015 [cited by applicant]
US 20150195406A1 · Dwyer · 2015 [cited by examiner]
US 20170011232A1 · Xue · 2017 [cited by examiner]
US 20170011233A1 · Xue · 2017 [cited by examiner]
US 20170013127A1 · Xue · 2017 [cited by examiner]
US 20170026514A1 · Dwyer · 2017 [cited by examiner]
US 20170242653A1 · Lang · 2017 [cited by applicant]
US 20170243576A1 · Millington · 2017 [cited by applicant]
US 20180165554A1 · Zhang · 2018 [cited by examiner]
US 20190080686A1 · Springer · 2019 [cited by examiner]
US 20190245971A1 · Dwyer · 2019 [cited by examiner]
US 20190245972A1 · Dwyer · 2019 [cited by examiner]
US 20190245973A1 · Dwyer · 2019 [cited by examiner]
US 20190245974A1 · Dwyer · 2019 [cited by examiner]
US 20200357390A1 · Bromand · 2020 [cited by examiner]
US 20210120126A1 · Dwyer · 2021 [cited by examiner]
US 20210193126A1 · Springer · 2021 [cited by examiner]
US 20220272196A1 · Dwyer · 2022 [cited by examiner]
US 20220377174A1 · Dwyer · 2022 [cited by examiner]
US 20220377175A1 · Dwyer · 2022 [cited by examiner]
US 20230179709A1 · Dwyer · 2023 [cited by examiner]
US 20230274735A1 · Springer · 2023 [cited by examiner]
Project Common Voice, Mozilla, Sep. 11, 2017, available online at: https://voice/mozilla.org/. [cited by applicant]