IP Library Granted Patent US 10,902,847
Granted Patent B2
US 10,902,847 · App. 16/124,697 · Granted Jan 26, 2021

System and method for assessing and correcting potential underserved content in natural language understanding applications

Inventors: Aaron Springer (Santa Cruz, CA); Henriette Cramer (San Francisco, CA); Sravana Reddy (Cambridge, MA)
Assignee: Spotify AB
G10L15/187G06F40/295G10L15/1815G10L15/22G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,902,847
App. No.
16/124,697
Granted
Jan 26, 2021
Kind
B2
Abstract

Methods, systems, and related products that provide detection of media content items that are under-locatable by machine voice-driven retrieval of uttered requests for retrieval of the media items. For a given media item, a resolvability value and/or an utterance resolve frequency is calculated by a number of playbacks of the media item by a speech retrieval modality to a total number of playbacks of the media item regardless of retrieval modality. In some examples, the methods, systems and related products also provide for improvement in the locatability of an under-locatable media item by collecting and/or generating one or more pronunciation aliases for the under-locatable item.

Claims (126)

1. A method for detecting an under-locatable media content item comprising:

ascertaining a first number of playbacks of a media content item;

ascertaining a second number of playbacks of the media content item, where each of the second number of playbacks corresponds to one of the first number of playbacks triggered by a machine voice-driven retrieval of a playback command utterance; and

comparing the first number of playbacks to the second number of playbacks to: 1) generate a resolvability value for the media content item and compare the resolvability value to a predefined threshold resolvability value and, if the resolvability value is less than the predefined threshold resolvability value, determine that the media content item is under-locatable; or 2) generate an utterance resolve frequency for the media content item and compare the utterance resolve frequency to a predefined threshold frequency and, if the utterance resolve frequency is less than the predefined threshold frequency, determine that the media content item is under-locatable.

2. The method of claim 1 , wherein each of the first number of playbacks of the media content item corresponds to a playback command provided by any of a plurality of command modalities, wherein one of the command modalities is natural speech and another of the command modalities is text.

3. The method of claim 1 ,

wherein the media content item is associated with one or more name entities, and the method further comprises:

classifying the one or more name entities with one or more tags, each of the one or more tags corresponding to a class of name entities that are under-locatable by machine voice-driven retrieval.

4. The method of claim 3 , further comprising:

generating at least one pronunciation alias for the one or more name entities;

wherein the generating is based at least partially on the one or more tags.

5. The method of claim 4 , further comprising:

locating the media content item using a machine voice-driven retrieval of an utterance containing the at least one pronunciation alias; and

playing back the located media content item.

6. The method of claim 1 , further comprising:

collecting at least one pronunciation alias for one or more name entities; and

associating a transcription of at least one pronunciation alias with the media content item.

7. The method of claim 6 , further comprising:

locating the media content item using a machine voice-driven retrieval of an utterance containing the at least one pronunciation alias; and

playing back the located media content item.

8. A system for detecting an under-locatable media content item, comprising:

one or more processors adapted to:

ascertain a first number of playbacks of a media content item;

ascertain a second number of playbacks of the media content item, where each of the second number of playbacks corresponds to one of the first number of playbacks triggered by a machine voice-driven retrieval of a playback command utterance; and

compare the first number of playbacks to the second number of playbacks to: 1) generate a resolvability value for the media content item and compare the resolvability value to a predefined threshold resolvability value and, if the resolvability value is less than the predefined threshold resolvability value, determine that the media content item is under-locatable; or 2) generate an utterance resolve frequency for the media content item and compare the utterance resolve frequency to a predefined threshold frequency and, if the utterance resolve frequency is less than the predefined threshold frequency, determine that the media content item is under-locatable.

9. The system of claim 8 , wherein each of the first number of playbacks of the media content item corresponds to a playback command provided by any of a plurality of command modalities, wherein one of the command modalities is natural speech and another of the command modalities is text.

10. The system of claim 8 ,

wherein the media content item is associated with one or more name entities, and the one or more processors are further adapted to:

classify the one or more name entities with one or more tags, each of the one or more tags corresponding to a class of name entities that are under-locatable by machine voice-driven retrieval.

11. The system of claim 10 , wherein the one or more processors are further adapted to:

generate at least one pronunciation alias for the one or more name entities;

wherein the generate is based at least partially on the one or more tags.

12. The system of claim 10 , wherein the one or more processors are further adapted to:

locate the media content item using a machine voice-driven retrieval of an utterance containing at least one pronunciation alias; and

play back the located media content item.

13. The system of claim 8 , wherein the one or more processors are further adapted to:

collect at least one pronunciation alias for the one or more name entities; and

associate a transcription of at least one pronunciation alias with the media content item.

14. The system of claim 13 , wherein the one or more processors are further adapted to:

locate the media content item using a machine voice-driven retrieval of an utterance containing the at least one pronunciation alias; and

play back the located media content item.

15. A non-transitory computer-readable medium, comprising:

one or more sequences of instructions that, when executed by one or more processors, causes the one or more processors to detect an under-locatable media content item by:

ascertaining a first number of playbacks of a media content item;

ascertaining a second number of playbacks of the media content item, where each of the second number of playbacks corresponds to one of the first number of playbacks triggered by a machine voice-driven retrieval of a playback command utterance; and

comparing the first number of playbacks to the second number of playbacks to: 1) generate a resolvability value for the media content item and compare the resolvability value to a predefined threshold resolvability value and, if the resolvability value is less than the predefined threshold resolvability value, determine that the media content item is under-locatable; or 2) generate an utterance resolve frequency for the media content item and compare the utterance resolve frequency to a predefined threshold frequency and, if the utterance resolve frequency is less than the predefined threshold frequency, determine that the media content item is under-locatable.

16. The non-transitory computer-readable medium of claim 15 , wherein each of the first number of playbacks of the media content item corresponds to a playback command provided by any of a plurality of command modalities, wherein one of the command modalities is natural speech and another of the command modalities is text.

17. The non-transitory computer-readable medium of claim 15 , wherein the media content item is associated with one or more name entities, and wherein the one or more sequences of instructions, when executed by the one or more processors, causes the one or more processors to:

classify the one or more name entities with one or more tags, each of the one or more tags corresponding to a class of name entities that are under-locatable by machine voice-driven retrieval.

18. The non-transitory computer-readable medium of claim 17 , wherein the one or more sequences of instructions, when executed by the one or more processors, causes the one or more processors to:

generate at least one pronunciation alias for the one or more name entities;

wherein the generate is based at least partially on the one or more tags.

19. The non-transitory computer-readable medium of claim 18 , wherein the one or more sequences of instructions, when executed by the one or more processors, causes the one or more processors to:

locate the media content item using a machine voice-driven retrieval of an utterance containing the at least one pronunciation alias; and

play back the located media content item.

20. The non-transitory computer-readable medium of claim 15 , wherein the one or more sequences of instructions, when executed by the one or more processors, causes the one or more processors to:

collect at least one pronunciation alias for one or more name entities; and

associate a transcription of the at least one pronunciation alias with the media content item.

21. The non-transitory computer-readable medium of claim 20 , wherein the one or more sequences of instructions, when executed by the one or more processors, causes the one or more processors to:

locate the media content item using a machine voice-driven retrieval of an utterance containing the at least one pronunciation alias; and

play back the located media content item.

22. A method for detecting an under-locatable media content item comprising:

ascertaining a first number of playbacks of a media content item, wherein the media content item is associated with one or more name entities;

ascertaining a second number of playbacks of the media content item, where each of the second number of playbacks corresponds to one of the first number of playbacks triggered by a machine voice-driven retrieval of a playback command utterance;

comparing the first number of playbacks to the second number of playbacks to generate: 1) a resolvability value for the media content item; or 2) an utterance resolve frequency for the media content item; and

classifying the one or more name entities with one or more tags, each of the one or more tags corresponding to a class of name entities that are under-locatable by machine voice-driven retrieval.

23. The method of claim 22 , wherein each of the first number of playbacks of the media content item corresponds to a playback command provided by any of a plurality of command modalities, wherein one of the command modalities is natural speech and another of the command modalities is text.

24. The method of claim 22 , further comprising:

1) comparing the utterance resolve frequency to a predefined threshold frequency and, if the utterance resolve frequency is less than the predefined threshold frequency, determining that the media content item is under-locatable; or

2) comparing the resolvability value to a predefined threshold resolvability value and, if the resolvability value is less than the predefined threshold resolvability value, determine that the media content item is under-locatable.

25. The method of claim 22 , further comprising:

generating at least one pronunciation alias for the one or more name entities;

wherein the generating is based at least partially on the one or more tags.

26. The method of claim 25 , further comprising:

locating the media content item using a machine voice-driven retrieval of an utterance containing the at least one pronunciation alias; and

playing back the located media content item.

27. The method of claim 22 , further comprising:

collecting at least one pronunciation alias for the one or more name entities; and

associating a transcription of the at least one pronunciation alias with the media content item.

28. The method of claim 27 , further comprising:

locating the media content item using a machine voice-driven retrieval of an utterance containing the at least one pronunciation alias; and

playing back the located media content item.

29. A system for detecting an under-locatable media content item, comprising:

one or more processors adapted to:

ascertain a first number of playbacks of a media content item, wherein the media content item is associated with one or more name entities;

ascertain a second number of playbacks of the media content item, where each of the second number of playbacks corresponds to one of the first number of playbacks triggered by a machine voice-driven retrieval of a playback command utterance;

compare the first number of playbacks to the second number of playbacks to generate: 1) a resolvability value for the media content item; or 2) an utterance resolve frequency for the media content item; and

classify the one or more name entities with one or more tags, each of the one or more tags corresponding to a class of name entities that are under-locatable by machine voice-driven retrieval.

30. The system of claim 29 , wherein each of the first number of playbacks of the media content item corresponds to a playback command provided by any of a plurality of command modalities, wherein one of the command modalities is natural speech and another of the command modalities is text.

31. The system of claim 29 , wherein the one or more processors are further adapted to:

1) compare the utterance resolve frequency to a predefined threshold frequency and, if the utterance resolve frequency is less than the predefined threshold frequency, determine that the media content item is under-locatable; or

2) compare the resolvability value to a predefined threshold resolvability value and, if the resolvability value is less than the predefined threshold resolvability value, determine that the media content item is under-locatable.

32. The system of claim 29 , wherein the one or more processors are further adapted to:

generate at least one pronunciation alias for the one or more name entities;

wherein the generate is based at least partially on the one or more tags.

33. The system of claim 29 , wherein the one or more processors are further adapted to:

locate the media content item using a machine voice-driven retrieval of an utterance containing at least one pronunciation alias; and

play back the located media content item.

34. The system of claim 29 , wherein the one or more processors are further adapted to:

collect at least one pronunciation alias for the one or more name entities; and

associate a transcription of the at least one pronunciation alias with the media content item.

35. The system of claim 34 , wherein the one or more processors are further adapted to:

locate the media content item using a machine voice-driven retrieval of an utterance containing the at least one pronunciation alias; and

play back the located media content item.

36. A non-transitory computer-readable medium, comprising:

one or more sequences of instructions that, when executed by one or more processors, causes the one or more processors to detect an under-locatable media content item by:

ascertaining a first number of playbacks of a media content item, wherein the media content item is associated with one or more name entities;

ascertaining a second number of playbacks of the media content item, where each of the second number of playbacks corresponds to one of the first number of playbacks triggered by a machine voice-driven retrieval of a playback command utterance;

comparing the first number of playbacks to the second number of playbacks to generate: 1) a resolvability value for the media content item; or 2) an utterance resolve frequency for the media content item; and

classifying the one or more name entities with one or more tags, each of the one or more tags corresponding to a class of name entities that are under-locatable by machine voice-driven retrieval.

37. The non-transitory computer-readable medium of claim 36 , wherein each of the first number of playbacks of the media content item corresponds to a playback command provided by any of a plurality of command modalities, wherein one of the command modalities is natural speech and another of the command modalities is text.

38. The non-transitory computer-readable medium of claim 36 , wherein the one or more sequences of instructions, when executed by the one or more processors, causes the one or more processors to:

1) compare the utterance resolve frequency to a predefined threshold frequency and, if the utterance resolve frequency is less than the predefined threshold frequency, determine that the media content item is under-locatable; or

2) compare the resolvability value to a predefined threshold resolvability value and, if the resolvability value is less than the predefined threshold resolvability value, determine that the media content item is under-locatable.

39. The non-transitory computer-readable medium of claim 36 , wherein the one or more sequences of instructions, when executed by the one or more processors, causes the one or more processors to:

generate at least one pronunciation alias for the one or more name entities;

wherein the generate is based at least partially on the one or more tags.

40. The non-transitory computer-readable medium of claim 36 , wherein the one or more sequences of instructions, when executed by the one or more processors, causes the one or more processors to:

locate the media content item using a machine voice-driven retrieval of an utterance containing at least one pronunciation alias; and

play back the located media content item.

41. The non-transitory computer-readable medium of claim 36 , wherein the one or more sequences of instructions, when executed by the one or more processors, causes the one or more processors to:

collect at least one pronunciation alias for the one or more name entities; and

associate a transcription of the at least one pronunciation alias with the media content item.

42. The non-transitory computer-readable medium of claim 41 , wherein the one or more sequences of instructions, when executed by the one or more processors, causes the one or more processors to:

locate the media content item using a machine voice-driven retrieval of an utterance containing the at least one pronunciation alias; and

play back the located media content item.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 6, 2019
From: SPRINGER, AARON; CRAMER, HENRIETTE; REDDY, SRAVANA
To: SPOTIFY AB
Reel/Frame 050296/0642 →
Continuity (3)
Provisional Application 62557265 · Sep 12, 2017
Provisional Application 62567582 · Oct 3, 2017
Related Publication 20190080686A1 · Mar 14, 2019
Cited By (1)
US 12,525,232