IP Library Granted Patent US 9,772,993
Granted Patent B2
US 9,772,993 · App. 15/215,114 · Granted Sep 26, 2017

System and method of recording utterances using unmanaged crowds for natural language processing

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,772,993
App. No.
15/215,114
Granted
Sep 26, 2017
Kind
B2
Abstract

A system and method of recording utterances for building Named Entity Recognition (“NER”) models, which are used to build dialog systems in which a computer listens and responds to human voice dialog. Utterances to be uttered may be provided to users through their mobile devices, which may record the user uttering (e.g., verbalizing, speaking, etc.) the utterances and upload the recording to a computer for processing. The use of the user's mobile device, which is programmed with an utterance collection application (e.g., configured as a mobile app), facilitates the use of crowd-sourcing human intelligence tasking for widespread collection of utterances from a population of users. As such, obtaining large datasets for building NER models may be facilitated by the system and method disclosed herein.

Claims (75)

1. A computer-implemented method of recording utterances from unmanaged crowds for natural language processing, the method being implemented in a user device having one or more physical processors programmed with computer program instructions that, when executed by the one or more physical processors, cause the user device to perform the method, the method comprising:

obtaining, by the user device, a token;

transmitting, by the user device, the token to a remote device via a network;

receiving, at the user device, from the remote device, one or more utterances to be uttered by a user and one or more campaign configuration parameters based on the token, wherein the one or more utterances and the one or more campaign configuration parameters are associated with a campaign that is associated with a natural language processing data collection effort;

configuring, by the user device, the computer program instructions based on the one or more campaign configuration parameters;

presenting to the user, by the user device, the one or more utterances to be uttered by the user; and

recording, by the user device, audio of the user uttering the one or more utterances.

2. The method of claim 1 , wherein the token comprises an alphanumeric string.

3. The method of claim 2 , wherein the alphanumeric string includes a campaign identifier that is associated with the one or more campaign configuration parameters.

4. The method of claim 2 , wherein the alphanumeric string includes a session script identifier that is associated with the one or more utterances to be uttered.

5. The method of claim 3 , wherein the one or more campaign configuration parameters include a configuration parameter that indicates that an audit check should be performed, the method further comprising:

configuring, by the user device, the computer program instructions to determine whether a number of failed audits associated with the campaign identifier with which the user device was involved exceeds a predetermined threshold number of failed audits; and

responsive to a determination that the number of failed audits with which the user device was involved exceeds the predetermined threshold number of failed audits, preventing the user device from participating in additional utterance recordings associated with the campaign identifier.

6. The method of claim 1 , wherein the one or more campaign configuration parameters include a calibration parameter that indicates that a calibration test should be performed, the method further comprising:

configuring, by the user device, the computer program instructions to perform a calibration test prior to presenting to the user the one or more utterances, by:

determining, by the user device, a level of ambient noise; and

determining, by the user device, whether the level of ambient noise satisfies the calibration test; and

presenting to the user, by the user device, the one or more utterances to be uttered by the user responsive to a determination that the ambient level of audio satisfies the calibration test.

7. The method of claim 6 , wherein the calibration parameter specifies a minimum level of ambient noise to satisfy the calibration test.

8. The method of claim 6 , wherein the calibration parameter specifies a maximum level of ambient noise to satisfy the calibration test.

9. The method of claim 1 , further comprising:

generating, by the user device, a completion code responsive to a determination that audio has been recorded of the user uttering each of the one or more utterances.

10. The method of claim 1 , further comprising:

causing, by the user device, the recorded audio to be provided to the remote device via the network.

11. The method of claim 1 , wherein obtaining a token comprises:

capturing, by the user device, an image of a machine-readable code, wherein the token is encoded in the machine-readable code.

12. The method of claim 1 , wherein obtaining a token comprises:

receiving, by the user device, a user input comprising an alphanumeric string that comprises the token.

13. The method of claim 1 , wherein transmitting the token to a remote device comprises:

transmitting, by the user device, the token to a transcription system that manages multiple campaigns associated with a natural language processing data collection effort.

14. The method of claim 1 , wherein transmitting the token to a remote device comprises:

transmitting, by the user device, the token to a crowd-sourcing service.

15. The method of claim 1 , further comprising:

causing, by the user device, recorded audio of the one or more utterances to be provided to the remote device via the network after the user has uttered all of the one or more utterances.

16. The method of claim 1 , further comprising:

causing, by the user device, recorded audio of each of the one or more utterances to be provided to the remote device individually via the network after the user has uttered each of the one or more utterances.

17. A system for recording utterances from unmanaged crowds for natural language processing, the system comprising:

a user device having one or more physical processors programmed with computer program instructions that, when executed by the one or more physical processors, cause the user device to:

obtain a token;

transmit the token to a remote device via a network;

receive, from the remote device, one or more utterances to be uttered by a user and one or more campaign configuration parameters based on the token, wherein the one or more utterances and the one or more campaign configuration parameters are associated with a campaign that is associated with a natural language processing data collection effort;

configure the computer program instructions based on the one or more campaign configuration parameters;

present to the user the one or more utterances to be uttered by the user; and

record audio of the user uttering the one or more utterances.

18. The system of claim 17 , wherein the token comprises an alphanumeric string.

19. The system of claim 18 , wherein

the alphanumeric string includes a campaign identifier that is associated with the one or more campaign configuration parameters.

20. The system of claim 18 , wherein

the alphanumeric string includes a session script identifier that is associated with the one or more utterances to be uttered.

21. The system of claim 19 , wherein the one or more campaign configuration parameters include a configuration parameter that indicates that an audit check should be performed, and wherein the user device is further programmed to:

configure the computer program instructions to determine whether a number of failed audits associated with the campaign identifier with which the user device was involved exceeds a predetermined threshold number of failed audits; and

responsive to a determination that the number of failed audits with which the user device was involved exceeds the predetermined threshold number of failed audits, prevent the user device from participating in additional utterance recordings associated with the campaign identifier.

22. The system of claim 17 , wherein the one or more campaign configuration parameters include a calibration parameter that indicates that a calibration test should be performed, and wherein the user device is further programmed to:

configure the computer program instructions to perform a calibration test prior to presenting to the user the one or more utterances, wherein to perform the calibration test, the user device is further programmed to:

determine a level of ambient noise; and

determine whether the level of ambient noise satisfies the calibration test; and

present to the user the one or more utterances to be uttered by the user responsive to a determination that the ambient level of audio satisfies the calibration test.

23. The system of claim 22 , wherein the calibration parameter specifies a minimum level of ambient noise to satisfy the calibration test.

24. The system of claim 22 , wherein the calibration parameter specifies a maximum level of ambient noise to satisfy the calibration test.

25. The system of claim 17 , wherein the user device is further programmed to:

generate a completion code responsive to a determination that audio has been recorded of the user uttering each of the one of one or more utterances.

26. The system of claim 17 , wherein the user device is further programmed to:

cause the recorded audio to be provided to the remote device via the network.

27. The system of claim 17 , wherein, to obtain the token, the user device is further programmed to:

capture an image of a machine-readable code, wherein the token is encoded in the machine-readable code.

28. The system of claim 17 , wherein, to obtain the token, the user device is further programmed to:

receive a user input comprising an alphanumeric string that comprises the token.

29. The system of claim 17 , wherein, to transmit the token, the user device is further programmed to:

transmit the token to a transcription system that manages multiple campaigns associated with a natural language processing data collection effort.

30. The system of claim 17 , wherein, to transmit the token, the user device is further programmed to:

transmit the token to a crowd-sourcing service.

31. The system of claim 17 , wherein the user device is further programmed to:

cause the recorded audio to be provided to the remote device via the network after the user has uttered all of the one or more utterances.

32. The system of claim 17 , wherein the user device is further programmed to:

cause recorded audio of each of the one or more utterances to be provided to the remote device individually via the network after the user has uttered each of the one or more utterances.

Assignments (8)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 24, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050818/0001 →
RELEASE OF SECURITY INTEREST Recorded Apr 5, 2018
From: ORIX GROWTH CAPITAL, LLC
To: VOICEBOX TECHNOLOGIES CORPORATION
Reel/Frame 045581/0630 →
SECURITY INTEREST Recorded Dec 22, 2017
From: VOICEBOX TECHNOLOGIES CORPORATION
To: ORIX GROWTH CAPITAL, LLC
Reel/Frame 044949/0948 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 20, 2016
From: BRAGA, DANIELA; KENNEWICK, MICHAEL; ROMANI, FARAZ; ELSHENAWY, AHMAD KHAMIS; ROTHWELL, SPENCER JOHN; CARTER, STEPHEN STEELE
To: VOICEBOX TECHNOLOGIES CORPORATION
Reel/Frame 039201/0238 →