IP Library Granted Patent US 9,448,993
Granted Patent B1
US 9,448,993 · App. 14/846,926 · Granted Sep 20, 2016

System and method of recording utterances using unmanaged crowds for natural language processing

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,448,993
App. No.
14/846,926
Granted
Sep 20, 2016
Kind
B1
Abstract

A system and method of recording utterances for building Named Entity Recognition (“NER”) models, which are used to build dialog systems in which a computer listens and responds to human voice dialog. Utterances to be uttered may be provided to users through their mobile devices, which may record the user uttering (e.g., verbalizing, speaking, etc.) the utterances and upload the recording to a computer for processing. The use of the user's mobile device, which is programmed with an utterance collection application (e.g., configured as a mobile app), facilitates the use of crowd-sourcing human intelligence tasking for widespread collection of utterances from a population of users. As such, obtaining large datasets for building NER models may be facilitated by the system and method disclosed herein.

Claims (51)

1. A computer implemented method of recording utterances from unmanaged crowds for natural language processing, the method being implemented in an end user device having one or more physical processors programmed with computer program instructions that, when executed by the one or more physical processors, cause the end user device to perform the method, the method comprising:

obtaining, by the end user device, a token;

obtaining, by the end user device, one or more campaign configuration parameters based on the token;

configuring, by the end user device, the computer program instructions based on the one or more campaign configuration parameters;

obtaining, by the end user device, one or more utterances to be uttered by a user based on the token;

displaying, by the end user device, the one or more utterances to be uttered by the user;

generating, by the end user device, an audio recording of the one or more utterances; and

causing, by the end user device, the audio recording to be provided to a remote device via a network.

2. The method of claim 1 , wherein the token comprises an alphanumeric string.

3. The method of claim 2 , wherein obtaining the one or more campaign configuration parameters comprises parsing the alphanumeric string to obtain a campaign identifier that is associated with the one or more campaign configuration parameters.

4. The method of claim 3 , wherein obtaining the one or more utterances to be uttered comprises parsing the alphanumeric string to obtain a session script identifier that is associated with the one or more utterances to be uttered.

5. The method of claim 3 , wherein the one or more campaign configuration parameters comprises a configuration parameter that indicates that an audit check should be performed, the method further comprising:

configuring, by the end user device, the computer program instructions to determine whether a number of failed audits associated with the campaign identifier with which the end user device was involved exceeds a predetermined threshold number of failed audits; and

responsive to a determination that the number of failed audits with which the end user device was involved exceeds the predetermined threshold number of failed audits, preventing the end user device from participating in additional utterance recordings associated with the campaign identifier.

6. The method of claim 1 , wherein the one or more campaign configuration parameters comprises a calibration parameter that indicates that a calibration test should be performed, the method further comprising:

configuring, by the end user device, the computer program instructions to perform a calibration test prior to generating the audio recording, by:

determining, by the end user device, a level of ambient noise;

determining, by the end user device, whether the level of ambient noise satisfies the calibration test, wherein the one or more utterances to be uttered are displayed responsive to a determination that the ambient level of audio satisfies the calibration test.

7. The method of claim 6 , wherein the calibration parameter specifies a minimum level of ambient noise to satisfy the calibration test.

8. The method of claim 6 , wherein the calibration parameter specifies a maximum level of ambient noise to satisfy the calibration test.

9. The method of claim 1 , the method further comprising:

generating, by the end user device, a completion code responsive to a determination that a corresponding audio recording has been generated for each one of one or more utterances, wherein the completion code is configured to provide information that validates the one or more utterances have been uttered and recorded.

10. The method of claim 1 , wherein causing the audio recording to be provided to a remote device via a network comprises:

immediately uploading, by the end user device, the audio recording to the remote device instead of through a batch process.

11. A system of recording utterances from unmanaged crowds for natural language processing, the system comprising:

an end user device having one or more physical processors programmed with computer program instructions that, when executed by the one or more physical processors, cause the end user device to:

obtain a token;

obtain one or more campaign configuration parameters based on the token;

configure the computer program instructions based on the one or more campaign configuration parameters;

obtain one or more utterances to be uttered by a user based on the token;

display the one or more utterances to be uttered by the user;

generate an audio recording of the one or more utterances; and

cause the audio recording to be provided to a remote device via a network.

12. The system of claim 11 , wherein the token comprises an alphanumeric string.

13. The system of claim 12 , wherein to obtain the one or more campaign configuration parameters, the end user device is further programmed to:

parse the alphanumeric string to obtain a campaign identifier that is associated with the one or more campaign configuration parameters.

14. The system of claim 13 , wherein to obtain the one or more utterances to be uttered, the end user device is further programmed to:

parse the alphanumeric string to obtain a session script identifier that is associated with the one or more utterances to be uttered.

15. The system of claim 13 , wherein the one or more campaign configuration parameters comprises a configuration parameter that indicates that an audit check should be performed, and wherein the end user device is further programmed to:

configure the computer program instructions to determine whether a number of failed audits associated with the campaign identifier with which the end user device was involved exceeds a predetermined threshold number of failed audits; and

responsive to a determination that the number of failed audits with which the end user device was involved exceeds the predetermined threshold number of failed audits, prevent the end user device from participating in additional utterance recordings associated with the campaign identifier.

16. The system of claim 11 , wherein the one or more campaign configuration parameters comprises a calibration parameter that indicates that a calibration test should be performed, and wherein the end user device is further programmed to:

configure the computer program instructions to perform a calibration test prior to generating the audio recording, wherein to perform the calibration test, the end user device is further programmed to:

determine a level of ambient noise;

determine whether the level of ambient noise satisfies the calibration test, wherein the one or more utterances to be uttered are displayed responsive to a determination that the ambient level of audio satisfies the calibration test.

17. The system of claim 16 , wherein the calibration parameter specifies a minimum level of ambient noise to satisfy the calibration test.

18. The system of claim 6 , wherein the calibration parameter specifies a maximum level of ambient noise to satisfy the calibration test.

19. The system of claim 11 , wherein the end user device is further programmed to:

generate a completion code responsive to a determination that a corresponding audio recording has been generated for each one of one or more utterances, wherein the completion code is configured to provide information that validates the one or more utterances have been uttered and recorded.

20. The system of claim 11 , wherein to cause the audio recording to be provided to a remote device via a network, the end user device is further programmed to:

immediately upload the audio recording to the remote device instead of through a batch process.

Assignments (8)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 24, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050818/0001 →
RELEASE OF SECURITY INTEREST Recorded Apr 5, 2018
From: ORIX GROWTH CAPITAL, LLC
To: VOICEBOX TECHNOLOGIES CORPORATION
Reel/Frame 045581/0630 →
SECURITY INTEREST Recorded Dec 22, 2017
From: VOICEBOX TECHNOLOGIES CORPORATION
To: ORIX GROWTH CAPITAL, LLC
Reel/Frame 044949/0948 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 29, 2015
From: BRAGA, DANIELA; KENNEWICK, MICHAEL; ROMANI, FARAZ; ELSHENAWY, AHMAD KHAMIS; ROTHWELL, SPENCER JOHN; CARTER, STEPHEN STEELE
To: VOICEBOX TECHNOLOGIES CORPORATION
Reel/Frame 036910/0944 →