IP Library Granted Patent US 11,043,215
Granted Patent B2
US 11,043,215 · App. 16/587,078 · Granted Jun 22, 2021

Method and system for generating textual representation of user spoken utterance

Inventors: Sergey Surenovich Galustyan (p. Sovkhoznyy, RU); Fedor Aleksandrovich Minkin (Moskovskaya obl., RU)
Assignee: YANDEX EUROPE AG
G10L15/197G10L15/02G10L15/063G10L15/22G10L25/63G10L25/84G10L25/90
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,043,215
App. No.
16/587,078
Granted
Jun 22, 2021
Kind
B2
Abstract

A method and a system for generating textual representation of user spoken utterance is disclosed. The method comprises receiving an indication of the user spoken utterance; generating, at least two hypotheses; generating, by the electronic device, from the at least two hypotheses a set of paired hypotheses, a given one of the set of paired hypotheses including a first hypothesis paired with a second hypothesis; determining, for the given one of the set of paired hypotheses, a pair score; generating a set of utterance features, the set of utterance features being indicative of one or more characteristics associated with the user spoken utterance; ranking, the first hypothesis and the second hypothesis based at least on the pair score and the set of utterance features; and in response to the first hypothesis being a highest ranked hypothesis, selecting the first hypothesis as the textual representation of user spoken utterance.

Claims (103)

1. A computer-implemented method for generating a textual representation of a user spoken utterance, the method being executable by an electronic device, the method comprising:

receiving, by the electronic device from a user, an indication of the user spoken utterance, the user spoken utterance being expressed in a natural language;

generating, by the electronic device, at least two hypotheses based on the user spoken utterance, each of the at least two hypotheses corresponding to a possible textual representation of the user spoken utterance;

generating, by the electronic device, from the at least two hypotheses a set of paired hypotheses, a given one of the set of paired hypotheses including a first hypothesis paired with a second hypothesis;

generating a first hypothesis profile, the first hypothesis profile corresponding to a first set of vector values representing one or more context-specific characteristics of the first hypothesis;

generating a second hypothesis profile, the second hypothesis profile corresponding to a second set of vector values representing one or more context-specific characteristics of the second hypothesis;

determining, using a pairwise classifier executable by the electronic device, for the given one of the set of paired hypotheses, a pair score, the pair score being indicative of a respective probability of the first hypothesis and the second hypothesis corresponding to a correct representation of the user spoken utterance, wherein the generating the pair score comprises:

generating the pair score based at least on the first hypothesis profile and the second hypothesis profile; and wherein

the pairwise classifier is a Price, Kner, Personnaz and Dreyfus (PKPD) algorithm previously trained using a training set of data prior to the receiving the user spoken utterance, the training set of data comprising at least:

a training pair of hypothesis including a first training hypothesis paired with a second training hypothesis, each of the first training hypothesis and the second training hypothesis having been generated in response to a training utterance;

a first training hypothesis profile, the first training hypothesis profile corresponding to a first training set of vector values representing one or more context-specific features of the first training hypothesis;

a second training hypothesis profile, the second training hypothesis profile corresponding to a second training set of vector values representing one or more context-specific features of the second training hypothesis;

a difference-in-profile score, the difference-in-profile score corresponding to a difference between the first training set of vector values and the second training set of vector values;

an aggregated profile score, the aggregated profile score corresponding to a quotient of the first training set of vector values and the second training set of vector values; and

a label indicative of one of the first training hypothesis and the second training hypothesis corresponding to the training utterance;

generating a set of utterance features, the set of utterance features being indicative of one or more characteristics associated with the user spoken utterance;

ranking, by a ranking algorithm executable by the electronic device, the first hypothesis and the second hypothesis based at least on the pair score and the set of utterance features; and

in response to the first hypothesis being a highest ranked hypothesis, selecting the first hypothesis as the textual representation of the user spoken utterance.

2. The method of claim 1 , wherein

the at least two hypotheses include the first hypothesis, the second hypothesis and at least one additional hypothesis;

the set of paired hypotheses includes a plurality of paired hypotheses comprising each of the at least two hypotheses paired with each of a remaining one of the at least two hypotheses; and wherein

determining the pair score comprises:

determining, a set of paired scores comprising one or more paired scores associated with each of the plurality of paired hypothesis.

3. The method of claim 2 , wherein

the ranking by the ranking algorithm comprises ranking the at least two hypotheses based at least on an entirety of the set of paired scores and the set of utterance features.

4. The method of claim 1 , wherein determining the pair score includes:

determining a first score, the first score being indicative of a relative probability of the first hypothesis corresponding to the correct representation of the user spoken utterance than the second hypothesis;

determining a second score, the second score being indicative of the relative probability of the second hypothesis corresponding to the correct representation of the user spoken utterance than the first hypothesis;

determining a first normalized score and a second normalized score;

the first normalized score being based on the first score;

the second normalized score being based on the second score; and

such that the first normalized score and the second normalized score add up together to a pre-determined total score.

5. The method of claim 1 , wherein generating the first hypothesis profile comprises:

analyzing, the first hypothesis using one or more context-specific models, each of the one or more context-specific models being trained on a plurality of context-specific words;

assigning, by each of the one or more context-specific models, a respective context-specific score corresponding to a vector value, a given vector value representing a proportion of context-specific words associated with a given context-specific model within the first hypothesis; and

aggregating the one or more vector values assigned by each of the one or more context-specific models.

6. The method of claim 1 , wherein the set of utterance features includes user-specific features, the user-specific features comprising at least one of:

an age of the user;

a gender of the user; and

an interest profile of the user.

7. The method of claim 6 , wherein the electronic device is further communicatively coupled to a database comprising one of a browsing log and a search log associated with the user, and wherein the user-specific features are generated based on at least one of:

a browsing history associated with the user; and

a search history associated with the user.

8. The method of claim 6 , wherein the electronic device is a user device, and wherein the user-specific features are generated based on previous interactions of the user with the user device.

9. The method of claim 1 , wherein the set of utterance features includes acoustic features, the acoustic features comprising at least one of:

a tone of the user spoken utterance;

a pitch of the user spoken utterance; and

a noise-to-signal ratio.

10. The method of claim 1 , wherein

the electronic device is a server; wherein

the server is coupled to a user device associated with the user; and wherein

receiving the user spoken utterance from the user comprises receiving the user spoken utterance from the user device.

11. The method of claim 1 , wherein

the ranking algorithm is a neural network; and

the method further comprises training the neural network using a training set of data prior to receiving the natural language input.

12. A computer-implemented method for generating a textual representation of a user spoken utterance, the method being executable by an electronic device, the method comprising:

receiving, by the electronic device from a user, an indication of the user spoken utterance, the user spoken utterance being expressed in a natural language;

generating, by the electronic device, at least two hypotheses based on the user spoken utterance, each of the at least two hypotheses corresponding to a possible textual representation of the user spoken utterance;

generating, by the electronic device, a set of paired hypotheses, the set of paired hypotheses comprising (i) each given one of the at least two hypotheses paired with (ii) each of a remaining one of the at least two hypotheses;

determining, using a pairwise classifier executable by the electronic device, a set of pair scores, the set of pair scores including a pair score for each of the paired hypothesis within the set of paired hypotheses, a given pair score for a given paired hypothesis being indicative of a respective probability of a first hypothesis and a second hypothesis of the given paired hypothesis corresponding to a correct representation of the user spoken utterance, wherein the pairwise classifier is a Price, Kner, Personnaz and Dreyfus (PKPD) algorithm;

ranking the at least two hypotheses by a ranking algorithm executable by the electronic device, the ranking algorithm being configured to rank the at least two hypotheses based on an entirety of the set of pair scores;

in response to the first hypothesis being a highest ranked hypothesis, selecting the first hypothesis as the textual representation of the user spoken utterance.

13. A system for generating a textual representation of a user spoken utterance, the system comprising an electronic device, the electronic device comprising a processor configured to:

receive, by the electronic device from a user, an indication of the user spoken utterance, the user spoken utterance being expressed in a natural language;

generate, by the electronic device, at least two hypotheses based on the user spoken utterance, each of the at least two hypotheses corresponding to a possible textual representation of the user spoken utterance;

generate, by the electronic device, from the at least two hypotheses a set of paired hypotheses, a given one of the set of paired hypotheses including a first hypothesis paired with a second hypothesis;

generate a first hypothesis profile, the first hypothesis profile corresponding to a first set of vector values representing one or more context-specific characteristics of the first hypothesis;

generate a second hypothesis profile, the second hypothesis profile corresponding to a second set of vector values representing one or more context-specific characteristics of the second hypothesis;

determine, using a pairwise classifier executable by the electronic device, a set of pair scores, the set of pair scores including comprising a pair score associated with the given one of the set of paired hypotheses, the pair score being indicative of a respective probability of the first hypothesis and the second hypothesis corresponding to a correct representation of the user spoken utterance, wherein to generate the pair score the processor is configured to:

generate the pair score based at least on the first hypothesis profile and the second hypothesis profile; and wherein

the pairwise classifier is a Price, Kner, Personnaz and Dreyfus (PKPD) algorithm previously trained using a training set of data prior to the receiving the user spoken utterance, the training set of data comprising at least:

a training pair of hypothesis including a first training hypothesis paired with a second training hypothesis, each of the first training hypothesis and the second training hypothesis having been generated in response to a training utterance;

a first training hypothesis profile, the first training hypothesis profile corresponding to a first training set of vector values representing one or more context-specific features of the first training hypothesis;

a second training hypothesis profile, the second training hypothesis profile corresponding to a second training set of vector values representing one or more context-specific features of the second training hypothesis;

a difference-in-profile score, the difference-in-profile score corresponding to a difference between the first training set of vector values and the second training set of vector values;

an aggregated profile score, the aggregated profile score corresponding to a quotient of the first training set of vector values and the second training set of vector values; and

a label indicative of one of the first training hypothesis and the second training hypothesis corresponding to the training utterance;

generate a set of utterance features, the set of utterance features being indicative of one or more characteristics associated with the user spoken utterance;

rank, by a ranking algorithm executable by the electronic device, the first hypothesis and the second hypothesis based at least on the entirety of the set of pair scores and the set of utterance features; and

in response to the first hypothesis being a highest ranked hypothesis, select the first hypothesis as the textual representation of the user spoken utterance.

14. The system of claim 13 , wherein

the at least two hypotheses include the first hypothesis, the second hypothesis and at least one additional hypothesis;

the set of paired hypotheses includes a plurality of paired hypotheses comprising each of the at least two hypotheses paired with each of a remaining one of the at least two hypotheses; and wherein

to determine the pair score, the processor is configured to:

determine, a set of paired scores comprising one or more paired scores associated with each of the plurality of paired hypothesis.

15. The system of claim 13 , wherein to determine the pair score, the processor is configured to:

determine a first score, the first score being indicative of a relative probability of the first hypothesis corresponding to the correct representation of the user spoken utterance than the second hypothesis;

determine a second score, the second score being indicative of the relative probability of the second hypothesis corresponding to the correct representation of the user spoken utterance than the first hypothesis;

determine a first normalized score and a second normalized score;

the first normalized score being based on the first score;

the second normalized score being based on the second score; and

such that the first normalized score and the second normalized score add up together to a pre-determined total score.

16. The system of claim 13 , wherein the set of utterance features includes user-specific features, the user-specific features comprising at least one of:

an age of the user;

a gender of the user; and

an interest profile of the user.

17. The system of claim 13 , wherein the set of utterance features includes acoustic features, the acoustic features comprising at least one of:

a tone of the user spoken utterance;

a pitch of the user spoken utterance; and

a noise-to-signal ratio.

18. The system of claim 13 , wherein

the ranking algorithm is a neural network; and

the processor is further configured to train the neural network using a training set of data prior to receiving the natural language input.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 9, 2024
From: DIRECT CURSUS TECHNOLOGY L.L.C
To: Y.E. HUB ARMENIA LLC
Reel/Frame 068534/0537 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 15, 2023
From: YANDEX EUROPE AG
To: DIRECT CURSUS TECHNOLOGY L.L.C
Reel/Frame 065692/0720 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 8, 2020
From: GALUSTYAN, SERGEY SURENOVICH; MINKIN, FEDOR ALEKSANDROVICH
To: YANDEX.TECHNOLOGIES LLC
Reel/Frame 051448/0375 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 8, 2020
From: YANDEX.TECHNOLOGIES LLC
To: YANDEX LLC
Reel/Frame 051448/0447 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 8, 2020
From: YANDEX LLC
To: YANDEX EUROPE AG
Reel/Frame 051448/0520 →