IP Library Granted Patent US 11,514,916
Granted Patent B2
US 11,514,916 · App. 16/992,943 · Granted Nov 29, 2022

Server that supports speech recognition of device, and operation method of the server

Inventors: Chanwoo Kim (Suwon-si, KR); Sichen Jin (Suwon-si, KR); Kyungmin Lee (Suwon-si, KR); Dhananjaya N. Gowda (Suwon-si, KR); Kwangyoun Kim (Suwon-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G10L15/30G10L15/02G10L15/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,514,916
App. No.
16/992,943
Granted
Nov 29, 2022
Kind
B2
Abstract

A server for supporting speech recognition of a device and an operation method of the server. The server and method identify a plurality of estimated character strings from the first character string and obtain a second character string, based on the plurality of estimated character strings, and transmit the second character string to the device. The first character string is output from a speech signal input to the device, via speech recognition.

Claims (51)

1. A server comprising:

a memory storing one or more computer-readable instructions;

a processor configured to execute the one or more computer-readable instructions stored in the memory; and

a communication interface configured to receive from a device a first character string of speech recognition by the device of a speech signal input to the device,

wherein the processor when executing the one or more computer-readable instructions is configured to:

identify an estimated character string to replace a portion of the first character string, based on the first character string; and

control the communication interface to transmit a second character string to the device, the second character string comprising the portion of the first character string replaced with the estimated character string, and

wherein the processor when executing the one or more computer-readable instructions is further configured to:

calculate likelihood matrices relating to replacement characters of the estimated character string that are to replace each character of the first character string, based on characters of the first character string accumulated prior to each character of the first character string; and

identify the estimated character string based on likelihood values within the likelihood matrices.

2. The server of claim 1 , wherein the processor when executing the one or more computer-readable instructions is further configured to:

obtain, the second character string, by replacing the portion of the first character string with the estimated character string based on the replacement characters, and

wherein the replacement characters are characters having pronunciations similar to each character within the first character string.

3. The server of claim 1 , wherein the processor when executing the one or more computer-readable instructions is further configured to:

calculate a likelihood of the estimated character string, based on the likelihood values within the likelihood matrices; and

select the estimated character string from among a plurality of estimated character strings, based on the likelihood, dictionary information, and a language model.

4. The server of claim 1 , wherein the likelihood matrices obtained for each character of the first character string are calculated based on posterior probabilities calculated based on characters of the first character string accumulated prior to each character of the first character string, and a character sequence probability calculated based on the characters of the first character string accumulated prior to each character of the first character string.

5. The server of claim 4 , wherein the posterior probabilities are calculated using an artificial intelligence recurrent neural network (RNN) including a plurality of long-short term memory (LSTM) layers and a softmax layer.

6. The server of claim 1 , wherein the likelihood matrices obtained for each character of the first character string are calculated based on a pre-determined confusion matrix.

7. The server of claim 1 , wherein the first character string includes characters respectively corresponding to speech signal frames obtained by splitting the speech signal at intervals of a preset time.

8. The server of claim 1 , wherein the processor when executing the one or more computer-readable instructions is further configured to provide a service associated with the speech signal input to the device, based on the second character string.

9. A device comprising:

a memory storing one or more computer-readable instructions;

a processor configured to execute the one or more computer-readable instructions stored in the memory; and

a communication interface configured to communicate with a server,

wherein the processor when executing the one or more computer-readable instructions is further configured to:

obtain a first character string by performing speech recognition on a speech signal;

determine whether to replace a portion of the first character string with another character string;

control the communication interface to transmit the first character string to the server, based on the determination; and

control the communication interface to receive, from the server, a second character string obtained by the server by replacing the portion included in the first character string with an estimated character string.

10. An operation method of a server, the operation method comprising:

receiving from a device a first character string of speech recognition by the device of a speech signal input to the device;

identifying an estimated character string to replace a portion of the first character string, based on the first character string;

transmitting a second character string to the device, the second character string comprising the portion of the first character string replaced with the estimated character string,

wherein the identifying comprises:

calculating likelihood matrices relating to replacement characters of the estimated character string that are to replace each character of the first character string, based on characters of the first character string accumulated prior to each character of the first character string; and

identifying the estimated character string based on likelihood values within the likelihood matrices.

11. The operation method of claim 10 ,

wherein the obtaining of the second character string, based on a plurality of estimated character strings, comprises obtaining, the second character string, by replacing the portion of the first character string with the estimated character string based on the replacement characters, and

the replacement characters are characters having pronunciations similar to each character within the first character string.

12. The operation method of claim 10 , wherein the obtaining of the second character string comprises:

calculating a likelihood of the estimated character string, based on the likelihood values within the likelihood matrices; and

selecting the estimated character string from among a plurality of estimated character strings, based on the likelihood, dictionary information, and a language model.

13. The operation method of claim 10 , wherein the likelihood matrices obtained for each character of the first character string are calculated based on posterior probabilities calculated based on characters of the first character string accumulated prior to each character of the first character string, and a character sequence probability calculated based on the characters of the first character string accumulated prior to each character of the first character string.

14. The operation method of claim 10 , wherein the first character string includes characters respectively corresponding to speech signal frames obtained by splitting the speech signal at intervals of a preset time.

15. The operation method of claim 10 , further comprising providing a service associated with the speech signal input to the device, based on the second character string.

16. An operation method of a device, the operation method comprising:

obtaining a first character string by performing speech recognition on a speech signal;

determining whether to replace a portion of the first character string with another character string;

transmitting the first character string to a server, based on the determination; and

receiving, from the server, a second character string obtained by the server by replacing the portion included in the first character string with an estimated character string.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 13, 2020
From: KIM, CHANWOO; JIN, SICHEN; LEE, KYUNGMIN; GOWDA, DHANANJAYA N.; KIM, KWANGYOUN
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 053491/0167 →
Priority Claims (2)
KR 10-2019-0133259 · Oct 24, 2019 · national
KR 10-2020-0018574 · Feb 14, 2020 · national
Continuity (2)
Provisional Application 62886027 · Aug 13, 2019
Related Publication 20210050018A1 · Feb 18, 2021