IP Library Granted Patent US 10,504,510
Granted Patent B2
US 10,504,510 · App. 15/578,523 · Granted Dec 10, 2019

Motion adaptive speech recognition for enhanced voice destination entry

Inventors: Munir Nikolai Alexander Georges (Kehl, DE); Josef Damianus Anastasiadis (Aachen, DE); Oliver Bender (Aachen, DE)
Assignee: Cerence Operating Company
G10L15/22G01C21/3608G10L15/065G10L15/30G10L15/32G10L2015/223G10L2015/225G10L2015/227G10L2015/228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,504,510
App. No.
15/578,523
Granted
Dec 10, 2019
Kind
B2
Abstract

A method or associated system for motion adaptive speech processing includes dynamically estimating a motion profile that is representative of a user's motion based on data from one or more resources, such as sensors and non-speech resources, associated with the user. The method includes effecting processing of a speech signal received from the user, for example, while the user is in motion, the processing taking into account the estimated motion profile to produce an interpretation of the speech signal. Dynamically estimating the motion profile can include computing a motion weight vector using the data from the one or more resources associated with the user, and can further include interpolating a plurality of models using the motion weight vector to generate a motion adaptive model. The motion adaptive model can be used to enhance voice destination entry for the user and re-used for other users who do not provide motion profiles.

Claims (21)

1. A method for motion adaptive speech processing for voice destination entry, the method comprising:

dynamically estimating a motion profile that is representative of a user's motion based on data from one or more resources associated with the user, the motion profile being a collection of snap shots of the user's position and motion, the data from the one or more resources including sensor data and data from a non-speech resource associated with the user;

wherein dynamically estimating the motion profile includes computing a motion weight vector using the data from the one or more resources associated with the user and interpolating a plurality of models using the motion weight vector to generate a motion adaptive model, wherein at least one of the models is associated with a geographic area, and wherein interpolating the models results in a probability that the user is or will be located in the geographic area; and

effecting processing of a speech signal received from the user, the processing taking into account the estimated motion profile by constraining a search space to user-relevant destinations based on the motion adaptive model to produce an interpretation of the speech signal.

2. The method according to claim 1 , wherein the sensor data include at least one member selected from the group consisting of position, speed, acceleration, direction, and a combination thereof.

3. The method according to claim 2 , wherein the data from the non-speech resource includes at least one member selected from the group consisting of navigation system data, address book data, calendar data, motion history data, crowd sourced data, configuration data, and a combination thereof.

4. The method according to claim 1 , wherein computing the motion weight vector includes determining a relation between the non-speech resource and a language resource associated with the plurality of models.

5. The method according to claim 1 , wherein dynamically estimating the motion profile further includes interpolating the motion adaptive model with a background model.

6. The method according to claim 1 , wherein the speech signal includes at least one of a voice audio signal, a video signal and data from gestures or text entry.

7. A system for motion adaptive speech processing for voice destination entry, the system comprising:

a motion profile estimator at a client, the estimator configured to estimate a motion profile that is representative of a user's motion dynamically based on data from one or more resources associated with the user, the motion profile being a collection of snap shots of the user's position and motion, the data from the one or more resources including sensor data and data from a non-speech resource associated with the user;

wherein the motion profile estimator is configured to compute a motion weight vector using the data from the one or more resources associated with the user and to interpolate a plurality of models using the motion weight vector to generate a motion adaptive model, wherein at least one of the models is associated with a geographic area, and wherein the motion profile estimator interpolates the models to produce a probability that the user is or will be located in the geographic area; and

a processor configured to effect processing of a speech signal received from the user, the processing taking into account the estimated motion profile by constraining a search space to user-relevant destinations based on the motion adaptive model to produce an interpretation of the speech signal.

8. The system according to claim 7 , wherein the motion profile estimator is further configured to interpolate the motion adaptive model with a background model.

9. The system according to claim 7 , wherein the processor is configured to perform automatic speech recognition (ASR) of the speech signal at the client using the estimated motion profile, the ASR producing the interpretation of the speech signal.

10. The system according to claim 7 , wherein the processor is configured to send the speech signal and estimated motion profile to a cloud service to perform automatic speech recognition (ASR) of the speech signal using the estimated motion profile, the ASR producing the interpretation of the speech signal.

11. The system according to claim 7 , wherein the processor is configured to send the speech signal to a cloud service for automatic speech recognition (ASR), receive results of the ASR from the cloud service, and re-rank the results using the estimated motion profile to produce the interpretation of the speech signal.

12. A computer program product comprising a non-transitory computer readable medium storing instructions for performing a method for motion adaptive speech processing for voice destination entry, the instructions, when executed by a processor, cause the processor to:

dynamically estimate a motion profile that is representative of a user's motion based on data from one or more resources associated with the user, the motion profile being a collection of snap shots of the user's position and motion, the data from the one or more resources including sensor data and data from a non-speech resource associated with the user;

wherein dynamically estimating the motion profile includes computing a motion weight vector using the data from the one or more resources associated with the user and interpolating a plurality of models using the motion weight vector to generate a motion adaptive model, wherein at least one of the models is associated with a geographic area, and wherein interpolating the models results in a probability that the user is or will be located in the geographic area; and

effect processing of a speech signal received from the user, the processing taking into account the estimated motion profile by constraining a search space to user-relevant destinations based on the motion adaptive model, to produce an interpretation of the speech signal.

Assignments (8)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REPLACE THE CONVEYANCE DOCUMENT WITH THE NEW ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 19, 2022
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 059804/0186 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY AGREEMENT. Recorded Oct 29, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 050871/0001 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 23, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050836/0191 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 12, 2018
From: GEORGES, MUNIR NIKOLAI ALEXANDER; ANASTASIADIS, JOSEF DAMIANUS; BENDER, OLIVER
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 044895/0454 →
Continuity (1)
Related Publication 20180158455A1 · Jun 7, 2018