IP Library Granted Patent US 11,024,297
Granted Patent B2
US 11,024,297 · App. 16/171,093 · Granted Jun 1, 2021

Method for using pauses detected in speech input to assist in interpreting the input during conversational interaction for information retrieval

Inventors: Murali Aravamudan (Andover, MA); Daren Gill (Concord, MA); Sashikumar Venkataraman (Andover, MA); Vineet Agarwal (Andover, MA); Ganesh Ramamoorthy (Salem, NH)
Assignee: Veveo, Inc.
G10L15/187G06F16/683G10L15/1822G06F16/433G10L15/1815G10L15/19G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,024,297
App. No.
16/171,093
Granted
Jun 1, 2021
Kind
B2
Abstract

A method for using speech disfluencies detected in speech input to assist in interpreting the input is provided. The method includes providing access to a set of content items with metadata describing the content items, and receiving a speech input intended to identify a desired content item. The method further includes detecting a speech disfluency in the speech input and determining a measure of confidence of a user in a portion of the speech input following the speech disfluency. If the confidence measure is lower than a threshold value, the method includes determining an alternative query input based on replacing the portion of the speech input following the speech disfluency with another word or phrase. The method further includes selecting content items based on comparing the speech input, the alternative query input (when the confidence measure is low), and the metadata associated with the content items.

Claims (64)

1. A method for using speech disfluencies detected in speech input to assist in interpreting the input, the method comprising:

providing access to a set of content items, each of the content items being associated with metadata that describes the corresponding content item;

receiving a speech input from a user;

detecting an absence of a period of silence in a beginning of the speech input;

in response to detecting the absence of the period of silence in the beginning of the speech input

determining that the speech input excludes a first portion of a first word of the speech input;

in response to determining that the speech input excludes the first portion of the first word of the speech input, determining a measure of confidence of the user in the speech input following the first word;

in response to determining that the measure of confidence of the user in the speech input following the first word exceeds a threshold value:

identifying a second portion of the first word included in the speech input;

identifying a plurality of words having a suffix matching the second portion of the first word;

identifying a subset of the plurality of words matching a user preference signature;

determining a plurality of interpretations for the speech input based on the subset of the plurality of words;

selecting a subset of content items, from the set of content items, wherein each content item of the subset of content items matches an interpretation of the plurality of interpretations of the speech input; and

in response to determining that the measure of confidence of the user in the speech input following the first word does not exceed the threshold value:

replacing a portion of the speech input with an alternative input and

selecting the subset of content items, from the set of content items, wherein each content item of the subset of content items matches the alternative input and

generating for display the subset of content items to the user.

2. The method of claim 1 , further comprising:

providing the user preference signature, the user preference signature comprising preferences of the user for at least one of (i) particular content items and (ii) metadata associated with the content items;

ranking the subset of content items based on the user preference signature; and

ordering the display of the subset of content based on the ranking.

3. The method of claim 1 , wherein the first portion is excluded from the speech input due to front end clipping of the first word.

4. The method of claim 3 , wherein the front-end clipping results in an incomplete detection of the first word.

5. The method of claim 1 , wherein detecting the absence of the period of silence in the beginning of the speech input comprises:

measuring a duration of the period of silence in the beginning of the speech input; and

determining that the duration of the period of silence is less than a threshold maximum duration.

6. The method of claim 1 wherein detecting the absence of the period of silence in the beginning of the speech input comprises:

measuring a sound level at the beginning of the speech input; and

determining that the sound level exceeds a threshold minimum value.

7. The method of claim 1 , further comprising:

determining the second portion of the first word as a subset of the speech input commencing at the beginning of the speech input and ending at a first period of silence in the speech input.

8. The method of claim 1 , further comprising receiving a user input indicating the beginning of the speech input.

9. A system for using speech disfluencies detected in speech input to assist in interpreting the input, the system comprising control circuitry configured to:

provide access to a set of content items, each of the content items being associated with metadata that describes the corresponding content item;

receive a speech input from a user;

detect an absence of a period of silence in a beginning of the speech input;

in response to detecting the absence of the period of silence in the beginning of the speech input

determine that the speech input excludes a first portion of a first word of the speech input;

in response to determining that the speech input excludes the first portion of the first word of the speech input, determining a measure of confidence of the user in the speech input following the first word;

in response to determining that the measure of confidence of the user in the speech input following the first word exceeds a threshold value:

identify a second portion of the first word included in the speech input;

identify a plurality of words having a suffix matching the second portion of the first word;

identify a subset of the plurality of words matching a user preference signature;

determine a plurality of interpretations for the speech input based on the subset of the plurality of words;

select a subset of content items, from the set of content items, wherein each content item of the subset of content items matches an interpretation of the plurality of interpretations of the speech input; and

in response to determining that the measure of confidence of the user in the speech input following the first word does not exceed the threshold value:

replace a portion of the speech input with an alternative input and

select the subset of content items, from the set of content items, wherein each content item of the subset of content items matches the alternative input and

generate for display the subset of content items to the user.

10. The system of claim 9 , wherein the control circuitry is further configured to:

provide the user preference signature, the user preference signature comprising preferences of the user for at least one of (i) particular content items and (ii) metadata associated with the content items;

rank the subset of content items based on the user preference signature; and

order the display of the subset of content based on the ranking.

11. The system of claim 9 , wherein the first portion is excluded from the speech input due to front end clipping of the first word.

12. The system of claim 11 , wherein the front-end clipping results in an incomplete detection of the first word.

13. The system of claim 9 , wherein the control circuitry is further configured, when detecting the absence of the period of silence in the beginning of the speech input, to:

measure a duration of the period of silence in the beginning of the speech input; and

determine that the duration of the period of silence is less than a threshold maximum duration.

14. The system of claim 9 , wherein the control circuitry is further configured, when detecting the absence of the period of silence in the beginning of the speech input, to:

measure a sound level at the beginning of the speech input; and

determine that the sound level exceeds a threshold minimum value.

15. The system of claim 9 , wherein the control circuitry is further configured to:

determine the second portion of the first word as a subset of the speech input commencing at the beginning of the speech input and ending at a first period of silence in the speech input.

16. The system of claim 9 , wherein the control circuitry is further configured to receive a user input indicating the beginning of the speech input.

Assignments (9)
CHANGE OF NAME Recorded Sep 24, 2024
From: VEVEO, INC.
To: VEVEO LLC
Reel/Frame 069036/0068 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 24, 2024
From: VEVEO LLC
To: ROVI GUIDES, INC.
Reel/Frame 069036/0130 →
CHANGE OF NAME Recorded Sep 24, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069036/0209 →
RELEASE OF SECURITY INTEREST Recorded Jun 5, 2020
From: HPS INVESTMENT PARTNERS, LLC
To: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
Reel/Frame 053458/0749 →
RELEASE OF SECURITY INTEREST Recorded Jun 5, 2020
From: MORGAN STANLEY SENIOR FUNDING, INC.
To: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
Reel/Frame 053481/0790 →
SECURITY INTEREST Recorded Jun 1, 2020
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS INC.; VEVEO, INC.; INVENSAS CORPORATION; INVENSAS BONDING TECHNOLOGIES, INC.; TESSERA, INC.; TESSERA ADVANCED TECHNOLOGIES, INC.; DTS, INC.; PHORUS, INC.; IBIQUITY DIGITAL CORPORATION
To: BANK OF AMERICA, N.A.
Reel/Frame 053468/0001 →
PATENT SECURITY AGREEMENT Recorded Nov 25, 2019
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
To: MORGAN STANLEY SENIOR FUNDING, INC., AS COLLATERAL AGENT
Reel/Frame 051110/0006 →
SECURITY INTEREST Recorded Nov 22, 2019
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
To: HPS INVESTMENT PARTNERS, LLC, AS COLLATERAL AGENT
Reel/Frame 051143/0468 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 5, 2019
From: ARAVAMUDAN, MURALI; GILL, DAREN; VENKATARAMAN, SASHIKUMAR; AGARWAL, VINEET; RAMAMOORTHY, GANESH
To: VEVEO, INC.
Reel/Frame 048241/0437 →
Continuity (4)
Continuation 15693162 · Aug 31, 2017
Continuation 13801837 · Mar 13, 2013
Provisional Application 61679184 · Aug 3, 2012
Related Publication 20190130899A1 · May 2, 2019
Cited By (1)
US 12,598,256