IP Library Granted Patent US 10,431,236
Granted Patent B2
US 10,431,236 · App. 15/813,832 · Granted Oct 1, 2019

Dynamic pitch adjustment of inbound audio to improve speech recognition

Inventor: Carly Gloge (Boulder, CO)
Assignee: Sphero, Inc.
G10L21/007G10L15/22G10L21/003G10L15/26G10L25/90
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,431,236
App. No.
15/813,832
Granted
Oct 1, 2019
Kind
B2
Abstract

Aspects of the present disclosure relate to dynamic pitch adjustment of inbound audio to improve speech recognition. Inbound audio may be received. Upon receiving the inbound audio, clusters of speech input may be detected within the received inbound audio. An average pitch may be detected from the inbound audio, using either subparts of the inbound audio or one or more of the detected speech clusters. A determination may be made using, among other things, the average pitch. Based on this determination, the pitch of the inbound audio may be adjusted. The adjusted input may then be passed to a speech recognition component.

Claims (62)

1. A system for dynamically adjusting the pitch of inbound audio, comprising:

at least one processor; and

memory encoding computer executable instructions that, when executed by the at least one processor, perform a method comprising:

receiving an input audio segment;

detecting one or more clusters of speech input within the input audio segment;

detecting an average pitch for at least one of the one or more clusters of speech input;

determining, based on at least the average pitch and an expected content for the input audio segment, whether the pitch of the input audio segment should be adjusted;

based on determining that the pitch should be adjusted, adjusting the pitch of at least one of the one or more speech clusters to generate an adjusted audio segment; and

transmitting the adjusted audio segment to a speech recognition component.

2. The system of claim 1 , wherein detecting one or more clusters of speech input comprises:

detecting an initial input unit and a subsequent input unit for the input audio segment;

comparing the initial input unit and the subsequent input unit using a threshold to determine whether a relationship between the initial input unit and the subsequent input unit meets the threshold; and

when it is determined that the relationship meets the threshold, associating the initial input unit and the subsequent input unit as a speech cluster of the one or more clusters of speech input.

3. The system of claim 1 , wherein determining whether the pitch of the input audio segment should be adjusted further comprises evaluating a duration of the input audio segment.

4. The system of claim 1 , wherein the method further comprises:

transmitting the input audio segment to the speech recognition component;

generating an original recognition result using the input audio segment;

generating an adjusted recognition result using the adjusted audio segment;

determining whether the adjusted recognition result is more accurate than the original recognition result; and

when it is determined that the adjusted recognition result is more accurate, using the adjusted recognition result.

5. The system of claim 1 , wherein determining whether the pitch of the input audio segment should be adjusted comprises analyzing at least one of the one or more clusters of speech input within the input audio segment.

6. The system of claim 1 , wherein the input audio segment is received from an audio capture device.

7. The system of claim 1 , wherein the input audio segment is received from an audio processing component.

8. A system for dynamically adjusting the pitch of inbound audio, comprising:

at least one processor; and

memory encoding computer executable instructions that, when executed by the at least one processor, perform a method comprising:

receiving an input audio segment, wherein the input audio segment is comprised of one or more clusters of speech input;

detecting an average pitch for at least one of the one or more clusters of speech input;

determining, based on at least the average pitch and an expected content for the input audio segment, whether the pitch of at least a part of the input audio segment should be adjusted; and

based on determining that the pitch should be adjusted, adjusting the pitch of at least one of the one or more speech clusters to generate an adjusted audio segment.

9. The system of claim 8 , wherein determining whether the pitch of the input audio segment should be adjusted further comprises evaluating a duration of the input audio segment.

10. The system of claim 8 , wherein the method further comprises:

transmitting the adjusted audio segment to a speech recognition component;

transmitting the input audio segment to the speech recognition component;

generating an original recognition result using the input audio segment;

generating an adjusted recognition result using the adjusted audio segment;

determining whether the adjusted recognition result is more accurate than the original recognition result; and

when it is determined that the adjusted recognition result is more accurate, using the adjusted recognition result.

11. The system of claim 8 , wherein determining whether the pitch of the input audio segment should be adjusted comprises analyzing at least one of the one or more clusters of speech input within the input audio segment.

12. The system of claim 8 , wherein the input audio segment is received from an audio capture device.

13. The system of claim 8 , wherein the input audio segment is received from an audio processing component.

14. A computer-implemented method for dynamically adjusting the pitch of inbound audio, comprising:

receiving an input audio segment;

detecting one or more clusters of speech input within the input audio segment;

detecting an average pitch for at least one of the one or more clusters of speech input;

determining, based on at least the average pitch and an expected content for the input audio segment, whether the pitch of the input audio segment should be adjusted;

based on determining that the pitch should be adjusted, adjusting the pitch of at least one of the one or more speech clusters to generate an adjusted audio segment; and

transmitting the adjusted audio segment to a speech recognition component.

15. The computer-implemented method of claim 14 , wherein detecting one or more clusters of speech input comprises:

detecting an initial input unit and a subsequent input unit for the input audio segment;

comparing the initial input unit and the subsequent input unit using a threshold to determine whether a relationship between the initial input unit and the subsequent input unit meets the threshold; and

when it is determined that the relationship meets the threshold, associating the initial input unit and the subsequent input unit as a speech cluster of the one or more clusters of speech input.

16. The computer-implemented method of claim 14 , wherein determining whether the pitch of the input audio segment should be adjusted further comprises evaluating a duration of the input audio segment.

17. The computer-implemented method of claim 14 , further comprising:

transmitting the input audio segment to the speech recognition component;

generating an original recognition result using the input audio segment;

generating an adjusted recognition result using the adjusted audio segment;

determining whether the adjusted recognition result is more accurate than the original recognition result; and

when it is determined that the adjusted recognition result is more accurate, using the adjusted recognition result.

18. The computer-implemented method of claim 14 , wherein determining whether the pitch of the input audio segment should be adjusted comprises analyzing at least one of the one or more clusters of speech input within the input audio segment.

19. The computer-implemented method of claim 14 , wherein the input audio segment is received from an audio capture device.

20. The computer-implemented method of claim 14 , wherein the input audio segment is received from an audio processing component.

Assignments (2)
SECURITY INTEREST Recorded May 11, 2020
From: SPHERO, INC.
To: SILICON VALLEY BANK
Reel/Frame 052623/0705 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 15, 2017
From: GLOGE, CARLY
To: SPHERO, INC.
Reel/Frame 044137/0508 →
Continuity (2)
Provisional Application 62422458 · Nov 15, 2016
Related Publication 20180137874A1 · May 17, 2018
Cited By (1)
US 12,505,830