IP Library Granted Patent US 10,269,371
Granted Patent B2
US 10,269,371 · App. 15/720,498 · Granted Apr 23, 2019

Techniques for decreasing echo and transmission periods for audio communication sessions

Inventors: Erik Kay (Emerald Hills, CA); Jonas Erik Lindberg (Stockholm, SE); Serge Lachapelle (Vallentuna, SE); Henrik Lundin (Sollentuna, SE)
Assignee: Google LLC
G10L21/043G10L21/0208H04L65/1069H04L65/80G10L15/20G10L19/005G10L21/0232G10L25/78
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,269,371
App. No.
15/720,498
Granted
Apr 23, 2019
Kind
B2
Abstract

A computer-implemented technique can include establishing an audio communication session between first and second computing devices and obtaining, by the first computing device, an audio input signal using audio data captured by a microphone. The first computing device can analyze the audio input signal to detect a speech input by its first user and can determine a duration of a detection period from when the audio input signal was obtained until the analyzing has completed. The first computing device can then transmit, to the second computing device, (i) a portion of the audio input signal beginning at a start of the speech input and (ii) the detection period duration, wherein receipt of the portion of the audio input signal and the detection period duration causes the second computing device to accelerate playback of the portion of the audio input signal to compensate for the detection period duration.

Claims (32)

1. A computer-implemented method to reduce echo, comprising:

receiving a detected portion of an audio input signal for an audio communication session between a first computing device and a second computing device, wherein the detected portion of the audio input signal was captured by a microphone of the first computing device and the detected portion of the audio input signal includes a pitch period;

identifying the pitch period for removal by cross-correlating the audio input signal with itself;

obtaining a modified portion of the audio input signal by removing the pitch period; and

outputting the modified portion of the audio input signal, wherein the modified portion of the audio input signal is output for playback on a speaker.

2. The computer-implemented method of claim 1 , wherein identifying the pitch period for removal by cross-correlating the audio input signal with itself includes identifying a peak in an autocorrelation signal based on application of a threshold.

3. The computer-implemented method of claim 1 , wherein the modified portion of the audio input signal has a shorter length than the detected portion of the audio input signal.

4. The computer-implemented method of claim 2 , wherein the threshold is between 0.5 and 0.9.

5. The computer-implemented method of claim 1 , further comprising transmitting the modified portion of the audio input signal to the second computing device.

6. The computer-implemented method of claim 1 , wherein the detected portion of the audio input signal was determined by applying a voice activity detection (VAD) technique to the audio input signal, and wherein a duration of the detection period is a delay associated with the VAD technique.

7. The computer-implemented method of claim 6 , wherein the VAD technique distinguishes speech by a first user from noise and speech by another user.

8. A computing system having one or more processors and a non-transitory memory storing a set of instructions that, when executed by the one or more processors, causes the computing system to perform operations comprising:

receiving a detected portion of an audio input signal for an audio communication session between a first computing device and a second computing device, wherein the detected portion of the audio input signal was captured by a microphone of the first computing device and the detected portion of the audio input signal includes a pitch period;

identifying a pitch period for removal from a detected portion of the audio input signal by cross-correlating the audio input signal with itself at one or more temporal points;

obtaining a modified portion of the audio input signal by removing the pitch period; and

outputting the modified portion of the audio input signal, wherein the modified portion of the audio input signal is output for playback on a speaker.

9. The computing system of claim 8 , wherein identifying the pitch period for removal identifying a peak in an autocorrelation signal based on application of a threshold.

10. The computing system of claim 8 , wherein the modified portion of the audio input signal has a shorter length than the detected portion of the audio input signal.

11. The computing system of claim 9 , wherein the threshold is between 0.5 and 0.9.

12. The computing system of claim 8 , wherein the audio input signal is part of a video communication session.

13. The computing system of claim 8 , wherein the detected portion of the audio input signal was determined by applying a voice activity detection (VAD) technique to the audio input signal, and wherein a duration of the detection period is a delay associated with the VAD technique.

14. The computer-implemented method of claim 13 , wherein the VAD technique distinguishes speech by a first user from noise and speech by another user.

15. A non-transitory computer-readable medium having a set of instructions stored thereon that, when executed by the one or more processors of a computing system, causes the computing system to perform operations comprising:

receiving a detected portion of an audio input signal for an audio communication session between a first computing device and a second computing device, wherein the detected portion of the audio input signal was captured by a microphone of the first computing device and the detected portion of the audio input signal includes a pitch period;

identifying the pitch period for removal by cross-correlating the audio input signal with itself;

obtaining a modified portion of the audio input signal by removing the pitch period; and

outputting the modified portion of the audio input signal, wherein the modified portion of the audio input signal is output for playback on a speaker.

16. The computer-readable medium of claim 15 , wherein identifying the pitch period for removal by cross-correlating the audio input signal with itself includes identifying a peak in an autocorrelation signal based on application of a threshold.

17. The computer-readable medium of claim 15 , wherein the modified portion of the audio input signal has a shorter length than the detected portion of the audio input signal.

18. The computer-readable medium of claim 16 , wherein the threshold is between 0.5 and 0.9.

19. The computer-readable medium of claim 15 , wherein the operations further comprise transmitting the modified portion of the audio input signal to the second computing device.

20. The computer-readable medium of claim 15 , wherein the detected portion of the audio input signal was determined by applying a voice activity detection (VAD) technique to the audio input signal, wherein the VAD technique distinguishes speech by a first user from noise and speech by another user, and wherein a duration of the detection period is a delay associated with the VAD technique.

Assignments (2)
CHANGE OF NAME Recorded Oct 8, 2018
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 047203/0341 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 29, 2017
From: KAY, ERIK; LINDBERG, JONAS ERIK; LACHAPELLE, SERGE; LUNDIN, HENRIK
To: GOOGLE INC.
Reel/Frame 043742/0283 →
Continuity (2)
Continuation 15246950 · Aug 25, 2016
Related Publication 20180061437A1 · Mar 1, 2018