IP Library Granted Patent US 11,804,237
Granted Patent B2
US 11,804,237 · App. 17/474,077 · Granted Oct 31, 2023

Conference terminal and echo cancellation method for conference

Inventors: Po-Jen Tu (New Taipei, TW); Jia-Ren Chang (New Taipei, TW); Kai-Meng Tzeng (New Taipei, TW)
Assignee: Acer Incorporated
G10L21/0364G10L19/018G10L21/0208H04N7/15G10L2021/02082
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,804,237
App. No.
17/474,077
Granted
Oct 31, 2023
Kind
B2
Abstract

A conference terminal and an echo cancellation method for a conference are provided. In the echo cancellation method, a synthetic speech signal is received. The synthetic speech signal includes a user speech signal of a speaking party corresponding to a first conference terminal of multiple conference terminals and an audio watermark signal corresponding to the first conference terminal. One or more delay times corresponding to the audio watermark signal are detected in a received audio signal. The received audio signal is recorded through a sound receiver of a second conference terminal of the conference terminals. An echo in the received audio signal is canceled according to the delay time.

Claims (40)

1. An echo cancellation method for a conference, adapted to a plurality of conference terminals each comprising a sound receiver and a loudspeaker, the echo cancellation method comprising:

receiving a synthetic speech signal, wherein the synthetic speech signal comprises a user speech signal of a speaking party corresponding to a first conference terminal of the plurality of conference terminals and an audio watermark signal corresponding to the first conference terminal;

detecting at least one delay time corresponding to the audio watermark signal in a received audio signal relative to the synthetic speech signal, wherein the received audio signal is recorded through the sound receiver of a second conference terminal of the plurality of conference terminals, and detecting the at least one delay time of the audio watermark signal comprises:

generating at least one initial delay signal corresponding to the user speech signal according to the at least one initial delay time, wherein a delay time of the at least one initial delay signal relative to the user speech signal is the at least one initial delay time between the audio watermark signal in the received audio signal and the audio watermark signal in the synthetic speech signal; and

estimating an echo path according to the at least one initial delay signal, wherein the audio watermark signal is delayed by the at least one initial delay time after passing through the echo path, and the echo path is a channel between the sound receiver and the loudspeaker; and

canceling an echo in the received audio signal according to the at least one delay time.

2. The echo cancellation method for a conference according to claim 1 , wherein detecting the at least one delay time corresponding to the audio watermark signal in the received audio signal comprises:

determining the at least one initial delay time according to a correlation between the received audio signal and the audio watermark signal, wherein the at least one initial delay time corresponds to a relatively high degree of the correlation.

3. The echo cancellation method for a conference according to claim 1 , wherein the synthetic speech signal further comprises a second user speech signal of the speaking party corresponding to a third conference terminal of the plurality of conference terminals, and a second audio watermark signal corresponding to the third conference terminal, and the echo cancellation method further comprises:

detecting at least one delay time corresponding to the second audio watermark signal in the received audio signal.

4. The echo cancellation method for a conference according to claim 1 , wherein the audio watermark signal has a frequency of higher than 16 kilohertz (kHz).

5. The echo cancellation method for a conference according to claim 1 , wherein canceling the echo in the received audio signal comprises:

generating at least one second synthetic speech signal which is the synthetic speech signal with the at least one delay time; and

canceling the at least one second synthetic speech signal from the received audio signal.

6. The echo cancellation method for a conference according to claim 1 , wherein estimating the echo path comparing:

estimating an impulse response of the echo path by applying the at least one initial delay signal to an adaptive filter.

7. The echo cancellation method for a conference according to claim 1 , further comprising:

playing, through the loudspeaker, the synthetic speech signal received via a network.

8. A conference terminal comprising:

a sound receiver, configured to perform recording and obtain a received audio signal of a speaking party corresponding thereto;

a loudspeaker, configured to play a sound;

a communication transceiver, configured to transmit or receive data; and

a processor, coupled to the sound receiver, the loudspeaker and the communication transceiver, and configured to:

receive a synthetic speech signal through the communication transceiver, wherein the synthetic speech signal comprises a user speech signal of the speaking party corresponding to a second conference terminal and an audio watermark signal corresponding to the second conference terminal;

detect at least one delay time corresponding to the audio watermark signal in the received audio signal relative to the synthetic speech signal, and the processor is further configured to

generate at least one initial delay signal corresponding to the user speech signal according to at least one initial delay time, wherein a delay time of the at least one initial delay signal relative to the user speech signal is the at least one initial delay time between the audio watermark signal in the received audio signal and the audio watermark signal in the synthetic speech signal; and

estimate an echo path according to the at least one initial delay signal, wherein the audio watermark signal is delayed by the at least one initial delay time after passing through the echo path, and the echo path is a channel between the sound receiver and the loudspeaker; and

cancel an echo in the received audio signal according to the at least one delay time.

9. The conference terminal according to claim 8 , wherein the processor is further configured to:

determine the at least one initial delay time according to a correlation between the received audio signal and the audio watermark signal, wherein the at least one initial delay time corresponds to a relatively high degree of the correlation.

10. The conference terminal according to claim 8 , wherein the synthetic speech signal further comprises a second user speech signal of the speaking party corresponding to a third conference terminal, and a second audio watermark signal corresponding to the third conference terminal, and the processor is further configured to:

detect at least one delay time corresponding to the second audio watermark signal in the received audio signal.

11. The conference terminal according to claim 8 , wherein the audio watermark signal has a frequency of higher than 16 kHz.

12. The conference terminal according to claim 8 , wherein the processor is further configured to:

generate at least one second synthetic speech signal which is the synthetic speech signal with the at least one delay time; and

cancel the at least one second synthetic speech signal from the received audio signal.

13. The conference terminal according to claim 8 , wherein the processor is further configured to:

estimate an impulse response of the echo path by applying the at least one initial delay signal to an adaptive filter.

14. The conference terminal according to claim 8 , wherein the processor is further configured to:

play, through the loudspeaker, the synthetic speech signal received via a network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 14, 2021
From: TU, PO-JEN; CHANG, JIA-REN; TZENG, KAI-MENG
To: ACER INCORPORATED
Reel/Frame 057469/0970 →
Priority Claims (1)
TW 110130678 · Aug 19, 2021 · national
Continuity (1)
Related Publication 20230058981A1 · Feb 23, 2023
Cited By (1)
US 12,470,661