IP Library Granted Patent US 8,781,818
Granted Patent B2
US 8,781,818 · App. 13/141,710 · Granted Jul 15, 2014

Speech capturing and speech rendering

Inventors: Cornelis Pieter Janse (Eindhoven, NL); Leon C. A. Van Stuivenberg (Eindhoven, NL); Harm Jan Willem Belt (Eindhoven, NL); Bahaa Eddine Sarroukh (Eindhoven, NL); Mahdi Triki (Eindhoven, NL)
Assignee: Koninklijke Philips N.V.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,781,818
App. No.
13/141,710
Granted
Jul 15, 2014
Kind
B2
Abstract

The invention proposes extracting one or more speech signals ( 151 - 154 ) as well as one or more ambient signals ( 131 ) from sound signals captured by microphones, wherein each of the speech signals corresponds to a different speaker. The invention proposes to transmit both the one or more speech signals ( 151 - 154 ) and the one or more ambient signals ( 131 ) to a rendering side, as opposed to sending only speech signals. This enables to reproduce the speech and ambient signals in a spatially different way at the rendering side. By reproducing the ambient signals a feeling of “being together” is created. In an embodiment, the invention enables reproducing two or more speech signals spatially from each other and from the ambient signals so that speech intelligibility is increased despite the presence of the ambient signals.

Claims (31)

1. A speech capturing device-comprising:

a capturing circuit, wherein the capturing circuit includes a plurality of microphones for capturing a plurality of sound signals originating from different spatial locations;

one or more extracting circuits each for deriving a respective speech signal corresponding to a respective speaker from the plurality of the sound signals;

a residual extracting circuit for deriving one or more ambient signals from the plurality of sound signals each decreased by the one or more speech signals derived by the one or more extracting circuits; and

a transmitting circuit for transmitting the one or more speech signals and the one or more ambient signals, further comprising:

an audiovisual locator for (i) determining one or more locations of the speakers and (ii) providing one or more output signals of spatial information about the locations of the speakers to the one or more extracting circuits, respectively, wherein each extracting circuit derives the respective speech signal further in response to a respective output signal of spatial information directed to a location of a respective one of the speakers.

2. The speech capturing device according to claim 1 , wherein the transmitting circuit is further arranged for transmitting the output signals of spatial information of the one or more locations of the speakers.

3. The speech capturing device according to claim 1 , wherein each extracting circuit comprises a generalized side-lobe canceller for deriving a corresponding speech signal.

4. The speech capturing device according to claim 1 , wherein each extracting circuit further comprises a post-processor circuit for performing a further noise reduction in a corresponding speech signal.

5. The speech capturing device according to claim 1 , wherein the residual extracting circuit further comprises a multi-channel adaptive filter.

6. The speech capturing device according to claim 5 , wherein the multi-channel adaptive filter is coupled to receive a sound signal captured by one of the microphones as a reference signal.

7. A communication system for communicating speech signals, the communication system comprising:

a speech capturing device according to claim 1 ; and

a speech rendering device, wherein the speech rendering device comprises a receiving circuit for receiving one or more speech signals and one or more ambient signals, wherein each speech signal corresponds to a different speaker, and

a rendering circuit for spatially reproducing the one or more speech signals and the one or more ambient signals, wherein the rendering circuit spatially reproduces the one or more speech signals such that respective directions from which the one or more speech signals are perceived by a listener (a) are aligned to a respective spatial location of a speaker of the different speakers and (b) comprise different directions than perceived directions of the spatially reproduced one or more ambient signals.

8. A hands-free audio or audiovisual conferencing terminal comprising the speech capturing device according to claim 1 and a speech rendering device, wherein the speech rendering device comprises a receiving circuit for receiving one or more speech signals and one or more ambient signals, wherein each speech signal corresponds to a different speaker, and

a rendering circuit for spatially reproducing the one or more speech signals and the one or more ambient signals, wherein the rendering circuit spatially reproduces the one or more speech signals such that respective directions from which the one or more speech signals are perceived by a listener (a) are aligned to a respective spatial location of a speaker of the different speakers and (b) comprise different directions than perceived directions of the spatially reproduced one or more ambient signals.

9. A speech rendering device comprising:

a receiving circuit for receiving one or more speech signals and one or more ambient signals wherein each speech signal corresponds to a different speaker at a different spatial location, the receiving circuit further for receiving spatial information about locations of the different speakers; and

a rendering circuit for spatially reproducing (i) the one or more speech signals and (ii) the one or more ambient signals, wherein the rendering circuit, in response to the spatial information, spatially reproduces the one or more speech signals such that respective directions from which the one or more speech signals are perceived by a listener (a) are aligned to a respective different spatial location represented by the spatial information of a speaker of the different speakers in a visualization of that speaker and (b) comprise different directions than perceived directions of the spatially reproduced one or more ambient signals.

10. The speech rendering device according to claim 9 , wherein the rendering circuit is arranged for spatially reproducing two or more of speech signals, wherein respective directions of the spatially reproduced two or more speech signals perceived by the listener comprise mutually different directions.

11. The speech rendering device according to claim 9 , wherein the rendering circuit is further arranged for reducing amplitudes of the one or more ambient signals.

12. A speech capturing method comprising the steps of:

capturing, via a plurality of microphones, a plurality of sound signals originating from different spatial locations;

deriving, via one or more extracting circuits, one or more speech signals corresponding to one or more respective speakers from the plurality of the sound signals;

deriving, via a residual extracting circuit, one or more ambient signals from the plurality of sound signals each decreased by the one or more speech signals; and

transmitting, via a transmitting circuit, the one or more speech signals and the one or more ambient signals, further comprising:

determining, via an audiovisual locator, one or more locations of the speakers and providing, via the audiovisual locator, one or more output signals of spatial information about the locations of the speakers to the one or more extracting circuits, respectively, wherein deriving the respective speech signal further includes deriving in response to a respective output signal of spatial information directed to a location of a respective one of the speakers.

13. A speech rendering method comprising the steps of:

receiving, via a receiving circuit, one or more speech signals and one or more ambient signals, wherein each speech signal corresponds to a different speaker at a different spatial location; and

spatially reproducing, via a rendering circuit, the one or more speech signals and the one or more ambient signals, wherein the spatially reproducing spatially reproduces the one or more speech signals such that respective directions from which the one or more speech signals are perceived by a listener (a) are aligned to a respective spatial location of a speaker of the different speakers and (b) comprise different directions than perceived directions of the spatially reproduced one or more ambient signals.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 19, 2019
From: KONINKLIJKE PHILIPS ELECTRONICS N.V.
To: KONINKLIJKE PHILIPS N.V.
Reel/Frame 048634/0295 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 19, 2019
From: KONINKLIJKE PHILIPS N.V.
To: MEDIATEK INC.
Reel/Frame 048634/0357 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2011
From: JANSE, CORNELIS PIETER; VAN STUIVENBERG, LEON C.A.; BELT, HARM JAN WILLEM; SARROUKH, BAHAA EDDINE; TRIKI, MAHDI
To: KONINKLIJKE PHILIPS ELECTRONICS N.V.
Reel/Frame 026487/0549 →
Priority Claims (1)
EP 08172683 · Dec 23, 2008 · regional
Continuity (1)
Related Publication 20110264450A1 · Oct 27, 2011