IP Library Granted Patent US 10,475,465
Granted Patent B2
US 10,475,465 · App. 16/026,449 · Granted Nov 12, 2019

Method and system for enhancing a speech signal of a human speaker in a video using visual information

Inventors: Shmuel Peleg (Mevaseret Zion, IL); Asaph Shamir (Jerusalem, IL); Tavi Halperin (Tel-Aviv, IL); Aviv Gabbay (Jerusalem, IL); Ariel Ephrat (Jerusalem, IL)
Assignee: Yissum Research Development Company, of The Hebrew University of Jerusalem Ltd.
G10L21/0205G06K9/00228G06K9/00523G06K9/4628G06K9/627G06K9/6289G10L21/0216G10L21/0232G10L21/0272G10L25/18G10L25/30G10L25/57G10L25/78
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,475,465
App. No.
16/026,449
Granted
Nov 12, 2019
Kind
B2
Abstract

A method and system for enhancing a speech signal is provided herein. The method may include the following steps: obtaining an original video, wherein the original video includes a sequence of original input images showing a face of at least one human speaker, and an original soundtrack synchronized with said sequence of images; and processing, using a computer processor, the original video, to yield an enhanced speech signal of said at least one human speaker, by detecting sounds that are acoustically unrelated to the speech of the at least one human speaker, based on visual data derived from the sequence of original input images.

Claims (35)

1. A method of enhancing a speech signal of a target human speaker, the method comprising:

obtaining a video, wherein said video comprises a sequence of images showing a face or parts of a face of the target human speaker, and an original soundtrack corresponding with said video;

representing said original soundtrack by a two-dimensional (2D) discrete time-frequency (DTF) audio transform;

generating a 2D time-frequency filter, having same dimensions as said 2D DTF audio transform, by analyzing said video;

obtaining a filtered DTF audio transform by point-wise multiplying the 2D DTF audio transform by said 2D time-frequency filter; and

generating an enhanced speech signal based on said filtered DTF audio transform, wherein said enhanced speech signal exhibits a removal, from the original soundtrack, of sounds that are unrelated to the speech of said target human speaker.

2. The method according to claim 1 , wherein said DTF audio transform is a Short-Term Fourier Transform (STFT) or a spectrogram.

3. The method according to claim 1 , wherein said generating of the 2D time frequency filter is carried out, at least in part, using a neural network.

4. The method according to claim 1 , wherein said 2D time-frequency filter is generated using an articulatory-to-acoustic mapping having as an input said sequence of original input images.

5. The method according to claim 3 , wherein the neural network is trained on a set of videos having respective clean speech signals.

6. The method according to claim 1 , wherein said enhanced speech signal exhibits less noise compared with the original soundtrack.

7. The method according to claim 1 , wherein said enhanced speech signal exhibits a better speaker separation of said target human speaker from another speaker included in the original soundtrack, compared with the original soundtrack.

8. The method according to claim 1 , wherein the sounds that are unrelated to the speech of said target human speaker comprise background sounds.

9. The method according to claim 8 , wherein the background sounds are unrelated to movements of the face or the part of the face of the target human speaker.

10. The method according to claim 1 , further comprising using the enhanced speech signal for control by speech of a computer-controlled device.

11. The method according to claim 10 , wherein said computer-controlled device comprises a vehicle.

12. A system for enhancing a speech signal of a target human speaker, the system comprising:

a computer memory configured to obtain a video, wherein said video comprises a sequence of images showing a face or parts of a face of the target human speaker, and an original soundtrack corresponding with said video; and

a computer processor configured to:

represent said original soundtrack by a two-dimensional (2D) discrete time-frequency (DTF) audio transform;

generate a 2D time-frequency filter, having same dimensions as said 2D DTF audio transform, by analyzing said video;

obtain a filtered DTF audio transform by point-wise multiplying the 2D DTF audio transform by said 2D time-frequency filter; and

generate an enhanced speech signal based on said filtered DTF audio transform, wherein said enhanced speech signal exhibits a removal, from the original soundtrack, of sounds that are unrelated to the speech of said target human speaker.

13. The system according to claim 12 , wherein said DTF audio transform is a Short-Term Fourier Transform (STFT) or a spectrogram.

14. The system according to claim 12 , wherein said generating of the 2D time frequency filter is carried out, at least in part, using a neural network.

15. The system according to claim 12 , wherein said 2D time-frequency filter is generated using an articulatory-to-acoustic mapping having as an input said sequence of original input images.

16. The system according to claim 14 , wherein the neural network is trained on a set of videos having respective clean speech signals.

17. The system according to claim 12 , wherein said enhanced speech signal exhibits less noise compared with the original soundtrack.

18. The system according to claim 12 , wherein said enhanced speech signal exhibits a better speaker separation of said target human speaker from another speaker included in the original soundtrack, compared with the original soundtrack.

19. A non-transitory computer readable medium for enhancing a speech signal of a target human speaker, the non-transitory computer readable medium comprising a set of instructions that when executed cause at least one computer processor to: represent said original soundtrack by a two-dimensional (2D) discrete time-frequency (DTF) audio transform; generate a 2D time-frequency filter, having same dimensions as said 2D DTF audio transform, by analyzing said video; obtain a filtered DTF audio transform by point-wise multiplying the 2D DTF audio transform by said 2D time-frequency filter; and generate an enhanced speech signal based on said filtered DTF audio transform, wherein said enhanced speech signal exhibits a removal, from the original soundtrack, of sounds that are unrelated to the speech of said target human speaker.

20. The non-transitory computer readable medium according to claim 19 , wherein said DTF audio transform is a Short-Term Fourier Transform (STFT) or a spectrogram.

21. The non-transitory computer readable medium according to claim 19 , wherein said generating of the 2D time frequency filter is carried out, at least in part, using a neural network.

22. The non-transitory computer readable medium according to claim 19 , wherein said 2D time-frequency filter is generated using an articulatory-to-acoustic mapping having as an input said sequence of original input images.

23. The non-transitory computer readable medium according to claim 21 , wherein the neural network is trained on a set of videos having respective clean speech signals.

24. The non-transitory computer readable medium according to claim 19 , wherein said enhanced speech signal exhibits less noise compared with the original soundtrack.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 7, 2018
From: PELEG, SHMUEL; SHAMIR, ASAPH; HALPERIN, TAVI; GABBAY, AVIV; EPHRAT, ARIEL
To: YISSUM RESEARCH DEVELOPMENT COMPANY OF THE HEBREW UNIVERSITY OF JERUSALEM LTD.
Reel/Frame 047087/0262 →
Continuity (4)
Provisional Application 62528225 · Jul 3, 2017
Provisional Application 62586472 · Nov 15, 2017
Provisional Application 62590774 · Nov 27, 2017
Related Publication 20190005976A1 · Jan 3, 2019
Cited By (1)
US 12,676,161