IP Library › Granted Patent US 9,818,427
Granted Patent B2
US 9,818,427 · App. 14/977,911 · Granted Nov 14, 2017

Automatic self-utterance removal from multimedia files

Inventors: Niall Cahill (Galway, IE); Jakub Wenus (Maynooth, IE); Mark Kelly (Leixlip, IE)
Assignee: Intel Corporation
G10L21/028G10L15/063G10L25/51G10L25/78H04W88/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,818,427
App. No.
14/977,911
Granted
Nov 14, 2017
Kind
B2
Abstract

Embodiments of a system and method for removing speech by a user from audio frames are generally described herein. A method may include receiving a plurality of frames of audio data, extracting a set of frames of the plurality of frames, the set of frames including speech by a user with a set of remaining frames in the plurality of frames not in the set of frames, suppressing the speech by the user from the set of frames using a trained model to create a speech-suppressed set of frames, and recompiling the plurality of frames using the speech-suppressed set of frames and the set of remaining frames.

Claims (41)

1. A device for removing self-utterances from audio, the device comprising:

a microphone to record a plurality of frames of audio data;

processing circuitry to:

create a trained model from a plurality of audio training frames, the plurality of audio training frames including speech by a user during a telephone call at a device, wherein the trained model includes using a Gaussian Scale Mixture Model (GSMM) with parameters optimized with a learning algorithm using a modified Expectation Maximization (EM) technique and wherein the learning algorithm is scheduled to stall when the learning rate of the EM technique stabilizes;

extract, at the device, a set of frames of the plurality of frames, the set of frames including speech by the user with a set of remaining frames in the plurality of frames not in the set of frames;

suppress the speech by the user from the set of frames using the trained model to create a speech-suppressed set of frames; and

recompile, at the device, the plurality of frames using the speech-suppressed set of frames and the set of remaining frames.

2. The device claim 1 , further comprising a speaker to play back the recompiled plurality of frames.

3. The device of claim 1 , wherein the device is a mobile device.

4. The device of claim 1 , wherein the plurality of frames of audio data are extracted from a multimedia file.

5. The device of claim 1 , wherein to extract the set of frames including the speech, the processing circuitry is to:

convert the plurality of frames to a frequency domain file;

determine high-energy frames of the frequency domain file; and

compare the high-energy frames to the trained model to determine whether the high-energy frames include speech.

6. The device of claim 5 , wherein the set of frames corresponds to the high-energy frames that are determined to include speech.

7. At least one non-transitory machine readable medium including instructions that, when executed, cause the machine to:

create a trained model from a plurality of audio training frames, the plurality of audio training frames including speech by a user during a telephone call at a device, wherein the trained model includes using a Gaussian Scale Mixture Model (GSMM) with parameters optimized with a learning algorithm using a modified Expectation Maximization (EM) technique and wherein the learning algorithm is scheduled to stall when the learning rate of the EM technique stabilizes;

receive, at a device, a plurality of frames of audio data;

extract, at the device, a set of frames of the plurality of frames, the set of frames including speech by the user with a set of remaining frames in the plurality of frames not in the set of frames;

suppress, at the device, the speech by the user from the set of frames using the trained model to create a speech-suppressed set of frames; and

recompile, at the device, the plurality of frames using the speech-suppressed set of frames and the set of remaining frames.

8. The at least one non-transitory machine readable medium of claim 7 , further comprising instructions to play back the recompiled plurality of frames.

9. The at least one non-transitory machine readable medium of claim 7 , wherein the device is a mobile device.

10. The at least one non-transitory machine readable medium of claim 7 , further comprising instructions to record the plurality of frames.

11. A method for removing self-utterances from audio, the method comprising:

creating a trained model from a plurality of audio training frames, the plurality of audio training frames including speech by a user during a telephone call at a device, wherein the trained model includes using a Gaussian Scale Mixture Model (GSMM) with parameters optimized with a learning algorithm using a modified Expectation Maximization (EM) technique and wherein the learning algorithm is scheduled to stall when the learning rate of the EM technique stabilizes;

receiving, at the device, a plurality of frames of audio data;

extracting, at the device, a set of frames of the plurality of frames, the set of frames including speech by the user with a set of remaining frames in the plurality of frames not in the set of frames;

suppressing, at the device, the speech by the user from the set of frames using the trained model to create a speech-suppressed set of frames; and

recompiling, at the device, the plurality of frames using the speech-suppressed set of frames and the set of remaining frames.

12. The method of claim 11 , further comprising playing back the recompiled plurality of frames.

13. The method of claim 11 , wherein the device is a mobile device.

14. The method of claim 11 , further comprising recording the plurality of frames.

15. The method of claim 11 , wherein the plurality of frames of audio data are extracted from a multimedia file.

16. The method of claim 11 , wherein extracting the set of frames including the speech includes:

converting the plurality of frames to a frequency domain file;

determining high-energy frames of the frequency domain file; and

comparing the high-energy frames to the trained model to determine whether the high-energy frames include speech.

17. The method of claim 16 , wherein the set of frames corresponds to the high-energy frames that are determined to include speech.

18. The method of claim 11 , wherein the set of remaining frames do not include speech by the user.

19. The method of claim 11 , further comprising recording the plurality of frames at the device, and wherein recompiling the plurality of frames includes recompiling the frames with self-utterances of the user at the device during recording removed.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 19, 2016
From: CAHILL, NIALL; WENUS, JAKUB; KELLY, MARK
To: INTEL CORPORATION
Reel/Frame 037778/0096 →
Continuity (1)
Related Publication 20170178661A1 · Jun 22, 2017