IP Library Granted Patent US 12,327,560
Granted Patent B2
US 12,327,560 · App. 18/504,312 · Granted Jun 10, 2025

Voice command scrubbing

Inventor: Christopher Iain Parkinson (Richland, WA)
Assignee: RealWear, Inc.
G10L15/22G10L15/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,327,560
App. No.
18/504,312
Granted
Jun 10, 2025
Kind
B2
Abstract

The invention is directed towards a an audio scrubbing system that allows for scrubbing recognized voice commands from audio data and replacing the recognized voice commands with environment audio data. Specifically, as a user captures video and audio data via a HMD, audio data captured by the HMD may be processed by an audio scrubbing module to identify voice commands in the audio data that are used for controlling the HMD. When a voice command is identified in the audio data, timestamps corresponding to the voice command may be determined. Filler audio data may then be generated to imitate the environment by processing at least a portion of the audio data by a neural network of a machine learning model. The filler audio data may then be used to replace the audio data corresponding to the identified voice commands, thereby scrubbing the voice command from the audio data.

Claims (37)

1. A head-mounted computing device comprising:

a microphone;

one or more processors;

one or more memory devices storing programmable instructions thereon that, when executed by the one or more processors, cause the one or more processors to execute operations including:

receiving input data, wherein the received input data includes a recorded audio portion;

determining that the recorded audio portion includes an audible voice utterance that corresponds to a defined voice command, wherein the audible voice utterance has a first duration;

generating, via a machine learning model, a filler audio portion based on audio segments sampled from the recorded audio portion prior to and after the audible voice utterance, the filler audio portion having a second duration equal to the first duration and excluding the audible voice utterance; and

scrubbing the recorded audio portion based on the generated filler audio portion, the scrubbed recorded audio portion excluding the audible voice utterance.

2. The device of claim 1 , wherein scrubbing the audio portion further comprises identifying a voice pattern corresponding to the audible voice utterance.

3. The device of claim 2 , further comprising neutralizing the audible voice utterance that corresponds to the defined voice command based on combining a scrubbing voice pattern with the voice pattern corresponding to the audible voice utterance.

4. The device of claim 1 , further comprising a plurality of microphones; and

wherein scrubbing the audio portion further comprises parsing a set of audio signals from a plurality of microphones, wherein audio signals determined to be near a user mouth are muted.

5. The device of claim 1 , wherein based on determining that a boom arm of the device is rotated into a first position, activating, by an application of the device, a directional microphone configured to capture speech from a user wearing the device.

6. The device of claim 1 , further comprising identifying a portion of audio data corresponding to the defined voice command based on stored voice recognition characteristics.

7. The device of claim 1 , further comprising storing the scrubbed recorded audio portion including the filler audio portion and the recorded corresponding video portion to a memory of the head-mounted computing device.

8. A non-transitory computer storage medium storing computer-useable instructions that, when used by one or more computing devices, cause the one or more computing devices to perform operations comprising:

receiving input data, wherein the received input data includes a recorded audio portion and a recorded corresponding video portion;

determining that the recorded audio portion includes an audible voice utterance that corresponds to a defined voice command, wherein the audible voice utterance has a first duration; and

generating, via a machine learning model, a filler audio portion based on audio segments sampled from the recorded audio portion prior to and after the audible voice utterance, the filler audio portion having a second duration equal to the first duration and excluding the audible voice utterance.

9. The computer storage medium of claim 8 , further comprising scrubbing the recorded audio portion based on the generated filler audio portion, the scrubbed recorded audio portion excluding the audible voice utterance and retaining environmental audio data corresponding to the audio portion.

10. The computer storage medium of claim 9 , wherein scrubbing the audio portion further comprises identifying a voice pattern corresponding to the audible voice utterance.

11. The computer storage medium of claim 10 , further comprising neutralizing the audible voice utterance that corresponds to the defined voice command based on combining a scrubbing voice pattern with the voice pattern corresponding to the audible voice utterance.

12. The computer storage medium of claim 8 , wherein scrubbing the audio portion further comprises parsing a set of audio signals from a plurality of microphones, wherein audio signals determined to be near a user mouth are muted.

13. The computer storage medium of claim 8 , wherein based on determining that a boom arm of the device is rotated into a first position, activating, by an application of the device, a directional microphone configured to capture speech from a user wearing the device.

14. The computer storage medium of claim 8 , further comprising identifying a portion of audio data corresponding to the defined voice command based on stored voice recognition characteristics.

15. A computer-implemented method comprising:

receiving input data, wherein the received input data includes an audio portion and a corresponding video portion;

determining that the audio portion includes an audible voice utterance that corresponds to a defined voice command;

excluding the audio portion from the input data based on the determination;

generating, by a machine learning model, a filler audio portion of a duration equal to a duration of the audible voice utterance and excluding the audible voice utterance; and

embedding the filler audio portion into the received input data based on a timestamp corresponding to the audible voice utterance.

16. The computer-implemented method of claim 15 , further comprising storing the scrubbed recorded audio portion including the filler audio portion and the recorded corresponding video portion to a memory of the head-mounted computing device.

17. The computer-implemented method of claim 15 , wherein excluding the audio portion further comprises identifying a voice pattern corresponding to the audible voice utterance.

18. The computer-implemented method of claim 17 , further comprising neutralizing the audible voice utterance that corresponds to the defined voice command based on combining a scrubbing voice pattern with the voice pattern corresponding to the audible voice utterance.

19. The computer-implemented method of claim 15 , wherein excluding the audio portion further comprises parsing a set of audio signals from a plurality of microphones, wherein audio signals determined to be near a user mouth are muted.

20. The computer-implemented method of claim 15 , further comprising:

wherein based on determining that a boom arm of the device is rotated into a first position, activating, by an application of the device, a directional microphone configured to capture speech from a user wearing the device.

Assignments (2)
SECURITY INTEREST Recorded Jun 6, 2024
From: REALWEAR, INC.
To: WESTERN ALLIANCE BANK
Reel/Frame 067646/0492 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 13, 2023
From: PARKINSON, CHRISTOPHER IAIN
To: REALWEAR, INC.
Reel/Frame 065859/0057 →
Continuity (2)
Continuation 17060579 · Oct 1, 2020
Related Publication 20240079009A1 · Mar 7, 2024
References Cited (39)
US 5954834A · Hassan · 1999 [cited by examiner]
US 8274571B2 · Zhu · 2012 [cited by applicant]
US 8767035B2 · Baldwin · 2014 [cited by examiner]
US 9530410B1 · Lebeau et al. · 2016 [cited by applicant]
US 9544491B2 · Pryszo et al. · 2017 [cited by applicant]
US 9548053B1 · Basye et al. · 2017 [cited by applicant]
US 9584774B2 · Bekiares et al. · 2017 [cited by applicant]
US 9691378B1 · Meyers et al. · 2017 [cited by applicant]
US 9728188B1 · Rosen et al. · 2017 [cited by applicant]
US 10152966B1 · O'Malley et al. · 2018 [cited by applicant]
US 10313417B2 · Chen et al. · 2019 [cited by applicant]
US 10354651B1 · Yi et al. · 2019 [cited by applicant]
US 10395428B2 · Stafford et al. · 2019 [cited by applicant]
US 10477158B2 · Galvin et al. · 2019 [cited by applicant]
US 10489887B2 · El-Khamy et al. · 2019 [cited by applicant]
US 11373686B1 · Gilmour · 2022 [cited by examiner]
US 20070256105A1 · Tabe · 2007 [cited by applicant]
US 20080221882A1 · Bundock et al. · 2008 [cited by applicant]
US 20090060207A1 · Barry et al. · 2009 [cited by applicant]
US 20110228925A1 · Birch · 2011 [cited by applicant]
US 20120020490A1 · Leichter · 2012 [cited by applicant]
US 20120050012A1 · Alsina et al. · 2012 [cited by applicant]
US 20130044893A1 · Mauchly et al. · 2013 [cited by applicant]
US 20130266127A1 · Schachter et al. · 2013 [cited by applicant]
US 20140350926A1 · Schuster et al. · 2014 [cited by applicant]
US 20160127691A1 · Bokowski et al. · 2016 [cited by applicant]
US 20170084276A1 · Lebeau et al. · 2017 [cited by applicant]
US 20170256271A1 · Lyon et al. · 2017 [cited by applicant]
US 20190073090A1 · Parkinson et al. · 2019 [cited by applicant]
US 20190253611A1 · Wang et al. · 2019 [cited by applicant]
US 20190267010A1 · Li et al. · 2019 [cited by applicant]
US 20190307313A1 · Wade · 2019 [cited by applicant]
EP Communication received for European Application No. 21876499.1, mailed on Sep. 24, 2024, 1 page. [cited by applicant]
Extended European Search Report received for European Application No. 21876499.1, mailed on Sep. 5, 2024, 8 pages. [cited by applicant]
Cybulska, M., et al., “Structure of pauses in speech in the context of speaker verification and classification of speech type”, EURASIP Journal on Audio, Speech, and Music Processing, pp. 1-16 (2016). [cited by applicant]
International Preliminary Report on Patentability received for PCT Patent Application No. PCT/US2021/052935, mailed on Apr. 13, 2023, 12 pages. [cited by applicant]
International Search Report and Written Opinion received for PCT Patent Application No. PCT/US2021/052935, mailed on Jan. 10, 2022, 19 pages. [cited by applicant]
Patel, Z., “Image Segmentation Approach for Realizing Zoomable Streaming HEVC Video”, Master of Science in Electrical Engineering, p. 74 (May 2015). [cited by applicant]
Watkins, N., “A modest proposal to prevent false triggers on voice assistants”, Retrieved from Internet URL :https://towardsdatascience.com/a-modest-proposal-for-voice-assistants-91ee48ed1325, accessed on Jan. 25, 2021,… [cited by applicant]