IP Library Granted Patent US 11,513,767
Granted Patent B2
US 11,513,767 · App. 17/095,812 · Granted Nov 29, 2022

Method and system for recognizing a reproduced utterance

Inventor: Kanstantsin Yurievich Artsiom (Minsk, BY)
Assignee: YANDEX EUROPE AG
G06F3/167G10L15/22G10L15/26G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,513,767
App. No.
17/095,812
Granted
Nov 29, 2022
Kind
B2
Abstract

There is provided a method for operating a speaker device able to be activated by receiving and recognizing a predetermined wake up word. The method is executable at a server. The method comprises: capturing, by the speaker device, an audio signal having been generated in a vicinity of the speaker device; retrieving, by the speaker device, a processing filter, the processing filter being indicative of a pre-determined signal augmentation pattern representative of an excluded portion that has been excluded from an originating utterance having the wake up word, the originating utterance to be reproduced by an other electronic device; applying, by the speaker device, the processing filter to determine presence of the pre-determined signal augmentation pattern in the audio signal; based on determining the presence of the pre-determined signal augmentation pattern in the audio signal, determining that the audio signal has been produced by the other electronic device.

Claims (49)

1. A computer-implemented method for operating a speaker device, the speaker device being associated with a first operating mode and a second operating mode, the speaker device being further associated with a pre-determined wake up word, the pre-determined wake up word being configured, once recognized by the speaker device being in the first operating mode, to cause the speaker device to switch into the second operating mode;

the method being executable by the speaker device, the method comprising:

capturing, by the speaker device, an audio signal having been generated in a vicinity of the speaker device, the audio signal having been generated by one of a human user and an other electronic device;

retrieving, by the speaker device, a processing filter, the processing filter being indicative of a pre-determined signal augmentation pattern representative of an excluded portion that has been excluded from an originating utterance having the wake up word,

the originating utterance to be reproduced by the other electronic device;

applying, by the speaker device, the processing filter to determine presence of the pre-determined signal augmentation pattern in the audio signal;

in response to determining the presence of the pre-determined signal augmentation pattern in the audio signal, determining that the audio signal has been produced by the other electronic device.

2. The method of claim 1 , wherein the method further comprises, in response to determining the presence of the pre-determined signal augmentation pattern in the audio signal:

discarding the audio signal from further processing.

3. The method of claim 1 , wherein the method further comprises, in response to determining the presence of the pre-determined signal augmentation pattern in the audio signal:

executing a pre-determined additional action other than processing the audio signal to determine presence of the pre-determined wake up word therein.

4. The method of claim 1 , wherein in response to not determining the presence of the pre-determined signal augmentation pattern in the audio signal, the method further comprising:

determining that the audio signal has been produced by the human user;

applying, by the speaker device, a speech-to-text algorithm to the audio signal to generate a text representation thereof;

processing, by the speaker device, the text representation to determine presence of the wake up word therein;

in response to determining the presence of the wake up word, switching the speaker device onto the second operating mode.

5. The method of claim 1 , wherein the signal augmentation pattern is associated with one of pre-determined frequency levels.

6. The method of claim 5 , wherein the signal augmentation pattern is associated with a plurality of pre-determined frequency levels, the plurality of pre-determined frequency levels being selected from a spectrum recognizable by a human ear.

7. The method of claim 5 , wherein the plurality of pre-determined frequency levels are such that they are not divisible by each other.

8. The method of claim 5 , wherein the signal augmentation pattern is associated with a plurality of pre-determined frequency levels, the plurality of pre-determined frequency levels being selected from a spectrum recognizable by a human ear and being not divisible by each other.

9. The method of claim 5 , wherein the plurality of pre-determined frequency levels are randomly selected.

10. The method of claim 9 , wherein the plurality of pre-determined frequency levels are randomly pre-selected.

11. The method of claim 5 , wherein the plurality of pre-determined frequency levels comprises:

486 Hz

638 Hz

814 Hz

1355 Hz

2089 Hz

2635 Hz

3351 Hz

4510 Hz.

12. The method of claim 11 , wherein the pre-determined signal augmentation pattern is a first pre-determined signal augmentation pattern and wherein the method further comprises receiving an indication of a second pre-determined signal augmentation pattern, different from the first pre-determined signal augmentation pattern.

13. The method of claim 12 , wherein the second pre-determined signal augmentation pattern is for indicating a type of the other electronic device.

14. The method of claim 1 , wherein the first operating mode is associated with local speech to text processing, and the second operating mode is associated with a server-based speech to text processing.

15. The method of claim 1 , wherein the other electronic device is located in the vicinity of the speaker device.

16. The method of claim 1 , wherein exclusion of the excluded portion forms a sound gap when the originating utterance is reproduced by the other device, the sound gap being substantially un-recognizable by a human ear.

17. The method of claim 1 , wherein the applying the processing filter comprises first processing the audio signal into time-frequency representation thereof.

18. The method of claim 17 , wherein the processing the audio signal comprises applying a Fourier transformation.

19. The method of claim 18 , wherein the applying the Fourier transformation is executed by a stacked window approach.

20. The method of claim 19 , wherein the applying, by the speaker device, the processing filter is executed for each of the stacked windows.

21. The method of claim 20 , wherein the determining the presence of the pre-determined signal augmentation pattern in the audio signal is in response to the presence of the pre-determined signal augmentation pattern in at least one of the stacked windows.

22. The method of claim 1 , wherein the applying the processing filter to determine the presence of the pre-determined signal augmentation pattern in the audio signal comprises determining energy levels in a plurality of pre-determined frequencies where sound has been filtered out.

23. The method of claim 22 , wherein the determining energy levels comprises comparing an energy level at a given one of the plurality of pre-determined frequencies where sound has been filtered out to an energy level in an adjacent frequency where the sound has not been filtered out.

24. The method of claim 23 , wherein the presence of the pre-determined signal augmentation pattern is determined in response to a difference between energy levels being above a pre-determined threshold.

25. A computer-implemented method for generating an audio feed for transmitting to an electronic device for audio-processing thereof, the audio feed having a content that includes a pre-determined wake up word, the pre-determined wake up word being configured, once recognized by a speaker device being in a first operating mode, to cause the speaker device to switch into a second operating mode from the first operating mode, the method being executable by a production server, the method comprising:

receiving, by the production server, the audio feed having the content, the audio feed having been pre-recorded;

retrieving, by the production server, a processing filter, the processing filter being indicative of a pre-determined signal augmentation pattern representative of an excluded portion to be excluded from the audio feed to indicate to the speaker device to ignore the wake up word contained in the content;

excluding, by the production server, the excluded portion from the audio feed thereby forming a sound gap when the audio feed is reproduced by the electronic device;

causing transmission of the audio feed to the electronic device.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 9, 2024
From: DIRECT CURSUS TECHNOLOGY L.L.C
To: Y.E. HUB ARMENIA LLC
Reel/Frame 068534/0619 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 15, 2023
From: YANDEX EUROPE AG
To: DIRECT CURSUS TECHNOLOGY L.L.C
Reel/Frame 065692/0720 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 21, 2022
From: ARTSIOM, KANSTANTSIN YURIEVICH
To: YANDEXBEL LLC
Reel/Frame 061501/0152 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 21, 2022
From: YANDEXBEL LLC
To: YANDEX LLC
Reel/Frame 061501/0313 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 21, 2022
From: YANDEX LLC
To: YANDEX EUROPE AG
Reel/Frame 061501/0792 →