IP Library Granted Patent US 11,282,526
Granted Patent B2
US 11,282,526 · App. 16/852,383 · Granted Mar 22, 2022

Methods and systems for processing audio signals containing speech data

Inventor: Patricia Scanlon (Dublin, IE)
Assignee: SoapBox Labs Ltd.
G10L17/04G06F16/61G06F16/636G06Q30/0185G06Q50/265G10L15/25G10L17/10G06F21/32
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,282,526
App. No.
16/852,383
Granted
Mar 22, 2022
Kind
B2
Abstract

Methods and systems for processing audio signals containing speech data are disclosed. Biometric data associated with at least one speaker are extracted from an audio input. A match is determined between the extracted biometric data and stored biometric data associated with a consenting user profile, where a consenting user profile is a user profile associated with a record indicating consent to store biometric data. If a match is determined to exist with such a profile, the speech data is stored in an archive after processing. If no such match is determined, or if the extracted biometric data includes data from a speaker not having a consenting user profile, the speech data is discarded, optionally after having been processed. The system and method provides a safeguard against transferring to storage data of users, particularly minors or children, for whom a verified and valid consent has not been obtained from an authorised adult.

Claims (68)

1. A method comprising:

storing one or more user profiles that are each associated with one of one or more users of a computing system, wherein each user profile is associated with a voiceprint that was generated to uniquely characterize voice characteristics of a respective user of the one or more users of the computing system, wherein at least one of the stored one or more user profiles is a consenting user profile of a child, wherein each consenting user profile of a child is associated with a record indicating consent by a parent of a child to store biometric data of a child;

processing an audio signal containing speech data received from a plurality of speakers at the computing system, wherein the plurality of speakers comprise a first child speaker and a second child speaker, and wherein the processing of the audio signal comprises extracting first biometric data associated with the first child speaker and second biometric data associated with the second child speaker;

determining whether any of the first extracted biometric data associated with the first child speaker or the second extracted biometric data associated with the second child speaker corresponds to a voiceprint associated with a respective consenting user profile associated with a record indicating consent by a parent of a respective child to store biometric data of the respective child;

responsive to determining that the first extracted biometric data associated with the first child speaker corresponds to the voiceprint associated with the respective consenting user profile associated with the record indicating consent by the parent of the respective child to store the biometric data of the respective child, performing at least one of:

(i) processing speech data from the first child speaker; or

(ii) storing the speech data from the first child speaker in an archive;

responsive to determining that the second extracted biometric data associated with the second child speaker does not correspond to any voiceprint associated with any consenting user profile:

processing speech data from the second child speaker by identifying a command of the second child speaker using speech recognition operations and executing the command to respond to the command of the second child speaker; and

after processing the speech data from the second child speaker, deleting the speech data from the second child speaker within a predetermined time period.

2. The method of claim 1 , wherein responsive to determining that the second extracted biometric data associated with the second speaker does not correspond to any voiceprint associated with any consenting user profile, the speech data from the second speaker is deleted without being processed further.

3. The method of claim 2 , wherein said predetermined time period is immediately after determining that the second extracted biometric data associated with the second speaker does not correspond to any voiceprint associated with any consenting user profile.

4. The method of claim 1 , wherein responsive to determining that the second extracted biometric data associated with the second speaker does not correspond to any voiceprint associated with any consenting user profile, the speech data from the second speaker is processed before being deleted within said predetermined time period.

5. The method of claim 4 , wherein said predetermined time period is immediately after processing the speech data from the second speaker.

6. The method of claim 4 , wherein the speech data from the second speaker is processed locally at the computing system and is not transmitted to a remote location.

7. The method of claim 1 , further comprising creating a consenting user profile, wherein creating a consenting user profile comprises:

verifying credentials of a first user of the computing system against a data source to ensure that the first user is authorized to provide consent to store speech data;

initializing a user profile associated with a second user, on instruction of the first user, wherein the second user is a child and the first user is a parent of the second user;

receiving speech data of the second user;

extracting biometric data from said speech data of the second user;

storing said biometric data from said speech data of the second user and associating said biometric data from said speech data of the second user with said user profile associated with the second user; and

storing said user profile associated with the second user as a consenting user profile.

8. The method of claim 7 , further comprising matching additional biometric data acquired from the first child speaker or the second child speaker against stored biometric data associated with a respective consenting user profile, and storing non-speech biometric data during profile creation.

9. The method of claim 1 , further comprising matching additional biometric data acquired from the first child speaker or the second child speaker against stored biometric data associated with a respective consenting user profile.

10. The method of claim 9 , wherein said additional biometric data is selected from:

a. image data of the face of the first child speaker or the second child speaker;

b. iris pattern data;

c. fingerprint data;

d. hand geometry data;

e. palm blood vessel pattern data;

f. retinal blood vessel pattern data;

g. mouth movement data; or

h. behavioural data.

11. The method of claim 1 , wherein determining whether the first extracted biometric data associated with the first child speaker or the second extracted biometric data associated the second child speaker corresponds to a voiceprint associated with a respective consenting user profile comprises determining a match against a user profile of a logged-in user.

12. The method of claim 11 , wherein said logged-in user is one of said one or more users of said computing system, and wherein said logged-in user has been logged into said computing system in response to detection of biometric data associated with the logged-in user.

13. The method of claim 1 , wherein determining whether any of the first extracted biometric data associated with the first child speaker or the second extracted biometric data associated with the second child speaker corresponds to the voiceprint associated with the respective consenting user profile associated with the record indicating consent by the parent of the respective child to store biometric data of the respective child comprises:

determining a match for any of the first extracted biometric data associated with the first child speaker or the second extracted biometric data associated with the second child speaker against both consenting user profiles and non-consenting user profiles, wherein a non-consenting user profile is a user profile not associated with a record indicating consent to store biometric data.

14. The method of claim 13 , further comprising creating a non-consenting user profile, wherein creating a non-consenting user profile comprises:

initializing a third user profile associated with a third user;

receiving speech data of the third user;

extracting third biometric data from said speech data of the third user;

storing said third biometric data and associating said third biometric data with said third user profile; and

storing said third user profile as a non-consenting user profile.

15. The method of claim 1 , further comprising updating a voiceprint associated with a respective consenting user profile based on the first extracted biometric data associated with the first child speaker or the second extracted biometric data associated with the second child speaker.

16. A computing system comprising:

an audio input;

a visual input;

a data store storing one or more user profiles that are each associated with one of one or more users of a computing system, wherein each user profile is associated with a voiceprint that was generated to uniquely characterize voice characteristics of a respective user of the one or more users of the computing system, wherein at least one of the stored one or more user profiles is a consenting user profile of a child, wherein each consenting user profile of a child is associated with a record indicating consent by a parent of a child to store biometric data of a child;

an interface to a storage archive storing speech data; and

a hardware processor, coupled to the audio input, the visual input and the data store, to:

process an audio signal received via the audio input and containing speech data received from a plurality of speakers, wherein the plurality of speakers comprise a first child speaker and a second child speaker, and wherein the audio signal is processed to extract first biometric data associated with the first child speaker and second biometric data associated with the second child speaker;

determine whether any of the first extracted biometric data associated with the first child speaker or the second extracted biometric data associated with the second child speaker corresponds to a voiceprint associated with a respective consenting user profile associated with a record indicating consent by a parent of a respective child to store biometric data of the respective child;

responsive to determining that the first extracted biometric data associated with the first child speaker corresponds to the voiceprint associated with the respective consenting user profile associated with the record indicating consent by the parent of the respective child to store the biometric data of the respective child, perform at least one of:

(i) processing speech data from the first child speaker; or

(ii) storing the speech data from the first child speaker in an archive;

responsive to determining that the second extracted biometric data associated with the second child speaker does not correspond to any voiceprint associated with any consenting user profile:

process speech data from the second child speaker by identifying a command of the second child speaker using speech recognition operations and executing the command to respond to the command of the second child speaker; and

after processing the speech data from the second child speaker, delete the speech data from the second child speaker within a predetermined time period.

17. A non-transitory computer-readable medium comprising instructions, which when executed by a processor, cause the processor to perform operations comprising:

storing one or more user profiles that are each associated with one of one or more users of a computing system, wherein each user profile is associated with a voiceprint that was generated to uniquely characterize voice characteristics of a respective user of the one or more users of the computing system, wherein at least one of the stored one or more user profiles is a consenting user profile of a child, wherein each consenting user profile of a child is associated with a record indicating consent by a parent of a child to store biometric data of a child;

processing an audio signal containing speech data received from a plurality of speakers at the computing system, wherein the plurality of speakers comprise a first child speaker and a second child speaker, and wherein the processing of the audio signal comprises extracting first biometric data associated with the first child speaker and second biometric data associated with the second child speaker;

determining whether any of the first extracted biometric data associated with the first child speaker or the second extracted biometric data associated with the second child speaker corresponds to a voiceprint associated with a respective consenting user profile associated with a record indicating consent by a parent of a respective child to store biometric data of the respective child;

responsive to determining that the first extracted biometric data associated with the first child speaker corresponds to the voiceprint associated with the respective consenting user profile associated with the record indicating consent by the parent of the respective child to store the biometric data of the respective child, performing at least one of:

(i) processing speech data from the first child speaker; or

(ii) storing the speech data from the first child speaker in an archive;

responsive to determining that the second extracted biometric data associated with the second child speaker does not correspond to any voiceprint associated with any consenting user profile:

processing speech data from the second child speaker by identifying a command of the second child speaker using speech recognition operations and executing the command to respond to the command of the second child speaker; and

after processing the speech data from the second child speaker, deleting the speech data from the second child speaker within a predetermined time period.

Assignments (2)
PATENT SECURITY AGREEMENT Recorded May 30, 2024
From: SOAPBOX LABS LIMITED
To: GOLDMAN SACHS BANK USA
Reel/Frame 067577/0354 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 22, 2020
From: SCANLON, PATRICIA
To: SOAPBOX LABS LTD.
Reel/Frame 052738/0809 →
Priority Claims (1)
EP 17197187 · Oct 18, 2017 · regional
Continuity (2)
Continuation PCTEP2018078470 · Oct 18, 2018
Related Publication 20200286490A1 · Sep 10, 2020