IP Library Granted Patent US 12,475,894
Granted Patent B2
US 12,475,894 · App. 18/217,411 · Granted Nov 18, 2025

Methods and systems for processing audio signals containing speech data

Inventor: Patricia Scanlon (Dublin, IE)
Assignee: SoapBox Labs Ltd.
G10L17/04G06F16/61G06F16/636G06Q30/0185G06Q50/265G10L15/25G10L17/10G06F21/32
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,475,894
App. No.
18/217,411
Granted
Nov 18, 2025
Kind
B2
Abstract

Methods and systems for processing audio signals containing speech data are disclosed. Biometric data associated with at least one speaker are extracted from an audio input. A correspondence is determined between the extracted biometric data and stored biometric data associated with a consenting user profile, where a consenting user profile is a user profile indicates consent to store biometric data. If no correspondence is determined, the speech data is discarded, optionally after having been processed.

Claims (71)

1 . A method comprising:

storing one or more user profiles that are each associated with one of one or more users of a computing system, wherein each user profile is associated with biometric data that uniquely characterizes a respective user of the one or more users of the computing system, and wherein at least one of the one or more user profiles is a consenting user profile, wherein each consenting user profile of a respective user indicates consent to store biometric data of the respective user;

determining first biometric data associated with a first speaker based at least in part on speech content derived from first image data of the first speaker;

determining whether the first biometric data associated with the first speaker corresponds to any biometric data associated with any consenting user profile indicating consent to store biometric data of a respective user; and

responsive to determining that the first biometric data associated with the first speaker does not correspond to any voiceprint associated with any consenting user profile:

deleting the first image data within a time period.

2 . The method of claim 1 further comprising:

processing second image data of a second speaker, wherein processing the second image data of the second speaker comprises extracting, from the second image data of the second speaker, second biometric data associated with the second speaker;

determining whether the second biometric data associated with the second speaker corresponds to any biometric data associated with any consenting user profile indicating consent to store biometric data of a respective user; and

responsive to determining that the second biometric data extracted from the second image data of the second speaker corresponds to a voiceprint associated with a consenting user profile, performing at least one of:

processing the second image data; or

storing the second image data in an archive.

3 . The method of claim 1 , wherein determining the first biometric data associated with the first speaker comprises analysing the first image data to determine the speech content based on movements of one or more facial features of the first speaker.

4 . The method of claim 3 , wherein the one or more facial features comprise one or more of a mouth, lips or jaw.

5 . The method of claim 1 , wherein determining the first biometric data associated with the first speaker comprises extracting, from the first image data of the first speaker, image-based biometric data associated with the first speaker.

6 . The method of claim 5 , wherein determining the first biometric data associated with the first speaker is further based on first speech data of the first speaker.

7 . The method of claim 6 , wherein determining the first biometric data associated with the first speaker further comprises extracting, from the first speech data of the first speaker, speech-based biometric data associated with the first speaker.

8 . The method of claim 1 , further comprising creating a consenting user profile, wherein creating a consenting user profile comprises:

verifying credentials of a first user of the computing system against a data source to ensure that the first user is authorized to provide consent to store speech data;

initializing a user profile associated with a second user, on instruction of the first user, wherein the second user is a child and the first user is a parent of the second user;

receiving speech data of the second user;

extracting biometric data from the speech data of the second user;

storing the biometric data from the speech data of the second user and associating the biometric data from the speech data of the second user with the user profile associated with the second user; and

storing the user profile associated with the second user as a consenting user profile.

9 . The method of claim 1 , further comprising matching additional biometric data acquired from the first speaker against stored biometric data associated with a respective consenting user profile.

10 . The method of claim 9 , wherein the additional biometric data is selected from:

a. image data of a face of the first speaker;

b. iris pattern data;

c. fingerprint data;

d. hand geometry data;

e. palm blood vessel pattern data;

f. retinal blood vessel pattern data;

g. mouth movement data; or

h. behavioural data.

11 . The method of claim 1 , further comprising creating a non-consenting user profile, wherein creating a non-consenting user profile comprises:

initializing a second user profile associated with a second user;

receiving speech data of the second user;

extracting second biometric data from the speech data of the second user;

storing the second biometric data and associating the second biometric data with the second user profile; and

storing the second user profile as a non-consenting user profile.

12 . A non-transitory computer-readable medium comprising instructions, which when executed by a processor, cause the processor to perform operations comprising:

storing one or more user profiles that are each associated with one of one or more users of a computing system, wherein each user profile is associated with biometric data that uniquely characterizes a respective user of the one or more users of the computing system, and wherein at least one of the one or more user profiles is a consenting user profile, wherein each consenting user profile of a respective user indicates consent to store biometric data of the respective user;

determining first biometric data associated with a first speaker based at least in part on speech content derived from first image data of the first speaker;

determining whether the first biometric data associated with the first speaker corresponds to any biometric data associated with any consenting user profile indicating consent to store biometric data of a respective user; and

responsive to determining that the first biometric data associated with the first speaker does not correspond to any voiceprint associated with any consenting user profile:

deleting the first image data within a time period.

13 . The non-transitory computer-readable medium of claim 12 , the operations further comprising:

processing second image data of a second speaker, wherein processing the second image data of the second speaker comprises extracting, from the second image data of the second speaker, second biometric data associated with the second speaker;

determining whether the second biometric data associated with the second speaker corresponds to any biometric data associated with any consenting user profile indicating consent to store biometric data of a respective user; and

responsive to determining that the second biometric data extracted from the second image data of the second speaker corresponds to a voiceprint associated with a consenting user profile, performing at least one of:

processing the second image data; or

storing the second image data in an archive.

14 . The non-transitory computer-readable medium of claim 12 , wherein determining the first biometric data associated with the first speaker comprises analysing the first image data to determine the speech content based on movements of one or more facial features of the first speaker, wherein the one or more facial features comprise one or more of a mouth, lips or jaw.

15 . The non-transitory computer-readable medium of claim 12 , wherein determining the first biometric data associated with the first speaker comprises extracting, from the first image data of the first speaker, image-based biometric data associated with the first speaker.

16 . The non-transitory computer-readable medium of claim 15 , wherein determining the first biometric data associated with the first speaker is further based on first speech data of the first speaker.

17 . The non-transitory computer-readable medium of claim 16 , wherein determining the first biometric data associated with the first speaker further comprises extracting, from the first speech data of the first speaker, speech-based biometric data associated with the first speaker.

18 . A system comprising:

a memory; and

a processing device coupled to the memory, to perform operations comprising:

storing one or more user profiles that are each associated with one of one or more users of a computing system, wherein each user profile is associated with biometric data that uniquely characterizes a respective user of the one or more users of the computing system, and wherein at least one of the one or more user profiles is a consenting user profile, wherein each consenting user profile of a respective user indicates consent to store biometric data of the respective user;

determining first biometric data associated with a first speaker based at least in part on speech content derived from first image data of the first speaker;

determining whether the first biometric data associated with the first speaker corresponds to any biometric data associated with any consenting user profile indicating consent to store biometric data of a respective user; and

responsive to determining that the first biometric data associated with the first speaker does not correspond to any voiceprint associated with any consenting user profile:

deleting the first image data within a time period.

19 . The system of claim 18 , the operations further comprising:

processing second image data of a second speaker, wherein processing the second image data of the second speaker comprises extracting, from the second image data of the second speaker, second biometric data associated with the second speaker;

determining whether the second biometric data associated with the second speaker corresponds to any biometric data associated with any consenting user profile indicating consent to store biometric data of a respective user; and

responsive to determining that the second biometric data extracted from the second image data of the second speaker corresponds to a voiceprint associated with a consenting user profile, performing at least one of:

processing the second image data; or

storing the second image data in an archive.

20 . The system of claim 18 , wherein determining the first biometric data associated with the first speaker comprises analysing the first image data to determine the speech content based on movements of one or more facial features of the first speaker, wherein the one or more facial features comprise one or more of a mouth, lips or jaw.

Assignments (2)
PATENT SECURITY AGREEMENT Recorded May 30, 2024
From: SOAPBOX LABS LIMITED
To: GOLDMAN SACHS BANK USA
Reel/Frame 067577/0354 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 14, 2024
From: SCANLON, PATRICIA
To: SOAPBOX LABS LTD.
Reel/Frame 066773/0731 →
Priority Claims (1)
EP 17197187 · Oct 18, 2017 · regional
Continuity (4)
Continuation 17700369 · Mar 21, 2022
Continuation 16852383 · Apr 17, 2020
Continuation PCTEP2018078470 · Oct 18, 2018
Related Publication 20240185860A1 · Jun 6, 2024
References Cited (110)
US 4091242A · Carrubba et al. · 1978 [cited by applicant]
US 5029214A · Hollander · 1991 [cited by applicant]
US 5623539A · Bassenyemukasa et al. · 1997 [cited by applicant]
US 5832100A · Lawton et al. · 1998 [cited by applicant]
US 5873061A · Hab-Umbach et al. · 1999 [cited by applicant]
US 5991429A · Coffin · 1999 [cited by examiner]
US 6067521A · Ishii et al. · 2000 [cited by applicant]
US 6219639B1 · Bakis et al. · 2001 [cited by applicant]
US 6314401B1 · Abbe et al. · 2001 [cited by applicant]
US 6554705B1 · Cumbers · 2003 [cited by applicant]
US 6567775B1 · Maali et al. · 2003 [cited by applicant]
US 6826306B1 · Lewis et al. · 2004 [cited by applicant]
US 7343553B1 · Kaye · 2008 [cited by applicant]
US 8054969B2 · Adhikari et al. · 2011 [cited by applicant]
US 8463488B1 · Hart · 2013 [cited by applicant]
US 8688306B1 · Nemec et al. · 2014 [cited by applicant]
US 8700392B1 · Hart et al. · 2014 [cited by applicant]
US 9070367B1 · Hoffmeister et al. · 2015 [cited by applicant]
US 9443514B1 · Taubman · 2016 [cited by applicant]
US 9881613B2 · Weinstein · 2018 [cited by applicant]
US 10178301B1 · Welbourne et al. · 2019 [cited by applicant]
US 11138334B1 · Garrod et al. · 2021 [cited by applicant]
US 20020007278A1 · Traynor · 2002 [cited by applicant]
US 20020010588A1 · Fujimori · 2002 [cited by applicant]
US 20020111809A1 · McIntosh · 2002 [cited by applicant]
US 20020116197A1 · Erten · 2002 [cited by applicant]
US 20020120866A1 · Mitchell et al. · 2002 [cited by applicant]
US 20030097353A1 · Gutta et al. · 2003 [cited by applicant]
US 20040003142A1 · Yokota et al. · 2004 [cited by applicant]
US 20050060412A1 · Chebolu · 2005 [cited by examiner]
US 20050096926A1 · Eaton et al. · 2005 [cited by applicant]
US 20050185779A1 · Toms · 2005 [cited by applicant]
US 20050193093A1 · Mathew et al. · 2005 [cited by applicant]
US 20050240582A1 · Hatonen et al. · 2005 [cited by applicant]
US 20060173793A1 · Glass · 2006 [cited by applicant]
US 20060222210A1 · Sundaram · 2006 [cited by examiner]
US 20060259305A1 · Pietruszka · 2006 [cited by applicant]
US 20070211921A1 · Popp et al. · 2007 [cited by applicant]
US 20070218955A1 · Cook et al. · 2007 [cited by applicant]
US 20070255564A1 · Yee · 2007 [cited by examiner]
US 20080004876A1 · He et al. · 2008 [cited by applicant]
US 20080235162A1 · Spring · 2008 [cited by applicant]
US 20080269958A1 · Filev · 2008 [cited by examiner]
US 20090122198A1 · Thorn · 2009 [cited by applicant]
US 20090173786A1 · Hatkoff · 2009 [cited by applicant]
US 20090185723A1 · Kurtz et al. · 2009 [cited by applicant]
US 20100088096A1 · Parsons · 2010 [cited by applicant]
US 20100233660A1 · Skala et al. · 2010 [cited by applicant]
US 20110235870A1 · Ichikawa · 2011 [cited by examiner]
US 20110243449A1 · Hannuksela et al. · 2011 [cited by applicant]
US 20120047560A1 · Underwood et al. · 2012 [cited by applicant]
US 20120253811A1 · Breslin et al. · 2012 [cited by applicant]
US 20120297284A1 · Matthews et al. · 2012 [cited by applicant]
US 20120323694A1 · Lita et al. · 2012 [cited by applicant]
US 20130162752A1 · Herz et al. · 2013 [cited by applicant]
US 20130185220A1 · Good et al. · 2013 [cited by applicant]
US 20130204607A1 · Forrest · 2013 [cited by applicant]
US 20130218566A1 · Qian et al. · 2013 [cited by applicant]
US 20130317817A1 · Ganong et al. · 2013 [cited by applicant]
US 20140004826A1 · Addy et al. · 2014 [cited by applicant]
US 20140136215A1 · Dai et al. · 2014 [cited by applicant]
US 20140254778A1 · Zeppenfeld · 2014 [cited by examiner]
US 20140278366A1 · Jacob et al. · 2014 [cited by applicant]
US 20140288932A1 · Yeracaris et al. · 2014 [cited by applicant]
US 20140337949A1 · Hoyos · 2014 [cited by applicant]
US 20140376786A1 · Johnson et al. · 2014 [cited by applicant]
US 20150073810A1 · Nishio et al. · 2015 [cited by applicant]
US 20150088515A1 · Beaumont et al. · 2015 [cited by applicant]
US 20150149174A1 · Gollan et al. · 2015 [cited by applicant]
US 20150194155A1 · Tsujikawa et al. · 2015 [cited by applicant]
US 20150220716A1 · Aronowitz et al. · 2015 [cited by applicant]
US 20150269943A1 · VanBlon et al. · 2015 [cited by applicant]
US 20150278922A1 · Isaacson et al. · 2015 [cited by applicant]
US 20150301796A1 · Visser et al. · 2015 [cited by applicant]
US 20150332041A1 · Maruyama · 2015 [cited by applicant]
US 20150340025A1 · Shima · 2015 [cited by applicant]
US 20150371637A1 · Neubacher et al. · 2015 [cited by applicant]
US 20160019372A1 · Clark · 2016 [cited by applicant]
US 20160063235A1 · Tussy · 2016 [cited by applicant]
US 20160063314A1 · Samet · 2016 [cited by applicant]
US 20160203699A1 · Mulhern et al. · 2016 [cited by applicant]
US 20160225048A1 · Zoldi et al. · 2016 [cited by applicant]
US 20160234024A1 · Mozer · 2016 [cited by examiner]
US 20160361663A1 · Watry · 2016 [cited by applicant]
US 20170061959A1 · Lehman et al. · 2017 [cited by applicant]
US 20170069327A1 · Heigold · 2017 [cited by examiner]
US 20170133018A1 · Parker · 2017 [cited by examiner]
US 20180129795A1 · Katz-Oz · 2018 [cited by examiner]
US 20180144742A1 · Ye et al. · 2018 [cited by applicant]
US 20180204209A1 · Kohli · 2018 [cited by examiner]
US 20180204577A1 · DeMerchant et al. · 2018 [cited by applicant]
US 20180210874A1 · Fuxman et al. · 2018 [cited by applicant]
US 20180349684A1 · Bapat · 2018 [cited by examiner]
US 20180366125A1 · Liu et al. · 2018 [cited by applicant]
US 20190095867A1 · Nishijima et al. · 2019 [cited by applicant]
US 20190213278A1 · Min · 2019 [cited by applicant]
US 20190220662A1 · Dawoud · 2019 [cited by applicant]
US 20190259388A1 · Robley et al. · 2019 [cited by applicant]
US 20190378171A1 · Bhat et al. · 2019 [cited by applicant]
US 20200110443A1 · Leong · 2020 [cited by examiner]
US 20200335098A1 · Saito · 2020 [cited by applicant]
US 20210266734A1 · Kapinos et al. · 2021 [cited by applicant]
JP 2009237256A · 2009 [cited by applicant]
JP 2020030739A · 2020 [cited by applicant]
WO 2005042314A1 · 2005 [cited by applicant]
WO 2015161240A2 · 2015 [cited by applicant]
Kepuska, V.,Z., & Rojanasthien, P., “Speech corpus generation from DVDs of movies and TV series”, 2011, Journal of International Technology and Information Management, 20(1), 49-82 (Year: 2011). [cited by applicant]
Google Patents Translation of JP-2020030739-A, https://patents.google.com/patent/JP2020030739A/en?oq=JP-2020030739-A (Year: 2020). [cited by applicant]
Google Patents Translation of JP-2009237256-A, https://patents.google.com/patent/JP2009237256A/en?oq=JP-2009237256-A (Year: 2009). [cited by applicant]
PCT International Search Report and Written Opinion for International Application No. PCT/US2018/078470 mailed Apr. 12, 2019, 17 pages. [cited by applicant]