IP Library Granted Patent US 12,229,313
Granted Patent B1
US 12,229,313 · App. 18/778,183 · Granted Feb 18, 2025

Systems and methods for analyzing speech data to remove sensitive data

Inventors: Tejas Shastry (Chicago, IL); Anthony Tassone (St. Charles, IL); Corey Burrows (Lisle, IL)
Assignee: Truleo, Inc.
G06F21/6245G06V20/41G10L17/02G10L17/04G10L25/57G11B27/031
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,229,313
App. No.
18/778,183
Granted
Feb 18, 2025
Kind
B1
Abstract

According to an embodiment, a method includes receiving audio data and providing the audio data as input to a first machine learning model to produce transcription data. The audio data is provided as input to a second machine learning model to produce speaker separation data, and the transcription data is segmented based on the speaker separation data to produce speaker separated transcription data. A portion of the speaker separated transcription data is provided as input to a third machine learning model to identify personal identifiable information (PII) text in the portion of the speaker separated transcription data, the portion being associated with a speaker from a plurality of speakers. The method also includes replacing the PII text with redaction text in the portion of the speaker separated transcription data and causing display of the transcription data including the redaction text at a user compute device.

Claims (72)

1. An apparatus, comprising:

a processor; and

a memory operably coupled to the processor, the memory storing instructions to cause the processor to:

receive video data that includes audio data,

provide the audio data as input to a first machine learning model to produce transcription data,

provide the audio data as input to a second machine learning model to produce speaker separation data,

segment the transcription data based on the speaker separation data to produce speaker separated transcription data,

provide a first portion of the speaker separated transcription data from a plurality of portions of the speaker separated transcription data as input to a context window of a third machine learning model to produce an attribute indication, the first portion (1) being associated with a first speaker from a plurality of speakers and (2) not being associated with remaining speakers from the plurality of speakers,

in response to producing the attribute indication, cause display of video data at a user compute device, the video data being associated with the first portion of the speaker separated transcription data and not remaining portions of the speaker separated transcription data from the plurality of portions of the speaker separated transcription data,

provide at least one of the first portion of the speaker separated transcription data or a second portion of the speaker separated transcription data as input to a fourth machine learning model to identify personal identifiable information (PII) text in the at least one of the first portion of the speaker separated transcription data or the second portion of the speaker separated transcription data, the second portion being associated with a second speaker from the plurality of speakers, and

in response to identifying the PII text:

replace the PII text with redaction text in the at least one of the first portion of the speaker separated transcription data or the second portion of the speaker separated transcription data, and

cause display of the transcription data including the redaction text at the user compute device.

2. The apparatus of claim 1 , wherein:

the first speaker is a peace officer; and

the second speaker is a civilian.

3. The apparatus of claim 1 , further comprising an interface configured to communicate with a body camera, the video data being received from the body camera.

4. The apparatus of claim 1 , wherein:

the third machine learning model includes (1) a first large language model (LLM) configured to receive the first portion of the speaker separated transcription data as input to produce a speaker specific attribute indication and (2) a second LLM configured to receive the first portion of the speaker separated transcription data and the second portion of the speaker separated transcription data as input to produce a conversation attribute.

5. The apparatus of claim 1 , wherein the memory further stores instructions to cause the processor to:

receive a confirmation indication from the user compute device in response to the processor causing display of the video data; and

generate a label for the video data, the label being associated with the attribute indication.

6. The apparatus of claim 5 , wherein the memory further stores instructions to cause the processor to cause a training signal to be sent to retrain the third machine learning model based on the confirmation indication.

7. The apparatus of claim 1 , wherein the memory further stores instructions to cause the processor to:

receive a rejection indication from the user compute device in response to the processor causing display of the video data; and

cause a training signal to be sent to retrain the third machine learning model based on the rejection indication.

8. The apparatus of claim 1 , wherein the memory further stores instructions to cause the processor to:

cause display of the first portion of the speaker separated transcription data at the user compute device;

in response to producing the attribute indication, cause display of a tag that (1) is associated with the attribute indication and (2) highlights the first portion of the speaker separated transcription data at the user compute device;

cause display of an identifier associated with the first speaker; and

cause display of a selectable element associated with at least one of a confirmation indication or a rejection indication.

9. The apparatus of claim 1 , wherein the instructions to cause the processor to cause display of the video data at the user compute device include instructions to cause the processor to add the video data to a selectable queue that is associated with the attribute indication.

10. A method, comprising:

receiving, via a processor, audio data;

providing, via the processor, the audio data as input to a first machine learning model to produce transcription data;

providing, via the processor, the audio data as input to a second machine learning model to produce speaker separation data;

segmenting, via the processor, the transcription data based on the speaker separation data to produce speaker separated transcription data;

providing a first portion of the speaker separated transcription data from a plurality of portions of the speaker separated transcription data as input to a context window of a third machine learning model to produce an attribute indication, the first portion (1) being associated with a first speaker from a plurality of speakers and (2) not being associated with remaining speakers from the plurality of speakers;

in response to producing the attribute indication, causing display of video data at a user compute device, the video data being associated with the first portion of the speaker separated transcription data and not the remaining portions of the speaker separated transcription data;

providing, via the processor, at least the first portion of the speaker separated transcription data or a second portion of the speaker separated transcription data as input to a fourth machine learning model to identify personal identifiable information (PII) text in the at least one of the first portion of the speaker separated transcription data or the second portion of the speaker separated transcription data, the second portion being associated with a second speaker from the plurality of speakers;

replacing, via the processor, the PII text with redaction text in the at least one of the first portion of the speaker separated transcription data or the second portion of the speaker separated transcription data; and

causing, via the processor, display of the transcription data including the redaction text at the user compute device.

11. The method of claim 10 , wherein the redaction text is first redaction text, the method further comprising:

receiving, via the processor and from the user compute device, a redaction request associated with a portion of the transcription data; and

replacing, via the processor, the portion of the transcription data with second redaction text.

12. The method of claim 10 , wherein the causing display of the transcription data including the redaction text at the user compute device includes causing display of the transcription data such that the user compute device is unable to revert the redaction text to view the PII text via the user compute device.

13. The method of claim 10 , wherein the providing the at least one of the first portion of the speaker separated transcription data or the second portion of the speaker separated transcription data as input to the fourth machine learning model includes:

generating, via the processor, a plurality of segments based on the at least one of the first portion of the speaker separated transcription data or the second portion of the speaker separated transcription data; and

iteratively providing, via the processor, each segment from the plurality of segments as input to a context window of the fourth machine learning model to identify the PII text.

14. The method of claim 10 , wherein the audio data is included in video data that is generated by a body worn camera.

15. A method, comprising:

receiving, via a processor, audio data;

providing, via the processor, the audio data as input to a first machine learning model to produce transcription data;

providing, via the processor, the audio data as input to a second machine learning model to produce speaker separation data;

segmenting, via the processor, the transcription data based on the speaker separation data to produce speaker separated transcription data;

providing, via the processor, a first portion of the speaker separated transcription data from a plurality of portions of the speaker separated transcription data as input to a third machine learning model to produce an attribute indication, the first portion of the speaker separated transcription data (1) being attributable to a law enforcement officer from a plurality of speakers and (2) not being associated with remaining speakers from the plurality of speakers;

causing, via the processor, display of video data at a user compute device, the video data being associated with the first portion of the speaker separated transcription data and not remaining portions of the speaker separated transcription data from the plurality of portions of the speaker separated transcription data;

providing, via the processor, at least the first portion of the speaker separated transcription data or a second portion of the speaker separated transcription data as input to a fourth machine learning model to identify personal identifiable information (PII) text in the at least one of the first portion of the speaker separated transcription data or the second portion of the speaker separated transcription data, the second portion being attributable to a community member;

replacing, via the processor, the PII text with redaction text in the at least one of the first portion of the speaker separated transcription data or the second portion of the speaker separated transcription data; and

causing, via the processor, display of the transcription data including the redaction text at the user compute device.

16. The method of claim 15 , wherein the attribute indication is associated with a speaker specific attribute.

17. The method of claim 15 , further comprising:

receiving, via the processor and from the user compute device, an invalidation request;

deleting, via the processor, the attribute indication in response to the receiving the invalidation request; and

generating, via the processor, audit data that includes an indication of the invalidation request.

18. The method of claim 15 , further comprising:

receiving, via the processor, a confirmation indication from the user compute device in response to the causing display of the video data; and

associating, via the processor, a label with the video data, the label being associated with the attribute indication.

19. The method of claim 15 , further comprising causing, via the processor, a training signal to be sent to retrain the third machine learning model based on a user confirmation of the attribute indication.

20. The method of claim 15 , wherein:

the providing the first portion of the speaker separated transcription data as input to the third machine learning model includes providing the second portion of the speaker separated transcription data as input to the third machine learning model to produce the attribute indication, the attribute indication being a conversation attribute indication.

21. The method of claim 20 , wherein the third machine learning model includes (1) a first large language model (LLM) configured to receive the first portion of the speaker separated transcription data as input to produce a speaker specific attribute indication and (2) a second LLM configured to receive the second portion of the speaker separated transcription data as input to produce the conversation attribute indication.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 13, 2024
From: SHASTRY, TEJAS; TASSONE, ANTHONY; BURROWS, COREY
To: TRULEO, INC.
Reel/Frame 068266/0927 →
Continuity (1)
Provisional Application 63514515 · Jul 19, 2023
References Cited (153)
US 3296374A · Clapper · 1967 [cited by applicant]
US 7409349B2 · Wang et al. · 2008 [cited by applicant]
US 7447654B2 · Ben-Levy et al. · 2008 [cited by applicant]
US 7475046B1 · Foley et al. · 2009 [cited by applicant]
US 7533054B2 · Hausman et al. · 2009 [cited by applicant]
US 7536335B1 · Weston et al. · 2009 [cited by applicant]
US 7610547B2 · Wang et al. · 2009 [cited by applicant]
US 7643998B2 · Yuen et al. · 2010 [cited by applicant]
US 7685048B1 · Hausman et al. · 2010 [cited by applicant]
US 7689495B1 · Kim et al. · 2010 [cited by applicant]
US 7774247B2 · Hausman et al. · 2010 [cited by applicant]
US 7822672B2 · Hausman · 2010 [cited by applicant]
US 8019665B2 · Hausman · 2011 [cited by applicant]
US 8099352B2 · Berger et al. · 2012 [cited by applicant]
US 8321465B2 · Farber et al. · 2012 [cited by applicant]
US 8332384B2 · Kemp · 2012 [cited by applicant]
US 8438158B2 · Kemp · 2013 [cited by applicant]
US 8442823B2 · Jeon et al. · 2013 [cited by applicant]
US 8473396B2 · Hausman et al. · 2013 [cited by applicant]
US 8489587B2 · Kemp · 2013 [cited by applicant]
US 8504483B2 · Foley et al. · 2013 [cited by applicant]
US 8676679B2 · Hausman et al. · 2014 [cited by applicant]
US 8788397B2 · Berger et al. · 2014 [cited by applicant]
US 8799140B1 · Toffee et al. · 2014 [cited by applicant]
US 8878853B2 · Baransky et al. · 2014 [cited by applicant]
US 8909516B2 · Medero et al. · 2014 [cited by applicant]
US 8965765B2 · Zweig et al. · 2015 [cited by applicant]
US 9313173B2 · Davis et al. · 2016 [cited by applicant]
US 9330659B2 · Ju et al. · 2016 [cited by applicant]
US 9613026B2 · Hodson · 2017 [cited by applicant]
US 9728184B2 · Xue et al. · 2017 [cited by applicant]
US 9760566B2 · Heck et al. · 2017 [cited by applicant]
US 9824698B2 · Jerauld · 2017 [cited by applicant]
US 9870424B2 · Neystadt et al. · 2018 [cited by applicant]
US 9906474B2 · Robarts et al. · 2018 [cited by applicant]
US 10002520B2 · Bohlander et al. · 2018 [cited by applicant]
US 10108306B2 · Khoo et al. · 2018 [cited by applicant]
US 10185989B2 · Ritter et al. · 2019 [cited by applicant]
US 10192277B2 · Hanchett et al. · 2019 [cited by applicant]
US 10210869B1 · King et al. · 2019 [cited by applicant]
US 10237716B2 · Bohlander et al. · 2019 [cited by applicant]
US 10237822B2 · Hanchett et al. · 2019 [cited by applicant]
US 10298875B2 · Klein et al. · 2019 [cited by applicant]
US 10354169B1 · Law et al. · 2019 [cited by applicant]
US 10354350B2 · Nakfour et al. · 2019 [cited by applicant]
US 10368225B2 · Hassan et al. · 2019 [cited by applicant]
US 10372755B2 · Blanco · 2019 [cited by applicant]
US 10381024B2 · Tan et al. · 2019 [cited by applicant]
US 10405786B2 · Sahin · 2019 [cited by applicant]
US 10417340B2 · Applegate et al. · 2019 [cited by applicant]
US 10419312B2 · Alazraki et al. · 2019 [cited by applicant]
US 10460746B2 · Costa et al. · 2019 [cited by applicant]
US 10477375B2 · Bohlander et al. · 2019 [cited by applicant]
US 10509988B2 · Woulfe et al. · 2019 [cited by applicant]
US 10534497B2 · Khoo et al. · 2020 [cited by applicant]
US 10586556B2 · Caskey et al. · 2020 [cited by applicant]
US 10594795B2 · Hanchett et al. · 2020 [cited by applicant]
US 10630560B2 · Adylov et al. · 2020 [cited by applicant]
US 10657962B2 · Zhang et al. · 2020 [cited by applicant]
US 10685075B2 · Blanco et al. · 2020 [cited by applicant]
US 10713497B2 · Womack et al. · 2020 [cited by applicant]
US 10720169B2 · Reitz et al. · 2020 [cited by applicant]
US 10728384B1 · Channakeshava · 2020 [cited by examiner]
US 10755729B2 · Dimino, Jr. et al. · 2020 [cited by applicant]
US 10779022B2 · MacDonald · 2020 [cited by applicant]
US 10779152B2 · Bohlander et al. · 2020 [cited by applicant]
US 10785610B2 · Bohlander et al. · 2020 [cited by applicant]
US 10796393B2 · Messerges et al. · 2020 [cited by applicant]
US 10805576B2 · Hanchett et al. · 2020 [cited by applicant]
US 10825479B2 · Hershfield et al. · 2020 [cited by applicant]
US 10848717B2 · Hanchett et al. · 2020 [cited by applicant]
US 10853435B2 · Reitz et al. · 2020 [cited by applicant]
US 10872636B2 · Smith et al. · 2020 [cited by applicant]
US 11120199B1 · Bachtiger · 2021 [cited by examiner]
US 11138970B1 · Han · 2021 [cited by examiner]
US 11423911B1 · Fu et al. · 2022 [cited by applicant]
US 11706391B1 · Heywood et al. · 2023 [cited by applicant]
US 11947872B1 · Mahler-Haug · 2024 [cited by examiner]
US 11948555B2 · Christie et al. · 2024 [cited by applicant]
US 12014750B2 · Shastry et al. · 2024 [cited by applicant]
US 12062368B1 · Arora · 2024 [cited by examiner]
US 20070117073A1 · Walker et al. · 2007 [cited by applicant]
US 20070167689A1 · Ramadas et al. · 2007 [cited by applicant]
US 20090155751A1 · Paul et al. · 2009 [cited by applicant]
US 20090292638A1 · Hausman · 2009 [cited by applicant]
US 20100121880A1 · Ursitti et al. · 2010 [cited by applicant]
US 20100332648A1 · Bohus et al. · 2010 [cited by applicant]
US 20110270732A1 · Ritter et al. · 2011 [cited by applicant]
US 20120004914A1 · Strom et al. · 2012 [cited by applicant]
US 20130156175A1 · Bekiares et al. · 2013 [cited by applicant]
US 20130173247A1 · Hodson · 2013 [cited by applicant]
US 20130300939A1 · Chou et al. · 2013 [cited by applicant]
US 20140006248A1 · Toffee · 2014 [cited by applicant]
US 20140081823A1 · Phadnis et al. · 2014 [cited by applicant]
US 20140101739A1 · Li et al. · 2014 [cited by applicant]
US 20140187190A1 · Schuler et al. · 2014 [cited by applicant]
US 20140207651A1 · Toffey et al. · 2014 [cited by applicant]
US 20150310729A1 · Lampert et al. · 2015 [cited by applicant]
US 20150310730A1 · Miller et al. · 2015 [cited by applicant]
US 20150310862A1 · Dauphin et al. · 2015 [cited by applicant]
US 20150381933A1 · Cunico et al. · 2015 [cited by applicant]
US 20160066085A1 · Chang et al. · 2016 [cited by applicant]
US 20160180737A1 · Clark et al. · 2016 [cited by applicant]
US 20170132703A1 · Oomman et al. · 2017 [cited by applicant]
US 20170316775A1 · Le et al. · 2017 [cited by applicant]
US 20170346904A1 · Fortna et al. · 2017 [cited by applicant]
US 20170364602A1 · Reitz et al. · 2017 [cited by applicant]
US 20180107943A1 · White et al. · 2018 [cited by applicant]
US 20180197548A1 · Palakodety · 2018 [cited by examiner]
US 20180233139A1 · Finkelstein et al. · 2018 [cited by applicant]
US 20180350389A1 · Garrido et al. · 2018 [cited by applicant]
US 20190019297A1 · Lim et al. · 2019 [cited by applicant]
US 20190042988A1 · Brown et al. · 2019 [cited by applicant]
US 20190096428A1 · Childress · 2019 [cited by examiner]
US 20190108270A1 · Dunne et al. · 2019 [cited by applicant]
US 20190121907A1 · Brunn et al. · 2019 [cited by applicant]
US 20190139438A1 · Tu et al. · 2019 [cited by applicant]
US 20190188814A1 · Kreitzer et al. · 2019 [cited by applicant]
US 20190258700A1 · Beaver et al. · 2019 [cited by applicant]
US 20190318725A1 · Le Roux et al. · 2019 [cited by applicant]
US 20200104698A1 · Ladvocat Cintra · 2020 [cited by applicant]
US 20200195726A1 · Hanchett et al. · 2020 [cited by applicant]
US 20200210907A1 · Ulizio et al. · 2020 [cited by applicant]
US 20200302043A1 · Vachon · 2020 [cited by applicant]
US 20200342857A1 · Moreno et al. · 2020 [cited by applicant]
US 20200365136A1 · Candelore et al. · 2020 [cited by applicant]
US 20210092224A1 · Rule et al. · 2021 [cited by applicant]
US 20210337307A1 · Wexler · 2021 [cited by examiner]
US 20210374601A1 · Liu et al. · 2021 [cited by applicant]
US 20220115022A1 · Sharifi et al. · 2022 [cited by applicant]
US 20220122615A1 · Chen et al. · 2022 [cited by applicant]
US 20220189501A1 · Shastry et al. · 2022 [cited by applicant]
US 20220310109A1 · Donsbach et al. · 2022 [cited by applicant]
US 20230103060A1 · Chaudhuri et al. · 2023 [cited by applicant]
US 20230186950A1 · Vanciu · 2023 [cited by examiner]
US 20230215439A1 · Kanda · 2023 [cited by examiner]
US 20230223038A1 · Shastry et al. · 2023 [cited by applicant]
US 20230419950A1 · Khare · 2023 [cited by examiner]
US 20240169854A1 · Shastry et al. · 2024 [cited by applicant]
US 20240256592A1 · O'Neill · 2024 [cited by examiner]
US 20240331721A1 · Shastry et al. · 2024 [cited by applicant]
WO WO2022133125A1 · 2022 [cited by applicant]
WO WO2024097345A1 · 2024 [cited by applicant]
Co-pending U.S. Appl. No. 18/738,819 inventor Tejas Shastry et al., filed Jun. 10, 2024. [cited by applicant]
Final Office Action for U.S. Appl. No. 18/180,652 mailed Oct. 13, 2023, 34 pages. [cited by applicant]
International Preliminary Report on Patentability for International Application No. PCT/US2021/063873 mailed Jun. 29, 2023, 10 pages. [cited by applicant]
International Search Report and Written Opinion for PCT Application No. PCT/US2021/063873 mailed Mar. 10, 2022, 11 pages. [cited by applicant]
International Search Report and Written Opinion for PCT Application No. PCT/US2023/036686 mailed Feb. 22, 2024, 8 pages. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 17/553,482 mailed Dec. 12, 2023, 12 pages. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 18/180,652 mailed Jul. 20, 2023, 19 pages. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 18/180,652 mailed Mar. 13, 2024, 15 pages. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 18/180,652 mailed Apr. 11, 2024, 14 pages. [cited by applicant]
Voigt, R. et al., “Language from police body camera footage shows racial disparities in officer respect,” Proceedings of the National Academy of Sciences, Jun. 20, 2017, 114(25), pp. 6521-6526. [cited by applicant]
Cited By (5)
US 12,489,778 US 12,499,243 US 12,537,750 US 12,609,968 US 12,694,106