IP Library Granted Patent US 12,051,397
Granted Patent B2
US 12,051,397 · App. 17/176,697 · Granted Jul 30, 2024

Audio privacy protection for surveillance systems

Inventors: Shaomin Xiong (Newark, CA); Toshiki Hirano (San Jose, CA); Pritam Das (Dublin, CA); Ramy Ayad (East Brunswick, NJ); Rajeev Nagabhirava (San Jose, CA)
Assignee: Western Digital Technologies, Inc.
G10K11/1754G06F21/32G06V20/40G06V40/172G10L17/06G10L21/028G10L25/57H04N5/76G06V20/44
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,051,397
App. No.
17/176,697
Granted
Jul 30, 2024
Kind
B2
Abstract

Systems and methods for audio privacy in network video surveillance systems are described. A video camera may include an image sensor and a microphone to generate a video stream. Responsive to detecting a human speaking condition in the video stream, the audio data may be selectively modified to mask a human voice component of the audio data for storing and/or displaying the surveillance video stream.

Claims (119)

1. A system, comprising:

a video camera comprising:

an image sensor configured to capture image data for a video stream; and

a microphone configured to capture audio data for the video stream; and

a controller configured to:

receive the video stream from the video camera;

determine a human speaking condition from the video stream;

determine a human voice component of the audio data during the human speaking condition;

process, responsive to the human speaking condition and using a facial recognition algorithm, the image data in the video stream to identify a speaker identity corresponding to the human voice component;

determine, based on comparing the speaker identity to a protected speaker data structure, that the speaker identity has a protected speaker status;

process, responsive to determining that the speaker identity has the protected speaker status and using a voice recognition algorithm, the audio data corresponding to the speaker identity in the video stream to identify an audio consent pattern, wherein the audio consent pattern comprises at least one word in the human voice component of the audio data configured to indicate consent to be recorded;

selectively modify, responsive to determining that the speaker identity has the protected speaker status and a lack of the audio consent pattern in the audio data corresponding to the speaker identity, the audio data in the video stream during the human speaking condition to mask the human voice component; and

store the video stream with the modified audio data.

2. The system of claim 1 , wherein:

the controller is further configured to process, responsive to determining that the speaker identity has the protected speaker status and using a behavior recognition algorithm, the image data corresponding to the speaker identity in the video stream to identify a visual consent pattern;

the visual consent pattern comprises a gesture of a speaker corresponding to the speaker identity and indicates consent to be recorded;

the controller processes the audio data for the audio consent pattern and the image data for the visual consent pattern in parallel; and

selectively modifying the audio data is further responsive to a lack of the visual consent pattern.

3. The system of claim 1 , wherein the audio consent pattern includes a phrase selected from a plurality of phrases indicative of consent to be recorded.

4. The system of claim 1 , wherein processing the audio data to mask the human voice component of the audio data includes:

separating the human voice component of the audio data from a non-voice remainder component of the audio data;

selectively masking the human voice component of the audio data; and

leaving the non-voice remainder component of the audio data unchanged.

5. The system of claim 1 , wherein processing the audio data to mask the human voice component of the audio data further includes:

separating the human voice component of the audio data into a plurality human voice components corresponding to individual speakers of a plurality of speakers;

identifying at least one individual speaker of the plurality of speakers; and

selectively masking, responsive to identifying the at least one individual speaker, a portion of the human voice component of the audio data corresponding to the at least one individual speaker.

6. The system of claim 1 , wherein:

the controller is further configured to:

separate the human voice component of the audio data into a plurality of human voice components corresponding to individual speakers of a plurality of speakers detect; and

process, using the voice recognition algorithm, the audio data in the video stream to identify the separated human voice component corresponding to the speaker identity that has protected speaker status; and

selectively modifying the audio data in the video stream during the human speaking condition comprises:

masking the separated human voice component corresponding to the speaker identity that has protected speaker status; and

leaving other human voice components of the plurality of human voice components in the audio data unchanged.

7. The system of claim 1 , wherein:

the controller is further configured to process, using a behavior recognition algorithm, the image data in the video stream to identify a visual consent pattern;

the visual consent pattern comprises at least one non-verbal behavior indicative of consent; and

selectively modifying the audio data in the video stream during the human speaking condition is further responsive to an absence of the visual consent pattern in the image data.

8. The system of claim 1 , wherein:

the controller is further configured to detect a video event from the image data in the video stream; and

determining the human speaking condition from the video stream is responsive to detecting the video event.

9. The system of claim 1 , wherein:

the controller is further configured to process, using the voice recognition algorithm, the audio data in the video stream to identify the speaker identity corresponding to the human voice component; and

determining that the speaker identity has protected speaker status comprises:

comparing a set of audio factors from the audio data to known voiceprints in the protected speaker data structure; and

comparing a set of facial vectors from the image data to known facial vectors in the protected speaker data structure.

10. The system of claim 1 , further comprising:

a storage device configured to store:

an unmodified video stream from the video camera; and

the video stream with the modified audio data from the controller;

an analytics engine configured to:

process the unmodified video stream to determine the human speaking condition from the video stream; and

notify the controller of the human speaking condition; and

a user device configured to:

display, using a graphical user interface and a speaker of the user device, the video stream with the modified audio data to a user;

determine a security credential for the user;

verify an audio access privilege corresponding to the security credential; and

display, using the graphical user interface and the speaker of the user device, the unmodified video stream to the user.

11. A computer-implemented method, comprising:

receiving a video stream from a video camera, wherein the video camera comprises:

an image sensor configured to capture image data for the video stream; and

a microphone configured to capture audio data for the video stream;

determining a human speaking condition from the video stream;

determining a human voice component of the audio data during the human speaking condition;

processing, responsive to the human speaking condition and using a facial recognition algorithm, the image data in the video stream to identify a speaker identity corresponding to the human voice component;

determining, based on comparing the speaker identity to a protected speaker data structure, that the speaker identity has a protected speaker status;

processing, responsive to determining that the speaker identity has the protected speaker status and using a voice recognition algorithm, the audio data corresponding to the speaker identity in the video stream to identify an audio consent pattern, wherein the audio consent pattern comprises at least one word in the human voice component of the audio data configured to indicate consent to be recorded;

selectively modifying, responsive to determining that the speaker identity has the protected speaker status and a lack of the audio consent pattern in the audio data corresponding to the speaker identity, the audio data in the video stream during the human speaking condition to mask the human voice component; and

storing the video stream with the modified audio data.

12. The computer-implemented method of claim 11 , further comprising:

processing, responsive to determining that the speaker identity has the protected speaker status and using a behavior recognition algorithm, the image data corresponding to the speaker identity in the video stream to identify a visual consent pattern, wherein:

the visual consent pattern comprises a gesture of a speaker corresponding to the speaker identity and indicates consent to be recorded;

processing the audio data for the audio consent pattern and the image data for the visual consent pattern occurs in parallel; and

selectively modifying the audio data is further responsive to a lack of the visual consent pattern.

13. The computer-implemented method of claim 11 , wherein the audio consent pattern includes a phrase selected from a plurality of phrases indicative of consent to be recorded.

14. The computer-implemented method of claim 11 , wherein selectively modifying the audio data to mask the human voice component of the audio data includes:

separating the human voice component of the audio data from a non-voice remainder component of the audio data;

selectively masking the human voice component of the audio data; and

leaving the non-voice remainder component of the audio data unchanged.

15. The computer-implemented method of claim 11 , wherein selectively modifying the audio data to mask the human voice component of the audio data includes:

separating the human voice component of the audio data into a plurality human voice components corresponding to individual speakers of a plurality of speakers;

identifying at least one individual speaker of the plurality of speakers; and

selectively masking, responsive to identifying the at least one individual speaker, a portion of the human voice component of the audio data corresponding to the at least one individual speaker.

16. The computer-implemented method of claim 11 , further comprising:

separating the human voice component of the audio data into a plurality of human voice components corresponding to individual speakers of a plurality of speakers; and

processing, using the voice recognition algorithm, the audio data in the video stream to identify the separated human voice component corresponding to the speaker identity that has protected speaker status, wherein selectively modifying the audio data in the video stream during the human speaking condition comprises:

masking the separated human voice component corresponding to the speaker identity that has protected speaker status; and

leaving other human voice components of the plurality of human voice components in the audio data unchanged.

17. The computer-implemented method of claim 11 , further comprising:

processing, using a behavior recognition algorithm, the image data in the video stream to identify a visual consent pattern, wherein:

the visual consent pattern comprises at least one non-verbal behavior indicative of consent; and

selectively modifying the audio data in the video stream during the human speaking condition is further responsive to an absence of the visual consent pattern in the image data.

18. The computer-implemented method of claim 11 , further comprising:

processing, using the voice recognition algorithm, the audio data in the video stream to identify the speaker identity corresponding to the human voice component, wherein determining that the speaker identity has protected speaker status comprises:

comparing a set of audio factors from the audio data to known voiceprints in the protected speaker data structure; and

comparing a set of facial vectors from the image data to known facial vectors in the protected speaker data structure.

19. The computer-implemented method of claim 18 , further comprising:

storing, in a storage device:

an unmodified video stream from the video camera; and

the video stream with the modified audio data;

processing the unmodified video stream to determine the human speaking condition from the video stream;

displaying, using a graphical user interface and a speaker of a user device, the video stream with the modified audio data to a user;

determining a security credential for the user;

verifying an audio access privilege corresponding to the security credential; and

displaying, using the graphical user interface and the speaker of the user device, the unmodified video stream to the user.

20. A surveillance system, comprising:

a video camera comprising:

an image sensor configured to capture image data for a video stream; and

a microphone configured to capture audio data for the video stream;

a processor;

a memory;

means for receiving the video stream from the video camera;

means for automatically determining a human speaking condition from the video stream;

means for determining a human voice component of the audio data during the human speaking condition;

means for processing, responsive to the human speaking condition and using a facial recognition algorithm, the image data in the video stream to identify a speaker identity corresponding to the human voice component;

means for determining, based on comparing the speaker identity to a protected speaker data structure, that the speaker identity has a protected speaker status;

means for processing, responsive to determining that the speaker identity has the protected speaker status and using a voice recognition algorithm, the audio data corresponding to the speaker identity in the video stream to identify a consent pattern, wherein the consent pattern comprises at least one word in the human voice component of the audio data configured to indicate consent to be recorded;

means for selectively modifying, responsive to determining that the speaker identity has the protected speaker status and a lack of the consent pattern in the audio data corresponding to the speaker identity, the audio data in the video stream during the human speaking condition to mask the human voice component; and

means for displaying the video stream using the modified audio data.

Assignments (10)
PARTIAL RELEASE OF SECURITY INTERESTS Recorded Apr 25, 2025
From: JPMORGAN CHASE BANK, N.A., AS AGENT
To: SANDISK TECHNOLOGIES, INC.
Reel/Frame 071382/0001 →
SECURITY AGREEMENT Recorded Apr 25, 2025
From: SANDISK TECHNOLOGIES, INC.
To: JPMORGAN CHASE BANK, N.A., AS COLLATERAL AGENT
Reel/Frame 071050/0001 →
PATENT COLLATERAL AGREEMENT Recorded Aug 23, 2024
From: SANDISK TECHNOLOGIES, INC.
To: JPMORGAN CHASE BANK, N.A., AS THE AGENT
Reel/Frame 068762/0494 →
CHANGE OF NAME Recorded Jun 27, 2024
From: SANDISK TECHNOLOGIES, INC.
To: SANDISK TECHNOLOGIES, INC.
Reel/Frame 067982/0032 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 29, 2024
From: WESTERN DIGITAL TECHNOLOGIES, INC.
To: SANDISK TECHNOLOGIES, INC.
Reel/Frame 067567/0682 →
PATENT COLLATERAL AGREEMENT - DDTL LOAN AGREEMENT Recorded Aug 21, 2023
From: WESTERN DIGITAL TECHNOLOGIES, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 067045/0156 →
PATENT COLLATERAL AGREEMENT - A&R LOAN AGREEMENT Recorded Aug 21, 2023
From: WESTERN DIGITAL TECHNOLOGIES, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 064715/0001 →
RELEASE OF SECURITY INTEREST AT REEL 056285 FRAME 0292 Recorded Feb 8, 2022
From: JPMORGAN CHASE BANK, N.A.
To: WESTERN DIGITAL TECHNOLOGIES, INC.
Reel/Frame 058982/0001 →
SECURITY INTEREST Recorded May 19, 2021
From: WESTERN DIGITAL TECHNOLOGIES, INC.
To: JPMORGAN CHASE BANK, N.A., AS AGENT
Reel/Frame 056285/0292 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 17, 2021
From: XIONG, SHAOMIN; HIRANO, TOSHIKI; DAS, PRITAM; AYAD, RAMY; NAGABHIRAVA, RAJEEV
To: WESTERN DIGITAL TECHNOLOGIES, INC.
Reel/Frame 055296/0064 →
Continuity (2)
Provisional Application 63119865 · Dec 1, 2020
Related Publication 20220172700A1 · Jun 2, 2022