IP Library Patent Application 17553482
Patent Application
App. No. 17/553,482

AUDIO ANALYSIS OF BODY WORN CAMERA

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
17/553,482
Abstract

Machine natural language processing to analyze language in apparatus, systems, and methods of using are provided. Audio from camera footage can be transcribed in one exemplary method includes extracting at least one audio segment from a body camera video track, detecting voice activity to identify starting and ending timestamps of voice, transcribing the at least one audio segment to identify and separate audio of at least a first speaker, and scoring the audio of the first speaker to identify interactions of interest. Audio could be analyzed and scored to record verbal performance, respectfulness, wellness, etc. and speakers from the audio can be detected.

Claims (35)

1 . A method of using machine natural language processing to analyze language in transcribed camera footage comprising:

extracting at least one audio segment from a body camera video track;

detecting voice activity to identify starting and ending timestamps of voice;

transcribing the at least one audio segment to identify and separate audio of at least one speaker;

scoring the audio of the at least one speaker to identify interactions of interest.

2 . The method of claim 1 wherein the at least one speaker is a figure of authority, including one of a: police officer, emergency technician, guard, soldier, doctor, or first responder.

3 . The method of claim 2 wherein the interactions of interest include whether the figure of authority is escalating or de-escalating a situation.

4 . The method of claim 2 wherein the interactions of interest include whether the figure of authority is using respectful language or negative language.

5 . The method of claim 2 wherein the scoring includes analyzing for word disfluencies or filler words to analyze speaker confidence.

6 . The method of claim 2 wherein the figure of authority is identified based on voice quality.

7 . The method of claim 6 wherein the transcribing identifies whether audio of at least an other speaker is included on the at least one audio segment.

8 . The method of claim 1 wherein the method further includes:

identifying events that may have occurred in the body camera video track based on language cues in the at least one audio segment.

9 . The method of claim 8 wherein the method further includes:

compressing the body camera video track based on the events.

10 . The method of claim 1 wherein the method is performed in real-time.

11 . A system of using machine natural language processing to analyze language in transcribed camera footage comprising:

an audio and language analyzer, operable to:

extract at least one audio segment from a body camera video track;

detect voice activity to identify starting and ending timestamps of voice;

transcribe the at least one audio segment to identify and separate audio of at least one speaker;

score the audio of the at least one speaker to identify interactions of interest.

12 . The system of claim 11 wherein the at least one speaker is a figure of authority, including one of a: police officer, emergency technician, guard, soldier, doctor, or first responder.

13 . The system of claim 12 wherein the interactions of interest include whether the figure of authority is escalating or de-escalating a situation.

14 . The system of claim 12 wherein the interactions of interest include whether the figure of authority is using respectful language or negative language.

15 . The system of claim 12 further comprising at least one other speaker and wherein the interactions of interest include whether the at least one other speaker is using negative language.

16 . The system of claim 12 wherein the score includes an analysis for word disfluencies or filler words to analyze speaker confidence.

17 . The system of claim 12 wherein the figure of authority is anonymously identified based on voice quality.

18 . The system of claim 16 wherein the transcription identifies whether audio of at least an other speaker is included on the at least one audio segment.

19 . The system of claim 18 wherein the audio of the at least other speaker is either selectively removed or analyzed by the system.

20 . The system of claim 11 wherein the system is further operable to:

identify events that may have occurred in the body camera video track based on language cues in the at least one audio segment.

21 . The system of claim 20 wherein the system is further operable to:

compress the body camera video track based on the events.

22 . The system of claim 11 wherein the system operates in real-time.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 15, 2022
From: SHASTRY, TEJAS; VERGUN, SVYATOSLAV; BROCHTRUP, COLIN; GOLDEY, MATTHEW
To: TRULEO, INC.
Reel/Frame 059273/0766 →