IP Library Granted Patent US 9,202,469
Granted Patent B1
US 9,202,469 · App. 14/487,361 · Granted Dec 1, 2015

Capturing noteworthy portions of audio recordings

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,202,469
App. No.
14/487,361
Granted
Dec 1, 2015
Kind
B1
Abstract

A technique for recording dictation, meetings, lectures, and other events includes automatically segmenting an audio recording into portions by detecting speech transitions within the recording and selectively identifying certain portions of the recording as noteworthy. Noteworthy audio portions are displayed to a user for selective playback. The user can navigate to different noteworthy audio portions while ignoring other portions. Each noteworthy audio portion starts and ends with a speech transition. Thus, the improved technique typically captures noteworthy topics from beginning to end, thereby reducing or avoiding the need for users to have to search for the beginnings and ends of relevant topics manually.

Claims (63)

1. A method of recording human speech, the method comprising:

acquiring, from a microphone coupled to electronic circuitry, an audio signal that conveys human speech;

identifying, by the electronic circuitry and in real time as the audio signal is being acquired, (i) a set of speech transitions in the audio signal that mark boundaries between respective portions of human speech, and (ii) a set of noteworthy audio segments, each noteworthy audio segment being one of the portions of human speech and meeting a noteworthiness criterion, the noteworthiness criterion providing a standard for evaluating noteworthiness of portions of human speech; and

after recording the audio signal, displaying a list of the identified set of noteworthy audio segments, the list enabling a user selectively to play back any of the noteworthy audio segments,

wherein identifying the set of speech transitions in the audio signal includes (i) detecting Pauses in human speech in the audio signal that exceed a predetermined interval of time and (ii) marking speech transitions at times relative to the audio signal when the Pauses occur, and

wherein detecting Pauses in human speech in the audio signal includes:

acquiring multiple power samples of the audio signal;

computing a set of statistics of the power samples;

computing a power threshold based on the set of statistics;

counting numbers of consecutive Silences, each Silence being a power sample whose power falls below the power threshold; and

identifying a Pause in the audio signal, in response to counting a predetermined number of consecutive Silences.

2. The method of claim 1 , wherein each noteworthy audio segment begins at a speech transition preceding a time when the noteworthiness criterion is met and wherein each noteworthy audio segment ends at a speech transition following the time when the noteworthiness criterion is met.

3. The method of claim 2 , wherein displaying the list of the identified set of noteworthy audio segments presents the noteworthy audio segments in order of time that the noteworthy audio segments were recorded, and wherein the method further comprises:

accepting, from the user, a selection of any noteworthy audio segment from the displayed list of noteworthy audio segments; and

in response to accepting the selection, playing back the selected noteworthy audio segment from a beginning of the selected noteworthy audio segment.

4. The method of claim 3 , further comprising, when displaying the list of noteworthy audio segments, displaying no portions of human speech besides the noteworthy audio segments, such that the user is unable selectively to play back any portions of human speech that fail to meet the noteworthiness criterion.

5. The method of claim 3 , wherein the noteworthiness criterion is met, for at least one of the portions of human speech, in response to the electronic circuitry detecting, while each such portion of human speech is being acquired, a predetermined manual operation performed by the user.

6. The method of claim 5 , wherein detecting the predetermined manual operation performed by the user includes the electronic circuitry detecting at least one of (i) a predetermined user input, (ii) a triggering of a proximity detector, (iii) a change in output of a light sensor, (vi) a change in output of an accelerometer, and (vii) a change in output of a gyroscope.

7. The method of claim 3 , further comprising, for at least one of the portions of human speech:

processing a corresponding portion of human speech as it is being acquired, wherein processing the corresponding portion of human speech generates a set of audio characteristics of the corresponding portion of human speech;

wherein the noteworthiness criterion is met for the corresponding portion of human speech when the generated set of audio characteristics is consistent with audio characteristics of noteworthy audio content.

8. The method of claim 7 , wherein the set of audio characteristics generated for each of the portions of human speech includes at least one of (i) a duration of an associated portion of human speech, (ii) an average power of the associated portion of human speech, (iii) one or more keywords transcribed from the associated portion of human speech, and (iv) a voice pattern assessment.

9. The method of claim 3 , wherein the electronic circuitry is provided as part of a computing device connected to a computing network, and wherein the method further comprises:

storing (i) metadata that identifies the set of speech transitions and (ii) the set of noteworthy audio segments; and

uploading the recorded audio signal and the stored metadata to a server connected to the computing device over the network, wherein uploading the recorded audio enables the set of noteworthy audio segments to be shared over the network with other users.

10. The method of claim 9 , wherein the computing device is a smartphone, and wherein uploading the recorded audio signal and the stored metadata to the server is performed, by an app running on the smartphone, in response to registering a single tap on a button displayed on the smartphone.

11. The method of claim 1 , wherein detecting pauses in human speech in the audio signal further includes:

acquiring, over time, new power samples; and

adapting to changes in background noise by recomputing the power threshold as new power samples are acquired.

12. The method of claim 11 , further comprising receiving the predetermined number of consecutive Silences as a user-adjustable parameter.

13. The method of claim 1 ,

wherein the electronic circuitry and the microphone are embodied together in a mobile computing device, the mobile computing device running an app, the app directing the acts of acquiring, identifying, and displaying, and

wherein the noteworthiness criterion is met, for a particular portion of the human speech, in response to (i) the app displaying a button on a display of the mobile computing device and while the particular portion of human speech is being acquired, and (ii) the app registering, while the particular portion of human speech is being acquired, an activation of the button.

14. A computing device, comprising:

electronic circuitry including processing circuitry and memory, the memory coupled to the processing circuitry and storing executable instructions which, when executed by the processing circuitry, cause the processing circuitry to:

acquire, from a microphone coupled to the electronic circuitry, an audio signal that conveys human speech;

identify, by the electronic circuitry and in real time as the audio signal is being acquired, (i) a set of speech transitions in the audio signal that mark boundaries between respective portions of human speech, and (ii) a set of noteworthy audio segments, each noteworthy audio segment being one of the portions of human speech and meeting a noteworthiness criterion, the noteworthiness criterion providing a standard for evaluating noteworthiness of portions of human speech; and

after recording the audio signal, display a list of the identified set of noteworthy audio segments, the list enabling a user selectively to play back any of the noteworthy audio segments,

wherein, when caused to identify the set of speech transitions in the audio signal, the processing circuitry is further caused to (i) detect Pauses in human speech in the audio signal that exceed a predetermined interval of time and (ii) mark speech transitions at times relative to the audio signal when the Pauses occur, and

wherein, when caused to detect Pauses in human speech in the audio signal, the processing circuit is further caused to:

acquire multiple power samples of the audio signal;

compute a set of statistics of the power samples;

compute a power threshold based on the set of statistics;

count numbers of consecutive Silences, each Silence being a power sample whose power falls below the power threshold; and

identify a Pause in the audio signal, in response to counting a predetermined number of consecutive Silences.

15. A computer program product including a non-transitory computer-readable medium having instructions which, when executed by processing circuitry, cause the processing circuitry to perform a method for recording human speech, the method comprising:

acquiring, from a microphone, an audio signal that conveys human speech;

identifying, by the electronic circuitry and in real time as the audio signal is being acquired, (i) a set of speech transitions in the audio signal that mark boundaries between respective portions of human speech, and (ii) a set of noteworthy audio segments, each noteworthy audio segment being one of the portions of human speech and meeting a noteworthiness criterion, the noteworthiness criterion providing a standard for evaluating noteworthiness of portions of human speech; and

after recording the audio signal, displaying a list of the identified set of noteworthy audio segments, the list enabling a user selectively to play back any of the noteworthy audio segments,

wherein identifying the set of speech transitions in the audio signal includes (i) detecting Pauses in human speech in the audio signal that exceed a predetermined interval of time and (ii) marking speech transitions at times relative to the audio signal when the Pauses occur, and

wherein detecting Pauses in human speech in the audio signal includes:

acquiring multiple power samples of the audio signal;

computing a set of statistics of the power samples;

computing a power threshold based on the set of statistics;

counting numbers of consecutive Silences, each Silence being a power sample whose power falls below the power threshold; and

identifying a Pause in the audio signal, in response to counting a predetermined number of consecutive Silences.

16. The computer program product of claim 15 , wherein displaying the list of the identified set of noteworthy audio segments presents the noteworthy audio segments in order of time that the noteworthy audio segments were recorded, and wherein the method further comprises:

accepting, from the user, a selection of any noteworthy audio segment from the displayed list of noteworthy audio segments; and

in response to accepting the selection, playing back the selected noteworthy audio segment from a beginning of the selected noteworthy audio segment.

17. The computer program product of claim 16 , wherein the noteworthiness criterion is met, for at least one of the portions of human speech, in response to the electronic circuitry detecting, while each such portion of human speech is being acquired, a predetermined manual operation performed by the user.

18. The computer program product of claim 16 , wherein the method further comprises, for at least one of the portions of human speech:

processing a corresponding portion of human speech as it is being acquired, wherein processing the corresponding portion of human speech generates a set of audio characteristics of the corresponding portion of human speech;

wherein the noteworthiness criterion is met for the corresponding portion of human speech when the generated set of audio characteristics is consistent with audio characteristics of noteworthy audio content.

Assignments (14)
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS (REEL/FRAME 053667/0169, REEL/FRAME 060450/0171, REEL/FRAME 063341/0051) Recorded Mar 15, 2024
From: BARCLAYS BANK PLC, AS COLLATERAL AGENT
To: GOTO GROUP, INC. (F/K/A LOGMEIN, INC.)
Reel/Frame 066800/0145 →
SECURITY INTEREST Recorded Feb 16, 2024
From: GOTO COMMUNICATIONS, INC.; GOTO GROUP, INC.; LASTPASS US LP
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION, AS THE NOTES COLLATERAL AGENT
Reel/Frame 066614/0355 →
SECURITY INTEREST Recorded Feb 16, 2024
From: GOTO COMMUNICATIONS, INC.,; GOTO GROUP, INC., A; LASTPASS US LP,
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION, AS THE NOTES COLLATERAL AGENT
Reel/Frame 066614/0402 →
SECURITY INTEREST Recorded Feb 7, 2024
From: GOTO GROUP, INC.,; GOTO COMMUNICATIONS, INC.; LASTPASS US LP
To: BARCLAYS BANK PLC, AS COLLATERAL AGENT
Reel/Frame 066508/0443 →
CHANGE OF NAME Recorded Apr 8, 2022
From: LOGMEIN, INC.
To: GOTO GROUP, INC.
Reel/Frame 059644/0090 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS (SECOND LIEN) Recorded Feb 16, 2021
From: BARCLAYS BANK PLC, AS COLLATERAL AGENT
To: LOGMEIN, INC.
Reel/Frame 055306/0200 →
SECOND LIEN PATENT SECURITY AGREEMENT Recorded Sep 1, 2020
From: LOGMEIN, INC.
To: BARCLAYS BANK PLC, AS COLLATERAL AGENT
Reel/Frame 053667/0079 →
NOTES LIEN PATENT SECURITY AGREEMENT Recorded Sep 1, 2020
From: LOGMEIN, INC.
To: U.S. BANK NATIONAL ASSOCIATION, AS NOTES COLLATERAL AGENT
Reel/Frame 053667/0032 →
FIRST LIEN PATENT SECURITY AGREEMENT Recorded Sep 1, 2020
From: LOGMEIN, INC.
To: BARCLAYS BANK PLC, AS COLLATERAL AGENT
Reel/Frame 053667/0169 →
RELEASE OF SECURITY INTEREST RECORDED AT REEL/FRAME 041588/0143 Recorded Aug 31, 2020
From: JPMORGAN CHASE BANK, N.A.
To: LOGMEIN, INC.; GETGO, INC.
Reel/Frame 053650/0978 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 4, 2019
From: GETGO, INC.
To: LOGMEIN, INC.
Reel/Frame 049843/0833 →
SECURITY INTEREST Recorded Feb 1, 2017
From: GETGO, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 041588/0143 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 8, 2016
From: CITRIX SYSTEMS, INC.
To: GETGO, INC.
Reel/Frame 039970/0670 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 8, 2014
From: MOORJANI, YOGESH; KASPER, RYAN WARREN; THAPLIYAL, ASHISH V.; KUMAR, AJAY; KURUVADI RAMESH BABU, ABHINAV; THAPLIYAL, ELIZABETH; KALBACH, JAMES; CRAMER, MARGARET DIANNE
To: CITRIX SYSTEMS, INC.
Reel/Frame 033913/0143 →