IP Library Granted Patent US 12,159,621
Granted Patent B2
US 12,159,621 · App. 17/834,355 · Granted Dec 3, 2024

Application software and services with register classification capabilities

Inventors: Huakai Liao (Redmond, WA); Ana Parra (San Jose, CA); Gaurav Vinayak Tendolkar (San Jose, CA); Amit Srivastava (San Jose, CA); Siliang Kang (Redwood City, CA)
Assignee: MICROSOFT TECHNOLOGY LICENSING, LLC
G10L15/02G10L15/04G10L15/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,159,621
App. No.
17/834,355
Granted
Dec 3, 2024
Kind
B2
Abstract

A computing apparatus comprises one or more computer readable storage media, one or more processors operatively coupled with the one or more computer readable storage media, and program instructions stored on the one or more computer readable storage media. The program instructions, when executed by the one or more processors, direct the computing apparatus to at least generate an audio recording of speech, extract features from the audio recording indicative of vocal patterns in the speech, determine a register classification of the speech based at least on the features, and display an indication of the register classification in a user interface.

Claims (45)

1. A computing apparatus comprising:

one or more computer readable storage media;

one or more processors operatively coupled with the one or more computer readable storage media; and

program instructions stored on the one or more computer readable storage media that, when executed by the one or more processors, direct the computing apparatus to at least:

generate an audio recording of speech;

divide the audio recording into chunks, wherein each chunk comprises a sequence of frames;

extract, from frames of the sequences of frames, features indicative of vocal patterns in the speech;

classify each frame, of the frames, as belonging to a register classification of a plurality of register classifications based on the extracted features;

classify each chunk, of the chunks, as belonging to the register classification of the plurality of register classifications based on a threshold number of frames of the respective chunk classified as belonging to the register classification;

determine the register classification of the speech based at least on a threshold number of consecutive chunks classified as belonging to the register classification; and

display an indication of the register classification in a user interface.

2. The computing apparatus of claim 1 wherein the program instructions further direct the computing apparatus to determine a register classification of each of the chunks based at least on features extracted from each chunk.

3. The computing apparatus of claim 2 wherein the program instructions further direct the computing apparatus to divide each of the chunks into a sequence of overlapping frames, and wherein to extract the features from the frames of the sequences of frames, the program instructions further direct the computing apparatus to extract the features from each of the sequence of overlapping frames.

4. The computing apparatus of claim 3 wherein, to classify each frame of the frames as belonging to the register classification of the plurality of register classifications, the program instructions direct the computing apparatus to classify each frame, of the sequence of overlapping frames, as belonging to a specific range of a vocal register if a vocal pattern expressed by a feature extracted for the frame matches a vocal pattern of the specific range.

5. The computing apparatus of claim 4 wherein the program instructions further direct the computing apparatus to classify each chunk, of the chunks, as belonging to the specific range of the vocal register if a subset of the overlapping frames belonging to the chunk includes at least a threshold number of frames classified as belonging to the specific range.

6. The computing apparatus of claim 5 wherein the program instructions further direct the computing apparatus to classify the speech as belonging to the specific range of the vocal register if the audio recording includes a threshold number of consecutive ones of the chunks classified as belonging to the specific range.

7. The computing apparatus of claim 6 wherein each of the chunks comprises a non-overlapping portion of the audio recording having a duration of about 0.5 seconds, and wherein each of the overlapping frames has a duration of about 50 milliseconds that overlaps with a preceding frame for a duration of about 40 milliseconds.

8. The computing apparatus of claim 7 wherein the specific range of the vocal register comprises a range associated with vocal fry.

9. One or more computer readable storage media having program instructions stored thereon that, when executed by one or more processors, direct a computing device to at least:

generate an audio recording of speech;

divide the audio recording into chunks, wherein each chunk comprises a sequence of frames;

extract, from frames of the sequence of frames, features indicative of vocal patterns in the speech;

classify each frame of the frames as belonging to a register classification of a plurality of register classifications based on the extracted features;

classify each chunk, of the chunks, as belonging to the register classification of the plurality of register classifications based on a threshold number of frames of the respective chunk classified as belonging to the register classification;

determine the register classification of the speech based at least on a threshold number of consecutive chunks classified as belonging to the register classification; and

display an indication of the register classification in a user interface.

10. The one or more computer readable storage media apparatus of claim 9 wherein the program instructions further direct the computing device to determine a register classification of each of the chunks based at least on features extracted from each chunk.

11. The one or more computer readable storage media of claim 10 wherein the program instructions further direct the computing device to divide each of the chunks into a sequence of overlapping frames, and wherein to extract the features from the frames of the sequence of frames, the program instructions further direct the computing device to extract the features from each of the sequence of overlapping frames.

12. The one or more computer readable storage media of claim 11 wherein, to classify each frame of the frames as belonging to the register classification of the plurality of register classifications, the program instructions direct the computing device to classify each frame, of the sequence of overlapping frames, as belonging to a specific range of a vocal register if a vocal pattern expressed by a feature extracted for the frame matches a vocal pattern of the specific range.

13. The one or more computer readable storage media of claim 12 wherein the program instructions further direct the computing device to classify each chunk, of the chunks, as belonging to the specific range of the vocal register if a subset of the overlapping frames belonging to the chunk includes at least a threshold number of frames classified as belonging to the specific range.

14. The one or more computer readable storage media of claim 13 wherein the program instructions further direct the computing device to classify the speech as belonging to the specific range of the vocal register if the audio recording includes a threshold number of consecutive ones of the chunks classified as belonging to the specific range.

15. The one or more computer readable storage media of claim 14 wherein each of the chunks comprises a non-overlapping portion of the audio recording having a duration of about 0.5 seconds, and wherein each of the overlapping frames has a duration of about 50 milliseconds that overlaps with a preceding frame for a duration of about 40 milliseconds.

16. The one or more computer readable storage media of claim 15 wherein the specific range of the vocal register comprises a range associated with vocal fry.

17. A method of operating a computing device, the method comprising:

the computing device generating an audio recording of speech;

the computing device dividing the audio recording into chunks;

the computing device dividing each of the chunks into a sequence of overlapping frames;

the computing device extracting, from frames of the sequences of overlapping frames, features indicative of vocal patterns in the speech;

the computing device classifying the frames, of the sequences of overlapping frames, as belonging to a register classification of a plurality of register classifications based on the extracted features;

the computing device classifying each chunk, of the chunks, as belonging to the register classification of the plurality of register classifications based on a threshold number of frames of the respective chunk classified as belonging to the register classification;

the computing device determining the register classification of the speech based at least on a threshold number of consecutive chunks classified as belonging to the register classification; and

the computing device displaying an indication of the register classification in a user interface.

18. The method of claim 17 further comprising the computing device classifying the speech as belonging to a specific range of a vocal register if the audio recording includes a threshold number of consecutive ones of the chunks classified as belonging to the specific range.

19. The method of claim 18 further comprising the computing device classifying each chunk, of the chunks, as belonging to the specific range of the vocal register if a subset of the overlapping frames belonging to the chunk includes at least a threshold number of frames classified as belonging to the specific range.

20. The method of claim 19 further comprising the computing device classifying each frame, of the sequences of overlapping frames, as belonging to the specific range of the vocal register if a vocal pattern indicated by a feature extracted for the frame matches a vocal pattern of the specific range.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 3, 2023
From: LIAO, HUAKAI; PARRA, ANA; TENDOLKAR, GAURAV VINAYAK; KANG, SILIANG
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 063211/0267 →
Continuity (1)
Related Publication 20230395064A1 · Dec 7, 2023