IP Library › Granted Patent US 12,444,421
Granted Patent B2
US 12,444,421 · App. 18/842,368 · Granted Oct 14, 2025

Computer-implemented method for punctuation of text from audio input

Inventors: Ville Ruutu (Helsinki, FI); Jussi Ruutu (Helsinki, FI); Honain Derrar (Helsinki, FI)
Assignee: Elisa Oyj
G10L15/26G10L15/05G10L15/08G10L15/14G10L25/93G10L15/02G10L15/04G10L25/03G10L25/18G10L25/78G10L25/87
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,444,421
App. No.
18/842,368
Granted
Oct 14, 2025
Kind
B2
Abstract

Disclosed herein is a computer-implemented method for punctuation of text from audio. The method includes obtaining an audio input comprising speech data; identifying a plurality of silent sections in the audio input; grouping the plurality of silent sections into a plurality of groups, where each group in the plurality of groups corresponds to a punctuation mark or a space without a punctuation mark; and associating each silent section in the plurality of silent sections with a punctuation mark or a space according to the grouping of the silent sections, thus obtaining punctuation information.

Claims (26)

1. A computer-implemented method for punctuation of text from an audio input, the method comprising:

obtaining an audio input comprising speech data;

identifying a plurality of silent sections in the audio input;

obtaining a type input indicating a type of the speech data in the audio input, wherein the type input indicates that the speech data is a customer service call, a public speech, or a lecture;

choosing an expected distribution, indicating an expected relative frequency of each group in a plurality of groups, according to the type input;

grouping the plurality of silent sections into the plurality of groups, wherein each group in the plurality of groups corresponds to a punctuation mark or a space without a punctuation mark, wherein the grouping the plurality of silent sections into the plurality of groups is done at least partially by using the expected distribution; and

associating each silent section in the plurality of silent sections with a punctuation mark or a space according to the group of the silent section, thus obtaining punctuation information,

performing a speech-to-text conversion on the audio input, thus obtaining a transcript of the speech data; and

punctuating the transcript according to the punctuation information by associating each silent section in the plurality of silent sections with the corresponding punctuation mark or the corresponding space without the punctuation mark;

wherein the expected distribution is based on statistical information about the plurality of silent sections between spoken words to apply punctuation.

2. The computer-implemented method according to claim 1 , wherein each group in the plurality of groups corresponds to a range of silent section temporal duration.

3. The computer-implemented method according to claim 1 , wherein the grouping the plurality of silent sections into the plurality of groups is done using a clustering algorithm.

4. The computer-implemented method according to claim 3 , wherein the clustering algorithm comprises k-means clustering.

5. The computer-implemented method according to claim 1 , wherein the expected distribution indicating an expected relative frequency of each group in the plurality of groups is at least partially based on an expected distribution of punctuation marks.

6. The computer-implemented method according to claim 5 , wherein the grouping the plurality of silent sections into the plurality of groups comprises:

determining at least one threshold temporal duration based on the expected distribution, wherein the at least one threshold temporal duration corresponds to a threshold between two groups in the plurality of groups; and

grouping the plurality of silent sections into the plurality of groups by comparing the temporal duration of each silent section in the plurality of silent sections to the at least one threshold temporal duration.

7. The computer-implemented method according to claim 5 , wherein a silent section in the plurality of silent sections is grouped into the plurality of groups using the expected distribution at least when the silent section cannot be grouped based on silent section temporal duration.

8. The computer-implemented method according to claim 1 , further comprising:

performing a speech-to-text conversion on the audio input, thus obtaining a transcript of the speech data; and

punctuating the transcript according to the punctuation information.

9. The computer-implemented method according to claim 8 , wherein the speech-to-text conversion further produces time stamp information of the plurality of silent sections, and wherein the plurality of silent sections is identified based on the time stamp information.

10. A computing device, comprising:

at least one processor; and

at least one memory including computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the computing device to perform the method according to claim 1 .

11. A non-transitory computer program product comprising program code configured to perform the method according to claim 1 when the computer program product is executed on a computer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 30, 2024
From: RUUTU, VILLE; RUUTU, JUSSI; DERRAR, HONAIN
To: ELISA OYJ
Reel/Frame 068456/0841 →
Priority Claims (1)
FI 20225351 · Apr 27, 2022 · national
Continuity (1)
Related Publication 20250111854A1 · Apr 3, 2025
References Cited (25)
US 6067514A · Chen · 2000 [cited by examiner]
US 7680657B2 · Shi · 2010 [cited by examiner]
US 9607613B2 · Buchanan · 2017 [cited by examiner]
US 10242669B1 · Sandler · 2019 [cited by examiner]
US 20010056344A1 · Ramaswamy · 2001 [cited by examiner]
US 20080300875A1 · Yao · 2008 [cited by examiner]
US 20130018656A1 · White · 2013 [cited by examiner]
US 20140350939A1 · Liu · 2014 [cited by examiner]
US 20230181093A1 · Shinkawa · 2023 [cited by examiner]
CN 102231278B · 2013 [cited by examiner]
CN 114512118A · 2022 [cited by examiner]
WO 2014187069A1 · 2014 [cited by applicant]
WO WO2022166218A1 · 2022 [cited by examiner]
Translation of CN-114512118-A (Year: 2022). [cited by examiner]
IP.com translation of WO2022166218A1. (Year: 2022). [cited by examiner]
IP.com translation of CN102231278B. (Year: 2013). [cited by examiner]
EP23719429 written opinions dated Jun. 14, 2023. (Year: 2023). [cited by examiner]
EP23719429 written opinions dated Jul. 9, 2024 and Jan. 11, 2024. (Year: 2024). [cited by examiner]
International Search Report and Written Opinion for International Application No. PCT/FI2023/050208 , mailing date of Jun. 14, 2023. [cited by applicant]
Search Report and Office Action for FI Application No. 20225351, dated Nov. 16, 2022. [cited by applicant]
Written Opinion of IPEA for International Application No. PCT/FI2023/050208 , mailing date of Jan. 11, 2024. [cited by applicant]
Office Action for FI Application No. 20225351, dated Mar. 8, 2024. [cited by applicant]
International Preliminary Report on Patentability (IPRP) Chapter II International Application No. PCT/FI2023/050208, mailing date of Jul. 9, 2024. [cited by applicant]
Levy et al., “The Effect of Pitch, Intensity and Pause Duration in Punctuation Detection”, 2012 IEEE 27th Convention of Electrical and Electronics Engineers in Israel, 2012, pp. 1-4, Eilat, Israel, doi: 10.1109/EEEI.201… [cited by applicant]
Rabiner, L.R., “A tutorial on hidden Markov models and selected applications in speech recognition,” in Proceedings of the IEEE, Feb. 1989, vol. 77, No. 2, pp. 257-286, doi: 10.1109/5.18626. [cited by applicant]