IP Library Granted Patent US 12,079,588
Granted Patent B2
US 12,079,588 · App. 18/345,663 · Granted Sep 3, 2024

Systems and methods for automatic speech translation

Inventor: Claudio Fantinuoli (Speyer, DE)
Assignee: KUDO, INC.
G06F40/58G10L13/033G10L15/04G10L15/22G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,079,588
App. No.
18/345,663
Granted
Sep 3, 2024
Kind
B2
Abstract

A method for providing automatic interpretation may include receiving, by a processor, audible speech from a speech source, generating, by the processor, in real-time, a speech transcript by applying an automatic speech recognition model on the speech, segmenting, by the processor, the speech transcript into speech segments based on a content of the speech by applying a segmenter model on the speech transcript, compressing, by the processor, the speech segments based on the content of the speech by applying a compressor model on the speech segments, generating, by the processor, a translation of the speech by applying a machine translation model on the compressed speech segments, and generating, by the processor, audible translated speech based on the translation of the speech by applying a text to speech model on the translation of the speech.

Claims (38)

1. A non-transitory, computer-readable medium including instructions which, when executed by at least one processor, cause the at least one processor to:

apply an automatic speech recognition model to audible speech to generate a speech transcript, the automatic speech recognition model using one or more known terms to generate the speech transcript;

apply a segmenter model on the speech transcript to segment the speech transcript into speech segments based on a content of the speech, each speech segment being a subset of the speech transcript segmented in accordance with the content and a number of words within each segment, wherein the segmenter model is trained to identify the speech segments in accordance with each speech segment's latency value corresponding to a time required to pronounce each segment;

apply a machine translation model on the speech segments to generate a translation of the speech; and

apply a text to speech model on the translation of the speech to generate audible translated speech based on the translation of the speech.

2. The medium of claim 1 , wherein the instructions further cause the at least one processor to apply a compressor model to compress speech segments based on the content of the speech.

3. The medium of claim 2 , wherein the instructions further cause the at least one processor to remove at least one segment or combine at least two segments into a compressed speech segment.

4. The medium of claim 2 , wherein the instructions further cause the at least one processor to adjust a compression level of the compressor model.

5. The medium of claim 4 , wherein the at least one processor adjusts the compression level based on a length of the speech transcript.

6. The medium of claim 1 , wherein the instructions further cause the at least one processor to adjust a speed of the text to speech model.

7. The medium of claim 6 , wherein the at least one processor adjusts the speed of the text to speech model based on a latency of the audible translated speech relative to the audible speech.

8. The medium of claim 1 , wherein segmenting the speech transcript into speech segments comprises selecting a number of words to be included in each speech segment.

9. A method comprising:

applying, by at least one processor, an automatic speech recognition model to audible speech to generate a speech transcript, the automatic speech recognition model using one or more known terms to generate the speech transcript;

applying, by the at least one processor, a segmenter model on the speech transcript to segment the speech transcript into speech segments based on a content of the speech, each speech segment being a subset of the speech transcript segmented in accordance with the content and a number of words within each segment, wherein the segmenter model is trained to identify the speech segments in accordance with each speech segment's latency value corresponding to a time required to pronounce each segment;

applying, by the at least one processor, a machine translation model on the speech segments to generate a translation of the speech; and

applying, by the at least one processor, a text to speech model on the translation of the speech to generate audible translated speech based on the translation of the speech.

10. The method of claim 9 , further comprising:

applying, by the at least one processor, a compressor model to compress speech segments based on the content of the speech.

11. The method of claim 10 , further comprising:

removing, by the at least one processor, at least one segment or combine at least two segments into a compressed speech segment.

12. The method of claim 10 , further comprising:

adjusting, by the at least one processor, a compression level of the compressor model.

13. The method of claim 12 , wherein the at least one processor adjusts the compression level based on a length of the speech transcript.

14. The method of claim 9 , further comprising:

adjusting, by the at least one processor, a speed of the text to speech model.

15. The method of claim 14 , wherein the at least one processor adjusts the speed of the text to speech model based on a latency of the audible translated speech relative to the audible speech.

16. The method of claim 9 , wherein segmenting the speech transcript into speech segments comprises selecting a number of words to be included in each speech segment.

17. A system comprising:

a first processor configured to receive audible speech;

a second processor in communication with the first processor, the second processor configured to:

apply an automatic speech recognition model to audible speech received from the first processor to generate a speech transcript, the automatic speech recognition model using one or more known terms to generate the speech transcript;

apply a segmenter model on the speech transcript to segment the speech transcript into speech segments based on a content of the speech, each speech segment being a subset of the speech transcript segmented in accordance with the content and a number of words within each segment, wherein the segmenter model is trained to identify the speech segments in accordance with each speech segment's latency value corresponding to a time required to pronounce each segment;

apply a machine translation model on the speech segments to generate a translation of the speech; and

apply a text to speech model on the translation of the speech to generate audible translated speech based on the translation of the speech.

18. The system of claim 17 , wherein the second processor is configured to apply a compressor model to compress speech segments based on the content of the speech.

19. The system of claim 18 , wherein the second processor is configured to remove at least one segment or combine at least two segments into a compressed speech segment.

20. The system of claim 18 , wherein the second processor is configured to adjust a compression level of the compressor model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 30, 2023
From: FANTINUOLI, CLAUDIO
To: KUDO, INC.
Reel/Frame 064132/0920 →
Continuity (2)
Continuation 17977555 · Oct 31, 2022
Related Publication 20240143947A1 · May 2, 2024