IP Library › Granted Patent US 11,200,881
Granted Patent B2
US 11,200,881 · App. 16/523,507 · Granted Dec 14, 2021

Automatic translation using deep learning

Inventors: Shonda A. Witherspoon (White Plains, NY); Shalisha Witherspoon (White Plains, NY); Bong Jun Ko (Harrington Park, NJ)
Assignee: International Business Machines Corporation
G10L13/00G06F40/58G06N3/0454G06N3/08G10H1/366
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,200,881
App. No.
16/523,507
Granted
Dec 14, 2021
Kind
B2
Abstract

Audio data of an original work is received. Text in the audio data is translated to a target language. The audio data is passed to a first deep learning model to learn voice features in the audio data. The audio data is passed to a second deep learning model to learn audio properties in the audio data. The translated text is synchronized to play in the position of original text of the original work in a synthesized voice. A translated audio data of the original work is created by combining the synchronized translated text in the synthesized voice with music of the audio data.

Claims (42)

1. A computer-implemented method comprising:

receiving audio data of an original work;

translating text in the audio data to a target language;

passing the audio data to a first deep learning model to learn voice features in the audio data;

passing the audio data to a second deep learning model to learn audio properties in the audio data;

synchronizing the translated text to play in the position of original text of the original work in a synthesized voice of the learned voice features; and

creating a translated audio data of the original work by combining the synchronized translated text in the synthesized voice with music of the audio data,

wherein the translating the text takes into account local connotation retention, syllable count, and rhyming scheme.

2. The method of claim 1 , further comprising:

separating the audio data into vocal and music portion, wherein the vocal portion is passed to the second deep learning model to learn audio properties, the audio properties including at least lyrics, notes and rhythm.

3. The method of claim 2 , wherein the translating text in the audio data to a target language comprises translating the lyrics to the target language.

4. The method of claim 1 , further comprising configuring at least one of the local connotation retention, syllable count and rhyming scheme to be considered more dominantly in the translating.

5. The method of claim 4 , wherein the configuring is performed based on user input.

6. The method of claim 5 , further comprising learning a user preference based on the user input.

7. A system comprising:

a hardware processor;

a memory device operatively coupled with the hardware processor;

the hardware processor operable to:

receive audio data of an original work;

translate text in the audio data to a target language;

pass the audio data to a first deep learning model to learn voice features in the audio data;

pass the audio data to a second deep learning model to learn audio properties in the audio data;

synchronize the translated text to play in the position of original text of the original work in a synthesized voice; and

create a translated audio data of the original work by combining the synchronized translated text in the synthesized voice with music of the audio data,

wherein the hardware processor is operable to take into account at least local connotation retention, syllable count, and rhyming scheme in translating the text.

8. The system of claim 7 , wherein the hardware processor is further operable to separate the audio data into vocal and music portion, wherein the vocal portion is passed to the second deep learning model to learn audio properties, the audio properties including at least lyrics, notes and rhythm.

9. The system of claim 8 , wherein the text includes the lyrics.

10. The system of claim 7 , wherein the hardware processor is further operable to configure at least one of the local connotation retention, syllable count and rhyming scheme to be considered more dominantly in the translating.

11. The system of claim 7 , wherein the synthesized voice is a synthesized voice of a selected singer.

12. The system of claim 7 , wherein the synthesized voice is a synthesized voice of the learned voice features from the audio file.

13. A computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a device to cause the device to:

receive audio data of an original work;

translate text in the audio data to a target language;

pass the audio data to a first deep learning model to learn voice features in the audio data;

pass the audio data to a second deep learning model to learn audio properties in the audio data;

synchronize the translated text to play in the position of original text of the original work in a synthesized voice; and

create a translated audio data of the original work by combining the synchronized translated text in the synthesized voice with music of the audio data,

wherein the device is further caused to take into account at least local connotation retention, syllable count, and rhyming scheme in the received audio data of the original work in translating the text.

14. The computer program product of claim 13 , wherein the device is further caused to separate the audio data into vocal and music portion, wherein the vocal portion is passed to the second deep learning model to learn audio properties, the audio properties including at least lyrics, notes and rhythm.

15. The computer program product of claim 14 , wherein the text includes the lyrics.

16. The computer program product of claim 13 , wherein the device is further caused to configure at least one of the local connotation retention, syllable count and rhyming scheme to be considered more dominantly in the translating.

17. The computer program product of claim 16 , wherein the device is further caused to configure at least one of the local connotation retention, syllable count and rhyming scheme to be considered more dominantly in the translating based on user input.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 26, 2019
From: WITHERSPOON, SHONDA A.; WITHERSPOON, SHALISHA; KO, BONG JUN
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 049874/0142 →
Continuity (1)
Related Publication 20210027761A1 · Jan 28, 2021
Cited By (1)
US 12,210,848