IP Library › Granted Patent US 11,551,568
Granted Patent B2
US 11,551,568 · App. 17/132,988 · Granted Jan 10, 2023

System and method for dual mode presentation of content in a target language to improve listening fluency in the target language

Inventor: Daniel Paul Raynaud (Austin, TX)
Assignee: JIVEWORLD, SPC
G09B5/06G09B19/06G10L21/10G09B19/04G10L15/08G10L15/26G10L2015/088H04N21/4394H04N21/4542
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,551,568
App. No.
17/132,988
Filed
Dec 23, 2020
Granted
Jan 10, 2023
Kind
B2
Examiner
YIP, JACK
Art Unit
3715
USPC
434/PCA.018
Abstract

Embodiments of a language learning system and method for implementing or assisting in self-study for improving listening fluency in a target language are disclosed. Such embodiments may simultaneously present the same piece of content in an auditory presentation and a corresponding visual presentation of a transcript of the auditory presentation, where the two presentations are adapted to work in tandem to increase the effectiveness of language learning for users.

Claims (69)

1. A system for language learning, comprising:

a server, comprising a processor and a non-transitory computer readable medium comprising instructions to:

receive an audio file comprising content in a target language;

obtain a transcript of the words of the content of the audio file in the target language;

generate a timestamp file based on the audio file and the transcript of the words of the content, the timestamp file including a word level timestamp for each word of the transcript, the word level timestamp for each word of the transcript corresponding to a time in the audio file associated with where that word occurs in the content;

generate a transcript and timing file corresponding to the audio file based on the transcript and the timestamp file, wherein the transcript and timing file comprises each word of the transcript of the content of the audio file and the associated word level timestamp for each word of the transcript of the content of the audio file; and

a client device, comprising a processor and a client application comprising instructions to:

obtain the audio file and the corresponding transcript and timing file;

auditorily present the content of the audio file using an audio player at the client device;

simultaneously with the auditory presentation of the content of the audio file, dynamically generate an interface using the transcript and timing file, the interface including a visual display of a visual transcript of the content in the target language, wherein:

the visual display of the visual transcript of the content is synchronized with the auditory presentation of the content by the audio player and includes a visual transcript of a set of words of the content,

the visual transcript of the set of words of the content includes a set of redacted words and a set of unredacted words,

each of the set of redacted words in the visual display are redacted by presenting the redacted word as a corresponding lozenge in the visual display, the lozenge sized according to the corresponding redacted word,

each of the set of unredacted words are presented in the visual display in text of the target language, and

synchronizing the visual display of the visual transcript of the content with the auditory presentation of the content comprises:

determining that a word is being presented in the auditory presentation of the content based on the word level timestamp associated with that word in the transcript and timing file and a state of the audio player;

highlighting the presentation of that word in the visual display substantially simultaneously with the auditory presentation of that word in the auditory presentation, wherein if the word is in the set of redacted words the lozenge corresponding to that word is highlighted, and if the word is in the set of unredacted words the textual presentation of the word is highlighted.

2. The system of claim 1 , wherein the set of unredacted words include one or more word group types, the word group types including vocabulary, incorrect usage, tricky bits or annotated words.

3. The system of claim 2 , wherein the one or more word group types is selected by a user using the interface.

4. The system of claim 1 , wherein the set of unredacted words includes one or more words determined during dynamic generation of the interface based on a user interaction with the presentation of the one or more words in the interface.

5. The system of claim 1 , wherein the client application comprises instructions for altering a ratio of the set of redacted words to unredacted words based on a desired amount of assistance.

6. The system of claim 1 , wherein the lozenge is shaped based on the content of the corresponding word.

7. The system of claim 1 , wherein the instructions of the non-transitory computer readable medium of the server, or the instructions of the client application, include instructions to:

determine a set of pauses in the auditory presentation of the content of the audio file, wherein the set of pauses are natural pauses associated with the auditory presentation; and

lengthen the determined set of pauses in the auditory presentation of the content.

8. A method for language learning, comprising:

obtaining an audio file comprising content in a target language, and a corresponding transcript and timing file, wherein the transcript and timing file was generated by:

obtaining a transcript of the words of the content of the audio file in the target language;

generating a timestamp file based on the audio file and the transcript of the words of the content, the timestamp file including a word level timestamp for each word of the transcript, the word level timestamp for each word of the transcript corresponding to a time in the audio file associated with where that word occurs in the content; and

generating the transcript and timing file corresponding to the audio file based on the transcript and the timestamp file, wherein the transcript and timing file comprises each word of the transcript of the content of the audio file and the associated word level timestamp for each word of the transcript of the content of the audio file;

auditorily presenting the content of the audio file using an audio player;

simultaneously with the auditory presentation of the content of the audio file, dynamically generating an interface using the transcript and timing file, the interface including a visual display of a visual transcript of the content in the target language, wherein:

the visual display of the visual transcript of the content is synchronized with the auditory presentation of the content by the audio player and includes a visual transcript of a set of words of the content,

the visual transcript of the set of words of the content includes a set of redacted words and a set of unredacted words,

each of the set of redacted words in the visual display are redacted by presenting the redacted word as a corresponding lozenge in the visual display, the lozenge sized according to the corresponding redacted word,

each of the set of unredacted words are presented in the visual display in text of the target language, and

synchronizing the visual display of the visual transcript of the content with the auditory presentation of the content comprises:

determining that a word is being presented in the auditory presentation of the content based on the word level timestamp associated with that word in the transcript and timing file and a state of the audio player;

highlighting the presentation of that word in the visual display substantially simultaneously with the auditory presentation of that word in the auditory presentation, wherein if the word is in the set of redacted words the lozenge corresponding to that word is highlighted, and if the word is in the set of unredacted words the textual presentation of the word is highlighted.

9. The method of claim 8 , wherein the set of unredacted words include one or more word group types, the word group types including vocabulary, incorrect usage, tricky bits or annotated words.

10. The method of claim 9 , wherein the one or more word group types is selected by a user using the interface.

11. The method of claim 8 , wherein the set of unredacted words includes one or more words determined during dynamic generation of the interface based on a user interaction with the presentation of the one or more words in the interface.

12. The method of claim 8 , wherein a ratio of the set of redacted words to unredacted words is altered based on a desired amount of assistance.

13. The method of claim 8 , wherein the lozenge is shaped based on the content of the corresponding word.

14. The method of claim 8 , further comprising:

determining a set of pauses in the auditory presentation of the content of the audio file, wherein the set of pauses are natural pauses associated with the auditory presentation; and

lengthening the determined set of pauses in the auditory presentation of the content.

15. A non-transitory computer readable medium, comprising instructions for:

obtaining an audio file comprising content in a target language, and a corresponding transcript and timing file, wherein the transcript and timing file was generated by:

obtaining a transcript of the words of the content of the audio file in the target language;

generating a timestamp file based on the audio file and the transcript of the words of the content, the timestamp file including a word level timestamp for each word of the transcript, the word level timestamp for each word of the transcript corresponding to a time in the audio file associated with where that word occurs in the content; and

generating the transcript and timing file corresponding to the audio file based on the transcript and the timestamp file, wherein the transcript and timing file comprises each word of the transcript of the content of the audio file and the associated word level timestamp for each word of the transcript of the content of the audio file;

auditorily presenting the content of the audio file using an audio player;

simultaneously with the auditory presentation of the content of the audio file, dynamically generating an interface using the transcript and timing file, the interface including a visual display of a visual transcript of the content in the target language, wherein:

the visual display of the visual transcript of the content is synchronized with the auditory presentation of the content by the audio player and includes a visual transcript of a set of words of the content,

the visual transcript of the set of words of the content includes a set of redacted words and a set of unredacted words,

each of the set of redacted words in the visual display are redacted by presenting the redacted word as a corresponding lozenge in the visual display, the lozenge sized according to the corresponding redacted word,

each of the set of unredacted words are presented in the visual display in text of the target language, and

synchronizing the visual display of the visual transcript of the content with the auditory presentation of the content comprises:

determining that a word is being presented in the auditory presentation of the content based on the word level timestamp associated with that word in the transcript and timing file and a state of the audio player;

highlighting the presentation of that word in the visual display substantially simultaneously with the auditory presentation of that word in the auditory presentation, wherein if the word is in the set of redacted words the lozenge corresponding to that word is highlighted, and if the word is in the set of unredacted words the textual presentation of the word is highlighted.

16. The non-transitory computer readable medium of claim 15 , wherein the set of unredacted words include one or more word group types, the word group types including vocabulary, incorrect usage, tricky bits or annotated words.

17. The non-transitory computer readable medium of claim 16 , wherein the one or more word group types is selected by a user using the interface.

18. The non-transitory computer readable medium of claim 15 , wherein the set of unredacted words includes one or more words determined during dynamic generation of the interface based on a user interaction with the presentation of the one or more words in the interface.

19. The non-transitory computer readable medium of claim 15 , wherein a ratio of the set of redacted words to unredacted words is altered based on a desired amount of assistance.

20. The non-transitory computer readable medium of claim 15 , wherein the lozenge is shaped based on the content of the corresponding word.

21. The non-transitory computer readable medium of claim 15 , further comprising instructions for:

determining a set of pauses in the auditory presentation of the content of the audio file, wherein the set of pauses are natural pauses associated with the auditory presentation; and

lengthening the determined set of pauses in the auditory presentation of the content.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 3, 2021
From: RAYNAUD, DANIEL PAUL
To: JIVEWORLD, SPC
Reel/Frame 055482/0536 →
Continuity (3)
Continuation 16844252 · Apr 9, 2020
Provisional Application 62831380 · Apr 9, 2019
Related Publication 20210183260A1 · Jun 17, 2021
Cited By (2)
US 12,190,752 US 12,424,117