IP Library Granted Patent US 12,452,390
Granted Patent B2
US 12,452,390 · App. 17/850,860 · Granted Oct 21, 2025

Word flow annotation

Inventors: Jeffrey Scott Sommers (Mountain View, CA); Jennifer M. R. Devine (Plantation, FL); Joseph Wayne Seuck (Tamarac, FL); Adrian Kaehler (Los Angeles, CA)
Assignee: MAGIC LEAP, INC.
H04N7/157G06F40/169G06F40/242G06F40/58G06V20/20G10L15/1815G10L15/26G10L25/84G10L2015/088H04N7/142
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,452,390
App. No.
17/850,860
Granted
Oct 21, 2025
Kind
B2
Abstract

An augmented reality (AR) device can be configured to monitor ambient audio data. The AR device can detect speech in the ambient audio data, convert the detected speech into text, or detect keywords such as rare words in the speech. When a rare word is detected, the AR device can retrieve auxiliary information (e.g., a definition) related to the rare word from a public or private source. The AR device can display the auxiliary information for a user to help the user better understand the speech. The AR device may perform translation of foreign speech, may display text (or the translation) of a speaker's speech to the user, or display statistical or other information associated with the speech.

Claims (65)

1. A method of identifying a thread in a text stream, the method comprising:

identifying a first audio stream at an augmented reality device associated with a first user;

receiving a second audio stream at the augmented reality device;

generating a first text stream from the first audio stream and a second text stream from the second audio stream;

identifying one or more keywords associated with both the first and second text streams, wherein the identifying the one or more keywords associated with the both the first and second text streams includes:

detecting a verbal repeating of the first word by the first user from the first or second text stream and designating the first word as one of the one or more keywords;

identifying an identity of a speaker of the second audio stream;

in response to identifying a plurality of keywords associated with both the first and second text streams, identifying a topic associated with each of the one or more keywords;

in response to identifying more than one unique topic, determining one or more groups of topics based at least in part on whether any of the more than one unique topics are related;

in response to determining more than one group of topics, generating a thread associated with each of the group of topics;

causing at least one of the generated threads to be rendered by the augmented reality device;

determining a context of each of the plurality of keywords based at least in part on the identity of the speaker of the second audio stream;

retrieving auxiliary information associated with the keywords based on the respective context; and

rendering the auxiliary information with the at least one of the generated threads for the keyword associated with topic of the at least one of the generated threads.

2. The method of claim 1 , wherein the first audio stream or the second audio stream is from at least one of: a person or an audio-visual content.

3. The method of claim 1 , wherein the method further comprises:

the tracking of the one or more eyes of the first user with the eye tracking sensors of the augmented reality device to collect the eye tracking information of the first user; and

selectively deemphasize or dismiss at least one of the generated threads based on the eye tracking information of the first user.

4. A method of identifying a thread in a text stream, the method comprising:

identifying a first audio stream at an augmented reality device associated with a first user;

receiving a second audio stream at the augmented reality device;

generating a first text stream from the first audio stream and a second text stream from the second audio stream;

identifying an identity of a speaker of the second audio stream;

parsing the first text stream and the second text stream to identify a first keyword associated with a first topic and a second keyword associated with a second topic, wherein the parsing to identify the first keyword associated with the first topic includes: detecting a verbal repeating of the first word by the first user from the first text stream and designating the first word as the first keyword;

generating a first thread associated with the first topic and a second thread associated with the second topic;

causing at least one of the first thread or the second thread to be rendered by the augmented reality device;

determining a context of each of the first and second keywords based at least in part on the identity of the speaker of the second audio stream;

retrieving auxiliary information associated with the first and second keywords based on the respective context; and

rendering the auxiliary information with the at least one of the generated threads for the keyword associated with topic of the at least one of the generated threads.

5. The method of claim 4 , wherein the first text stream or the second text stream is from at least one of: a person or an audio-visual content.

6. The method of claim 4 , wherein the first text stream is from a first person and the second text stream is from a second person.

7. The method of claim 4 , wherein the first topic further comprises a plurality of sub-topics.

8. The method of claim 4 , wherein both the first thread and the second thread are rendered by the augmented reality device.

9. The method of claim 8 ,

wherein the first thread is rendered by the augmented reality device on the left side of a user's field of view, and the second thread is rendered by the augmented reality device on the right side of the user's field of view, and

wherein the method further comprises tracking one or more eyes of the first user with eye tracking sensors of the augmented reality device to collect eye tracking information of the first user; and selectively deemphasize or dismiss at least one of the first thread or the second thread based on the eye tracking information of the first user.

10. The method of claim 8 , wherein the first thread is rendered by the augmented reality device in a first color and the second thread is rendered by the augmented reality device in a second color.

11. The method of claim 4 , further comprising:

parsing the first text stream and the second text stream to identify a third keyword associated with a third topic; and

generating a third thread associated with the third topic.

12. The method of claim 4 , wherein the method further comprises:

the tracking of the one or more eyes of the first user with the eye tracking sensors of the augmented reality device to collect the eye tracking information of the first user; and

selectively deemphasize or dismiss at least one of the first thread or the second thread when the eye tracking information of the first user indicates that the eyes of first user has tracked an entire display area of the auxiliary information.

13. An augmented reality device comprising a hardware processor and an augmented reality display of an augmented reality device associated with a first user, wherein the hardware processor is programmed to:

identify a first audio stream with the augmented reality device;

detect ambient sounds with an audio sensor of the augmented reality device;

identify a second audio stream within the ambient sounds;

identify an identity of a speaker of the second audio stream;

generate a first text stream from the first audio stream and a second text stream from the second audio stream;

parse the first text stream and the second text stream to identify a first keyword associated with a first topic and a second keyword associated with a second topic, wherein the parsing to identify the first keyword associated with the first topic includes: detecting a verbal repeating of the first word by the first user from the first text stream and designating the first word as the first keyword;

generate a first thread associated with the first topic and a second thread associated with the second topic;

determining a context of the first keyword based at least in part on the identity of the speaker of the second audio stream;

retrieve auxiliary information associated with the first keyword based on the context;

render, on the augmented reality display, at least one of the first thread or the second thread; and

render, on the augmented reality display, the auxiliary information with at least one of the first thread or the second thread.

14. The augmented reality device of claim 13 , wherein the first topic further comprises a plurality of sub-topics.

15. The augmented reality device of claim 13 , wherein the hardware processor is programmed to render both the first thread and the second thread on the augmented reality display.

16. The augmented reality device of claim 15 , wherein the hardware processor is programmed to render the first thread on the left side of the augmented reality display and the second thread on the right side of the augmented reality display.

17. The augmented reality device of claim 15 , wherein the hardware processor is programmed to render the first thread in a first color and the second thread in a second color.

18. The augmented reality device of claim 13 , wherein the hardware processor is further programmed to:

parse the first text stream and the second text stream to identify a third keyword associated with a third topic; and

generate a third thread associated with the third topic.

19. The augmented reality device of claim 13 , further comprising the eye tracking sensors configured to collect the eye tracking information for the first user, and wherein the hardware processor is programmed to:

modify, after a time period, a presentation of the auxiliary information to notify the first user that the auxiliary information is preparing to be dismissed; and

selectively postpone the dismissal the auxiliary information from the augmented reality device when the eye tracking information indicates that the eyes of first user is reading the auxiliary information.

Assignments (4)
SECURITY INTEREST Recorded Oct 28, 2025
From: MAGIC LEAP, INC.; MENTOR ACQUISITION ONE, LLC; MOLECULAR IMPRINTS, INC.
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 073388/0027 →
SECURITY INTEREST Recorded Feb 7, 2023
From: MAGIC LEAP, INC.; MENTOR ACQUISITION ONE, LLC; MOLECULAR IMPRINTS, INC.
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 062681/0065 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 3, 2022
From: SEUCK, JOSEPH WAYNE; SOMMERS, JEFFREY; DEVINE, JENNIFER M.R.
To: MAGIC LEAP, INC.
Reel/Frame 061963/0848 →
PROPRIETARY INFORMATION AND INVENTIONS AGREEMENT Recorded Dec 3, 2022
From: KAEHLER, ADRIAN
To: MAGIC LEAP, INC.
Reel/Frame 062054/0205 →