IP Library Granted Patent US 12676139
Granted Patent B1
US 12676139 · App. 19/071,469 · Granted Jul 7, 2026

Mixed reality text narration with dynamic text detection

Inventor: Benjamin Joseph Smith (San Carlos, CA)
Assignee: Meta Platforms Technologies, LLC
G10L13/033G06F3/013G06F3/017G06T19/006G06V20/20G06V20/63
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12676139
App. No.
19/071,469
Granted
Jul 7, 2026
Kind
B1
Abstract

An embodiment includes receiving, via a video camera of a mixed reality headset, a first video frame comprising a first text portion. An embodiment includes narrating the first text portion, the narrating comprising converting, using a trained text-to-speech conversion model, the first text portion to corresponding audio. An embodiment includes detecting, in a second video frame received via the video camera, during the narrating, that the first text portion has been replaced by a second text portion. An embodiment includes adjusting, responsive to the detecting, the narrating, the adjusting comprising ceasing the narrating at a first word in the first text portion and starting the narrating at a second word in the second text portion.

Claims (76)

1 . A computer-implemented method comprising:

receiving a first video frame via a video camera of a mixed reality headset, wherein the first video frame comprises a first text portion;

providing the first text portion to a trained text-to-speech (TTS) model;

receiving a first audio corresponding to the first text portion from the TTS model;

narrating the first text portion based on the first audio;

while narrating the first text portion, (i) receiving a second video frame via the video camera, wherein the second video frame comprises a second text portion and (ii) determining that the first text portion has been replaced by the second text portion in a field of view of the mixed reality headset;

responsive to determining that the first text portion has been replaced by the second text portion, stopping narration of the first text portion;

providing the second text portion to the TTS model;

receiving a second audio corresponding to the second text portion from the TTS model;

after stopping narration of the first text portion, narrating the second text portion based on the second audio;

while narrating the second text portion, (i) receiving a third video frame via the video camera and (ii) determining based on the third video frame that the field of view is directed away from the second text portion; and

responsive to determining that the field of view is directed away from the second text portion, pausing narration of the second text portion.

2 . The computer-implemented method of claim 1 , further comprising:

determining that the second text portion has been in view of the video camera for a predetermined amount of time;

wherein stopping narration of the first text portion is performed responsive to determining that the second text portion has been in view of the video camera for the predetermined amount of time.

3 . The computer-implemented method of claim 1 , further comprising:

while narrating the first text portion, detecting a first gesture indicating a particular word in the first text portion; and

responsive to detecting the first gesture, (i) pausing narration of the first text portion and resuming narration of the first text portion at the particular word, or (ii) adjusting an appearance of the particular word in a display of the mixed reality headset.

4 . The computer-implemented method of claim 1 , further comprising:

while narrating the first text portion, detecting a second gesture; and

responsive to detecting the second gesture, adjusting a narration speed of the first text portion.

5 . The computer-implemented method of claim 1 , further comprising:

prior to providing the first text portion to the TTS model, translating the first text portion from a first human language to a second human language;

wherein the first audio is in the second human language.

6 . The computer-implemented method of claim 1 , further comprising:

while narrating the first text portion, determining an eye gaze location indicating a particular word in the first text portion; and

responsive to detecting the eye gaze location, pausing narration of the first text portion and resuming narration of the first text portion at the particular word.

7 . The computer-implemented method of claim 1 , further comprising:

while narrating the first text portion, detecting an audio input;

recognizing the audio input as a voice command with a trained speech-recognition model; and

responsive to the voice command, adjusting narration of the first text portion.

8 . The computer-implemented method of claim 1 , wherein stopping narration of the first text portion is performed prior to completing narration of the first text portion.

9 . The computer-implemented method of claim 1 , wherein pausing narration of the second text portion is performed prior to completing narration of the second text portion.

10 . A non-transitory, computer-readable medium storing instructions, which when executed by a processor of an electronic device, cause the electronic device to:

receive a first video frame via a video camera of a mixed reality headset, wherein the first video frame comprises a first text portion;

provide the first text portion to a trained TTS model;

receive a first audio corresponding to the first text portion from the TTS model;

narrate the first text portion based on the first audio;

while narrating the first text portion, (i) receive a second video frame via the video camera, wherein the second video frame comprises a second text portion and (ii) determine that the first text portion has been replaced by the second text portion in a field of view of the mixed reality headset;

responsive to determining that the first text portion has been replaced by the second text portion, stop narration of the first text portion;

provide the second text portion to the TTS model;

receive a second audio corresponding to the second text portion from the TTS model;

after stopping narration of the first text portion, narrate the second text portion based on the second audio;

while narrating the second text portion, (i) receive a third video frame via the video camera and (ii) determine based on the third video frame that the field of view is directed away from the second text portion; and

responsive to determining that the field of view is directed away from the second text portion, pause narration of the second text portion.

11 . The non-transitory, computer-readable medium of claim 10 , wherein the instructions further cause the electronic device to:

determine that the second text portion has been in view of the video camera for a predetermined amount of time;

wherein stopping narration of the first text portion is performed responsive to determining that the second text portion has been in view of the video camera for the predetermined amount of time.

12 . The non-transitory, computer-readable medium of claim 10 , wherein the instructions further cause the electronic device to:

while narrating the first text portion, detect a first gesture indicating a particular word in the first text portion; and

responsive to detecting the first gesture, (i) pause narration of the first text portion and resume narration of the first text portion at the particular word, or (ii) adjust an appearance of the particular word in a display of the mixed reality headset.

13 . The non-transitory, computer-readable medium of claim 10 , wherein the instructions further cause the electronic device to:

while narrating the first text portion, detect a second gesture; and

responsive to detecting the second gesture, adjust a narration speed of the first text portion.

14 . The non-transitory, computer-readable medium of claim 10 , wherein the instructions further cause the electronic device to:

prior to providing the first text portion to the TTS model, translate the first text portion from a first human language to a second human language;

wherein the first audio is in the second human language.

15 . The non-transitory, computer-readable medium of claim 10 , wherein the instructions further cause the electronic device to:

while narrating the first text portion, determine an eye gaze location indicating a particular word in the first text portion; and

responsive to detecting the eye gaze location, pause narration of the first text portion and resume narration of the first text portion at the particular word.

16 . The non-transitory, computer-readable medium of claim 10 , wherein stopping narration of the first text portion is performed prior to completing narration of the first text portion.

17 . The non-transitory, computer-readable medium of claim 10 , wherein pausing narration of the second text portion is performed prior to completing narration of the second text portion.

18 . A system comprising:

a processor; and

a non-transitory, computer-readable medium storing instructions, which when executed by the processor, cause the system to:

receive a first video frame via a video camera of a mixed reality headset, wherein the first video frame comprises a first text portion;

provide the first text portion to a trained TTS model;

receive a first audio corresponding to the first text portion from the TTS model;

narrate the first text portion based on the first audio;

while narrating the first text portion, (i) receive a second video frame via the video camera, wherein the second video frame comprises a second text portion and (ii) determine that the first text portion has been replaced by the second text portion in a field of view of the mixed reality headset;

responsive to determining that the first text portion has been replaced by the second text portion, stop narration of the first text portion;

provide the second text portion to the TTS model;

receive a second audio corresponding to the second text portion from the TTS model;

after stopping narration of the first text portion, narrate the second text portion based on the second audio;

while narrating the second text portion, (i) receive a third video frame via the video camera and (ii) determine based on the third video frame that the field of view is directed away from the second text portion; and

responsive to determining that the field of view is directed away from the second text portion, pause narration of the second text portion.