IP Library Granted Patent US 12,100,310
Granted Patent B2
US 12,100,310 · App. 18/534,524 · Granted Sep 24, 2024

Interactive reading assistant

Inventors: Barry-John Theobald (Sunnyvale, CA); Russell Y. Webb (San Jose, CA); Nicholas Elia Apostoloff (San Jose, CA)
Assignee: APPLE INC.
G09B19/00G06F3/013G06F3/167G06T11/00G09B5/06G10L15/02G10L15/22G10L25/51G06T2200/24
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,100,310
App. No.
18/534,524
Granted
Sep 24, 2024
Kind
B2
Abstract

A method includes obtaining a speech proficiency value indicator indicative of a speech proficiency value associated with a user of the electronic device. The method further includes in response to determining that the speech proficiency value satisfies a threshold proficiency value: displaying training text via the display device; obtaining, from the audio sensor, speech data associated with the training text, wherein the speech data is characterized by the speech proficiency value; determining, using a speech classifier, one or more speech characterization vectors for the speech data based on linguistic features within the speech data; and adjusting one or more operational values of the speech classifier based on the one or more speech characterization vectors and the speech proficiency value.

Claims (57)

1. A method comprising:

at an electronic device including one or more processors, a non-transitory memory, a display device, and an audio sensor:

obtaining a speech proficiency value indicator indicative of a speech proficiency value associated with a user of the electronic device; and

in response to determining that the speech proficiency value satisfies a threshold proficiency value:

displaying training text via the display device;

obtaining, from the audio sensor, speech data associated with the training text, wherein the speech data is characterized by the speech proficiency value;

determining, using a speech classifier, one or more speech characterization vectors for the speech data based on linguistic features within the speech data; and

adjusting one or more operational values of the speech classifier based on the one or more speech characterization vectors and the speech proficiency value.

2. The method of claim 1 , wherein a customized error threshold for the user is generated by adjusting the one or more operational values of the speech classifier.

3. The method of claim 2 , further comprising:

determining whether or not a difference between the linguistic features and expected linguistic features satisfies the customized error threshold; and

in accordance with a determination that the difference between the linguistic features and the expected linguistic features does not satisfy the customized error threshold, displaying, via the display device, a speaking prompt in order to obtain additional speech data from the audio sensor.

4. The method of claim 3 , further comprising:

in accordance with a determination that the difference between the linguistic features and the expected linguistic features satisfies the customized error threshold, distinguishing, via the display device, an appearance of a portion of text content from the remainder of the text content.

5. The method of claim 1 , further comprising determining whether or not the speech proficiency value satisfies the threshold proficiency value based on user profile data.

6. The method of claim 5 , wherein determining whether or not the speech proficiency value satisfies the threshold proficiency value includes:

obtaining, from an image sensor, image data associated with the user, wherein the image data is included within the user profile data; and

determining whether or not the image data satisfies the threshold proficiency value.

7. The method of claim 5 , wherein determining whether or not the speech proficiency value satisfies the threshold proficiency value includes:

obtaining context data associated with the user, wherein the context data is included within the user profile data; and

determining whether or not the context data satisfies the threshold proficiency value.

8. The method of claim 5 , wherein determining whether or not the speech proficiency value satisfies the threshold proficiency value includes:

obtaining, via the audio sensor, sample speech data associated with the user, wherein the sample speech data is included within the user profile data; and

determining whether or not the sample data satisfies the threshold proficiency value.

9. The method of claim 1 , wherein the one or more speech characterization vectors provide speech-style values characterizing the speech data.

10. The method of claim 1 , wherein the speech classifier corresponds to a neural network.

11. The method of claim 1 , wherein the speech classifier utilizes natural language processing (NLP).

12. The method of claim 1 , further comprising detecting the linguistic features within the speech data.

13. The method of claim 1 , further comprising:

generating one or more speech characterization values based on the speech proficiency value; and

comparing the one or more speech characterization vectors and the one or more speech characterization values in order to adjust the one or more operational values of the speech classifier.

14. An electronic device comprising:

one or more processors;

a non-transitory memory;

an audio sensor;

a display device; and

one or more programs, wherein the one or more programs are stored in the non-transitory memory and configured to be executed by the one or more processors, the one or more programs including instructions for:

obtaining a speech proficiency value indicator indicative of a speech proficiency value associated with a user of the electronic device; and

in response to determining that the speech proficiency value satisfies a threshold proficiency value:

displaying training text via the display device;

obtaining, from the audio sensor, speech data associated with the training text, wherein the speech data is characterized by the speech proficiency value;

determining, using a speech classifier, one or more speech characterization vectors for the speech data based on linguistic features within the speech data; and

adjusting one or more operational values of the speech classifier based on the one or more speech characterization vectors and the speech proficiency value.

15. The electronic device of claim 14 , the one or more programs including instructions for generating a customized error threshold for the user by adjusting the one or more operational values of the speech classifier.

16. The electronic device of claim 15 , the one or more programs including instructions for:

determining whether or not a difference between the linguistic features and expected linguistic features satisfies the customized error threshold; and

in accordance with a determination that the difference between the linguistic features and the expected linguistic features does not satisfy the customized error threshold, displaying, via the display device, a speaking prompt in order to obtain additional speech data from the audio sensor.

17. The electronic device of claim 16 , the one or more programs including instructions for, in accordance with a determination that the difference between the linguistic features and the expected linguistic features satisfies the customized error threshold, distinguishing, via the display device, an appearance of a portion of text content from the remainder of the text content.

18. The electronic device of claim 14 , wherein the one or more speech characterization vectors provide speech-style values characterizing the speech data.

19. The electronic device of claim 14 , wherein the speech classifier corresponds to a neural network.

20. A non-transitory computer readable storage medium storing one or more programs, the one or more programs comprising instructions, which, when executed by an electronic device with one or more processors, an audio sensor, and a display device, cause the electronic device to:

obtain a speech proficiency value indicator indicative of a speech proficiency value associated with a user of the electronic device; and

in response to determining that the speech proficiency value satisfies a threshold proficiency value:

display training text via the display device;

obtain, from the audio sensor, speech data associated with the training text, wherein the speech data is characterized by the speech proficiency value;

determine, using a speech classifier, one or more speech characterization vectors for the speech data based on linguistic features within the speech data; and

adjust one or more operational values of the speech classifier based on the one or more speech characterization vectors and the speech proficiency value.

Continuity (4)
Continuation 17750923 · May 23, 2022
Continuation 16798820 · Feb 24, 2020
Provisional Application 62824158 · Mar 26, 2019
Related Publication 20240105079A1 · Mar 28, 2024