IP Library Granted Patent US 11,941,345
Granted Patent B2
US 11,941,345 · App. 17/510,994 · Granted Mar 26, 2024

Voice instructed machine authoring of electronic documents

Inventors: Timo Mertens (San Francisco, CA); Vipul Raheja (San Francisco, CA); Chad Mills (Georgetown, TX); Ihor Skliarevskyi (Vancouver, CA); Ignat Blazhko (San Francisco, CA); Robyn Perry (San Mateo, CA); Nicholas Bern (San Francisco, CA); Dhruv Kumar (Vancouver, CA); Melissa Lopez (Ewing, NJ)
Assignee: Grammarly, Inc.
G06F40/166G06F3/0481G06F40/103G06F40/20G10L15/22G10L15/26G06F40/30G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,941,345
App. No.
17/510,994
Granted
Mar 26, 2024
Kind
B2
Abstract

A computer-implemented process is programmed to process a source input, determine text enhancements, and present the text enhancements to apply to the sentences dictated from the source input. A text processor may use machine-learning models to process an audio input to generate sentences in a presentable format. An audio input can be processed by an automatic speech recognition model to generate electronic text. The electronic text may be used to generate sentence structures using a normalization model. A comprehension model may be used to identify instructions associated with the sentence structures and generate sentences based on the instructions and the sentence structures. An enhancement model may be used to identify enhancements to apply to the sentences. The enhancements may be presented alongside sentences generated by the comprehension model to provide the user an option to select either the enhancements or the sentences.

Claims (50)

1. A computer-implemented method comprising, executed by one or more of a computer system and a text processor that is coupled to the one or more computer system via a network:

receiving, at a user interface of the one or more computer systems, a first input to initiate a dictation process;

receiving, in response to an initiation of the dictation process, an audio input from the one or more computer systems;

converting, using an automatic speech recognition model, the audio input to digitally stored electronic text;

generating, using a normalization model, one or more digitally stored sentence structures based on the digitally stored electronic text;

identifying, using a comprehension model, one or more instructions represented in the one or more digitally stored sentence structures;

generating, using the comprehension model, one or more sentences based on the one or more instructions and the one or more sentence structures;

identifying, using an enhancement model to analyze the one or more sentences, one or more enhancements to apply to the one or more sentences that alters the one or more sentences; and

presenting, at a display of the one or more computer systems, the one or more enhancements near the one or more sentences in the user interface of the one or more computer systems, each particular enhancement of the one or more enhancements being selectable via a second input from the one or more computer systems to apply the particular enhancement to the one or more sentences.

2. The method of claim 1 , further comprising formatting and transmitting to the one or more computer systems presentation instructions that are programmed to cause displaying the user interface with a selectable graphical widget to initiate the dictation process and receiving the first input by receiving a selection of the selectable graphical widget.

3. The method of claim 1 , further comprising: formatting and transmitting to the one or more computer systems presentation instructions that are programmed to cause the display of the user interface with a selectable graphical widget to reinitiate the dictation process; receiving a second input from the one or more computer systems specifying selection of the graphical widget; and, in response to the second input, visually erasing the previously generated one or more sentences and the one or more enhancements.

4. The method of claim 1 , further comprising updating the one or more enhancements and the one or more sentences in the user interface in real-time as the audio input is received from the one or more computer systems.

5. The method of claim 1 , further comprising identifying, in the one or more sentence structures using the comprehension model, one or more of a command to generate the one or more sentences, a command to access an email thread, or a command to generate a response to the email thread.

6. The method of claim 1 , further comprising generating the one or more sentences by one or more of rearranging the one or more sentence structures, removing one or more words from the sentence structures, or adding one or more additional words to the sentence structures.

7. The method of claim 1 , further comprising identifying, using the enhancement model to analyze the one or more sentences, one or more of inserting one or more transition words into the one or more sentences, deleting one or more transition words from the one or more sentences, merging two sentences of the one or more sentences, splitting one sentence of the one or more sentences into two separate sentences.

8. The method of claim 1 , further comprising:

receiving from the one or more computer systems a third input specifying a selection of the one or more sentences or one of the one or more enhancements;

accessing an email thread comprising a plurality of email messages directed between the first user and a second user; and

sending the selection as a response within the email thread.

9. The method of claim 8 , further comprising identifying, using the enhancement model to analyze the one or more sentences, the one or more enhancements to apply to the one or more sentences based on the email thread between the first user and the second user.

10. One or more computer-readable non-transitory storage media embodying software that is operable when executed to:

receive, at a user interface of a computer system, a first input to initiate a dictation process;

receive, in response to the initiation of the dictation process, an audio input from the computer system;

convert, using an automatic speech recognition model, the audio input to digitally stored electronic text;

generate, using a normalization model, one or more digitally stored sentence structures based on the text;

identify, using a comprehension model, one or more instructions represented in the one or more digitally stored sentence structures;

generate, using the comprehension model, one or more sentences based on the one or more instructions and the one or more digitally stored sentence structures;

identify, using an enhancement model to analyze the one or more sentences, one or more enhancements to apply to the one or more sentences that alters the one or more sentences; and

present, at a display of the computer system, the one or more enhancements near the one or more sentences in the user interface of the computer system, each particular enhancement of the one or more enhancements being selectable via a second input from the computer system to apply the particular enhancement to the one or more sentences.

11. The media of claim 10 , wherein the software is further operable when executed to format and transmit to the computer system presentation instructions that are programmed to cause displaying the user interface with a selectable graphical widget to initiate the dictation process, and receive the first input by receiving a selection of the selectable graphical widget.

12. The media of claim 10 , wherein the software is further operable when executed to format and transmit to the computer system presentation instructions that are programmed to cause displaying the user interface with a selectable graphical widget to reinitiate the dictation process; receive a second input from the computer system specifying selection of the graphical widget; in response to the second input, visually erase the previously generated one or more sentences and the one or more enhancements.

13. The media of claim 10 , wherein the software is further operable when executed to update the one or more enhancements and the one or more sentences in the user interface in real-time as the audio input is received from the computer system.

14. The media of claim 10 , wherein the software is further operable when executed to identify, in the one or more sentence structures using the comprehension model, one or more of a command to generate the one or more sentences, a command to access an email thread, or a command to generate a response to the email thread.

15. The media of claim 10 , wherein the software is further operable when executed to generate the one or more sentences by one or more of rearranging the one or more sentence structures, removing one or more words from the sentence structures, or adding one or more additional words to the sentence structures.

16. A system comprising: one or more processors; and a non-transitory memory coupled to the processors comprising instructions executable by the processors, the processors operable when executing the instructions to:

receive, at a user interface of a computer system, a first input to initiate a dictation process;

receive, in response to the initiation of the dictation process, an audio input from the computer system;

convert, using an automatic speech recognition model, the audio input to digitally stored electronic text;

generate, using a normalization model, one or more digitally stored sentence structures based on the text;

identify, using a comprehension model, one or more instructions represented in the one or more digitally stored sentence structures;

generate, using the comprehension model, one or more sentences based on the one or more instructions and the one or more digitally stored sentence structures;

identify, using an enhancement model to analyze the one or more sentences, one or more enhancements to apply to the one or more sentences that alters the one or more sentences; and

present, at a display of the computer system, the one or more enhancements near the one or more sentences in the user interface of the computer system, each particular enhancement of the one or more enhancements being selectable via a second input from the computer system to apply the particular enhancement to the one or more sentences.

17. The system of claim 16 , wherein the processors are further operable when executed to format and transmit to the computer system presentation instructions that are programmed to cause displaying the user interface with a selectable graphical widget to initiate the dictation process, and receive the first input by receiving a selection of the selectable graphical widget.

18. The system of claim 16 , wherein the processors are further operable when executed to format and transmit to the computer system presentation instructions that are programmed to cause displaying the user interface with a selectable graphical widget to reinitiate the dictation process; receive a second input from the computer system specifying selection of the graphical widget; in response to the second input, visually erase the previously generated one or more sentences and the one or more enhancements.

19. The system of claim 16 , wherein the processors are further operable when executed to update the one or more enhancements and the one or more sentences in the user interface in real-time as the audio input is received from the computer system.

20. The system of claim 16 , wherein the processors are further operable when executed to identify, in the one or more sentence structures using the comprehension model, one or more of a command to generate the one or more sentences, a command to access an email thread, or a command to generate a response to the email thread.

21. The method of claim 1 , the comprehension model and the enhancement model being the same model.

22. The media of claim 10 , the comprehension model and the enhancement model being the same model.

23. The system of claim 16 , the comprehension model and the enhancement model being the same model.

Assignments (2)
CHANGE OF NAME Recorded Nov 21, 2025
From: GRAMMARLY, INC.
To: SUPERHUMAN PLATFORM INC.
Reel/Frame 073655/0099 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 17, 2021
From: MERTENS, TIMO; RAHEJA, VIPUL; MILLS, CHAD; SKLIAREVSKYI, IHOR; BLAZHKO, IGNAT; PERRY, ROBYN; BERN, NICHOLAS; KUMAR, DHRUV; LOPEZ, MELISSA
To: GRAMMARLY INC.
Reel/Frame 058138/0599 →
Continuity (1)
Related Publication 20230125194A1 · Apr 27, 2023