IP Library Granted Patent US 11,871,148
Granted Patent B1
US 11,871,148 · App. 17/544,412 · Granted Jan 9, 2024

Artificial intelligence communication assistance in audio-visual composition

Inventors: Oleksiy Shevchenko (West Vancouver, CA); Ayan Mandal (Oakdale, CA); Bradley Jon Hoover (San Francisco, CA); Joel Tetreault (New York, NY); Maksym Lytvyn (West Vancouver, CA); Dmytro Lider (Kyiv, UA)
Assignee: Grammarly, Inc.
H04N7/147G06F9/453G06N20/00G10L15/197G10L15/22H04N7/148
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,871,148
App. No.
17/544,412
Granted
Jan 9, 2024
Kind
B1
Abstract

In embodiments of the present invention improved capabilities are described for artificial intelligence communication assistance for aiding in the audio-visual composition of electronic communications.

Claims (38)

1. A method of electronic communication assistance, the method comprising:

receiving an audio-visual electronic communication at an artificial intelligence assistant computing platform from a first user, the audio-visual electronic communication comprising an audio communication content and a video communication content, an intended recipient of the audio-visual electronic communication being a second user;

extracting an audio communication information from the audio communication content;

extracting a video communication information from the video communication content;

processing the audio-visual electronic communication with a processor using at least one of a machine learning model, deep learning model, or statistical learning algorithm, to generate feedback and suggestions associated with different target audience characteristics, communication conditions, and communication channel types, for a compositional change for communication content of the audio-visual electronic communication based on the extracted audio communication information and the extracted video communication information;

providing the feedback and the suggestions to the first user.

2. The method of claim 1 , at least one of the video communication information or the audio communication information comprising non-verbal signals.

3. The method of claim 1 , the video communication information comprising body language.

4. The method of claim 1 , the video communication information comprising one or more of facial expressions, postures, and gestures.

5. The method of claim 1 , at least one of the audio communication information or the video communication information comprising at least one of voice tone or communication environment.

6. The method of claim 1 , the audio communication information comprising background noise, the suggestions for the compositional change being generated based on processing the audio-visual electronic communication for an environmental state based on the background noise.

7. The method of claim 1 , the video communication information comprising visual background, the suggestions for the compositional change being made based on processing the audio-visual electronic communication based on the visual background.

8. The method of claim 1 , at least one of the feedback or the suggestions comprising one or more of incremental feedback, general feedback, or specific modifications to a message.

9. A method of electronic communication assistance, the method comprising:

receiving an audio-visual electronic communication at an artificial intelligence assistant computing platform from a first user, the audio-visual electronic communication comprising an audio communication content and a video communication content, an intended recipient of the audio-visual electronic communication being a second user;

extracting an audio communication information from the audio communication content;

extracting a video communication information from the video communication content;

processing the audio-visual electronic communication with a processor using at least one of a machine learning model, deep learning model, or statistical learning algorithm, to generate feedback and suggestions for a compositional change associated with different target audience characteristics, communication conditions, and communication channel types, for communication content of the audio-visual electronic communication using at least one of a sensor input from a wearable user device, a first user communication attribute retrieved from a communication profile for the first user, or a second user communication attribute retrieved from a communication profile for the second user;

providing the feedback and the suggestions to the first user.

10. The method of claim 9 , the suggestions for the compositional change being derived from representations of previous content and context from a plurality of user profiles stored in a communication profile database which are similar to at least one of the communication profile for the first user or the communication profile for the second user.

11. The method of claim 9 , the processor being trained on large-scale data mixed with prior communication and effective communications from a plurality of user profiles.

12. The method of claim 9 , the suggestions for the compositional change being generated using at least one of a machine learning language model or a natural language statistical algorithm.

13. The method of claim 9 , the processor generating the suggestions for the compositional change by optimizing generated language of the first user as determined by the processor from the first user communication attribute.

14. The method of claim 9 , the processor generating the suggestions for the compositional change by replicating a communication style of the first user as determined by the processor from the first user communication attribute.

15. The method of claim 9 , the audio-visual electronic communication further comprising a communication goal, and the processor generating the suggestions for the compositional change by optimizing in respect of the communication goal.

16. The method of claim 9 , further comprising providing a textual representation of the audio-visual electronic communication within a graphical user interface on a screen of a computing device of the first user and providing the suggestions for the compositional change on the screen.

17. The method of claim 9 , at least one of the feedback or the suggestions comprising one or more of incremental feedback, general feedback, or specific modifications to a message.

18. A storage medium storing program instructions capable of being executed by one or more processing devices and which, when executed by the one or more processing devices, cause the one or more processing devices to execute:

receiving an audio-visual electronic communication at an artificial intelligence assistant computing platform from a first user, the audio-visual electronic communication comprising an audio communication content and a video communication content, an intended recipient of the audio-visual electronic communication being a second user;

extracting an audio communication information from the audio communication content;

extracting a video communication information from the video communication content;

processing the audio-visual electronic communication with a processor using at least one of a machine learning model, deep learning model, or statistical learning algorithm, to generate feedback and suggestions associated with different target audience characteristics, communication conditions, and communication channel types, for a compositional change for communication content of the audio-visual electronic communication based on the extracted audio communication information and the extracted video communication information;

providing the feedback and the suggestions to the first user.

19. The storage medium of claim 18 , the video communication information comprising non-verbal signals, the non-verbal signals comprising body language.

20. The storage medium of claim 18 , at least one of the audio communication information or the video communication information comprising at least one of voice tone or communication environment.

21. The storage medium of claim 18 , the audio communication information comprising background noise, and the storage medium further comprising program instructions which when executed by the one or more processing devices cause the one or more processing devices to generate the suggestions for the compositional change based on processing the audio-visual electronic communication for an environmental state based on the background noise.

22. The storage medium of claim 18 , the video communication information comprising visual background, and the storage medium further comprising program instructions which when executed by the one or more processing devices cause the one or more processing devices to generate the suggestions for the compositional change based on processing the audio-visual electronic communication based on the visual background.

23. The storage medium of claim 18 at least one of the feedback or the suggestions comprising one or more of incremental feedback, general feedback, or specific modifications to a message.

Assignments (2)
CHANGE OF NAME Recorded Nov 21, 2025
From: GRAMMARLY, INC.
To: SUPERHUMAN PLATFORM INC.
Reel/Frame 073655/0099 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 7, 2021
From: SHEVCHENKO, OLEKSIY; MANDAL, AYAN; HOOVER, BRADLEY JON; TETREAULT, JOEL; LYTVYN, MAKSYM; LIDER, DMYTRO
To: GRAMMARLY, INC.
Reel/Frame 058327/0951 →
Continuity (3)
Continuation 16905050 · Jun 18, 2020
Continuation 16055036 · Aug 4, 2018
Provisional Application 62541203 · Aug 4, 2017
Cited By (1)
US 12,333,247