IP Library Granted Patent US 12,424,198
Granted Patent B2
US 12,424,198 · App. 18/069,380 · Granted Sep 23, 2025

Word replacement during poor network connectivity or network congestion

Inventor: Nick Swerdlow (Santa Clara, CA)
Assignee: Zoom Communications, Inc.
G10L13/08G10L13/027G10L15/22G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,424,198
App. No.
18/069,380
Granted
Sep 23, 2025
Kind
B2
Abstract

A server generates a continuous audio stream during periods of poor network connectivity or network congestion. The server obtains a first audio stream from a user device connected to a real-time communication session and detects speech data in the first audio stream. The server converts the speech data to text data that includes one or more words. The server determines that the text data is missing a word based on a context of the one or more words. The server synthesizes a predicted word for replacing the missing word in a voice of a user of the user device and combines the synthesized word with the first audio stream to generate the continuous audio stream. The server transmits the continuous audio stream to other user devices connected to the real-time communication session.

Claims (62)

1. A method, comprising:

obtaining a first audio stream from a user device connected to a real-time communication session;

detecting first speech data in the first audio stream;

converting the first speech data to first text data including one or more words;

determining that the first text data has a missing word based on a context of the one or more words;

synthesizing a predicted word for replacing the missing word in a voice of a user of the user device to obtain a synthesized word;

combining the synthesized word with the first audio stream to generate a second audio stream, wherein the first audio stream is stored in a buffer for a predetermined duration to combine the synthesized word with the first audio stream; and

transmitting the second audio stream to one or more other user devices connected to the real-time communication session.

2. The method of claim 1 , further comprising:

transmitting a notification to the one or more other user devices that indicates that a portion of the second audio stream is synthesized.

3. The method of claim 1 , wherein the second audio stream contains second speech data, the method further comprising:

converting the second speech data to second text data; and

transmitting the second text data to the one or more other user devices.

4. The method of claim 1 , wherein the first audio stream is stored in a buffer to combine the synthesized word with the first audio stream.

5. The method of claim 1 , further comprising:

determining a speech profile of the user based on real-time communication session recordings associated with the user; and

determining the missing word based on the speech profile.

6. The method of claim 1 , further comprising:

determining the context of the one or more words based on a neighboring word range for a keyword of the one or more words.

7. The method of claim 1 , further comprising:

transmitting a request to the user device; and

receiving, in response to the request, a message that indicates whether the synthesized word is accurate.

8. A system, comprising:

a first user device connected to a real-time communication session;

and

a server configured to:

obtain a first audio stream from a user device connected to the real-time communication session;

detect first speech data in the first audio stream;

convert the first speech data to first text data including one or more words;

determine that the first text data has a missing word based on a context of the one or more words;

synthesize a predicted word for replacing the missing word in a voice of a user of the user device to obtain a synthesized word;

combine the synthesized word with the first audio stream to generate a second audio stream, wherein the first audio stream is stored in a buffer for a predetermined duration to combine the synthesized word with the first audio stream; and

transmit the second audio stream to one or more other user devices connected to the real-time communication session.

9. The system of claim 8 , wherein the server is further configured to:

transmit a notification to the one or more other user devices that indicates that a portion of the second audio stream is synthesized.

10. The system of claim 8 , wherein the server further configured to synthesize the one or more predicted words using a deep neural network.

11. The system of claim 8 , wherein the missing word is determined based on a speech profile.

12. The system of claim 8 , wherein the server is further configured to:

determine a speech profile of the user based on real-time communication session recordings associated with the user.

13. The system of claim 8 , wherein the server is further configured to:

determine the context of the one or more words based on a speech profile of the user.

14. The system of claim 8 , wherein the server is further configured to:

transmit a feedback request to the user device; and

receive, in response to the feedback request, a message that indicates whether the synthesized word is accurate.

15. A non-transitory computer-readable medium comprising instructions stored in a memory, that when executed by a processor, cause the processor to perform operations comprising:

obtaining a first audio stream from a user device connected to a real-time communication session;

detecting first speech data in the first audio stream;

converting the first speech data to first text data including one or more words;

determining that the first text data has a missing word based on a context of the one or more words;

synthesizing a predicted word for replacing the missing word in a voice of a user of the user device to obtain a synthesized word;

combining the synthesized word with the first audio stream to generate a second audio stream, wherein the first audio stream is stored in a buffer for a predetermined duration to combine the synthesized word with the first audio stream; and

transmitting the second audio stream to one or more other user devices connected to the real-time communication session.

16. The non-transitory computer-readable medium of claim 15 , the operations further comprising:

transmitting a notification to the one or more other user devices that indicates that at least one word in the second audio stream is synthesized.

17. The non-transitory computer-readable medium of claim 15 , wherein the second audio stream contains second speech data, the operations further comprising:

converting the second speech data to second text data.

18. The non-transitory computer-readable medium of claim 15 , operations further comprising:

determining a speech profile of the user based on real-time communication session recordings; and

determining the missing word based on the speech profile.

19. The non-transitory computer-readable medium of claim 15 , wherein the predicted word is synthesized in the voice of the user using a vocal model.

20. The non-transitory computer-readable medium of claim 15 , the operations further comprising:

receiving a message that indicates whether the synthesized word is accurate.

Assignments (2)
CHANGE OF NAME Recorded Jan 7, 2025
From: ZOOM VIDEO COMMUNICATIONS, INC.
To: ZOOM COMMUNICATIONS, INC.
Reel/Frame 069839/0593 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 21, 2022
From: SWERDLOW, NICK
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 062169/0344 →