IP Library Granted Patent US 11,876,632
Granted Patent B2
US 11,876,632 · App. 17/723,459 · Granted Jan 16, 2024

Audio transcription for electronic conferencing

Inventors: Christopher Maury (Pittsburgh, PA); James A. Forrest (Pittsburgh, PA); Christopher M. Garrido (Santa Clara, CA); Patrick Miauton (Redwood City, CA)
Assignee: Apple Inc.
H04L12/1831G06F16/683H04L12/1822
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,876,632
App. No.
17/723,459
Granted
Jan 16, 2024
Kind
B2
Abstract

Aspects of the subject technology provide for transcription of audio content during a conferencing session, such as an audio conferencing session or a video conferencing session. The transcription can be generated by the device at which the audio input is received, and transmitted to a remote device at which the transcription is displayed. Video content can also be provided from the device that generates the transcription to the remote device that displays in the transcription. The transcription can be provided with time information corresponding to time information in the video content, for synchronized display of the transcription and the corresponding video content.

Claims (70)

1. A method, comprising:

during a conferencing session between at least a first device and a second device:

receiving, by the first device, a first audio input;

generating, by the first device using a voice model previously stored at the first device and having been trained on one or more voice inputs from a user of the first device, a first transcription of the first audio input; and

sending the first transcription from the first device to the second device; and

during the conferencing session and after sending the first transcription:

receiving, by the first device, a second audio input;

generating, by the first device, a second transcription of the second audio input; and

sending the second transcription from the first device to the second device.

2. The method of claim 1 , further comprising, during the conferencing session, sending a first audio stream corresponding to the first audio input from the first device to the second device with the first transcription.

3. The method of claim 1 , further comprising, during the conferencing session:

receiving, by the first device, a first video input; and

sending a first video stream corresponding to the first video input from the first device to the second device.

4. The method of claim 3 , further comprising, during the conferencing session, sending time information corresponding to the first transcription from the first device to the second device.

5. The method of claim 1 , further comprising, during the conferencing session:

receiving an audio stream at the first device from the second device; and

generating an audio output corresponding to the audio stream, wherein the first device does not generate a transcription of the received audio stream.

6. The method of claim 1 , wherein the first transcription is associated with a corresponding confidence score, the method further comprising:

after sending the first transcription from the first device to the second device, sending an update to the first transcription from the first device to the second device, the update associated with an updated corresponding confidence score.

7. The method of claim 1 , wherein sending the first transcription from the first device to the second device comprises sending the first transcription integrated into a video stream from the first device to the second device.

8. The method of claim 1 , further comprising:

receiving a transcription request at the first device from the second device; and

generating the first transcription and the second transcription based on receiving the transcription request.

9. The method of claim 1 , further comprising:

providing, from the first device to the second device, a request for transcription of audio input corresponding to the second device;

determining, by the first device, that the second device is unable to generate the transcription of the audio input corresponding to the second device;

receiving, at the first device from the second device, an audio stream corresponding to the audio input corresponding to the second device; and

generating, by the first device, a transcription of the audio stream received from the second device.

10. The method of claim 1 , further comprising:

receiving an audio stream at the first device from the second device; and

in accordance with one or more first criteria being met:

generating, by the first device, a third transcription corresponding to the audio stream from the second device; and

providing the third transcription to a third device.

11. The method of claim 10 , wherein the one or more first criteria includes a criterion that is based on computing capabilities of the first device and a fourth device.

12. The method of claim 1 , further comprising:

after sending the second transcription, receiving, by the first device, a request to end the conferencing session; and

ending the conferencing session responsive to the request to end the conferencing session.

13. A non-transitory computer-readable medium storing instructions which, when executed by one or more processors, cause the one or more processors to perform operations comprising:

during a conferencing session between at least a first device and a second device:

receiving, by the first device, a first audio input;

generating, by the first device using a voice model previously stored at the first device and having been trained on one or more voice inputs from a user of the first device, a first transcription of the first audio input; and

sending the first transcription from the first device to the second device; and

during the conferencing session and after sending the first transcription:

receiving, by the first device, a second audio input;

generating, by the first device, a second transcription of the second audio input; and

sending the second transcription from the first device to the second device.

14. The non-transitory computer-readable medium of claim 13 , the operations further comprising, during the conferencing session, sending a first audio stream corresponding to the first audio input from the first device to the second device with the first transcription.

15. The non-transitory computer-readable medium of claim 13 , the operations further comprising, during the conferencing session:

receiving, by the first device, a first video input; and

sending a first video stream corresponding to the first video input from the first device to the second device.

16. The non-transitory computer-readable medium of claim 15 , the operations further comprising, during the conferencing session, sending time information corresponding to the first transcription from the first device to the second device.

17. An electronic device, comprising:

memory; and

one or more processors configured to:

during a conferencing session between at least the electronic device and another device:

receive a first audio input;

generate, using a voice model previously stored at the electronic device and having been trained on one or more voice inputs from a user of the electronic device, a first transcription of the first audio input; and

send the first transcription to the other device; and

during the conferencing session and after sending the first transcription:

receive a second audio input;

generate a second transcription of the second audio input; and

send the second transcription to the other device.

18. The electronic device of claim 17 , wherein the one or more processors are further configured to, during the conferencing session, send a first audio stream corresponding to the first audio input to the other device with the first transcription.

19. The electronic device of claim 17 , wherein the one or more processors are further configured to, during the conferencing session:

receive a first video input; and

send a first video stream corresponding to the first video input to the other device; and

send time information corresponding to the first transcription to the other device.

20. The electronic device of claim 17 , wherein the one or more processors are further configured to, during the conferencing session:

receive an audio stream from the other device; and

generate an audio output corresponding to the audio stream, wherein the one or more processors do not generate a transcription of the received audio stream.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 25, 2022
From: MAURY, CHRISTOPHER; FORREST, JAMES A.; MIAUTON, PATRICK; GARRIDO, CHRISTOPHER M.
To: APPLE INC.
Reel/Frame 060020/0143 →
Continuity (2)
Provisional Application 63197485 · Jun 6, 2021
Related Publication 20220393898A1 · Dec 8, 2022
Cited By (1)
US 12,739,142