IP Library Granted Patent US 12,739,142
Granted Patent B2
US 12,739,142 · App. 18/534,558 · Granted Sep 15, 2026

Audio transcription for electronic conferencing

Inventors: Christopher Maury (Pittsburgh, PA); James A. Forrest (Pittsburgh, PA); Christopher M. Garrido (Santa Clara, CA); Patrick Miauton (Redwood City, CA)
Assignee: Apple Inc.
H04L12/1831G06F16/683H04L12/1822
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,739,142
App. No.
18/534,558
Granted
Sep 15, 2026
Kind
B2
Abstract

Aspects of the subject technology provide for transcription of audio content during a conferencing session, such as an audio conferencing session or a video conferencing session. The transcription can be generated by the device at which the audio input is received, and transmitted to a remote device at which the transcription is displayed. Video content can also be provided from the device that generates the transcription to the remote device that displays in the transcription. The transcription can be provided with time information corresponding to time information in the video content, for synchronized display of the transcription and the corresponding video content.

Claims (56)

1 . A method, comprising:

during a conferencing session between at least a first device and a second device:

receiving, by the first device, a first audio input;

generating, by the first device using a voice model previously stored at the first device and having been trained on one or more voice inputs from a user of the first device, a first transcription of the first audio input; and

sending the first transcription from the first device to the second device.

2 . The method of claim 1 , further comprising, during the conferencing session, sending a first audio stream corresponding to the first audio input from the first device to the second device with the first transcription.

3 . The method of claim 1 , further comprising, during the conferencing session:

receiving, by the first device, a first video input; and

sending a first video stream corresponding to the first video input from the first device to the second device.

4 . The method of claim 3 , further comprising, during the conferencing session, sending time information corresponding to the first transcription from the first device to the second device.

5 . The method of claim 1 , further comprising, during the conferencing session:

receiving an audio stream at the first device from the second device; and

generating an audio output corresponding to the audio stream, wherein the first device does not generate a transcription of the received audio stream.

6 . The method of claim 1 , wherein the first transcription is associated with a corresponding confidence score, the method further comprising:

after sending the first transcription from the first device to the second device, sending an update to the first transcription from the first device to the second device, the update associated with an updated corresponding confidence score.

7 . The method of claim 1 , wherein sending the first transcription from the first device to the second device comprises sending the first transcription integrated into a video stream from the first device to the second device.

8 . The method of claim 1 , further comprising:

receiving a transcription request at the first device from the second device; and

generating the first transcription based on receiving the transcription request.

9 . The method of claim 1 , further comprising:

providing, from the first device to the second device, a request for a transcription of an audio input corresponding to the second device;

determining, by the first device, that the second device is unable to generate the transcription of the audio input corresponding to the second device;

receiving, at the first device from the second device, an audio stream corresponding to the audio input corresponding to the second device; and

generating, by the first device, a transcription of the audio stream received from the second device.

10 . The method of claim 1 , further comprising:

receiving an audio stream at the first device from the second device; and

in accordance with one or more first criteria being met:

generating, by the first device, a third transcription corresponding to the audio stream from the second device; and

providing the third transcription to a third device.

11 . The method of claim 10 , wherein the one or more first criteria includes a criterion that is based on computing capabilities of the first device and a fourth device.

12 . The method of claim 1 , further comprising:

after sending the first transcription, receiving, by the first device, a request to end the conferencing session; and

ending the conferencing session responsive to the request to end the conferencing session.

13 . A method, comprising:

providing, from a first device to a second device during a conferencing session between at least the first device and the second device, a request for a transcription of an audio input corresponding to the second device;

determining, by the first device, that the second device is unable to generate the transcription of the audio input corresponding to the second device;

receiving, at the first device from the second device, an audio stream corresponding to the audio input corresponding to the second device; and

generating, by the first device, a transcription of the audio stream received from the second device.

14 . The method of claim 13 , wherein the audio input comprises a first audio input, and the transcription comprises a first transcription, the method further comprising, during the conferencing session:

displaying the first transcription at the first device;

receiving, by the first device, a second audio input;

generating, by the first device, a second transcription of the second audio input; and

sending the second transcription from the first device to the second device.

15 . The method of claim 14 , further comprising, during the conferencing session, sending another audio stream corresponding to the second audio input from the first device to the second device with the second transcription.

16 . The method of claim 13 , wherein determining that the second device is unable to generate the transcription of the audio input corresponding to the second device comprises receiving an indication from the second device or a server that the second device does not have a transcription capability.

17 . A method, comprising:

receiving an audio stream at a first device from a second device during a conferencing session between at least the first device, the second device, and a third device; and

in accordance with one or more criteria being met:

generating, by the first device, a transcription corresponding to the audio stream from the second device; and

providing the transcription to the third device.

18 . The method of claim 17 , wherein the one or more criteria includes a criterion that is based on computing capabilities of the first device and a fourth device.

19 . The method of claim 18 , wherein the computing capabilities of the first device and the fourth device comprise an audio transcription capability that is available at first device and the fourth device and that is unavailable at the second device.

20 . The method of claim 19 , wherein:

the computing capabilities of the first device and the fourth device further comprise, for each of the first device and the fourth device, one or more of: a processor speed, a memory size, a battery power, or a network connection quality; and

the method further comprises generating the transcription at the first device responsive to a nomination of the first device, from among the first device and the fourth device, based on the computing capabilities of the first device and the fourth device.

21 . The method of claim 1 , wherein the voice model has been trained locally at the first device to generate transcriptions of audio inputs.

Continuity (3)
Continuation 17723459 · Apr 18, 2022
Provisional Application 63197485 · Jun 6, 2021
Related Publication 20240113905A1 · Apr 4, 2024
References Cited (42)
US 5563804A · Mortensen · 1996 [cited by applicant]
US 5867654A · Ludwig · 1999 [cited by examiner]
US 7106381B2 · Molaro et al. · 2006 [cited by applicant]
US 7756923B2 · Caspi · 2010 [cited by applicant]
US 8407049B2 · Cromack · 2013 [cited by applicant]
US 8812510B2 · Romanov · 2014 [cited by examiner]
US 9071728B2 · Begeja · 2015 [cited by applicant]
US 9420227B1 · Shires et al. · 2016 [cited by applicant]
US 10019989B2 · Gauci · 2018 [cited by applicant]
US 10788963B2 · Quinn · 2020 [cited by examiner]
US 10992907B1 · Bran · 2021 [cited by examiner]
US 11165597B1 · Matsuguma · 2021 [cited by applicant]
US 11594227B2 · Hermanns · 2023 [cited by applicant]
US 11699456B2 · Donofrio · 2023 [cited by applicant]
US 11876632B2 · Maury · 2024 [cited by examiner]
US 12020708B2 · Bradley · 2024 [cited by examiner]
US 20020069069A1 · Kanevsky · 2002 [cited by applicant]
US 20030169330A1 · Ben-Shachar · 2003 [cited by applicant]
US 20060190250A1 · Saindon · 2006 [cited by applicant]
US 20070143103A1 · Asthana et al. · 2007 [cited by applicant]
US 20080066001A1 · Majors · 2008 [cited by examiner]
US 20080129864A1 · Stone et al. · 2008 [cited by applicant]
US 20080204587A1 · Takahara et al. · 2008 [cited by applicant]
US 20080295040A1 · Crinon · 2008 [cited by applicant]
US 20100268534A1 · Kishan Thambiratnam · 2010 [cited by applicant]
US 20100333175A1 · Cox · 2010 [cited by applicant]
US 20130311177A1 · Bastide · 2013 [cited by applicant]
US 20140108288A1 · Calman · 2014 [cited by applicant]
US 20150002611A1 · Thapliyal · 2015 [cited by applicant]
US 20170236532A1 · Reynolds · 2017 [cited by examiner]
US 20170317843A1 · Brieskorn · 2017 [cited by examiner]
US 20190147903A1 · Totzke · 2019 [cited by applicant]
US 20200036546A1 · Soni · 2020 [cited by examiner]
US 20210074298A1 · Coeytaux · 2021 [cited by applicant]
US 20220046066A1 · Rothschild · 2022 [cited by applicant]
US 20220115019A1 · Bradley · 2022 [cited by applicant]
US 20220191055A1 · Christensen · 2022 [cited by examiner]
US 20220198140A1 · Trim · 2022 [cited by applicant]
US 20220385758A1 · Tadesse · 2022 [cited by applicant]
Liu J et al. CN118379990 A, Voice data processing method, device and equipment, 2021, 16 pages (Year: 2021). [cited by examiner]
European Patent Application No. 22738111.8; Office Action dated Jul. 18, 2025, 6 pages. [cited by applicant]
Indian Patent Application No. 202317075805; Office Action dated Nov. 11, 2025, 5 pages with English translation. [cited by applicant]