IP Library Granted Patent US 12,700,330
Granted Patent B2
US 12,700,330 · App. 18/337,187 · Granted Aug 4, 2026

Augmented reality speech-to-text captioning

Inventors: Jill S. Dhillon (Jupiter, FL); Jennifer M. Hatfield (Portland, OR); Tushar Agrawal (West Fargo, ND); Jeremy R. Fox (Georgetown, TX)
Assignee: International Business Machines Corporation
G09B21/009G06F3/013G06F3/017G06F3/167G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,700,330
App. No.
18/337,187
Filed
Jun 19, 2023
Granted
Aug 4, 2026
Kind
B2
Art Unit
3715
USPC
434/PCA.01
Abstract

An approach for improving communications for hearing challenged individuals. The approach captures audio communication and video communication associated with a speaker presentation. The approach converts the audio communication to text. The approach sends the text to an augmented reality (AR) device, associated with a user having hearing challenges, as captions. The approach stores the audio communication, the video communication, and the captions for replay by the user.

Claims (36)

1 . A computer-implemented method for improving audio communications, the computer-implemented method comprising:

capturing audio communication and video communication associated with a speaker presentation to a user;

converting the audio communication to text, the converting performed responsive to a machine learning prediction that comprehension of the audio by the user will fall below a threshold level of comprehension;

sending the text to an augmented reality (AR) device as captions for the audio communication; and

displaying the captions along with the video communication on the AR device.

2 . The computer-implemented method of claim 1 , wherein the sending is based on receiving a request from the user to begin sending audio captions to the AR device, the request facilitating the machine learning prediction.

3 . The computer-implemented method of claim 2 , wherein the request is based on at least one of a predetermined eye movement or a predetermined hand swipe gesture.

4 . The computer-implemented method of claim 1 , wherein the converting comprises translating a language of the text generated from a spoken language associated with the speaker presentation to a proficient language associated with the user.

5 . The computer-implemented method of claim 1 , wherein the sending requires valid security credentials.

6 . The computer-implemented method of claim 1 , wherein the AR device comprises at least one of glasses, a mobile phone, a tablet computer or a smart watch.

7 . The computer-implemented method of claim 1 , wherein the converting is further performed based on factors comprising at least one of the audio communication, speaker sentiment, context of the audio communication, and environment of the audio communication.

8 . The computer-implemented method of claim 1 , further comprising:

storing the audio communication, the text, and the video communication for replay.

9 . The computer-implemented method of claim 8 , further comprising:

responsive to a request to replay the audio and video communications, rendering the stored text as audio captions superimposed on top of the video communication.

10 . A computer system for improving audio communications, the computer system comprising:

one or more computer processors;

one or more non-transitory computer readable storage media; and

program instructions stored on the one or more non-transitory computer readable storage media, the program instructions comprising:

program instructions to capture audio communication and video communication associated with a speaker presentation to a user;

program instructions to convert, responsive to a machine learning prediction that comprehension of the audio by the user will fall below a threshold level of comprehension, the audio communication to text;

program instructions to send the text to an augmented reality (AR) device as captions for the audio communication; and

program instructions to display the captions along with the video communication on the AR device.

11 . The computer system of claim 10 , wherein the sending is based on receiving a request from the user to begin sending audio captions to the AR device, the request facilitating the machine learning prediction.

12 . The computer system of claim 11 , wherein the request is based on at least one of a predetermined eye movement or a predetermined hand swipe gesture.

13 . The computer system of claim 10 , wherein the converting comprises translating a language of the text generated from a spoken language associated with the speaker presentation to a written language associated with the user according to the user comprehension of the spoken language.

14 . The computer system of claim 10 , wherein the sending requires valid security credentials.

15 . The computer system of claim 10 , wherein the AR device comprises at least one of glasses, a mobile phone, a tablet computer or a smart watch.

16 . The computer system of claim 10 , wherein the converting is further performed based on factors comprising at least one of the audio communication, speaker sentiment, context of the audio communication, and environment of the audio communication.

17 . A computer program product for improving audio communications, the computer program product comprising:

one or more computer readable storage media; and

program instructions stored on the one or more computer readable storage media, the program instructions comprising:

program instructions to capture audio communication and video communication associated with a speaker presentation to a user;

program instructions to convert, responsive to a machine learning prediction that comprehension of the audio by the user will fall below a threshold level of comprehension, the audio communication to text;

program instructions to send the text to an augmented reality (AR) device as captions for the audio communication; and

program instructions to display the captions along with the video communication on the AR device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 19, 2023
From: DHILLON, JILL S.; HATFIELD, JENNIFER M.; AGRAWAL, TUSHAR; FOX, JEREMY R.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 063985/0366 →
Continuity (1)
Related Publication 20240420591A1 · Dec 19, 2024
References Cited (27)
US 10187512B2 · Singh · 2019 [cited by examiner]
US 10347254B2 · Forutanpour · 2019 [cited by applicant]
US 10567569B2 · Singh · 2020 [cited by examiner]
US 10878819B1 · Chavez · 2020 [cited by applicant]
US 11122341B1 · Decrop · 2021 [cited by applicant]
US 20140129207A1 · Bailey · 2014 [cited by applicant]
US 20180033329A1 · Suleiman · 2018 [cited by applicant]
US 20180091643A1 · Singh · 2018 [cited by examiner]
US 20180098002A1 · Peterson · 2018 [cited by applicant]
US 20190132439A1 · Singh · 2019 [cited by examiner]
US 20200077136A1 · Kwatra · 2020 [cited by applicant]
US 20200228452A1 · Boss · 2020 [cited by applicant]
US 20210160583A1 · Hirtzel · 2021 [cited by applicant]
US 20220312128A1 · Rosenwein · 2022 [cited by examiner]
“Captioning on Glass”, downloaded from the Internet on Jan. 27, 2023, 4 pages, <https://cog.gatech.edu/>. [cited by applicant]
“IBM 5G and edge computing”, downloaded from the Internet on Jan. 27, 2023, 24 pages, <https://www.ibm.com/downloads/cas/0WOR6ORJ>. [cited by applicant]
“IBM accessibility requirements—IBM Accessibility”, downloaded from the Internet on Jan. 27, 2023, <https://www.ibm.com/able/requirements/requirements>, 8 pages. [cited by applicant]
“IBM Equal Access Toolkit—IBM Accessibility”, downloaded from the Internet on Jan. 27, 2023, <https://www.ibm.com/able/toolkit/>, 3 pages. [cited by applicant]
Desalvo et al., “Augmented reality Head Mounted Display for viewing Closed Captioned and/or subtitled streams,” An IP.com Prior Art Database Technical Disclosure, IP.com No. IPCOM000207866D, IP.com Electronic Publicatio… [cited by applicant]
Grossi, Alexandra, “Diversely Deaf IBMers Work to Create An Accessible Future”, IBM, Dec. 3, 2019, 6 pages, <https://www.ibm.com/blogs/age-and-ability/2019/12/03/diversely-deaf-ibmers-work-to-create-an-accessible-future… [cited by applicant]
IBM Careers Blog, “Providing an Inclusive Work Environment for the Deaf and Hearing Impaired”, Dec. 11, 2019, <https://www.ibm.com/blogs/jobs/2019/12/11/providing-a-work-environment-for-the-deaf-and-hearing-impaired/>, … [cited by applicant]
IBM, “Artificial Intelligence”, IBM Research—Israel, downloaded from the Internet on Jan. 27, 2023, <https://research.ibm.com/haifa/dept/imt/index.shtml>, 9 pages. [cited by applicant]
IBM, “Natural Language Processing”, IBM Research, downloaded from the Internet on Jan. 27, 2023, <https://research.ibm.com/topics/natural-language-processing>, 8 pages. [cited by applicant]
IBM, “New to accessibility”, IBM Accessibility Central, downloaded from the Internet on Jan. 27, 2023, <https://pages.github.ibm.com/IBMa/able/Get_Started/New_to_Accessibility/>, 7 pages. [cited by applicant]
IBM, “Tools & automation”, IBM Accessibility Central, downloaded from the Internet on Jan. 27, 2023, <https://pages.github.ibm.com/IBMa/able/Find_accessibility_tools/Find_accessibility_tools#speech-recognition>, 14 page… [cited by applicant]
IBM, “Watson Speech to Text”, downloaded from the Internet on Jan. 27, 2023, <https://www.ibm.com/cloud/watson-speech-to-text>, 9 pages. [cited by applicant]
IBM, “Watson Text to Speech”, downloaded from the Internet on Jan. 27, 2023, 9 pages, <https://www.ibm.com/cloud/watson-text-to-speech>. [cited by applicant]