IP Library › Granted Patent US 12,190,886
Granted Patent B2
US 12,190,886 · App. 17/448,967 · Granted Jan 7, 2025

Selective inclusion of speech content in documents

Inventors: Sushain Pandit (Austin, TX); Sarbajit K. Rakshit (Kolkata, IN)
Assignee: International Business Machines Corporation
G10L15/26H04W4/029H04W4/33G06T19/006
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,190,886
App. No.
17/448,967
Granted
Jan 7, 2025
Kind
B2
Abstract

In an approach for enabling a user to visualize a transcript of a discussion via a head-mounted AR device and to selectively copy one or more parts of the transcript that is contextually relevant for inclusion in a new document file or a previously created document file, a processor captures audio of a spoken content of a first participant of a discussion via an AR device worn by a user. A processor analyzes the audio of the spoken content of the first participant. A processor converts the audio of the spoken content of the first participant to text to create a transcript. A processor creates a visualization of the transcript. A processor presents the visualization of the transcript to the user via the AR device. A processor enables the user to copy one or more parts of the transcript into a document file via a selection support.

Claims (90)

1. A computer-implemented method comprising:

capturing, by one or more processors, audio of a spoken content of a first participant of a discussion via an augmented reality (AR) device worn by a user;

prior to capturing the audio of the spoken content of the first participant of the discussion via the AR device worn by the user, identifying, by the one or more processors, a plurality of participants, including the first participant, who are each physically present at a physical location to participate in the discussion, wherein each of the plurality of participants are identified based on the AR device that they are wearing;

analyzing, by the one or more processors, the audio of the spoken content of the first participant;

converting, by the one or more processors, the audio of the spoken content of the first participant to text to create a transcript;

creating, by the one or more processors, a visualization of the transcript;

presenting, by the one or more processors, the visualization of the transcript to the user via the AR device; and

enabling, by the one or more processors, the user to copy one or more parts of the transcript into a document file via a selection support,

wherein enabling the user to copy the one or more parts of the transcript into the document file via the selection support further comprises:

enabling, by the one or more processors, the user to point at a location of the first participant; and

selecting, by the one or more processors, one or more parts of the transcript associated with the first participant,

wherein the AR device provides a view of the physical location of the discussion and includes a computer-generated graphic including the visualization of the transcript overlaid on the view of the physical location.

2. The computer-implemented method of claim 1 , further comprising:

assigning, by the one or more processors, a first unique identification number to each participant of the plurality of participants;

authenticating, by the one or more processors, the AR device of each participant of the plurality of participants;

identifying, by the one or more processors, the physical location where the discussion is occurring;

assigning, by the one or more processors, a second unique identification number to the physical location where the discussion is occurring; and

creating, by the one or more processors, an indoor positioning system that represents the physical location where the discussion is occurring.

3. The computer-implemented method of claim 1 , further comprising:

prior to capturing the audio of the spoken content of the first participant of the discussion via the AR device worn by the user, enabling, by the one or more processors, the first participant to assign one or more permissions to each participant of the plurality of participants, wherein the one or more permissions determine whether each participant of the plurality of participants can copy the one or more parts of the transcript into the document file.

4. The computer-implemented method of claim 1 , wherein analyzing the audio of the spoken content of the first participant further comprises:

identifying, by the one or more processors, the first participant by a tone of voice;

determining, by the one or more processors, a direction from which the audio of the spoken content of the first participant originated via a beam forming sensor on the AR device worn by the user and the indoor positioning system;

creating, by the one or more processors, a direction vector representing the direction from which the audio of the spoken content of the first participant originated;

determining, by the one or more processors, a time when the audio of the spoken content of the first participant originated;

creating, by the one or more processors, a time stamp representing the time when the audio of the spoken content of the first participant originated; and

preparing, by the one or more processors, a time scale to organize the audio of the spoken content of the first participant according to when the audio was captured.

5. The computer-implemented method of claim 4 , wherein the transcript includes an identifying factor of the first participant, the direction vector representing the direction from which the audio of the spoken content of the first participant originated, and the time stamp representing the time when the audio of the spoken content of the first participant originated.

6. The computer-implemented method of claim 1 , wherein the document file is a new document file or a previously created document file, and wherein the one or more parts of the transcript are placed in a user-defined position.

7. A computer program product comprising:

one or more computer readable storage media and program instructions stored on the one or more computer readable storage media, the program instructions comprising:

program instructions to capture audio of a spoken content of a first participant of a discussion via an AR device worn by a user;

prior to capturing the audio of the spoken content of the first participant of the discussion via the AR device worn by the user, program instructions to identify, by the one or more processors, a plurality of participants, including the first participant, who are each physically present in a physical location to participate in the discussion, wherein each of the plurality of participants are identified based on the AR device that they are wearing;

program instructions to analyze the audio of the spoken content of the first participant;

program instructions to convert the audio of the spoken content of the first participant to text to create a transcript;

program instructions to create a visualization of the transcript;

program instructions to present the visualization of the transcript to the user via the AR device; and

program instructions to enable the user to copy one or more parts of the transcript into a document file via a selection support,

wherein enabling the user to copy the one or more parts of the transcript into the document file via the selection support further comprises:

program instructions to enable the user to point at a location of the first participant; and

program instructions to select one or more parts of the transcript associated with the first participant,

wherein the AR device provides a view of the physical location of the discussion and includes a computer-generated graphic including the visualization of the transcript overlaid on the view of the physical location.

8. The computer program product of claim 7 , further comprising:

program instructions to assign a first unique identification number to each participant of the plurality of participants;

program instructions to authenticate the AR device of each participant of the plurality of participants;

program instructions to identify the physical location where the discussion is occurring;

program instructions to assign a second unique identification number to the physical location where the discussion is occurring; and

program instructions to create an indoor positioning system that represents the physical location where the discussion is occurring.

9. The computer program product of claim 7 , further comprising:

prior to capturing the audio of the spoken content of the first participant of the discussion via the AR device worn by the user, program instructions to enable the first participant to assign one or more permissions to each participant of the plurality of participants, wherein the one or more permissions determine whether each participant of the plurality of participants can copy the one or more parts of the transcript into the document file.

10. The computer program product of claim 7 , wherein analyzing the audio of the spoken content of the first participant further comprises:

program instructions to identify the first participant by a tone of voice;

program instructions to determine a direction from which the audio of the spoken content of the first participant originated via a beam forming sensor on the AR device worn by the user and the indoor positioning system;

program instructions to create a direction vector representing the direction from which the audio of the spoken content of the first participant originated;

program instructions to determine a time when the audio of the spoken content of the first participant originated;

program instructions to create a time stamp representing the time when the audio of the spoken content of the first participant originated; and

program instructions to prepare a time scale to organize the audio of the spoken content of the first participant according to when the audio was captured.

11. The computer program product of claim 10 , wherein the transcript includes an identifying factor of the first participant, the direction vector representing the direction from which the audio of the spoken content of the first participant originated, and the time stamp representing the time when the audio of the spoken content of the first participant originated.

12. The computer program product of claim 7 , wherein the document file is a new document file or a previously created document file, and wherein the one or more parts of the transcript are placed in a user-defined position.

13. A computer system comprising:

one or more computer processors;

one or more computer readable storage media;

program instructions collectively stored on the one or more computer readable storage media for execution by at least one of the one or more computer processors, the stored program instructions comprising:

program instructions to capture audio of a spoken content of a first participant of a discussion via an AR device worn by a user;

prior to capturing the audio of the spoken content of the first participant of the discussion via the AR device worn by the user, program instructions to identify, by the one or more processors, a plurality of participants, including the first participant, who are each physically present in a physical location to participate in the discussion, wherein each of the plurality of participants are identified based on the AR device that they are wearing;

program instructions to analyze the audio of the spoken content of the first participant;

program instructions to convert the audio of the spoken content of the first participant to text to create a transcript;

program instructions to create a visualization of the transcript;

program instructions to present the visualization of the transcript to the user via the AR device; and

program instructions to enable the user to copy one or more parts of the transcript into a document file via a selection support,

wherein enabling the user to copy the one or more parts of the transcript into the document file via the selection support further comprises:

program instructions to enable the user to point at a location of the first participant; and

program instructions to select one or more parts of the transcript associated with the first participant,

wherein the AR device provides a view of the physical location of the discussion and includes a computer-generated graphic including the visualization of the transcript overlaid on the view of the physical location.

14. The computer system of claim 13 , further comprising:

program instructions to assign a first unique identification number to each participant of the plurality of participants;

program instructions to authenticate the AR device of each participant of the plurality of participants;

program instructions to identify the physical location where the discussion is occurring;

program instructions to assign a second unique identification number to the physical location where the discussion is occurring; and

program instructions to create an indoor positioning system that represents the physical location where the discussion is occurring.

15. The computer system of claim 13 , further comprising:

prior to capturing the audio of the spoken content of the first participant of the discussion via the AR device worn by the user, program instructions to enable the first participant to assign one or more permissions to each participant of the plurality of participants, wherein the one or more permissions determine whether each participant of the plurality of participants can copy the one or more parts of the transcript into the document file.

16. The computer system of claim 13 , wherein analyzing the audio of the spoken content of the first participant further comprises:

program instructions to identify the first participant by a tone of voice;

program instructions to determine a direction from which the audio of the spoken content of the first participant originated via a beam forming sensor on the AR device worn by the user and the indoor positioning system;

program instructions to create a direction vector representing the direction from which the audio of the spoken content of the first participant originated;

program instructions to determine a time when the audio of the spoken content of the first participant originated;

program instructions to create a time stamp representing the time when the audio of the spoken content of the first participant originated; and

program instructions to prepare a time scale to organize the audio of the spoken content of the first participant according to when the audio was captured.

17. The computer system of claim 13 , wherein the document file is a new document file or a previously created document file, and wherein the one or more parts of the transcript are placed in a user-defined position.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 27, 2021
From: PANDIT, SUSHAIN; RAKSHIT, SARBAJIT K.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 057608/0395 →
Continuity (1)
Related Publication 20230267933A1 · Aug 24, 2023
References Cited (38)
US 8954330B2 · Koenig · 2015 [cited by examiner]
US 9135231B1 · Barra · 2015 [cited by examiner]
US 10614828B1 · Cowburn · 2020 [cited by applicant]
US 10878819B1 · Chavez · 2020 [cited by applicant]
US 20050171926A1 · Thione · 2005 [cited by examiner]
US 20120059651A1 · Delgado · 2012 [cited by examiner]
US 20140081634A1 · Forutanpour · 2014 [cited by applicant]
US 20140129207A1 · Bailey · 2014 [cited by applicant]
US 20160261673A1 · Cardonha · 2016 [cited by examiner]
US 20170133007A1 · Drewes · 2017 [cited by examiner]
US 20170263248A1 · Gruber · 2017 [cited by examiner]
US 20180067717A1 · Limaye · 2018 [cited by examiner]
US 20180101550A1 · Calcaterra · 2018 [cited by examiner]
US 20180197624A1 · Robaina · 2018 [cited by examiner]
US 20180307303A1 · Powderly · 2018 [cited by examiner]
US 20180350144A1 · Rathod · 2018 [cited by examiner]
US 20180374483A1 · Florexil · 2018 [cited by applicant]
US 20190317606A1 · Jain · 2019 [cited by examiner]
US 20190362312A1 · Platt · 2019 [cited by examiner]
US 20190377753A1 · Goldberg · 2019 [cited by examiner]
US 20200368616A1 · Delamont · 2020 [cited by examiner]
US 20210173480A1 · Osterhout · 2021 [cited by examiner]
US 20220068276A1 · Munetomo · 2022 [cited by examiner]
US 20220247824A1 · Springer · 2022 [cited by examiner]
US 20220374645A1 · Santoro · 2022 [cited by examiner]
US 20230247418A1 · Sun · 2023 [cited by examiner]
US 20230271083A1 · Cai · 2023 [cited by examiner]
US 20240031261A1 · Mukherjee · 2024 [cited by examiner]
CN 102116941A · 2011 [cited by applicant]
CN 104464389A · 2015 [cited by applicant]
CN 105554662A · 2016 [cited by applicant]
CN 109361527A · 2019 [cited by applicant]
CN 109923462A · 2019 [cited by applicant]
CN 110288861A · 2019 [cited by applicant]
KR 20110024880A · 2011 [cited by applicant]
WO 2019237428A1 · 2019 [cited by applicant]
WO 2021026617A1 · 2021 [cited by applicant]
“Patent Cooperation Treaty PCT International Search Report”, Applicant's File Reference: F22W2440, International Application No. PCT/CN2022/107340, International Filing Date: Jul. 22, 2022, Date of Mailing: Sep. 26, 202… [cited by applicant]