IP Library Granted Patent US 10,923,121
Granted Patent B2
US 10,923,121 · App. 16/101,130 · Granted Feb 16, 2021

Method, apparatus, and computer program product for searchable real-time transcribed audio and visual content within a group-based communication system

Inventors: Andrew Locascio (San Francisco, CA); Lynsey Haynes (San Francisco, CA); Jahanzeb Sherwani (San Francisco, CA); Jason DiCioccio (Santa Clara, CA)
Assignee: SlackTechnologies, Inc.
G10L15/22G06F16/685G06K9/00228G06K9/00288G10L15/1822G10L15/26G10L15/30H04L12/1831H04L51/066H04L12/1822H04M2203/50H04N7/15
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,923,121
App. No.
16/101,130
Granted
Feb 16, 2021
Kind
B2
Abstract

Embodiments of the present disclosure provide methods, systems, apparatuses, and computer program products for generating a searchable transcript of a group-based audio/video connection within a group-based communication system.

Claims (93)

1. An apparatus for generating a searchable transcript of a group-based audio feed for display within a group-based communication channel interface of a group-based communication system, the group-based communication system comprising a plurality of users organized among a plurality of group-based communication channels, the apparatus comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus to:

receive a group-based audio feed comprised of a plurality of sub-feeds, each sub-feed received from a respective client device, and each sub-feed comprising a plurality of sequential audio snippets, the group-based audio feed associated with a group-based communication channel and a group-based communication connection, wherein the group-based communication channel comprises a communications feed configured to display messaging communications transmitted and viewable by channel members according to access controls associated with the group-based communication channel;

for each sub-feed received from a client device,

for each sequential audio snippet,

using a speech recognition engine, convert the sequential audio snippet to a final assembled text string;

assign a connection sequence number to the final assembled text string; and

assign a group-based communication channel identifier, a user identifier, and a group-based audio feed identifier to the final assembled text string, wherein the user identifier is associated with the group-based communication channel identifier and the group-based communication channel identifier is associated with the group-based communication channel;

transmit, while the group-based communication connection is occurring, to each of the respective client devices, instructions for rendering a group-based communication channel interface comprising the final assembled text strings arranged according to their respective connection sequence number into a real-time transcript that is simultaneously displayed at each of the respective client devices, the group-based communication channel interface associated with the group-based communication channel identifier; and

upon completion of the group-based communication connection, index the real-time transcript into a store the searchable transcript for storage in a group-based communication repository and searching within the group-based communication system.

2. The apparatus of claim 1 , wherein the at least one memory and the computer program code are further configured to, with the at least one processor, cause the apparatus to:

transmit to each of the respective client devices a group-based communication channel interface comprising the temporary assembled text strings in a temporary format; and

upon assembly of the final assembled text strings, transmit to each of the respective client devices a group-based communication channel interface comprising the final assembled text strings in a final format.

3. The apparatus of claim 1 , wherein the at least one memory and the computer program code are further configured to, with the at least one processor, cause the apparatus to:

for each final assembled text string,

parse the final assembled text string to identify a spoken informality;

remove the spoken informality from the final assembled text string.

4. The apparatus of claim 3 , wherein the at least one memory and the computer program code are further configured to, with the at least one processor, cause the apparatus to: identify the spoken informality by comparing the parsed final assembled text string to a spoken informality store.

5. The apparatus of claim 4 , wherein the spoken informality store is generated based on a machine learning model.

6. The apparatus of claim 1 , wherein the at least one memory and the computer program code are further configured to, with the at least one processor, cause the apparatus to: determine, using voice recognition, the user identifier associated with one of the sub-feed or sequential audio snippet.

7. The apparatus of claim 1 , wherein the at least one memory and the computer program code are further configured to, with the at least one processor, cause the apparatus to:

determine, based on a client device associated with the sub-feed, the user identifier associated with the sub-feed.

8. The apparatus of claim 1 , wherein each client device comprises a video capturing mechanism, and wherein the at least one memory and the computer program code are further configured to, with the at least one processor, cause the apparatus to:

receive a video signal associated with a sub-feed of the group-based audio feed; and

determine, using facial recognition, the user identifier associated with the sub-feed.

9. The apparatus of claim 1 , wherein the at least one memory and the computer program code are further configured to, with the at least one processor, cause the apparatus to: using the speech recognition engine, determine that the sequential audio snippet does not include speech.

10. The apparatus of claim 1 , wherein the at least one memory and the computer program code are further configured to, with the at least one processor, cause the apparatus to:

receive a video signal associated with a sub-feed of the group-based audio feed; and

determine, using voice recognition and facial recognition, a user associated with the sub-feed and the video signal.

11. The apparatus of claim 1 , wherein the at least one memory and the computer program code are further configured to, with the at least one processor, cause the apparatus to:

detect a spoken name within the group-based audio feed, the spoken name associated with a notification request; and

transmit a notification to a client device associated with a user identifier in the notification request that the spoken name has been detected.

12. The apparatus of claim 1 , wherein the at least one memory and the computer program code are further configured to, with the at least one processor, cause the apparatus to:

detect a topic with the group-based audio feed, the topic associated with a notification request; and

transmit a notification to a client device associated with a user identifier in the notification request that the topic has been detected.

13. A system for generating a searchable transcript of a group-based audio feed for display within a group-based communication channel interface of a group-based communication system, the group-based communication system comprising a plurality of users organized among a plurality of group-based communication channels, the system comprising at least one repository and at least one server comprising at least one processor and at least one memory including computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the system to:

receive a group-based audio feed comprised of a plurality of sub-feeds, each sub-feed received from a respective client device, and each sub-feed comprising a plurality of sequential audio snippets, the group-based audio feed associated with a group-based communication channel and a group-based communication connection, wherein the group-based communication channel comprises a communications feed configured to display messaging communications transmitted and viewable by channel members according to access controls associated with the group-based communication channel;

for each sub-feed received from a client device,

for each sequential audio snippet,

using a speech recognition engine, convert the sequential audio snippet to a final assembled text string;

assign a connection sequence number to the final assembled text string; and

assign a group-based communication channel identifier, a user identifier, and a group-based audio feed identifier to the final assembled text string, wherein the user identifier is associated with the group-based communication channel identifier and the group-based communication channel identifier is associated with the group-based communication channel;

transmit, while the group-based communication connection is occurring, to each of the respective client devices, instructions for rendering a group-based communication channel interface comprising the final assembled text strings arranged according to their respective connection sequence number into a real-time transcript that is simultaneously displayed at each of the respective client devices, the group-based communication channel interface associated with the group-based communication channel identifier; and

upon completion of the group-based communication connection, index the real-time transcript into a searchable transcript for storage in a group-based communication repository and searching within the group-based communication system.

14. The system of claim 13 , wherein the at least one memory and the computer program code are further configured to, with the at least one processor, cause the system to:

for each final assembled text string,

parse the final assembled text string to identify a spoken informality;

remove the spoken informality from the final assembled text string.

15. The system of claim 14 , wherein the at least one memory and the computer program code are further configured to, with the at least one processor, cause the system to: identify the spoken informality by comparing the parsed final assembled text string to a spoken informality store, wherein the spoken informality store is generated based on a machine learning model.

16. The system of claim 13 , wherein the at least one memory and the computer program code are further configured to, with the at least one processor, cause the system to:

determine, using voice recognition, the user identifier associated with one of the sub-feed or sequential audio snippet.

17. The system of claim 13 , wherein the at least one memory and the computer program code are further configured to, with the at least one processor, cause the system to: determine, based on a client device associated with the sub-feed, the user identifier associated with the sub-feed.

18. The system of claim 13 , wherein each client device comprises a video capturing mechanism, and wherein the at least one memory and the computer program code are further configured to, with the at least one processor, cause the apparatus to:

receive a video signal associated with a sub-feed of the group-based audio feed; and

determine, using facial recognition, the user identifier associated with the sub-feed.

19. The system of claim 13 , wherein the at least one memory and the computer program code are further configured to, with the at least one processor, cause the system to: using the speech recognition engine, determine that the sequential audio snippet does not include speech.

20. The system of claim 13 , wherein the at least one memory and the computer program code are further configured to, with the at least one processor, cause the system to:

receive a video signal associated with a sub-feed of the group-based audio feed; and

determine, using voice recognition and facial recognition, a user associated with the sub-feed and the video signal.

21. The system of claim 13 , wherein the at least one memory and the computer program code are further configured to, with the at least one processor, cause the system to:

detect a spoken name or a topic within the group-based audio feed, the spoken name or topic associated with a notification request; and

transmit a notification to a client device associated with a user identifier in the notification request that the detection occurred.

22. The system of claim 13 , wherein the at least one memory and the computer program code are further configured to, with the at least one processor, cause the system to:

for each sub-feed received from a client device,

for each sequential audio snippet,

using a speech recognition engine, convert the sequential audio snippet to a temporary converted text string;

assign a temporary sequence number to the temporary converted text string;

assemble a plurality of temporary converted text strings into a temporary assembled text string according to the temporary sequence number assigned to each temporary converted text string;

for each temporary assembled text string,

parse the temporary assembled text string to identify a conversion error; and

based on any identified conversion errors, assign final sequence numbers to the temporary converted text strings of the temporary assembled text string;

assemble the plurality of temporary text strings of the temporary assembled text strings into a final assembled text string according to the final sequence number assigned to each temporary converted text strings; and

assign a connection sequence number to the final assembled text string.

23. The system of claim 22 , wherein the at least one memory and the computer program code are further configured to, with the at least one processor, cause the system to:

transmit to each of the respective client devices a group-based communication channel interface comprising the temporary assembled text strings in a temporary format; and

upon assembly of the final assembled text strings, transmit to each of the respective client devices a group-based communication channel interface comprising the final assembled text strings in a final format.

24. A computer-implemented method for generating a searchable transcript of a group-based audio feed for outputting to a group-based communication system, the method comprising:

receiving, by a group-based communication server, a group-based audio feed comprised of a plurality of sub-feeds, each sub-feed received from a respective client device, and each sub-feed comprising a plurality of sequential audio snippets;

for each sub-feed received from a client device,

for each sequential audio snippet,

using, by the group-based communication server, a speech recognition engine, converting the sequential audio snippet to a temporary converted text string; and

assigning, by the group-based communication server, a temporary sequence number to the temporary converted text string;

assembling, by the group-based communication server, a plurality of temporary converted text strings into a temporary assembled text string according to the temporary sequence number assigned to each temporary converted text string;

for each temporary assembled text string,

parsing, by the group-based communication server, the temporary assembled text string to identify a conversion error; and

based on any identified conversion errors, assigning, by the group-based communication server, final sequence numbers to the temporary converted text strings of the temporary assembled text string;

assembling, by the group-based communication server, the plurality of temporary text strings of the temporary assembled text strings into a final assembled text string according to the final sequence number assigned to each temporary converted text strings;

assigning, by the group-based communication server, a connection sequence number to the final assembled text string;

assigning, by the group-based communication server, a group-based communication channel identifier, a user identifier, and a group-based audio feed identifier to the final assembled text string, wherein the user identifier is associated with the group-based communication channel identifier;

transmitting, by the group-based communication server, to each of the respective client devices, a group-based communication channel interface comprising the final assembled text strings arranged according to their respective connection sequence number into a searchable transcript; and

storing, by the group-based communication server, the searchable transcript in a group-based communication repository, wherein the searchable transcript is indexed for searching within the group-based communication system.

25. The method of claim 24 , further comprising:

receiving, by the group-based communication server, a search query from a client device;

retrieving, by the group-based communication server and from the group-based communication repository, search results comprising a plurality of searchable transcripts based on parameters extracted from the search query.

Assignments (7)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 4, 2026
From: SLACK TECHNOLOGIES, LLC
To: SALESFORCE, INC.
Reel/Frame 075518/0986 →
CORRECTIVE ASSIGNMENT TO CORRECT THE NEWLY MERGED ENTITY'S NEW NAME, AND TO REMOVE THE PERIOD PREVIOUSLY RECORDED AT REEL: 057254 FRAME: 0738. ASSIGNOR(S) HEREBY CONFIRMS THE MERGER AND CHANGE OF NAME. Recorded Sep 9, 2021
From: SKYLINE STRATEGIES II LLC; SLACK TECHNOLOGIES, INC.
To: SLACK TECHNOLOGIES, LLC
Reel/Frame 057514/0930 →
MERGER AND CHANGE OF NAME Recorded Aug 2, 2021
From: SKYLINE STRATEGIES II LLC; SLACK TECHNOLOGIES, INC.; SLACK TECHNOLOGIES, LLC
To: SLACK TECHNOLOGIES, LLC.
Reel/Frame 057254/0738 →
MERGER Recorded Aug 2, 2021
From: SLACK TECHNOLOGIES, INC.; SKYLINE STRATEGIES I INC.
To: SLACK TECHNOLOGIES, INC.
Reel/Frame 057254/0693 →
RELEASE OF SECURITY INTEREST IN PATENT COLLATERAL AT REEL/FRAME NO. 49332/0349 Recorded Jul 19, 2021
From: MORGAN STANLEY SENIOR FUNDING, INC.
To: SLACK TECHNOLOGIES, INC.
Reel/Frame 057649/0882 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 6, 2020
From: LOCASCIO, ANDREW; HAYNES, LYNSEY; SHERWANI, JAHANZEB; DICIOCCIO, JASON
To: SLACK TECHNOLOGIES, INC.
Reel/Frame 052039/0178 →
PATENT SECURITY AGREEMENT Recorded May 30, 2019
From: SLACK TECHNOLOGIES, INC.
To: MORGAN STANLEY SENIOR FUNDING, INC., AS COLLATERAL AGENT
Reel/Frame 049332/0349 →