IP Library › Granted Patent US 11,714,595
Granted Patent B1
US 11,714,595 · App. 17/359,227 · Granted Aug 1, 2023

Adaptive audio for immersive individual conference spaces

Inventors: Phil Libin (San Francisco, CA); Leonid Kitainik (San Jose, CA)
Assignee: mmhmm inc.
G06F3/165G06V40/174G06V40/20G10L15/22G10L15/26H04N7/15H04R3/005H04S7/302H04S7/305G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,714,595
App. No.
17/359,227
Granted
Aug 1, 2023
Kind
B1
Abstract

Adapting an audio portion of a video conference includes a presenter providing content for the video conference by delivering live content, prerecorded content, or combining live content with prerecorded content, at least one additional co-presenter provides content for the video conference, and untangling overlapping audio streams of the presenter and the co-presenter by replaying individual audio streams from the presenter and/or the at least one co-presenter or separating the audio streams by diarization. Adapting an audio portion of a video conference may also include recording the presenter to provide a recorded audio stream, using speech-to-text conversion to convert the recorded audio stream to text, correlating the text to the recorded audio stream, retrieving a past portion of the recorded audio stream using a keyword search of the text, and replaying the past portion of the recorded audio stream. The keyword may be entered using a voice recognition system.

Claims (43)

1. A method of adapting an audio portion of a video conference, comprising:

a presenter providing content for the video conference by delivering live content, prerecorded content, or combining live content with prerecorded content;

at least one additional co-presenter provides content for the video conference;

untangling overlapping audio streams of the presenter and the co-presenter by at least one of: replaying individual audio streams from the presenter and the at least one co-presenter or separating the audio streams by diarization;

recording the presenter to provide a recorded audio stream;

using speech-to-text conversion to convert the recorded audio stream to text;

correlating the text to the recorded audio stream;

retrieving a past portion of the recorded audio stream using a keyword search of the text; and

replaying the past portion of the recorded audio stream.

2. A method, according to claim 1 , wherein a corresponding video stream is replayed along with the past portion of the audio stream.

3. A method, according to claim 1 , wherein the keyword is entered using a voice recognition system.

4. A method of adapting an audio portion of a video conference, comprising:

a presenter providing content for the video conference by delivering live content, prerecorded content, or combining live content with prerecorded content;

at least one additional co-presenter provides content for the video conference;

untangling overlapping audio streams of the presenter and the co-presenter by at least one of: replaying individual audio streams from the presenter and the at least one co-presenter or separating the audio streams by diarization;

eliminating background noise by applying filters thereto; and

generating background sounds as a productivity and attention booster.

5. A method, according to claim 4 , wherein background sounds are based on at least one of: audience reaction and presentation specifics.

6. A method, according to claim 5 , wherein audience feedback is acoustically and visually enhanced by changing spatial acoustic properties to emulate acoustic properties of a larger conference room or hall and by zooming out a scene to show the presenter and participants in a virtual conference room, a hall or other shared space using special video features.

7. A method of adapting an audio portion of a video conference, comprising:

a presenter providing content for the video conference by delivering live content, prerecorded content, or combining live content with prerecorded content;

at least one additional co-presenter provides content for the video conference;

untangling overlapping audio streams of the presenter and the co-presenter by at least one of: replaying individual audio streams from the presenter and the at least one co-presenter or separating the audio streams by diarization; and

emulating audience feedback, wherein emulating audience feedback includes providing sounds corresponding to at least one of: a laugh, a sigh, applause, happy exclamations, or angry exclamations.

8. A method, according to claim 7 , wherein emulated audience feedback is controlled by at least one of: a facial recognition component, a gesture recognition component, a speech recognition component, and an expression/emotion recognition component and wherein the recognition components are applied to a visual appearance and an audio stream of the presenter.

9. A method of adapting an audio portion of a video conference, comprising:

a presenter providing content for the video conference by delivering live content, prerecorded content, or combining live content with prerecorded content;

at least one additional co-presenter provides content for the video conference;

untangling overlapping audio streams of the presenter and the co-presenter by at least one of: replaying individual audio streams from the presenter and the at least one co-presenter or separating the audio streams by diarization; and further comprising at least one of:

altering acoustic properties of the audio portion according to at least one of: a number of participants in the video conference and characteristics of a presentation space being emulated for the video conference; or

altering at least one of: pitch, timbre, and expression of at least one of the audio streams provided by the presenter and the co-presenter.

10. A method, according to claim 9 , wherein altering acoustic properties includes varying echo and reverberation levels and intensities.

11. A method of adapting an audio portion of a video conference, comprising:

a presenter providing content for the video conference by delivering live content, prerecorded content, or combining live content with prerecorded content; and

actuating audience microphones to select one of three modes: a first mode where sound from a corresponding audience member is broadcast in real time to all participants of the video conference, a second mode where each of the audience microphones is muted, and a third mode where audio tracks from the audience microphones are captured and broadcast at opportune periods of time, wherein the method further includes at least one of the following features:

the audio tracks are not broadcast to participants of the video conference while the audio tracks are being captured;

when the audience microphones are in the third mode, the audio tracks are captured at a particular one of the audience microphones in response to a corresponding one of the audience members providing a verbal command or actuating a control;

when the audience microphones are in the third mode, the audio tracks are captured at a particular one of the audience microphones in response to the presenter providing a verbal command or actuating a control;

captured, pre-processed, mixed and broadcast audio tracks from the audience microphones represent audience feedback;

the opportune periods of time correspond to pauses in presenter audio caused by seeking audience feedback;

voice direction and location of the presenter is adjusted based on relocation of an image of the presenter; or

in the third mode, audio tracks from the audience microphones are pre-processed and mixed.

12. A method, according to claim 11 , wherein audience feedback is acoustically and visually enhanced by changing spatial acoustic properties to emulate acoustic properties of a larger conference room or hall and by zooming out a scene to show the presenter and participants in a virtual conference room, a hall or other shared space using special video features.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 2, 2021
From: LIBIN, PHIL; KITAINIK, LEONID
To: MMHMM INC.
Reel/Frame 057051/0662 →
Continuity (1)
Provisional Application 63062504 · Aug 7, 2020
Cited By (6)
US 12,283,291 US 12,362,954 US 12,418,765 US 12,554,813 US 12,598,268 US 12,671,953