IP Library Granted Patent US 12,149,914
Granted Patent B2
US 12,149,914 · App. 17/669,599 · Granted Nov 19, 2024

Multi-channel speech compression system and method

Inventors: Dushyant Sharma (Mountain House, CA); Patrick A. Naylor (Reading, GB); Uwe Helmut Jost (Groton, MA)
Assignee: Microsoft Technology Licensing, LLC
H04S7/30G06T7/70G10L15/063G10L15/22G10L19/008G10L19/167G10L21/0208H04R1/406H04R3/005H04R5/027H04S3/008G10L2019/0001G10L2019/0002G10L2021/02166H04R2201/401H04S2400/01H04S2400/15
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,149,914
App. No.
17/669,599
Granted
Nov 19, 2024
Kind
B2
Abstract

A method, computer program product, and computing system for obtaining machine vision encounter information using one or more machine vision systems. Audio encounter information may be obtained using a plurality of audio acquisition devices of an audio recording system. The audio encounter information may be encoded using an audio codec. The encoding of the audio encounter information by the audio codec may be adapted based upon, at least in part, the machine vision encounter information.

Claims (28)

1. A computer-implemented method, executed on a computing device, comprising:

obtaining machine vision encounter information using one or more machine vision systems;

obtaining audio encounter information using a plurality of audio acquisition devices of an audio recording system;

encoding the audio encounter information using one or more codecs;

adapting the encoding of the audio encounter information by the one or more codecs based upon, at least in part, the machine vision encounter information; and

generating a plurality of acoustic relative transfer functions between the plurality of audio acquisition devices of the audio recording system,

wherein adapting the encoding of the audio encounter information by the one or more codecs based upon, at least in part, the machine vision encounter information includes adapting the one or more codecs to encode the audio encounter information using one or more acoustic relative transfer functions associated with a particular acoustic source when the machine vision encounter information detects the acoustic source.

2. The computer-implemented method of claim 1 , wherein the plurality of audio acquisition devices of the audio recording system are positioned within a fixed geometry relative to each other.

3. The computer-implemented method of claim 1 , wherein adapting the encoding of the audio encounter information by the one or more codecs based upon, at least in part, the machine vision encounter information includes adapting the one or more codecs to estimate one or more acoustic relative transfer functions when the machine vision encounter information indicates at least a threshold change in the acoustic environment.

4. The computer-implemented method of claim 1 , wherein adapting the encoding of the audio encounter information by the one or more codecs based upon, at least in part, the machine vision encounter information includes adapting the one or more codecs to selectively encode the audio encounter information based upon, at least in part, whether the machine vision encounter information indicates that an audio source is speaking.

5. The computer-implemented method of claim 1 , wherein adapting the encoding of the audio encounter information by the one or more codecs based upon, at least in part, the machine vision encounter information includes adapting the one or more codecs to generate the plurality of acoustic relative transfer functions between the plurality of audio acquisition devices of the audio recording system based upon, at least in part, location information associated with an acoustic source from the machine vision encounter information.

6. A computer program product residing on a non-transitory computer readable medium having a plurality of instructions stored thereon which, when executed by a processor, cause the processor to perform operations comprising:

obtaining machine vision encounter information using one or more machine vision systems;

obtaining audio encounter information using a plurality of audio acquisition devices of an audio recording system;

encoding the audio encounter information using one or more codecs;

adapting the encoding of the audio encounter information by the one or more codecs based upon, at least in part, the machine vision encounter information; and

generating a plurality of acoustic relative transfer functions between the plurality of audio acquisition devices of the audio recording system,

wherein adapting the encoding of the audio encounter information by the one or more codecs based upon, at least in part, the machine vision encounter information includes adapting the one or more codecs to encode the audio encounter information using one or more acoustic relative transfer functions associated with a particular acoustic source when the machine vision encounter information detects the acoustic source.

7. The computer program product of claim 6 , wherein the plurality of audio acquisition devices of the audio recording system are positioned within a fixed geometry relative to each other.

8. The computer program product of claim 6 , wherein adapting the encoding of the audio encounter information by the one or more codecs based upon, at least in part, the machine vision encounter information includes adapting the one or more codecs to estimate one or more acoustic relative transfer functions when the machine vision encounter information indicates at least a threshold change in the acoustic environment.

9. The computer program product of claim 6 , wherein adapting the encoding of the audio encounter information by the one or more codecs based upon, at least in part, the machine vision encounter information includes adapting the one or more codecs to selectively encode the audio encounter information based upon, at least in part, whether the machine vision encounter information indicates that an audio source is speaking.

10. The computer program product of claim 6 , wherein adapting the encoding of the audio encounter information by the one or more codecs based upon, at least in part, the machine vision encounter information includes adapting the one or more codecs to generate the plurality of acoustic relative transfer functions between the plurality of audio acquisition devices of the audio recording system based upon, at least in part, location information associated with an acoustic source from the machine vision encounter information.

11. A computing system comprising:

a memory; and

a processor configured to obtain machine vision encounter information using one or more machine vision systems, wherein the processor is further configured to obtain audio encounter information using a plurality of audio acquisition devices of an audio recording system, wherein the processor is further configured to encode the audio encounter information using one or more codecs, wherein the processor is further configured to adapt the encoding of the audio encounter information by the one or more codecs based upon, at least in part, the machine vision encounter information, wherein the processor is further configured to generate a plurality of acoustic relative transfer functions between the plurality of audio acquisition devices of the audio recording system, and wherein adapting the encoding of the audio encounter information by the one or more codecs based upon, at least in part, the machine vision encounter information includes adapting the one or more codecs to encode the audio encounter information using one or more acoustic relative transfer functions associated with a particular acoustic source when the machine vision encounter information detects the acoustic source.

12. The computing system of claim 11 , wherein the plurality of audio acquisition devices of the audio recording system are positioned within a fixed geometry relative to each other.

13. The computing system of claim 11 , wherein adapting the encoding of the audio encounter information by the one or more codecs based upon, at least in part, the machine vision encounter information includes adapting the one or more codecs to estimate one or more acoustic relative transfer functions when the machine vision encounter information indicates at least a threshold change in the acoustic environment.

14. The computing system of claim 11 , wherein adapting the encoding of the audio encounter information by the one or more codecs based upon, at least in part, the machine vision encounter information includes adapting the one or more codecs to selectively encode the audio encounter information based upon, at least in part, whether the machine vision encounter information indicates that an audio source is speaking.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065578/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 10, 2022
From: SHARMA, DUSHYANT; NAYLOR, PATRICK A.; JOST, UWE HELMUT
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 059225/0534 →