IP Library › Granted Patent US 12,567,408
Granted Patent B2
US 12,567,408 · App. 17/626,617 · Granted Mar 3, 2026

Multi-modal smart audio device system attentiveness expression

Inventors: Christopher Graham Hines (Sydney, AU); Rowan James Katekar (Redfern, AU); Glenn N. Dickins (Como, AU); Richard J. Cartwright (Sydney, AU); Jeremiha Emile Douglas (Mill Valley, CA); Mark R.P. Thomas (Walnut Creek, CA)
Assignee: Dolby Laboratories Licensing Corporation
G10L15/22G10L15/32G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,567,408
App. No.
17/626,617
Filed
Jan 12, 2022
Granted
Mar 3, 2026
Kind
B2
Art Unit
2659
USPC
704/251
Abstract

A method may involve receiving output signals from each microphone of a plurality of microphones in the environment, each of the plurality of microphones residing in a microphone location of the environment, the output signals corresponding to an utterance of a person. The method may involve determining, based at least in part on the output signals, a zone within the environment that has at least a threshold probability of including the person's location and generating a plurality of spatially-varying attentiveness signals within the zone. Each attentiveness signal may be generated by a device located within the zone. Each attentiveness signal may indicate that a corresponding device is in an operating mode in which the corresponding device is awaiting a command and may indicate a relevance metric of the corresponding device.

Claims (19)

1 . A method of controlling a system of devices in an environment, the method comprising:

receiving output signals from each microphone of a plurality of microphones in the environment, each microphone of the plurality of microphones residing in a microphone location of the environment, the output signals corresponding to an utterance of a person;

determining, based at least in part on the output signals, a zone within the environment that has at least a threshold probability of including a location of the person;

generating a plurality of spatially-varying attentiveness signals within the zone, each attentiveness signal of the plurality of attentiveness signals being generated by a device located within the zone, each attentiveness signal indicating that a corresponding device is in an operating mode in which the corresponding device is awaiting a command, wherein each attentiveness signal is spatially varied based on a relevance metric of the corresponding device, and wherein the relevance metric is based, at least in part, on an estimated distance from the location of the person to an acoustic centroid of a plurality of microphones within the zone, wherein the acoustic centroid is computed as a function of zone classification posteriors corresponding to a set of utterances by the person and an estimated position of the person during each utterance in the set.

2 . The method of claim 1 , wherein an attentiveness signal generated by a first device indicates a relevance metric of a second device, the second device being the corresponding device.

3 . The method of claim 1 , wherein the relevance metric is based, at least in part, on an estimated visibility of the corresponding device.

4 . The method of claim 1 , wherein the utterance comprises a wakeword.

5 . The method of claim 1 , wherein at least one of the plurality of spatially-varying attentiveness signals comprises a modulation of at least one previous signal generated by the device located within the zone prior to a time of the utterance.

6 . The method of claim 5 , wherein the at least one previous signal comprises a light signal and wherein the modulation comprises at least one of a color modulation, a color saturation modulation or a light intensity modulation.

7 . The method of claim 1 , wherein at least one microphone of the plurality of microphones is included in or configured for communication with a smart audio device.

8 . The method of claim 1 , further comprising an automated process of determining whether the device is in a device group.

9 . The method of claim 8 , wherein the automated process is based, at least in part, on sensor data corresponding to at least one of light or sound emitted by the device, or, wherein the automated process is based, at least in part, on communications between at least one of a source and an orchestrating hub device or a receiver and the orchestrating hub device, or, wherein the automated process is based, at least in part, on a light source or a sound source being switched on and off for a duration of time.

10 . The method of claim 8 , further comprising automatically updating the automated process according to implicit feedback based on one or more of a success of beamforming based on an estimated zone, a success of microphone selection based on the estimated zone, a determination that the person has terminated a response of a voice assistant abnormally, a command recognizer returning a low-confidence result or a second-pass retrospective wakeword detector returning low confidence that a wakeword was spoken.

11 . The method of claim 1 , further comprising selecting at least one speaker of the device located within the zone and controlling the at least one speaker to provide sound to the person.

12 . The method of claim 1 , further comprising selecting at least one microphone of the device located within the zone and providing signals output by the at least one microphone to a smart audio device.

13 . The method of claim 1 , wherein a first microphone of the plurality of microphones samples audio data according to a first sample clock and a second microphone of the plurality of microphones samples the audio data according to a second sample clock.

14 . An apparatus configured to perform the method of claim 1 .

15 . A system of devices configured to perform the method of claim 1 .

16 . One or more non-transitory media having software stored thereon, the software including instructions for controlling one or more devices to perform the method of claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 18, 2022
From: HINES, CHRISTOPHER GRAHAM; KATEKAR, ROWAN JAMES; DICKINS, GLENN N.; CARTWRIGHT, RICHARD J.; DOUGLAS, JEREMIHA EMILE; THOMAS, MARK R.P.
To: DOLBY LABORATORIES LICENSING CORPORATION
Reel/Frame 058683/0857 →
Continuity (5)
Provisional Application 62880110 · Jul 30, 2019
Provisional Application 62880112 · Jul 30, 2019
Provisional Application 62964018 · Jan 21, 2020
Provisional Application 63003788 · Apr 1, 2020
Related Publication 20220270601A1 · Aug 25, 2022
References Cited (52)
US 7297929B2 · Cernasov · 2007 [cited by applicant]
US 7843353B2 · Pan · 2010 [cited by applicant]
US 8279079B2 · Bergman · 2012 [cited by applicant]
US 8930841B2 · Huang · 2015 [cited by applicant]
US 9076450B1 · Sadek et al. · 2015 [cited by applicant]
US 9270139B2 · Rofougaran · 2016 [cited by applicant]
US 9275637B1 · Salvador · 2016 [cited by applicant]
US 9318107B1 · Sharifi · 2016 [cited by applicant]
US 9368105B1 · Freed · 2016 [cited by applicant]
US 9674781B2 · Szewczyk · 2017 [cited by applicant]
US 9691378B1 · Meyers · 2017 [cited by applicant]
US 9721586B1 · Bay · 2017 [cited by applicant]
US 9723689B2 · Łagutko · 2017 [cited by applicant]
US 9734830B2 · Lindahl · 2017 [cited by applicant]
US 9774507B2 · Britt · 2017 [cited by applicant]
US 9826601B2 · Vangeel · 2017 [cited by applicant]
US 9927249B2 · Barnard · 2018 [cited by applicant]
US 9996316B2 · Jorgovanovic · 2018 [cited by applicant]
US 9997070B1 · Komanduri · 2018 [cited by applicant]
US 10013981B2 · Ramprashad · 2018 [cited by applicant]
US 10026399B2 · Gopalan · 2018 [cited by applicant]
US 10043521B2 · Bocklet · 2018 [cited by applicant]
US 10134399B2 · Lang · 2018 [cited by applicant]
US 10140849B2 · Sloo · 2018 [cited by applicant]
US 10181323B2 · Beckhardt · 2019 [cited by applicant]
US 10192546B1 · Piersol · 2019 [cited by applicant]
US 10304450B2 · Tak · 2019 [cited by examiner]
US 10380852B2 · Horling · 2019 [cited by examiner]
US 10425781B1 · Devaraj · 2019 [cited by examiner]
US 20040141418A1 · Matsuo · 2004 [cited by applicant]
US 20050018861A1 · Tashev · 2005 [cited by applicant]
US 20140163978A1 · Basye · 2014 [cited by applicant]
US 20160259419A1 · Chatterjee · 2016 [cited by applicant]
US 20170083285A1 · Meyers · 2017 [cited by applicant]
US 20170090864A1 · Jorgovanovic · 2017 [cited by applicant]
US 20170330429A1 · Tak · 2017 [cited by applicant]
US 20180047394A1 · Tian · 2018 [cited by applicant]
US 20180108351A1 · Beckhardt et al. · 2018 [cited by applicant]
US 20180211665A1 · Park · 2018 [cited by applicant]
US 20180228006A1 · Baker · 2018 [cited by applicant]
US 20180330589A1 · Horling · 2018 [cited by applicant]
US 20180335903A1 · Coffman · 2018 [cited by applicant]
US 20180342151A1 · Vanblon · 2018 [cited by applicant]
US 20190066670A1 · White et al. · 2019 [cited by applicant]
US 20190096398A1 · Sereshki · 2019 [cited by applicant]
US 20190130914A1 · Sharifi · 2019 [cited by applicant]
US 20190179611A1 · Wojogbe · 2019 [cited by applicant]
WO 2014130463A2 · 2014 [cited by applicant]
WO 2019059939A1 · 2019 [cited by applicant]
WO 2021021814 · 2021 [cited by applicant]
Wolfram Mathworld, Correlation Coefficient http://mathworld.wolfram.com/CorrelationCoefficient.html. [cited by applicant]
Wikipedia: Dynamic Time Warping ; https://en.wikipedia.org/wiki/Dynamic_time_warping. [cited by applicant]