IP Library › Granted Patent US 11,722,731
Granted Patent B2
US 11,722,731 · App. 17/103,908 · Granted Aug 8, 2023

Integrating short-term context for content playback adaption

Inventors: Victor Carbune (Zürich, CH); Matthew Sharifi (Kilchberg, CH)
Assignee: Google LLC
H04N21/458H04N21/4396H04N21/44218H04N21/4532H04N21/466H04N21/47202H04N21/47217
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,722,731
App. No.
17/103,908
Granted
Aug 8, 2023
Kind
B2
Abstract

While an assistant-enabled device is playing back media content, a method includes receiving a contextual signal from an environment of the assistant-enabled device and executing an event recognition routine to determine whether the received contextual signal is indicative of an event that conflicts with the playback of the media content from the assistant-enabled device. When the event recognition routine determines that the received contextual signal is indicative of the event that conflicts with the playback of the media content, the method also includes adjusting content playback settings of the assistant-enabled device.

Claims (76)

1. A method comprising:

receiving, at data processing hardware of an assistant-enabled device while the assistant-enabled device is playing back media content in an environment of the assistant-enabled device, a contextual signal representing the environment, the contextual signal comprising at least one of audio present in the environment detected by a microphone of the assistant-enabled device or image data representing an image of the environment captured by an image capture device of the assistant-enabled device;

executing, by the data processing hardware of the assistant-enabled device while the assistant-enabled device is playing back the media content, an event recognition routine to determine whether the received contextual signal is indicative of an event present in the environment that conflicts with the playback of the media content in the environment from the assistant-enabled device; and

in response to the event recognition routine determining that the received contextual signal is indicative of the event present in the environment that conflicts with the playback of the media content in the environment, automatically adjusting, by the data processing hardware of the assistant-enabled device, content playback settings of the assistant-enabled device while continuing to receive the contextual signal that is indicative of the event present in the environment that conflicts with the playback of the media content in the environment from the assistant-enabled device,

wherein executing the event recognition routine comprises executing a neural network-based classification model configured to receive the contextual signal as input and generate, as output, a classification result indicating whether the received contextual signal is indicative of the event present in the environment that conflicts with the playback of the media content in the environment from the assistant-enabled device.

2. The method of claim 1 , wherein:

the contextual signal received at the neural network-based classification model as input comprises an audio stream; and

the classification result generated by the neural network-based classification model as output comprises an audio event that conflicts with the playback of the media content in the environment.

3. The method of claim 2 , wherein the classification result generated by the neural network-based classification model as output is further based on an audible level of the audio stream.

4. The method of claim 1 , wherein:

the contextual signal received at the neural network-based classification model as input comprises an image stream; and

the classification result generated by the neural network-based classification model as output comprises an activity event that conflicts with the playback of the media content in the environment.

5. The method of claim 1 , further comprising:

determining, by the data processing hardware of the assistant-enabled device, that the received contextual signal is indicative of an audio event;

obtaining, by the data processing hardware of the assistant-enabled device, an audible level associated with the audio event;

obtaining, by the data processing hardware of the assistant-enabled device, an audible level of the media content playing back in the environment from the assistant-enabled device; and

determining, by the data processing hardware of the assistant-enabled device, a likelihood score indicating a likelihood that the media content playing back in the environment from the assistant-enabled device interrupts an ability of a user associated with the assistant-enabled device to hear the audio event,

wherein adjusting the content playback settings of the assistant-enabled device comprises one of, based on the likelihood score:

lowering the audible level of the media content playing back in the environment from the assistant-enabled device; or

stopping/pausing the playback of the media content in the environment from the assistant-enabled device.

6. The method of claim 1 , further comprising, when the event recognition routine determines that the received contextual signal is indicative of the event that conflicts with the playback of the media content in the environment:

obtaining, by the data processing hardware of the assistant-enabled device, playback features associated with the media content playing back in the environment from the assistant-enabled device;

obtaining, by the data processing hardware of the assistant-enabled device, event-based features associated with the event; and

determining, by the data processing hardware of the assistant-enabled device, using a trained machine learning model configured to receive the playback features and the event-based features as input, a likelihood score indicating a likelihood that the media content playing back in the environment from the assistant-enabled device interrupts an ability of a user associated with the assistant-enabled device to recognize the event,

wherein adjusting the content playback settings of the assistant-enabled device is based on the likelihood score.

7. The method of claim 6 , wherein:

the event-based features comprise at least one of an audio level associated with the event, an event type, or event importance; and

the playback features comprise at least one of an audible level of the media content playing back in the environment from the assistant-enabled device, a media content type, or playback importance.

8. The method of claim 6 , further comprising, after adjusting the content playback settings of the assistant-enabled device:

obtaining, by the data processing hardware of the assistant-enabled device, user feedback indicating at least one of:

acceptance of the adjusted content playback settings; or

a subsequent manual adjustment to the content playback settings of the assistant-enabled device; and

executing, by the data processing hardware of the assistant-enabled device, a training process that re-trains the machine learning model on at least one of the obtained playback features, the obtained event-based features, the adjusted content playback settings, and the obtained user feedback.

9. The method of claim 1 , wherein adjusting the content playback settings of the assistant-enabled device comprises at least one of increasing/decreasing an audio level of the playback of the media content, stopping/pausing the playback of the media content, or instructing the assistant-enabled device to playback a different type of media content.

10. The method of claim 1 , further comprising:

receiving, at the data processing hardware of the assistant-enabled device, user-defined configuration settings indicating user preferences for adjusting the content playback settings of the assistant-enabled device,

wherein adjusting the content playback settings of the assistant-enabled device is based on the user-defined configuration settings.

11. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

receiving, while an assistant-enabled device is playing back media content in an environment of the assistant-enabled device, a contextual signal representing the environment, the contextual signal comprising at least one of audio present in the environment detected by a microphone of the assistant-enabled device or image data representing an image of the environment captured by an image capture device of the assistant-enabled device;

executing, while the assistant-enabled device is playing back the media content, an event recognition routine to determine whether the received contextual signal is indicative of an event present in the environment that conflicts with the playback of the media content in the environment from the assistant-enabled device; and

in response to the event recognition routine determining that the received contextual signal is indicative of the event present in the environment that conflicts with the playback of the media content in the environment, automatically adjusting content playback settings of the assistant-enabled device while continuing to receive the contextual signal that is indicative of the event present in the environment that conflicts with the playback of the media content in the environment from the assistant-enabled device,

wherein executing the event recognition routine comprises executing a neural network-based classification model configured to receive the contextual signal as input and generate, as output, a classification result indicating whether the received contextual signal is indicative of the event present in the environment that conflicts with the playback of the media content in the environment from the assistant-enabled device.

12. The system of claim 11 , wherein:

the contextual signal received at the neural network-based classification model as input comprises an audio stream; and

the classification result generated by the neural network-based classification model as output comprises an audio event that conflicts with the playback of the media content in the environment.

13. The system of claim 12 , wherein the classification result generated by the neural network-based classification model as output is further based on an audible level of the audio stream.

14. The system of claim 11 , wherein:

the contextual signal received at the neural network-based classification model as input comprises an image stream; and

the classification result generated by the neural network-based classification model as output comprises an activity event that conflicts with the playback of the media content in the environment.

15. The system of claim 11 , wherein the operations further comprise:

determining that the received contextual signal is indicative of an audio event;

obtaining an audible level associated with the audio event;

obtaining an audible level of the media content playing back in the environment from the assistant-enabled device; and

determining a likelihood score indicating a likelihood that the media content playing back in the environment from the assistant-enabled device interrupts an ability of a user associated with the assistant-enabled device to hear the audio event,

wherein adjusting the content playback settings of the assistant-enabled device comprises one of, based on the likelihood score:

lowering the audible level of the media content playing back in the environment from the assistant-enabled device; or

stopping/pausing the playback of the media content in the environment from the assistant-enabled device.

16. The system of claim 11 , wherein the operations further comprise, when the event recognition routine determines that the received contextual signal is indicative of the event that conflicts with the playback of the media content in the environment:

obtaining playback features associated with the media content playing back in the environment from the assistant-enabled device;

obtaining event-based features associated with the event; and

determining, using a trained machine learning model configured to receive the playback features and the event-based features as input, a likelihood score indicating a likelihood that the media content playing back in the environment from the assistant-enabled device interrupts an ability of a user associated with the assistant-enabled device to recognize the event,

wherein adjusting the content playback settings of the assistant-enabled device is based on the likelihood score.

17. The system of claim 16 , wherein:

the event-based features comprise at least one of an audio level associated with the event, an event type, or an event importance; and

the playback features comprise at least one of an audible level of the media content playing back in the environment from the assistant-enabled device, a media content type, or a playback importance.

18. The system of claim 16 , wherein the operations further comprise, after adjusting the content playback settings of the assistant-enabled device:

obtaining user feedback indicating at least one of:

acceptance of the adjusted content playback settings; or

a subsequent manual adjustment to the content playback settings of the assistant-enabled device; and

executing a training process that re-trains the machine learning model on at least one of the obtained playback features, the obtained event-based features, the adjusted content playback settings, and the obtained user feedback.

19. The system of claim 11 , wherein adjusting the content playback settings of the assistant-enabled device comprises at least one of increasing/decreasing an audio level of the playback of the media content, stopping/pausing the playback of the media content, or instructing the assistant-enabled device to playback a different type of media content.

20. The system of claim 11 , wherein the operations further comprise:

receiving user-defined configuration settings indicating user preferences for adjusting the content playback settings of the assistant-enabled device,

wherein adjusting the content playback settings of the assistant-enabled device is based on the user-defined configuration settings.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 25, 2020
From: SHARIFI, MATTHEW; CARBUNE, VICTOR
To: GOOGLE LLC
Reel/Frame 054466/0379 →
Continuity (1)
Related Publication 20220167049A1 · May 26, 2022