IP Library › Patent Application 16596756
Patent Application
App. No. 16/596,756

VOICE ASSISTANT WITH CONTEXTUALLY-ADJUSTED AUDIO OUTPUT

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
16/596,756
Abstract

A voice assistant has a contextually-adjusted audio output. The audio output can be adjusted, for example, based on media content characteristics.

Claims (42)

1 . A method for generating synthesized speech of a voice assistant having a contextually-adjusted audio output using a voice-enabled device, the method comprising:

identifying media content characteristics associated with media content;

identifying base characteristics of audio output;

generating contextually-adjusted characteristics of audio output based at least in part on the base characteristics and the media content characteristics; and

using the contextually-adjusted audio output characteristics to generate the synthesized speech.

2 . The method of claim 1 , wherein the contextually-adjusted characteristics of audio output are further based on user-specific adjustments to the base characteristics of audio output.

3 . The method of claim 1 , wherein using the contextually-adjusted audio output comprises receiving voice content and generating the synthesized speech to convey the voice content to the user according to the contextually-adjusted audio output.

4 . The method of claim 1 , wherein identifying the media content characteristics comprises:

analyzing audio of the media content to determine musical characteristics of the media content; and

analyzing media content metadata to determine metadata-based characteristics.

5 . The method of claim 4 , wherein generating a contextually-adjusted audio output is based at least in part upon the musical characteristics of the media content.

6 . The method of claim 5 , wherein generating the contextually-adjusted audio output comprises generating mood-related attributes that are compatible with the musical characteristics of the media content.

7 . The method of claim 5 , wherein generating the contextually-adjusted audio output comprises generating mood-related attributes that are compatible with metadata-based characteristics of the media content.

8 . The method of claim 1 , wherein the user-specific adjustments are based on the user's listening history.

9 . The method of claim 1 , wherein using the contextually-adjusted audio output to generate synthesize speech further comprises:

selecting words to be spoken by the voice assistant using a natural language generator based upon language adjustments associated with the contextually-adjusted audio output characteristics; and

determining a pronunciation and an emotion for speaking the words based upon speech adjustments associated with the contextually-adjusted audio output characteristics.

10 . The method of claim 1 , further comprising generating a mood associated with the contextually-adjusted audio output, the mood comprising:

the contextually-adjusted audio output;

one or more audio cues; and

one or more visual representations.

11 . A voice assistant system comprising:

at least one processing device; and

at least one computer readable storage device storing data instructions that, when executed by the at least one processing device, cause the at least one processing device to:

identify media content characteristics associated with media content;

identify base characteristics of audio output;

generate contextually-adjusted audio output characteristics based at least in part on the base characteristics of audio output and the media content characteristics; and

use the contextually-adjusted audio output characteristics to generate synthesized speech.

12 . The voice assistant system of claim 11 , further comprising a voice-enabled device configured for interaction with a user via voice, wherein the voice-enabled device comprises the at least one processing device and the at least one computer readable storage device.

13 . The voice assistant system of claim 11 , further comprising a media delivery system comprising at least one server computing device comprising the at least one processing device at the at least one computer readable storage device.

14 . The voice assistant system of claim 11 , wherein the base characteristics of audio output are user-specific characteristics of audio output generated based at least in part on a listening history of a user and brand characteristics of audio output.

15 . The voice assistant system of claim 11 , wherein the data instructions that cause the at least one processing device to identify media content characteristics associated with media content further comprises:

analyzing audio content of the media content to identify musical characteristics of the media content; and

analyzing media content metadata of the media content to identify metadata based characteristics of the media content; and

wherein the media content characteristics used to generate the contextually-adjusted audio output further comprise:

the musical characteristics of the media content; and

the metadata characteristics of the media content.

16 . The voice assistant system of claim 11 , wherein generating the contextually-adjusted audio output is performed by a contextual audio output adjuster, and wherein the contextual audio output adjuster further comprises data instructions that cause the at least one processing device to:

generate language adjustments based on the contextually-adjusted audio output;

send the language adjustments to a natural language generator to select words to be spoken by the voice assistant;

generate speech adjustments based on the contextually-adjusted audio output; and

send the speech adjustments to a text-to-speech engine, the speech adjustments defining pronunciation adjustments and emotion adjustments to be applied to the words when spoken by the voice assistant.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 27, 2023
From: SPOTIFY USA INC.
To: SPOTIFY AB
Reel/Frame 063105/0815 →
EMPLOYMENT AGREEMENT Recorded Dec 14, 2022
From: KUMAR, ROHIT
To: SPOTIFY USA INC.
Reel/Frame 062567/0015 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 14, 2022
From: MENNICKEN, SARAH; MOULTON, PAUL; STECKEL, MIRA; CRAMER, HENRIETTE SUSANNE MARTINE; LE LAY, FRANÇOIS
To: SPOTIFY AB
Reel/Frame 062088/0919 →