IP Library Granted Patent US 12,177,648
Granted Patent B2
US 12,177,648 · App. 17/849,137 · Granted Dec 24, 2024

Systems and methods for orientation-responsive audio enhancement

Inventor: Warren Keith Edwards (Atlanta, GA)
Assignee: Adeia Guides Inc.
H04S7/303G06F3/013H04S2400/01H04S2400/11
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,177,648
App. No.
17/849,137
Granted
Dec 24, 2024
Kind
B2
Abstract

Sound objects are identified within a content item, and location metadata is extracted from the content item for each sound object. A reference layout is generated, relative to a user position, for the sound objects based on the location metadata. A user's gaze is then determined using pupil tracking, body movement data, head orientation data, or other techniques. Using the reference layout, a sound object along a path defined by the gaze of the user is identified, and audio of the identified object is enhanced.

Claims (88)

1. A method for orientation-responsive audio enhancement, the method comprising:

identifying, within a content item, a plurality of virtual sound objects;

extracting, from the content item, location metadata for each respective virtual sound object;

generating a reference layout, relative to a user position, for the plurality of virtual sound objects based on the location metadata;

detecting a gaze of the user;

identifying, based on the reference layout, at least two virtual sound objects of the plurality of virtual sound objects, along a path defined by the gaze of the user;

analyzing audio of each of the at least two virtual sound objects along the path defined by the gaze of the user;

determining, based on the analyzing, that a first virtual sound object of the at least two virtual sound objects comprises a voice; and

based on determining that the first virtual sound object comprises a voice, enhancing the audio of the first virtual sound object.

2. The method of claim 1 , wherein enhancing audio of the virtual sound object further comprises modifying an amplitude of audio of at least one virtual sound object.

3. The method of claim 1 , further comprising:

identifying a plurality of users consuming the content item; and

for each user of the plurality of users, determining second location metadata corresponding to locations of each other user of the plurality of users;

wherein the plurality of virtual sound objects includes the plurality of users.

4. The method of claim 1 , further comprising:

receiving a negative input; and

in response to the negative input:

stopping enhancement of audio of the first virtual sound object;

selecting a second virtual sound object of the at least two virtual sound objects; and

enhancing audio of the second virtual sound object.

5. The method of claim 4 , wherein the negative input is a gesture.

6. The method of claim 4 , wherein the negative input is a speech input.

7. The method of claim 1 , further comprising:

generating for display a representation of each virtual sound object of the at least two virtual sound objects;

detecting a second gaze of the user; and

identifying a representation of a virtual sound object, of the representations of the at least two virtual sound objects, along a second path defined by the second gaze of the user.

8. The method of claim 1 , wherein the gaze is a second gaze and wherein determining the second gaze of the user comprises:

identifying a first gaze of the user;

receiving movement data from a device associated with the user; and

determining, based on the movement data, the second gaze of the user.

9. The method of claim 1 , wherein detecting the gaze of the user comprises:

tracking pupils of the user;

determining, based on the tracking, a direction in which the pupils of the user are focused; and

determining the gaze based on the direction.

10. The method of claim 1 , wherein detecting the gaze of the user comprises:

tracking an orientation of a head of the user; and

determining, based on the orientation of the head of the user, the gaze of the user.

11. The method of claim 1 , wherein the audio of the first virtual sound object is enhanced without enhancing audio of the other of the plurality of the virtual sound objects.

12. A system for orientation-responsive audio enhancement, the system comprising:

input/output circuitry configured to receive a content item; and

control circuitry configured to:

identify, within the content item, a plurality of virtual sound objects;

extract, from the content item, location metadata for each respective virtual sound object;

generate a reference layout, relative to a user position, for the plurality of virtual sound objects based on the location metadata;

detect a gaze of the user;

identify, based on the reference layout, at least two virtual sound objects of the plurality of virtual sound objects along a path defined by the gaze of the user;

analyze audio of each of the at least two virtual sound objects along the path defined by the gaze of the user;

determine, based on the analyzing, that a first virtual sound object of the at least two virtual sound objects comprises a voice; and

based on determining that the first virtual sound object comprises a voice, enhance the audio of the first virtual sound object.

13. The system of claim 12 , wherein the control circuitry configured to enhance audio of the virtual sound object is further configured to modify an amplitude of audio of at least one virtual sound object.

14. The system of claim 12 , wherein the control is further configured to:

identify a plurality of users consuming the content item; and

for each user of the plurality of users, determine second location metadata corresponding to locations of each other user of the plurality of users;

wherein the plurality of virtual sound objects includes the plurality of users.

15. The system of claim 12 , wherein the control circuitry is further configured to:

receive a negative input; and

in response to the negative input:

stop enhancement of audio of the first virtual sound object;

select a second virtual sound object of the subset of virtual sound objects; and

enhance audio of the second virtual sound object.

16. The system of claim 15 , wherein the negative input is a gesture.

17. The system of claim 15 , wherein the negative input is a speech input.

18. The system of claim 12 , wherein the control circuitry is further configured to:

generate for display a representation of each virtual sound object of the at least two virtual sound objects;

detect a second gaze of the user; and

identify a representation of a virtual sound object, of the representations of the at least two virtual sound objects, along a second path defined by the second gaze of the user.

19. The system of claim 12 , wherein the gaze is a second gaze and wherein the control circuitry configured to determine the second gaze of the user is further configured to:

identify a first gaze of the user;

receive movement data from a device associated with the user; and

determine, based on the movement data, the second gaze of the user.

20. The system of claim 12 , wherein the control circuitry configured to detect the gaze of the user is further configured to:

track pupils of the user;

determine, based on the tracking, a direction in which the pupils of the user are focused; and

determine the gaze based on the direction.

21. The system of claim 12 , wherein the control circuitry configured to detect the gaze of the user is further configured to:

track an orientation of a head of the user; and

determine, based on the orientation of the head of the user, the gaze of the user.

22. The system of claim 12 , wherein the audio of the first virtual sound object is enhanced without enhancing audio of the other of the plurality of the virtual sound objects.

23. A method for orientation-responsive audio enhancement, the method comprising:

identifying, within a content item, a plurality of virtual sound objects;

extracting, from the content item, location metadata for each respective virtual sound object;

generating a reference layout, relative to a user position, for the plurality of virtual sound objects based on the location metadata;

detecting a gaze of the user;

identifying, based on the reference layout, at least two virtual sound objects, of the plurality of virtual sound objects, along a path defined by the gaze of the user;

based on identifying the at least two virtual sound objects along the path defined by the gaze of the user, generating for display at least two representations respectively corresponding to the at least two virtual sound objects;

receiving a selection of a first representation of the at least two representations; and

based on the received selection, enhancing the audio of a first virtual sound object corresponding to the first representation.

24. The method of claim 23 , wherein the audio of the first virtual sound object is enhanced without enhancing audio of the other of the plurality of the virtual sound objects.

Assignments (4)
CHANGE OF NAME Recorded Sep 25, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069049/0212 →
SECURITY INTEREST Recorded May 3, 2023
From: ADEIA GUIDES INC.; ADEIA IMAGING LLC; ADEIA MEDIA HOLDINGS LLC; ADEIA MEDIA SOLUTIONS INC.; ADEIA SEMICONDUCTOR ADVANCED TECHNOLOGIES INC.; ADEIA SEMICONDUCTOR BONDING TECHNOLOGIES INC.; ADEIA SEMICONDUCTOR INC.; ADEIA SEMICONDUCTOR SOLUTIONS LLC; ADEIA SEMICONDUCTOR TECHNOLOGIES LLC; ADEIA SOLUTIONS LLC
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 063529/0272 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 6, 2022
From: EDWARDS, WARREN KEITH
To: ROVI GUIDES, INC.
Reel/Frame 060994/0312 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2022
From: EDWARDS, KEITH
To: ROVI GUIDES, INC.
Reel/Frame 060317/0732 →