IP Library Granted Patent US 11,606,508
Granted Patent B1
US 11,606,508 · App. 17/529,642 · Granted Mar 14, 2023

Enhanced representations based on sensor data

Inventors: Rahul B. Desai (Hoffman Estates, IL); Amit Kumar Agrawal (Bangalore, IN)
Assignee: Motorola Mobility LLC
H04N5/2628G06V20/40G06V40/176G06V40/20G10L25/51H04L65/403
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,606,508
App. No.
17/529,642
Granted
Mar 14, 2023
Kind
B1
Abstract

Techniques for generating enhanced representations based on sensor data are described and are implementable in a video conference setting. Generally, the described implementations enable an enhanced representation of a focal individual, for instance a speaker, to be generated based on sensor data, for instance audio and visual sensor data. The audio data can identify an individual as the speaker or determine a general location of a source of audio. Visual sensors can detect gestures of individuals located in the general location of the source of audio to identify gestures which indicate that one or more individuals are speaking or are about to speak.

Claims (38)

1. A method, comprising:

identifying, by a first device, at least one focal individual in a viewable region of a video capture device based on positional audio data and gesture information obtained from one or more sensors of the first device;

generating an enhanced representation of the at least one focal individual by processing content associated with the at least one focal individual, the enhanced representation containing enhanced visual content and enhanced audial content pertaining to the at least one focal individual and further including an orientation tag indicating a display orientation for the enhanced representation based on a spatial position of the at least one focal individual in relation to the first device; and

communicating the enhanced representation of the at least one focal individual for display in a user interface of a second device.

2. The method of claim 1 , wherein identifying the at least one focal individual includes validating the positional audio data against the gesture information to verify that the at least one focal individual is speaking.

3. The method of claim 1 , wherein identifying the at least one focal individual includes filtering location specific profile information associated with the at least one focal individual.

4. The method of claim 1 , wherein the gesture information includes gesture information from an individual other than the at least one focal individual.

5. The method of claim 1 , wherein said identifying the at least one focal individual comprises:

generating a position map of individuals present with the at least one focal individual; and

identifying the at least one focal individual by correlating the positional audio data and the gesture information to the position map.

6. The method of claim 1 , wherein the enhanced representation includes a zoomed in view of the at least one focal individual.

7. The method of claim 6 , further comprising utilizing one or more super-resolution techniques to generate the enhanced representation.

8. The method of claim 1 , wherein said generating the enhanced representation comprises utilizing beamforming to suppress audio that does not originate with the at least one focal individual in the enhanced representation.

9. The method of claim 1 , wherein the enhanced representation includes information from a user profile associated with the at least one focal individual.

10. The method of claim 1 , wherein the enhanced representation simulates a perspective view of the at least one focal individual in relation to the first device based on the orientation tag and utilizes spatialized audio to simulate an audial perspective relative to the first device.

11. An apparatus comprising:

a processing system implemented at least in part in hardware of the apparatus; and

a video conference module implemented at least in part in hardware of the apparatus and executable by the processing system to:

receive, by an audio sensor of a first device, positional audio data indicating a location of a source of audio from one or more individuals;

detect, by the first device, facial gestures of the one or more individuals positioned in the location;

identify, by the first device, at least one focal individual as speaking based on the positional audio data and detected facial gestures;

generate an enhanced representation of the at least one focal individual, containing visual and audial content, a display orientation of the enhanced representation based in part on a spatial position of the at least one focal individual; and

communicate the enhanced representation of the at least one focal individual for display in a user interface of a second device.

12. The apparatus of claim 11 , wherein to identify the at least one focal individual is based on an indication from the positional audio data that the at least one focal individual is within the location and the detected facial gestures verify that the at least one focal individual is speaking.

13. The apparatus of claim 11 , wherein to identify the at least one focal individual includes correlating the positional audio data and detected facial gestures to location specific profile information associated with the at least one focal individual.

14. The apparatus of claim 11 , wherein the enhanced representation includes information from a user profile associated with the at least one focal individual.

15. The apparatus of claim 11 , wherein the enhanced representation simulates a perspective view of the at least one focal individual in relation to the first device and utilizes spatialized audio to simulate an audial perspective relative to the first device.

16. A system comprising:

one or more processors; and

one or more computer-readable storage media storing instructions that are executable by the one or more processors to:

identify at least one focal individual as speaking within a viewable region of a video capture device of a first device;

determine a spatial position of the at least one focal individual in relation to the first device;

generate an enhanced representation of the at least one focal individual based on the spatial position, the enhanced representation containing enhanced visual content that simulates a perspective view of the at least one focal individual in relation to the first device and enhanced audial content to simulate an audial perspective relative to the first device; and

communicate the enhanced representation of the at least one focal individual for display in a user interface of a second device.

17. The system of claim 16 , wherein to identify the at least one focal individual includes verifying audio data obtained from one or more audio sensors of the first device against facial gestures detected by the first device to validate that the at least one focal individual is speaking.

18. The system of claim 16 , wherein the enhanced representation utilizes spatialized audio to simulate the audial perspective relative to the first device.

19. The system of claim 16 , wherein the enhanced representation includes information from a user profile associated with the at least one focal individual.

20. The system of claim 19 , wherein the information from the user profile associated with the at least one focal individual includes one or more of a name, job description, position, technical background, company designation, contact information, user photo, or expertise.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 21, 2021
From: DESAI, RAHUL B.; AMIT KUMAR, AMIT KUMAR
To: MOTOROLA MOBILITY LLC
Reel/Frame 058173/0559 →