IP Library › Granted Patent US 11,475,618
Granted Patent B2
US 11,475,618 · App. 16/871,901 · Granted Oct 18, 2022

Dynamic vision sensor for visual audio processing

Inventors: Xiaoyong Ye (San Mateo, CA); Yuichiro Nakamura (San Mateo, CA)
Assignee: Sony Interactive Entertainment Inc.
G06T13/40G06K9/6217G06T7/20G06V40/171G06V40/172G06V40/174G06V40/193G10L25/30G10L25/63
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,475,618
App. No.
16/871,901
Granted
Oct 18, 2022
Kind
B2
Abstract

To track certain difficult facial features during speech such as the corners of the mouth and the teeth, a camera sensor system generates RGB/IR images and the system also uses light intensity change signals from an event driven sensor (EDS), as well as voice analysis. In this way, the camera sensor system enables improved performance tracking (equivalent to using very high-speed camera) at lower bandwidth and power consumption.

Claims (57)

1. An assembly comprising:

at least one camera unit configured to generate red-green-blue (RGB) images of a face;

at least one event driven sensor (EDS) configured to output signals representing changes in illumination intensity of the face, the EDS providing an output of −1 responsive to intensity of light being sensed decreasing, +1 responsive to intensity of light being sensed increasing, and 0 responsive to a change in light intensity below a threshold;

at least one microphone configured to output signals representing speech; and

at least one processor configured with executable instructions to:

receive signals from the camera unit, the EDS, and the microphone;

execute at least one neural network to generate, based on the signals from the camera unit, the EDS, and the microphone, at least one of: emotion prediction, tracking of at least a portion of the face.

2. The assembly of claim 1 , wherein the camera unit is configured to generate infrared (IR) images.

3. The assembly of claim 1 , wherein the camera unit, processor, and EDS are disposed on a single chip.

4. The assembly of claim 1 , wherein the instructions are executable to:

execute the at least one neural network to generate, based on the signals from the camera unit, the EDS, and the microphone, emotion prediction.

5. The assembly of claim 1 , wherein the instructions are executable to:

execute the at least one neural network to generate, based on the signals from the camera unit, the EDS, and the microphone, tracking of at least a portion of the face.

6. The assembly of claim 5 , wherein the portion comprises at least one eye pupil.

7. The assembly of claim 5 , wherein the portion comprises corners of the mouth.

8. The assembly of claim 5 , wherein the portion comprises the interior of the mouth including teeth.

9. A system comprising:

at least one camera unit configured to generate images of a person;

at least one microphone;

at least one event driven sensor (EDS) configured to output signals representative of the person; and

at least one processor programmed with instructions to:

process output of the microphone using a short term Fourier transform (STFT);

process output of the STFT using at least one audio processing convolutional neural network (CNN);

process at least features in images from the camera unit using at least one visual processing CNN;

process representations of output signals from the EDS using at least one event processing CNN; and

fuse outputs of the CNNs in fully connected neural network layers to generate at least one of:

a prediction of emotion of the person,

tracking of at least a portion of the face of the person,

at least one virtual reality (VR) image of the person,

an identification of the person.

10. The system of claim 9 , wherein the processor is configured with instructions to:

process outputs of the CNNs using a recurrent neural network (RNN); and

process output of the RNN using the fully connected neural network layers to generate mouth tracking of the person.

11. The system of claim 9 , wherein the processor is configured with instructions to:

generate a prediction of emotion of the person.

12. The system of claim 9 , wherein the processor is configured with instructions to:

generate tracking of at least a portion of the face of the person.

13. The system of claim 9 , wherein the processor is configured with instructions to:

generate at least one virtual reality (VR) image of the person.

14. The system of claim 9 , wherein the processor is configured with instructions to:

generate an identification of the person.

15. A method comprising:

receiving signals from at least one camera unit;

receiving signals from at least one event-driven sensor (EDS);

receiving signals from at least one microphone; and

executing, using at least one computer, at least one neural network to generate, based on the signals from the camera unit, the EDS, and the microphone, at least one of: emotion prediction, tracking of at least a portion of the face, identification of a person, generating a virtual reality (VR) image of the person, and

presenting on at least one display the at least one of: emotion prediction, tracking of at least the portion of the face, identification of the person, the VR image of the person, wherein

the EDS is configured to provide an output of −1 responsive to intensity of light being sensed decreasing, +1 responsive to intensity of light being sensed increasing, and 0 responsive to a change in light intensity below a threshold.

16. The method of claim 15 , comprising:

executing at least one neural network to generate, based on the signals from the camera unit, the EDS, and the microphone, emotion prediction.

17. The method of claim 15 , comprising:

executing at least one neural network to generate, based on the signals from the camera unit, the EDS, and the microphone, tracking of at least a portion of the face.

18. The method of claim 15 , comprising:

executing at least one neural network to generate, based on the signals from the camera unit, the EDS, and the microphone, identification of a person.

19. The method of claim 15 , comprising:

executing at least one neural network to generate, based on the signals from the camera unit, the EDS, and the microphone, generating a virtual reality (VR) image of the person.

20. The method of claim 17 , wherein the portion of the face is the corners of the mouth and interior of the mouth.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 15, 2022
From: YE, XIAOYONG; NAKAMURA, YUICHIRO
To: SONY INTERACTIVE ENTERTAINMENT INC.
Reel/Frame 061108/0529 →
Continuity (1)
Related Publication 20210350602A1 · Nov 11, 2021