IP Library Granted Patent US 11,375,293
Granted Patent B2
US 11,375,293 · App. 16/177,232 · Granted Jun 28, 2022

Textual annotation of acoustic effects

Inventors: Naveen Kumar (San Mateo, CA); Justice Adams (San Mateo, CA); Arindam Jati (Los Angeles, CA); Masanori Omote (Half Moon Bay, CA)
Assignee: Sony Interactive Entertainment Inc.
H04N21/84G06F16/635G06F16/685G06N3/08G09B21/00H04N21/233
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,375,293
App. No.
16/177,232
Granted
Jun 28, 2022
Kind
B2
Abstract

Accommodation for color or visual impairments may be implemented by selective color substitution. A color accommodation module receives an image frame from a host system and generates a color-adapted version of the image frame. The color accommodation module may include a rule based filter that substitutes one or more colors within the image frame with one or more corresponding alternative colors.

Claims (16)

1. A system for enhancing the accessibility of Audio Visual content, the system comprising:

an acoustic effects annotation module configured classify primary audio events occurring within an audio segment to generate one or more tags describing the primary audio events occurring within the audio segment, wherein the acoustic effects annotation module includes two neural networks that together are configured to classify the primary acoustic effects occurring within the audio segment wherein a first neural network is trained with unsupervised learning techniques to predict the primary acoustic effects occurring within the audio segment and a second neural network trained with supervised learning techniques to classify the primary acoustic effects predicted by the first neural network and wherein the acoustic effect annotation module includes an audio tagging function that combines the one or more tags with video frame information wherein the tags appear on the corresponding video images.

2. The system of claim 1 wherein the one or more primary audio events include a top three most important sounds within the audio segment.

3. The system of claim 1 wherein the audio segment is a clip of video game audio having multiple sounds associated with multiple sources.

4. The system of claim 1 , further comprising a controller coupled to the acoustic effects annotation module, wherein the controller is configured to provide the one or more tags to a host system for display on a display screen and synchronize the output of the acoustic effects annotation module with one or more other neural network modules.

5. The system of claim 4 wherein the one or more other neural network modules includes a Graphical Style Modification module configured to apply a style adapted from a reference image frame to a source image frame wherein the source image frame is synchronized to appear during the audio segment.

6. The system of claim 1 , further comprising a controller coupled to the host system and the action description module, wherein the controller is configured to synchronize presentation of text corresponding to the one or more tags with display of a sequence of image frames associated with the audio segment.

7. A method for enhancing accessibility of Audio Visual content, comprising:

classifying primary acoustic effects occurring within an audio segment to generate one or more tags describing the primary acoustic effects occurring within the audio segment with an acoustic effects annotation module, wherein the acoustic effects annotation module includes two neural networks that together are configured to classify the primary acoustic effects occurring within the audio segment, wherein the first neural network is trained unsupervised learning techniques to predict the primary acoustic effects occurring within the audio segment and a second neural network trained with supervised learning techniques to classify the primary acoustic effects predicted by the first neural network and wherein the acoustic effect annotation module includes an audio tagging function that combines the one or more tags with video frame information wherein the tags appear on the corresponding video images.

8. The method of claim 7 wherein the one or more primary audio events include a top three most important sounds within the audio segment.

9. The method of claim 7 wherein the audio segment is a clip of video game audio having multiple sounds associated with multiple sources.

10. The method of claim 7 , further comprising providing the one or more tags to a host system for display on a display screen and synchronizing the output of the audio description module with one or more other neural network modules with a controller coupled to the audio description module.

11. The method of claim 10 wherein the one or more other neural network modules includes a Graphical Style Modification module configured to apply a style adapted from a reference image frame to a source image frame wherein the source image frame is synchronized to appear during the audio segment.

12. The method of claim 7 , further comprising a controller coupled to the host system and the action description module, wherein the controller is configured to synchronize presentation of text corresponding to the one or more tags with display of a sequence of image frames associated with the audio segment.

13. A non-transitory computer-readable medium having computer readable instructions embodied therein, the instructions being configured upon execution to implement a method for enhancing accessibility of Audio Visual content, the method comprising

classifying primary audio events occurring within an audio segment to generate one or more tags describing the primary audio events occurring within the audio segment with an audio description module, wherein the acoustic effects annotation module includes two neural networks that together are configured to classify the primary acoustic effects occurring within the audio segment, wherein a first neural network is trained unsupervised learning techniques to predict the primary acoustic effects occurring within the audio segment and a second neural network trained with supervised learning techniques to classify the primary acoustic effects predicted by the first neural network and wherein the acoustic effect annotation module includes an audio tagging function that combines the one or more tags with video frame information wherein the tags appear on the corresponding video images.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 2, 2022
From: KUMAR, NAVEEN; ADAMS, JUSTICE; JATI, ARINDAM; OMOTE, MASAONORI
To: SONY INTERACTIVE ENTERTAINMENT INC.
Reel/Frame 059777/0178 →
Continuity (1)
Related Publication 20200137463A1 · Apr 30, 2020
Cited By (2)
US 12,288,567 US 12,530,401