IP Library Granted Patent US 11,074,939
Granted Patent B1
US 11,074,939 · App. 16/853,861 · Granted Jul 27, 2021

Disambiguation of audio content using visual context

Inventors: Jason Malinowski (Mahopac, NY); Swaminathan Balasubramanian (Troy, MI); Cheranellore Vasudevan (Bastrop, TX); Thomas G. Lawless, III (Wallkil, NY)
Assignee: International Business Machines Corporation
G11B27/036G06K9/00711G10L15/183
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,074,939
App. No.
16/853,861
Granted
Jul 27, 2021
Kind
B1
Abstract

Provided is a method for disambiguating an audio component extracted from audiovisual content. Audiovisual content is identified. The audiovisual content includes an audio component and a video component. An ambiguous expression is detected in the audio component. An object referenced by the ambiguous expression is identified in the video component. A verbal description of the object is generated. The verbal description is injected into the audio component to generate a modified audio component.

Claims (71)

1. A method for disambiguating an audio component extracted from audiovisual content, the method comprising:

identifying audiovisual content that includes an audio component and a video component;

detecting an ambiguous expression in the audio component;

determining to disambiguate the audio component for a user that is consuming the audio component, wherein determining to disambiguate the audio component comprises one or more selected from the group consisting of:

determining that the user's screen is off;

determining that the user is receiving the audio component but not the video component;

determining, based on a gaze of the user, that the user is not watching the video component; and

determining, based on a user profile, that the user has a visual impairment;

identifying, in the video component, an object referenced by the ambiguous expression;

generating a verbal description of the object; and

inserting the verbal description into the audio component to generate a modified audio component.

2. The method of claim 1 , the method further comprising:

transmitting the modified audio component to a user device.

3. The method of claim 1 , wherein detecting the ambiguous expression in the audio component comprises:

comparing words and phrases in the audio component to a dictionary of ambiguous expressions; and

determining that at least one word or phrase in the audio component is in the dictionary of ambiguous expressions.

4. The method of claim 1 , the method further comprising:

modifying the verbal description of the object prior to inserting the verbal description into the audio component, wherein the modifying comprises:

determining a subject of the object;

determining, using a user profile for the user consuming the audio component, a familiarity score of the user, the familiarity score indicating a level of familiarity of the user with the subject; and

modifying the verbal description of the object based on the familiarity score of the user.

5. The method of claim 1 , wherein detecting the ambiguous expression in the audio component comprises:

detecting, using natural language processing, a phrase in the audio component that directs a focus of the user to a portion of the video component; and

determining that the portion of the video component is not visible to the user.

6. A system for disambiguating an audio component extracted from audiovisual content, the system comprising:

a memory; and

a processor communicatively coupled to the memory, wherein the processor is configured to perform a method comprising:

identifying audiovisual content that includes an audio component and a video component;

detecting an ambiguous expression in the audio component, wherein detecting the ambiguous expression in the audio component comprises:

detecting, using natural language processing, a phrase in the audio component that directs a focus of a user to a portion of the video component; and

determining that the portion of the video component is not visible to the user;

identifying, in the video component, an object referenced by the ambiguous expression;

generating a verbal description of the object; and

inserting the verbal description into the audio component to generate a modified audio component.

7. The system of claim 6 , wherein the method further comprises:

transmitting the modified audio component to a user device.

8. The system of claim 6 , wherein the method further comprises:

determining to disambiguate the audio component for the user.

9. The system of claim 8 , wherein the user is consuming the audio component, and wherein determining to disambiguate the audio component for the user comprises one or more selected from the group consisting of:

determining that the user's screen is off;

determining that the user is receiving the audio component but not the video component;

determining, based on a gaze of the user, that the user is not watching the video component; and

determining, based on a user profile, that the user has a visual impairment.

10. The system of claim 6 , wherein detecting the ambiguous expression in the audio component comprises:

comparing words and phrases in the audio component to a dictionary of ambiguous expressions; and

determining that at least one word or phrase in the audio component is in the dictionary of ambiguous expressions.

11. The system of claim 6 , wherein the method further comprises:

modifying the verbal description of the object prior to inserting the verbal description into the audio component, wherein the modifying comprises:

determining a subject of the object;

determining, using a user profile for the user consuming the audio component, a familiarity score of the user, the familiarity score indicating a level of familiarity of the user with the subject; and

modifying the verbal description of the object based on the familiarity score of the user.

12. A computer program product for disambiguating an audio component extracted from audiovisual content, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to perform a method comprising:

identifying audiovisual content that includes an audio component and a video component;

detecting an ambiguous expression in the audio component;

determining to disambiguate the audio component for a user that is consuming the audio component, wherein determining to disambiguate the audio component comprises one or more selected from the group consisting of:

determining that the user's screen is off;

determining that the user is receiving the audio component but not the video component;

determining, based on a gaze of the user, that the user is not watching the video component; and

determining, based on a user profile, that the user has a visual impairment identifying, in the video component, an object referenced by the ambiguous expression;

generating a verbal description of the object; and

inserting the verbal description into the audio component to generate a modified audio component.

13. The computer program product of claim 12 , wherein the method further comprises:

transmitting the modified audio component to a user device.

14. The computer program product of claim 12 , wherein detecting the ambiguous expression in the audio component comprises:

comparing words and phrases in the audio component to a dictionary of ambiguous expressions; and

determining that at least one word or phrase in the audio component is in the dictionary of ambiguous expressions.

15. The computer program product of claim 12 , wherein the method further comprises:

modifying the verbal description of the object prior to inserting the verbal description into the audio component, wherein the modifying comprises:

determining a subject of the object;

determining, using a user profile for the user consuming the audio component, a familiarity score of the user, the familiarity score indicating a level of familiarity of the user with the subject; and

modifying the verbal description of the object based on the familiarity score of the user.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 21, 2020
From: MALINOWSKI, JASON; BALASUBRAMANIAN, SWAMINATHAN; VASUDEVAN, CHERANELLORE; LAWLESS, THOMAS G., III
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 052449/0938 →
Cited By (2)
US 12,556,559 US 12,689,640