IP Library Granted Patent US 11,822,868
Granted Patent B2
US 11,822,868 · App. 15/906,388 · Granted Nov 21, 2023

Augmenting text with multimedia assets

Inventors: Emre Demiralp (San Jose, CA); Gavin Stuart Peter Miller (Los Altos, CA); Walter W. Chang (San Jose, CA); Grayson Squier Lang (Santa Clara, CA); Daicho Ito (San Jose, CA)
Assignee: ADOBE INC.
G06F40/134G06F16/00G06F40/279
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,822,868
App. No.
15/906,388
Granted
Nov 21, 2023
Kind
B2
Abstract

Systems and methods are provided for providing a navigation interface to access or otherwise use electronic content items. In one embodiment, an augmentation application identifies at least one entity referenced in a document. The entity can be referenced in at least two portions of the document by at least two different words or phrases. The augmentation application associates the at least one entity with at least one multimedia asset. The augmentation application generates a layout including at least some content of the document referencing the at least one entity and the at least one multimedia asset associated with the at least one entity. The augmentation application renders the layout for display.

Claims (122)

1. A method in which one or more processing devices perform operations comprising:

identifying an entity referenced in a document, wherein the entity is identified using a first word or phrase located in a first portion of the document, wherein identifying the entity comprises:

(i) executing a natural-language-processing algorithm that identifies entities in the document including the entity, wherein the natural-language-processing algorithm comprises identifying an n-gram in the document corresponding to the entity, wherein the n-gram comprises a contiguous sequence of n items from the document, and

wherein identifying the n-gram comprises:

splitting text in the document into token elements; and

applying, to the token elements, n-gram extraction;

(ii) accessing a feature ontology, wherein the feature ontology comprises a mapping between a plurality of entities identified by the natural-language-processing algorithm and a plurality of classifications for the entities, wherein the feature ontology includes a character class that maps character entities to corresponding character references in the document, a setting class that maps setting entities to corresponding setting references in the document, an emotion class that maps intangible emotion entities to a corresponding emotional state reference in the document; and

(iii) modifying the feature ontology to include the entity;

classifying the entity based in part on a perceptual property associated with the entity, wherein the perceptual property is a visual attribute;

associating the entity with a multimedia asset via the feature ontology, wherein the multimedia asset includes an image with the visual attribute;

determining that a second word or phrase in a second portion of the document refers to the entity;

generating a layout for the second portion of the document based on determining that the second word or phrase refers to the entity and retrieving the multimedia asset via the feature ontology, wherein the layout includes the multimedia asset associated with the entity, wherein generating the layout comprises:

determining a relationship between the entity and an additional entity referenced in the document, wherein determining the relationship comprises:

classifying the entity as a protagonist;

classifying the additional entity as an antagonist;

retrieving, from the feature ontology, an additional multimedia asset associated with the additional entity; and

based on the determined relationship between the entity and the additional entity, positioning the multimedia asset associated with the entity at a first position in the layout and positioning the additional multimedia asset associated with the additional entity at a second position in the layout facing the first position; and

rendering the layout with the second portion of the document for display.

2. The method of claim 1 , wherein identifying the n-gram further comprises:

converting the text in the document to a root form of the text;

removing, from the text, at least one of (i) an article or (ii) a preposition;

applying, to the token element, at least one of segmentation or part-of-speech tagging; and

modifying the feature ontology to include the n-gram corresponding to the entity.

3. The method of claim 2 , wherein converting the text in the document to the root form of the text comprises at least one of (i) converting a plural form of the text to a singular form of the text or (ii) converting a possessive version of the text to a non-possessive version of the text.

4. The method of claim 1 ,

wherein the entity comprises a plurality of entities referenced in the document and the multimedia asset comprises a plurality of multimedia assets,

wherein associating the entity with the multimedia asset comprises associating each entity from the plurality of entities with a respective one of the plurality of multimedia assets, and

wherein the operations further comprise:

identifying, for each entity from the plurality of entities, a respective portion of the document referencing the entity; and

generating a navigation interface comprising the multimedia assets associated with the plurality of entities, wherein each of the multimedia assets included in the navigation interface is configured to receive input representing a command to navigate to a respective portion of the document referencing a respective entity associated with the multimedia asset.

5. The method of claim 1 , wherein associating the entity with the multimedia asset comprises:

identifying a virtual community of clients to which the document is accessible;

generating a request to identify the multimedia asset;

providing the request to a client in the virtual community; and

receiving from the virtual community a recommended association of the entity with the multimedia asset, wherein the entity is associated with the multimedia asset based on the recommended association.

6. The method of claim 1 ,

wherein the additional entity is identified using a third word or phrase in the document that is different from the first word or phrase or the second word or phrase.

7. The method of claim 1 ,

wherein the entity comprises a plurality of entities referenced in the document and the multimedia asset comprises a plurality of multimedia assets,

wherein associating the entity with the multimedia asset comprises associating each of the plurality of entities with a respective one of the plurality of multimedia assets, and

wherein the operations further comprise:

determining a frequency with which the entity is referenced in the document;

generating, based on the frequency, analytics describing the document; and

providing an interface including the analytics.

8. The method of claim 1 , wherein the document comprises an electronic book and wherein the layout is rendered for display by an electronic book reader application.

9. A system comprising:

one or more processing devices; and

a non-transitory computer-readable medium communicatively coupled to the one or more processing devices,

wherein the one or more processing devices are configured to executed program code stored in the non-transitory computer-readable medium and thereby perform operations comprising:

identifying an entity referenced in a document, wherein the entity is identified using a first word or phrase located in a first portion of the document, wherein identifying the entity comprises:

(i) executing a natural-language-processing algorithm that identifies entities in the document including the entity, wherein the natural-language-processing algorithm comprises identifying an n-gram in the document corresponding to the entity, wherein the n-gram comprises a contiguous sequence of n items from the document, and wherein identifying the n-gram comprises:

splitting text in the document into token elements; and

applying, to the token elements, n-gram extraction;

(ii) accessing a feature ontology, wherein the feature ontology comprises a mapping between a plurality of entities identified by the natural-language-processing algorithm and a plurality of classifications for the entities, wherein the feature ontology includes a character class that maps character entities to corresponding character references in the document, a setting class that maps setting entities to corresponding setting references in the document, and an emotion class that maps intangible emotion entities to a corresponding emotional state reference in the document; and

(iii) modifying the feature ontology to include the entity;

classifying the entity based in part on a perceptual property associated with the entity, wherein the perceptual property is a visual attribute;

associating the entity with a multimedia asset via the feature ontology, wherein the multimedia asset includes an image with the visual attribute;

determining that a second word or phrase in a second portion of the document refers to the entity;

generating a layout for the second portion of the document based on determining that the second word or phrase refers to the entity and retrieving the multimedia asset via the feature ontology, wherein the layout includes the multimedia asset associated with the entity, wherein generating the layout comprises:

determining a relationship between the entity and an additional entity referenced in the document, wherein determining the relationship comprises:

classifying the entity as a protagonist;

classifying the additional entity as an antagonist;

retrieving, from the feature ontology, an additional multimedia asset associated with the additional entity; and

based on the determined relationship between the entity and the additional entity, positioning the multimedia asset associated with the entity at a first position in the layout and positioning the additional multimedia asset associated with the additional entity at a second position in the layout facing the first position; and

rendering the layout with the second portion of the document for display.

10. The system of claim 9 , wherein identifying the n-gram further comprises:

converting text in the document to a root form of the text;

removing, from the text, at least one of (i) an article or (ii) a preposition;

applying, to the token element, at least one of segmentation or part-of-speech tagging; and

modifying the feature ontology to include the n-gram corresponding to the entity.

11. The system of claim 10 , wherein converting the text in the document to the root form of the text comprises at least one of (i) converting a plural form of the text to a singular form of the text or (ii) converting a possessive version of the text to a non-possessive version of the text.

12. The system of claim 9 ,

wherein the entity comprises a plurality of entities referenced in the document and the multimedia asset comprises a plurality of multimedia assets,

wherein associating the entity with the multimedia asset comprises associating each entity from the plurality of entities with a respective one of the plurality of multimedia assets, and

wherein the operations further comprise:

identifying, for each entity from the plurality of entities, a respective portion of the document referencing the entity; and

generating a navigation interface comprising the multimedia assets associated with the plurality of entities, wherein each of the multimedia assets included in the navigation interface is configured to receive input representing a command to navigate to a respective portion of the document referencing a respective entity associated with the multimedia asset.

13. The system of claim 9 , wherein associating the entity with the multimedia asset comprises:

identifying a virtual community of clients to which the document is accessible;

generating a request to identify the multimedia asset;

providing the request to a client in the virtual community; and

receiving from the virtual community a recommended association of the entity with the multimedia asset, wherein the entity is associated with the multimedia asset based on the recommended association.

14. The system of claim 9 ,

wherein the additional entity is identified using a third word or phrase in the document that is different from the first word or phrase or the second word or phrase.

15. The system of claim 9 ,

wherein the entity comprises a plurality of entities referenced in the document and the multimedia asset comprises a plurality of multimedia assets,

wherein associating the entity with the multimedia asset comprises associating each of the plurality of entities with a respective one of the plurality of multimedia assets; and

wherein the operations further comprise:

determining a frequency with which the entity is referenced in the document;

generating, based on the frequency, analytics describing the document; and

providing an interface including the analytics.

16. A non-transitory computer-readable medium storing program code that when executed by one or more processing devices, causes the one or more processing devices to perform operations comprising:

identifying an entity referenced in a document, wherein the entity is identified using a first word or phrase located in a first portion of the document, wherein identifying the entity comprises:

(i) executing a natural-language-processing algorithm that identifies entities in the document including the entity, wherein the natural-language-processing algorithm comprises identifying an n-gram in the document corresponding to the entity, wherein the n-gram comprises a contiguous sequence of n items from the document, and wherein identifying the n-gram comprises:

splitting text in the document into token elements; and

applying, to the token elements, n-gram extraction;

(ii) accessing a feature ontology, wherein the feature ontology comprises a mapping between a plurality of entities identified by the natural-language-processing algorithm and a plurality of classifications for the entities, wherein the feature ontology includes a character class that maps character entities to corresponding character references in the document, a setting class that maps setting entities to corresponding setting references in the document, and an emotion class that maps intangible emotion entities to a corresponding emotional state reference in the document; and

(iii) modifying the feature ontology to include the entity;

classifying the entity based in part on a perceptual property associated with the entity, wherein the perceptual property is a visual attribute;

associating the entity with a multimedia asset via the feature ontology, wherein the multimedia asset includes an image with the visual attribute;

determining that a second word or phrase in a second portion of the document refers to the entity;

generating a layout for the second portion of the document based on determining that the second word or phrase refers to the entity and retrieving the multimedia asset via the feature ontology, wherein the layout includes the multimedia asset associated with the entity, wherein generating the layout comprises:

determining a relationship between the entity and an additional entity referenced in the document, wherein determining the relationship comprises:

classifying the entity as a protagonist;

classifying an additional entity as an antagonist;

retrieving, from the feature ontology, an additional multimedia asset associated with the additional entity; and

based on the determined relationship between the entity and the additional entity, positioning the multimedia asset associated with the entity at a first position in the layout and positioning the additional multimedia asset associated with the additional entity at a second position in the layout facing the first position; and

rendering the layout with the second portion of the document for display.

17. The non-transitory computer-readable medium of claim 16 , wherein identifying the n-gram comprises:

converting text in the document to a root form of the text;

removing, from the text, at least one of (i) an article or (ii) a preposition;

applying, to the token element, at least one of segmentation or part-of-speech tagging; and

modifying the feature ontology to include the n-gram corresponding to the entity.

18. The non-transitory computer-readable medium of claim 17 , wherein converting the text in the document to the root form of the text comprises at least one of (i) converting a plural form of the text to a singular form of the text or (ii) converting a possessive version of the text to a non-possessive version of the text.

19. The non-transitory computer-readable medium of claim 16 ,

wherein the entity comprises a plurality of entities referenced in the document and the multimedia asset comprises a plurality of multimedia assets,

wherein associating the entity with the multimedia asset comprises associating each entity from the plurality of entities with a respective one of the plurality of multimedia assets, and

wherein the operations further comprise:

identifying, for each entity from the plurality of entities, a respective portion of the document referencing the entity; and

generating a navigation interface comprising the multimedia assets associated with the plurality of entities, wherein each of the multimedia assets included in the navigation interface is configured to receive input representing a command to navigate to a respective portion of the document referencing a respective entity associated with the multimedia asset.

20. The non-transitory computer-readable medium of claim 16 ,

wherein the additional entity is identified using a third word or phrase in the document that is different from the first word or phrase or the second word or phrase.

Assignments (3)
CORRECTIVE ASSIGNMENT TO CORRECT THE THE 4TH ASSIGNOR DAICHI ITO'S NAME PREVIOUSLY RECORDED AT REEL: 45051 FRAME: 653. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded May 28, 2024
From: DEMIRALP, EMRE; MILLER, GAVIN STUART PETER; CHANG, WALTER W.; ITO, DAICHI; LANG, GRAYSON SQUIER
To: ADOBE SYSTEMS INCORPORATED
Reel/Frame 067959/0340 →
CHANGE OF NAME Recorded Mar 6, 2019
From: ADOBE SYSTEMS INCORPORATED
To: ADOBE INC.
Reel/Frame 048525/0042 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 27, 2018
From: DEMIRALP, EMRE; MILLER, GAVIN STUART PETER; CHANG, WALTER W.; ITO, DAICHO; LANG, GRAYSON SQUIER
To: ADOBE SYSTEMS INCORPORATED
Reel/Frame 045051/0653 →