IP Library Granted Patent US 10,805,647
Granted Patent B2
US 10,805,647 · App. 15/850,746 · Granted Oct 13, 2020

Automatic personalized story generation for visual media

Inventors: Ying Zhang (Palo Alto, CA); Shengbo Guo (San Jose, CA)
Assignee: FACEBOOK, INC.
H04N21/23418G06F40/10G06F40/20G06F40/56G06K9/00751G06Q50/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,805,647
App. No.
15/850,746
Granted
Oct 13, 2020
Kind
B2
Abstract

Exemplary embodiments relate to the automatic generation of captions for visual media, including photos, photo albums, non-live video, and live video. The visual media may be analyzed to determine contextual information (such as location information, people and objects in the video, time, etc.). A system may integrate this information with information from the user's social network and a personalized language model built using public-facing language from the user. The personalized language model captures the user's way of speaking to make the generated captions more detailed and personalized. The language model may account for the context in which the video was generated. The captions maybe used to simplify and encourage content generation, and may also be used to index visual media, rank the media, and recommend the media to users likely to engage with the media.

Claims (30)

1. A method, comprising: accessing visual media;

analyzing information associated with the visual media to identify a context of the visual media;

providing the context to a personalized language model configured to reflect a personal narrative style of the user; and

generating a caption for the visual media in the user's personal narrative style using the personalized language model, wherein generating the caption comprises applying the personalized language model in a first personal narrative style specific to the user for a first context and a second personal narrative style specific to the same user, differing from the first narrative style, in a second context differing from the first context.

2. The method of claim 1 , wherein the context includes location information, a person in the visual media, a recognized object in the visual media, a location of the visual media, or a time at which the visual media is recorded.

3. The method of claim 1 , wherein the personalized language model is configured using public-facing language collected from the user.

4. The method of claim 1 , wherein the user is connected to a second user in a social graph, and wherein the personalized language model is configured using public-facing language of the second user.

5. The method of claim 1 , wherein the caption serves as an index to the visual media for ranking or recommending the visual media.

6. The method of claim 1 , wherein the visual media is provided through an API call from a third-party source.

7. A non-transitory computer-readable medium storing instructions configured to cause one or more processors to:

access visual media;

analyze information associated with the visual media to identify a context of the visual media;

provide the context to a personalized language model configured to reflect a personal narrative style of the user; and

generate a caption for the visual media in the user's personal narrative style using the personalized language model, wherein generating the caption comprises applying the personalized language model in a first personal narrative style specific to the user for a first context and a second personal narrative style specific to the same user, differing from the first narrative style, in a second context differing from the first context.

8. The medium of claim 7 , wherein the context includes location information, a person in the visual media, a recognized object in the visual media, a location of the visual media, or a time at which the visual media is recorded.

9. The medium of claim 7 , wherein the personalized language model is configured using public-facing language collected from the user.

10. The medium of claim 7 , wherein the user is connected to a second user in a social graph, and wherein the personalized language model is configured using public-facing language of the second user.

11. The medium of claim 7 , wherein the caption serves as an index to the visual media for ranking or recommending the visual media.

12. The medium of claim 7 , wherein the visual media is provided through an API call from a third-party source.

13. An apparatus comprising:

a non-transitory computer readable medium configured to store instructions for interacting with visual media; and

a processor configured to execute the instructions, the instructions configured to cause the processor to:

access the visual media;

analyze information associated with the visual media to identify a context of the visual media;

provide the context to a personalized language model configured to reflect a personal narrative style of the user; and

generate a caption for the visual media in the user's personal narrative style using the personalized language model, wherein generating the caption comprises applying the personalized language model in a first personal narrative style specific to the user for a first context and a second personal narrative style specific to the same user, differing from the first narrative style, in a second context differing from the first context.

14. The apparatus of claim 13 , wherein the context includes location information, a person in the visual media, a recognized object in the visual media, a location of the visual media, or a time at which the visual media is recorded.

15. The apparatus of claim 13 , wherein the personalized language model is configured using public-facing language collected from the user.

16. The apparatus of claim 13 , wherein the user is connected to a second user in a social graph, and wherein the personalized language model is configured using public-facing language of the second user.

17. The apparatus of claim 13 , wherein the caption serves as an index to the visual media for ranking or recommending the visual media.

Assignments (2)
CHANGE OF NAME Recorded May 5, 2022
From: FACEBOOK, INC.
To: META PLATFORMS, INC.
Reel/Frame 059858/0387 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 22, 2018
From: ZHANG, YING; GUO, SHENGBO
To: FACEBOOK, INC.
Reel/Frame 044683/0728 →
Continuity (1)
Related Publication 20190200050A1 · Jun 27, 2019
Cited By (1)
US 12,382,139