IP Library Granted Patent US 10,108,707
Granted Patent B1
US 10,108,707 · App. 15/712,571 · Granted Oct 23, 2018

Data ingestion pipeline

Inventors: Yahui Chu (Cheswick, PA); Stephen Allen Whitney (Sunnyvale, CA)
Assignee: Amazon Technologies, Inc.
G06F17/30761G06F17/30516G06F17/30867G10L13/08G10L15/265
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,108,707
App. No.
15/712,571
Granted
Oct 23, 2018
Kind
B1
Abstract

Techniques for expanding system capabilities to execute user commands relating to trending topics (e.g., real-time news questions, trending questions, sports questions, game questions, politic questions, etc.) are described. The system gathers data from a variety of sources (e.g., news feeds, social media feeds, RSS feeds, news websites, etc.). The system segments gathered data corresponding to, for example, topic and or entity. The system may only store data corresponding to a topic or entity in a dedicated trending storage if the system receives data corresponding to the topic or entity from a number of different sources satisfying a threshold number of sources. Data in the dedicated trending storage may be maintained using decay models or algorithms. For example, the more often the system receives data corresponding to a topic or entity from one or more sources, the longer the data is maintained in the storage, and vice versa.

Claims (93)

1. A computer-implemented method comprising:

during a first period of time:

receiving, from a first remote device, first content data associated with a topic;

receiving, from a second remote device, second content data associated with the topic;

based on the first content data and the second content data both being associated with the topic, grouping the first content data and the second content data to generate first grouped data;

determining the first remote device and the second remote device correspond to a number of remote devices satisfying a threshold number of remote devices;

storing, based on the number satisfying the threshold number, the first grouped data as first stored data;

during a second period of time after the first period of time:

receiving, from a first device, input audio data corresponding to an utterance;

performing speech processing on the input audio data to determine a command corresponding to the topic;

determining, in a profile associated with the first device, a preferred content source associated with the topic;

determining the first remote device corresponds to the preferred content source;

performing text-to-speech (TTS) processing on the first content data to generate output audio data; and

causing the first device to emit audio corresponding to the output audio data.

2. The computer-implemented method of claim 1 , further comprising:

determining a number of sources that, within a time period, publish data corresponding to the topic;

determining that the number falls below a frequency threshold; and

determining the first stored data is no longer trending.

3. The computer-implemented method of claim 1 , further comprising:

determining, during a first time period, a first number of sources from which a first plurality of data corresponding to the topic is received; determining, during a second time period, a second number of sources from which a second plurality of data corresponding to the topic is received;

determining the second number is less than the first number; and

determining, based on the second number being less than the first number, a future time when the first stored data is to be deleted from a trending knowledge base.

4. The computer-implemented method of claim 3 , further comprising:

storing the first grouped data in a general knowledge base; and

permitting the first grouped data to persist in the general knowledge base after the future time.

5. A system comprising:

at least one processor; and

at least one memory including instructions that, when executed by the at least one processor, cause the system to:

receive input data;

perform speech processing on the input data to determine the input data corresponds to a topic;

determine, from profile data associated with a device, a preferred content source associated with the topic;

determine stored data corresponding to the topic, the stored data being received from a number of content sources satisfying a threshold number of content sources;

determine at least a portion of the stored data received from the preferred content source; and

cause the device to output content corresponding to the at least a portion.

6. The system of claim 5 , wherein the at least one memory further includes instructions that, when executed by the at least one processor, further cause the system to:

receive first text data from a first content source;

receive second text data from a second content source;

perform natural language processing on the first text data to determine the topic;

perform natural language processing on the second text data to determine the topic;

determine the number to include at least to the first content source and the second content source;

determine the number satisfies the threshold number; and

generate, based on the number satisfying the threshold number, the stored data to include the first text data and the second text data.

7. The system of claim 5 , wherein the at least one memory further includes instructions that, when executed by the at least one processor, further cause the system to:

determine a number of content sources that, within a time period, publish data corresponding to the topic;

determine that the number falls below a frequency threshold; and

based on the number falling below the frequency threshold, determine a future time that the stored data is to be deleted from a trending knowledge.

8. The system of claim 7 , wherein the at least one memory further includes instructions that, when executed by the at least one processor, further cause the system to:

store the stored data in a general knowledge base as second stored data; and

permit the second stored data to persist in the general knowledge base after the future time.

9. The system of claim 5 , wherein the at least one memory further includes instructions that, when executed by the at least one processor, further cause the system to:

determine, during a first time period, a first number of content sources from which a first plurality of data corresponding to the topic is received;

determine, during a second time period, a second number of content sources from which a second plurality of data corresponding to the topic is received;

determine the second number is less than the first number; and

determine, based on the second number being less than the first number, a future time when the stored data is to be deleted.

10. The system of claim 5 , wherein the at least one memory further includes instructions that, when executed by the at least one processor, further cause the system to:

perform natural language processing on the stored data to generate natural language results;

generate, based on the natural language results, a synopsis of the stored data; and

cause the device to output content corresponding to the synopsis.

11. The system of claim 5 , wherein the stored data corresponds to at least one of website content data, news data, news feed data, podcast data, or rich site summary (RSS) feed data.

12. The system of claim 5 , wherein the stored data includes a plurality of data files received from a single content source, the plurality of data files satisfying a threshold number of data files.

13. A computer-implemented method comprising:

receiving input data;

performing speech processing on the input data to determine the input data corresponds to a topic;

determining, from profile data associated with a device, a preferred content source associated with the topic;

determining stored data corresponding to the topic, the stored data being received from a number of content sources satisfying a threshold number of content sources;

determining at least a portion of the stored data received from the preferred content source; and

causing the device to output content corresponding to the at least a portion.

14. The computer-implemented method of claim 13 , further comprising

receiving first text data from a first content source;

receiving second text data from a second content source;

performing natural language processing on the first text data to determine the topic;

performing natural language processing on the second text data to determine the topic;

determining the number to include at least to the first content source and the second content source;

determining the number satisfies the threshold number; and

generating, based on the number satisfying the threshold number, the stored data to include the first text data and the second text data.

15. The computer-implemented method of claim 13 , further comprising:

determining a number of content sources that, within a time period, publish data corresponding to the topic;

determining that the number falls below a frequency threshold; and

based on the number falling below the frequency threshold, determine a future time that the stored data is to be deleted from a trending knowledge base.

16. The computer-implemented method of claim 15 , further comprising:

storing the stored data in a general knowledge base as second stored data; and

permitting the second stored data to persist in the general knowledge base after the future time.

17. The computer-implemented method of claim 13 , further comprising:

determining, during a first time period, a first number of content sources from which a first plurality of data corresponding to the topic is received;

determining, determine a second time period, a second number of content sources from which a second plurality of data corresponding to the topic is received;

determining the second number is less than the first number; and

determining, based on the second number being less than the first number, a future time when the stored data is to be deleted.

18. The computer-implemented method of claim 13 , further comprising:

performing natural language processing on the stored data to generate natural language results;

generating, based on the natural language results, a synopsis of the stored data; and

causing the device to output content corresponding to the synopsis.

19. The computer-implemented method of claim 13 , wherein the stored data corresponds to at least one of web site content data, news data, news feed data, podcast data, or rich site summary (RSS) feed data.

20. The computer-implemented method of claim 13 , wherein the stored data includes a plurality of data files received from a single content source, the plurality of data files satisfying a threshold number of data files.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 22, 2017
From: CHU, YAHUI; WHITNEY, STEPHEN ALLEN
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 043663/0770 →
Cited By (4)
US 12,198,413 US 12,374,097 US 12,406,316 US 12,475,698