IP Library › Granted Patent US 11,461,394
Granted Patent B2
US 11,461,394 · App. 17/039,880 · Granted Oct 4, 2022

Storing semi-structured data

Inventor: Martin Probst (Munich, DE)
Assignee: Google LLC
G06F16/86G06F16/213G06F16/83G06F16/835
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,461,394
App. No.
17/039,880
Granted
Oct 4, 2022
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for storing semi-structured data. One of the methods includes maintaining a plurality of schemas; receiving a first semi-structured data item; determining that the first semi-structured data item does not match any of the schemas in the plurality of schemas; and in response to determining that the first semi-structured data item does not match any of the schemas in the plurality of schemas: generating a new schema, encoding the first semi-structured data item in the first data format to generate the first new encoded data item in accordance with the new schema, storing the first new encoded data item in the data item repository, and associating the first new encoded data item with the new schema.

Claims (66)

1. A method comprising:

maintaining, at data processing hardware, a plurality of schemas each associated with one or more data items stored in a data format in a data item repository;

receiving, at the data processing hardware, a semi-structured data item in a semi-structured data format, the semi-structured data item comprising a plurality of key/value pairs;

determining, by the data processing hardware, that the semi-structured data item does not match any of the schemas in the plurality of schemas; and

in response to determining that the semi-structured data item does not match any of the schemas in the plurality of schemas:

generating, by the data processing hardware, a new schema for the semi-structured data item, the new schema specifying a corresponding maximum length for each value of at least one key/value pair in the plurality of key/value pairs of the received semi-structured data item;

determining, by the data processing hardware, whether each value of the at least one key/value pair in the plurality of key/value pairs exceeds the corresponding maximum length specified by the new schema; and

when each value of the at least one key/value pair in the plurality of key/value pairs fails to exceed the corresponding maximum length:

encoding, by the data processing hardware, each value as part of the semi-structured data item into an encoded semi-structured data item, the encoded semi-structured data item requiring less storage space than the semi-structured data item;

storing, by the data processing hardware, the encoded semi-structured data item in the data item repository;

assigning, by the data processing hardware, the new schema a unique identifier;

associating, by the data processing hardware, the encoded semi-structured data item with the new schema using the unique identifier; and

mapping, by the data processing hardware, a location of each value in the encoded semi-structured data item to a corresponding key using the new schema.

2. The method of claim 1 , wherein determining that the semi-structured data item does not match any of the schemas in the plurality of schemas comprises determining that the keys from the key/value pairs do not match the keys mapped to by any of the plurality of schemas.

3. The method of claim 1 , wherein a first schema of the plurality of schemas maps each of the keys from the key/value pairs to locations and identifies requirements for values of one or more of the keys from the key/value pairs, and wherein determining that the semi-structured data item does not match any of the schemas in the plurality of schemas comprises determining that the values from the key/value pairs do not satisfy the requirements identified in the first schema.

4. The method of claim 1 , further comprising:

receiving, at the data processing hardware, a second semi-structured data item, wherein the second semi-structured data item comprises one or more second key/value pairs;

determining, by the data processing hardware, that the second semi-structured data item matches a second schema from the plurality of schemas; and

in response to determining that the second semi-structured data item matches the second schema:

encoding, by the data processing hardware, the second semi-structured data item in the data format to generate a second new data item by storing values corresponding to the values from the second key/value pairs at respective locations in the second new data item in accordance with the second schema,

storing, by the data processing hardware, the second new data item in the data item repository, and

associating, by the data processing hardware, the second new data item with the second schema.

5. The method of claim 4 , wherein determining that the second semi-structured data item matches the second schema from the plurality of schemas comprises determining that the keys mapped to locations by the second schema match the keys from the second key/value pairs.

6. The method of claim 4 , wherein the second schema identifies requirements for values of one or more of the keys mapped to locations by the second schema.

7. The method of claim 6 , wherein determining that the second semi-structured data item matches the second schema from the plurality of schemas comprises determining that the values from the second key/value pairs satisfy the requirements identified in the second schema.

8. The method of claim 1 , further comprising:

receiving, by the data processing hardware, a query for semi-structured data items, wherein the query specifies requirements for values for one or more keys;

identifying, by the data processing hardware, schemas from the plurality of schemas that identify locations for values corresponding to each of the one or more keys;

for each identified schema, searching, by the data processing hardware, the data items associated with the schema to identify data items that satisfy the query; and

providing, by the data processing hardware, data identifying values from the data items that satisfy the query in response to the query.

9. The method of claim 8 , wherein searching the data items associated with the schema comprises searching, for each data item associated with the schema, the locations in the data item identified by the schema as storing values for the associated keys to identify whether the data item stores values for the associated keys that satisfy the requirements specified in the query.

10. The method of claim 1 , wherein the new schema further specifies a corresponding data type for each value of the at least one key/value pair in the plurality of key/value pairs of the received semi-structured data item.

11. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

maintaining a plurality of schemas each associated with one or more data items stored in a data format in a data item repository;

receiving a semi-structured data item in a semi-structured data format, the semi-structured data item comprising a plurality of key/value pairs;

determining that the semi-structured data item does not match any of the schemas in the plurality of schemas; and

in response to determining that the semi-structured data item does not match any of the schemas in the plurality of schemas:

generating a new schema for the semi-structured data item, the new schema specifying a corresponding maximum length for each value of at least one key/value pair in the plurality of key/value pairs of the received semi-structured data item;

determining whether each value of the at least one key/value pair in the plurality of key/value pairs exceeds the corresponding maximum length specified by the new schema; and

when each value of the at least one key/value pair in the plurality of key/value pairs fails to exceed the corresponding maximum length:

encoding each value as part of the semi-structured data item independent of a corresponding key into an encoded semi-structured data item, the encoded semi-structured data item requiring less storage space than the semi-structured data item;

storing the encoded semi-structured data item in the data item repository;

assigning the new schema a unique identifier;

associating the encoded semi-structured data item with the new schema using the unique identifier; and

mapping a location of each value in the encoded semi-structured data item to a corresponding key using the new schema.

12. The system of claim 11 , wherein determining that the semi-structured data item does not match any of the schemas in the plurality of schemas comprises determining that the keys from the key/value pairs do not match the keys mapped to by any of the plurality of schemas.

13. The system of claim 11 , wherein a first schema of the plurality of schemas maps each of the keys from the key/value pairs to locations and identifies requirements for values of one or more of the keys from the key/value pairs, and wherein determining that the semi-structured data item does not match any of the schemas in the plurality of schemas comprises determining that the values from the key/value pairs do not satisfy the requirements identified in the first schema.

14. The system of claim 11 , wherein the operations further comprise:

receiving a second semi-structured data item, wherein the second semi-structured data item comprises one or more second key/value pairs;

determining that the second semi-structured data item matches a second schema from the plurality of schemas; and

in response to determining that the second semi-structured data item matches the second schema:

encoding the second semi-structured data item in the data format to generate a second new data item by storing values corresponding to the values from the second key/value pairs at respective locations in the second new data item in accordance with the second schema,

storing the second new data item in the data item repository, and

associating the second new data item with the second schema.

15. The system of claim 14 , wherein determining that the second semi-structured data item matches the second schema from the plurality of schemas comprises determining that the keys mapped to locations by the second schema match the keys from the second key/value pairs.

16. The system of claim 14 , wherein the second schema identifies requirements for values of one or more of the keys mapped to locations by the second schema.

17. The system of claim 16 , wherein determining that the second semi-structured data item matches the second schema from the plurality of schemas comprises determining that the values from the second key/value pairs satisfy the requirements identified in the second schema.

18. The system of claim 11 , wherein the operations further comprise:

receiving a query for semi-structured data items, wherein the query specifies requirements for values for one or more keys;

identifying schemas from the plurality of schemas that identify locations for values corresponding to each of the one or more keys;

for each identified schema, searching the data items associated with the schema to identify data items that satisfy the query; and

providing data identifying values from the data items that satisfy the query in response to the query.

19. The system of claim 18 , wherein searching the data items associated with the schema comprises searching, for each data item associated with the schema, the locations in the data item identified by the schema as storing values for the associated keys to identify whether the data item stores values for the associated keys that satisfy the requirements specified in the query.

20. The system of claim 11 , wherein the new schema further specifies a corresponding data type for each value of the at least one key/value pair in the plurality of key/value pairs of the received semi-structured data item.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 1, 2020
From: PROBST, MARTIN
To: GOOGLE INC.
Reel/Frame 053945/0009 →
CONVERSION Recorded Oct 1, 2020
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 053962/0001 →
Continuity (3)
Continuation 15669603 · Aug 4, 2017
Continuation 14507690 · Oct 6, 2014
Related Publication 20210019291A1 · Jan 21, 2021
Cited By (1)
US 12,468,841