IP Library Granted Patent US 9,645,816
Granted Patent B2
US 9,645,816 · App. 14/866,123 · Granted May 9, 2017

Multi-language code search index

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,645,816
App. No.
14/866,123
Granted
May 9, 2017
Kind
B2
Abstract

A method and apparatus for generating a code index for multiple types of code is provided. The method comprises: analyzing a plurality of files that includes a first file that contains first code in a first programming language and a second file that contains second code in a second programming language; identifying a first plurality of tokens within the first file based on a first tokenizing approach; identifying a second plurality of tokens within the second file based on a second tokenizing approach that is different than the first tokenizing approach; storing the first plurality of tokens and the second plurality of tokens within a particular index.

Claims (53)

1. A method comprising:

analyzing a plurality of files that includes a first file that contains first code in a first programming language and a second file that contains second code in a second programming language;

identifying a first plurality of tokens within the first file based on a first tokenizing approach;

identifying a second plurality of tokens within the second file based on a second tokenizing approach that is different than the first tokenizing approach;

storing the first plurality of tokens and the second plurality of tokens within a particular index that comprises a plurality of language-specific fields and one or more language-specific values associated with each of the plurality of language-specific fields;

analyzing the particular index based on a query that specifies a language-specific field of the plurality of language-specific fields and one of the one or more language-specific values associated with the language-specific field;

returning, in response to the analyzing, an indication of a file comprising the one of the one or more language-specific values associated with the language-specific field;

wherein the method is performed by one or more computing devices.

2. The method of claim 1 , wherein:

the particular index includes a particular entry that stores a token that is in the first plurality of tokens and the second plurality of tokens:

the particular entry includes data that identifies the first programming language and the second programming language.

3. The method of claim 1 , further comprising:

receiving a search query that includes one or more terms;

in response to receiving the search query, determining, based on the particular index, a result of the search query, wherein the result indicates the first file and the second file.

4. The method of claim 3 , wherein the one or more terms of the search query comprise one or more Boolean operators.

5. The method of claim 3 , wherein the one or more terms of the search query comprise terms identifying the first programming language and/or the second programming language.

6. The method of claim 1 , further comprising:

using a first parsing technique to identify first tokens in the first plurality of tokens;

using a second parsing technique that is different than the first parsing technique to identify second tokens in the second plurality of tokens.

7. The method of claim 1 , wherein the first programming language and the second programming language are selected from a group consisting of: Java, SCALA, Pearl, C++, JSON, SCSS, and JSP.

8. The method of claim 1 , further comprising:

identifying, within the first file, a code block generated when the first file is executed;

storing, in the particular index, a code block identifier for the code block generated when the first file is executed; and

storing, in the particular index, an association between the code block identifier and the first file.

9. The method of claim 8 , wherein the code block generated when the first file is executed is generated in the second programming language.

10. The method of claim 8 , wherein the code block generated when the first file is executed is generated in the first programming language.

11. A data processing system, comprising:

one or more processors;

a non-transitory computer-readable medium having instructions embodied thereon, the instructions when executed by the one or more processors, cause performance of:

analyzing a plurality of files that includes a first file that contains first code in a first programming language and a second file that contains second code in a second programming language;

identifying a first plurality of tokens within the first file based on a first tokenizing approach;

identifying a second plurality of tokens within the second file based on a second tokenizing approach that is different than the first tokenizing approach;

storing the first plurality of tokens and the second plurality of tokens within a particular index that comprises a plurality of language-specific fields and one or more language-specific values associated with each of the plurality of language-specific fields;

analyzing the particular index based on a query that specifies a language-specific field of the plurality of language-specific fields and one of the one or more language-specific values associated with the language-specific field;

returning, in response to the analyzing, an indication of a file comprising the one of the one or more language-specific values associated with the language-specific field.

12. The data processing system of claim 11 , wherein:

the particular index includes a particular entry that stores a token that is in the first plurality of tokens and the second plurality of tokens:

the particular entry includes data that identifies the first programming language and the second programming language.

13. The data processing system of claim 11 , wherein the instructions, when executed by the one or more processors, further cause performance of:

receiving a search query that includes one or more terms;

in response to receiving the search query, determining, based on the particular index, a result of the search query, wherein the result indicates the first file and the second file.

14. The data processing system of claim 13 , wherein the one or more terms of the search query comprise one or more Boolean operators.

15. The data processing system of claim 13 , wherein the one or more terms of the search query comprise terms identifying the first programming language and the second programming language.

16. The data processing system of claim 11 , wherein the instructions, when executed by the one or more processors, further cause performance of:

using a first parsing technique to identify first keywords in the first plurality of tokens;

using a second parsing technique that is different than the first parsing technique to identify second keywords in the second plurality of tokens.

17. The data processing system of claim 11 , wherein the first programming language and the second programming language are selected from a group consisting of: Java, SCALA, Pearl, C++, JSON, SCSS, and JSP.

18. The data processing system of claim 11 , wherein the instructions, when executed by the one or more processors, further cause:

identifying, within the first file, a code block generated when the first file is executed;

storing, in the particular index, a code block identifier for the code block generated when the first file is executed; and

storing, in the particular index, an association between the code block identifier and the first file.

19. The data processing system of claim 18 , wherein the code block generated when the first file is executed is generated in the second programming language.

20. The data processing system of claim 18 , wherein the code block generated when the first file is executed is generated in the first programming language.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 1, 2017
From: LINKEDIN CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 044746/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 25, 2015
From: SHLOSBERG, VLAD
To: LINKEDIN CORPORATION
Reel/Frame 036661/0540 →