IP Library Granted Patent US 7,933,765
Granted Patent B2
US 7,933,765 · App. 11/692,777 · Granted Apr 26, 2011

Cross-lingual information retrieval

Assignee: Corbis Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,933,765
App. No.
11/692,777
Granted
Apr 26, 2011
Kind
B2
Abstract

Multi-lingual search and retrieval of digital content. Embodiments are generally directed to methods and systems for creating an English language database that associates non-English terms with English terms in multiple categories of metadata. Language experts use an interface to create equivalencies between non-English terms and English terms, Boolean expressions, synonyms, and other forms of search terms. Language dictionaries and other sources also create equivalencies. The database is used to evaluate non-English search terms submitted by a user, and to determine English search terms that can be used to perform a search for content. The multiple categories of metadata may comprise structured data, such as keywords of a structured vocabulary, and/or unstructured data, such as captions, titles, descriptions, etc. Weighting and/or prioritization can be applied to the search terms, to the process of searching the multiple categories, and/or to the search results, to rank the search results.

Claims (81)

1. A method for identifying digital content with a client computer, the method enabling operations, comprising:

generating an equivalency list with a translations generator in communication with the client computer, wherein the list is based on a secondary-language query term associated with at least one primary-language query term, and wherein each of the at least one primary-language query term is in a pre-selected language and the secondary-language term is in a language that is different from the pre-selected primary language;

receiving the secondary-language query term in a search request for a search engine that is in communication with the client computer;

selecting the at least one primary-language query term from the equivalency list, based on the received secondary-language query term;

identifying digital content that is associated with structured text metadata, if the at least one primary-language query term is included in the structured text metadata; and

identifying digital content that corresponds to unstructured free-text metadata, if the at least one primary-language query term is included in the corresponding unstructured free-text metadata and is not a unique identifier of a defined term in a controlled vocabulary.

2. The method of claim 1 , wherein the equivalency list is one of a plurality of equivalency lists, each of which comprises another query term in a language different from that of any other one of the plurality of equivalency lists, and wherein the other query term is associated with the at least one primary-language query term.

3. The method of claim 1 , wherein the at least one primary-language query term comprises at least one of the following; a controlled vocabulary keyword, the secondary-language query term, a synonym associated with the secondary-language query term, and a boolean expression of terms in the pre-selected language.

4. The method of claim 1 , wherein generating the equivalency list comprises:

providing an interface enabling a user to associate the at least one primary-language query term with at least one term in the language of the secondary-language term that is different from the pre-selected primary language;

receiving an indication through the interface that the secondary-language query term is a primary term; and

receiving an indication through the interface on whether a unique identifier shall be associated with a controlled vocabulary term in the at least one primary-language query term.

5. The method of claim 1 , wherein generating the equivalency list comprises determining whether to remove the unique identifier from the at least one primary-language query term.

6. The method of claim 1 , wherein the unique identifier indicates a limitation of a meaning of the at least one primary-language query term.

7. The method of claim 1 , wherein the structured text metadata comprises at least one keyword from a controlled vocabulary, wherein each of the at least one keyword is identified by a unique identifier associated with a precise concept.

8. The method of claim 1 , wherein the corresponding unstructured free-text metadata comprises at least one of the following; a caption, a title, a paragraph, and a date.

9. The method of claim 1 , further comprising weighting at least one of the following; the at least one primary-language query term, the structured text metadata, and the corresponding unstructured free-text metadata.

10. The method of claim 1 , further comprising prioritizing the identified content based on a weighting of at least one of the following; the at least one primary-language query term, the structured text metadata, and the corresponding unstructured free-text metadata.

11. The method of claim 1 , further comprising removing a noise word from the secondary-language query term prior to selecting the at least one primary-language query term from the equivalency list.

12. The method of claim 1 , wherein the pre-selected language comprises English.

13. A processor readable non-transitory storage medium that includes a plurality of executable instructions, wherein the execution of the instructions enables operations for performing the steps of claim 1 .

14. A system for identifying digital content with a client computer, comprising:

a translations generator that is in communication with the client computer, the translations generator is arranged to perform a plurality of operations including:

generating an equivalency list based on a secondary-language query term associated with at least one primary-language query term, wherein each of the at least one primary-language query term is in a pre-selected language and the secondary-language term is in a language that is different from the pre-selected primary language;

a translating machine that is in communication with the client computer, the translations generator, and a search engine, the translating machine is arranged to perform a plurality of operations including:

receiving the secondary-language query term in a search request for the search engine; and

selecting the at least one primary-language query term from the equivalency list, based on the secondary-language query term; and

the search engine that performs a plurality of operations, including:

identifying digital content that is associated with structured text metadata, if the at least one primary-language query term is included in the structured text metadata; and

identifying digital content that corresponds to unstructured free-text metadata, if the at least one primary-language query term is included in the corresponding unstructured free-text metadata and is not a unique identifier of a defined term in a controlled vocabulary.

15. The system of claim 14 , wherein the translations generator is further arranged to perform a plurality of operations, including:

providing an interface enabling a user to associate the at least one primary-language query term with at least one term in the language of the secondary-language term that is different from the pre-selected primary language;

receiving an indication through the interface that the secondary-language query term is a primary term; and

receiving an indication through the interface on whether a unique identifier shall be associated with a controlled vocabulary term in the at least one primary-language query term.

16. The system of claim 14 , wherein the search engine further performs the operation of prioritizing the identified content based on a weighting of at least one of the following; the at least one primary-language query term, the structured text metadata, and the corresponding unstructured free-text metadata.

17. A method for associating terms in an equivalency list for identifying digital content with a client computer in communication with a translations generator, the translations generator is arranged to perform a plurality of operations; comprising:

associating a secondary-language term with a controlled vocabulary keyword in a primary-language, if the secondary-language term has a unique meaning depending on a context;

indicating that the secondary-language term exists in the primary language, if the secondary-language term is identical in the primary language;

associating the secondary-language term with a synonym in the primary language, if the secondary-language term is synonymous with the synonym;

designating the secondary-language term as a primary translation based on a primary-language term; and

associating the secondary-language term with a Boolean expression, if a meaning of the secondary-language term can be expressed by a combination of primary-language terms.

18. The method of claim 17 , wherein the controlled vocabulary keyword includes a unique identifier indicating that the controlled vocabulary keyword has a meaning depending on the context.

19. The method of claim 17 , further comprising at least one of the following:

identifying digital content associated with the secondary-language term based on structured text metadata that is associated with the controlled vocabulary keyword; and

identifying digital content associated with the secondary-language term based on corresponding unstructured free-text metadata that is associated with at least one of the following; the secondary-language term that is identical in the primary language, the synonym in the primary language; and the Boolean expression.

20. A method for generating a list for identifying digital content with a client computer in communication with a translations generator, the translations generator is arranged to perform a plurality of operations, comprising:

receiving a subset of equivalencies comprising a plurality of secondary-language terms that are associated with a primary-language term;

parsing the subset into a list of equivalencies, wherein each equivalency comprises an association of at least one of the plurality of secondary-language terms with the primary-language term;

associating a unique identifier with the primary-language term in at least one equivalency of the list, if at least one of the plurality of secondary-language terms has a limited meaning that is associated with the unique identifier;

adding at least one of the secondary-language terms to at least one equivalency in the list, if the primary-language term is identical to the at least one secondary-language term, wherein the at least one secondary-language term is one of the plurality of secondary-language terms;

adding a primary-language lead-in term to at least one equivalency in the list, if at least one of the plurality of secondary-language terms is synonymous with the primary-language lead-in term;

designating one of the plurality of secondary-language terms as a primary translation based on the primary-language term; and

adding a Boolean expression to at least one equivalency in the list, if the least one of the plurality of secondary-language terms is associated with a combination of terms in the primary language.

21. The method of claim 20 , wherein the subset of equivalencies is received from an interface that enables, a user to designate associations between the plurality of secondary-language terms and the primary-language term.

22. The method of claim 20 , further comprising providing the list of equivalencies to a search engine that identifies digital content based on the list and at least one of the following; structured text metadata and corresponding unstructured free-text metadata.

23. A system for generating a list for identifying digital content with a client computer, comprising:

a parser that is in communication with the client computer, the parser is arranged to perform a plurality of operations, including:

receiving a subset of equivalencies comprising a plurality of secondary-language terms that are associated with a primary-language term; and

parsing the subset into a list of equivalencies based on each equivalency comprising an association of at least one of the plurality of secondary-language terms with the primary-language term; and

a list generator that is in communication with the client computer and the parser, the list generator is arranged to perform a plurality of operations, including:

associating a unique identifier with the primary-language term in at least one equivalency of the list, if at least one of the plurality of secondary-language terms has a limited meaning that is associated with the unique identifier;

adding a nonprimary-language term to at least one equivalency in the list, if the primary-language term is identical to the nonprimary-language term, wherein the nonprimary-language term is one of the plurality of secondary-language terms;

adding a primary-language lead-in term to at least one equivalency in the list, if at least one of the plurality of secondary-language terms is synonymous with the primary-language lead-in term;

designating one of the plurality of secondary-language terms as a primary translation based on the primary-language term; and

adding a Boolean expression to at least one equivalency in the list, if the least one of the plurality of secondary-language terms is associated with a combination of terms in the primary language.

24. The system of claim 23 wherein the list generator further performs the operation of providing the list of equivalencies to a search engine that identifies digital content based on the list and at least one of the following: structured text metadata and corresponding unstructured free-text metadata.

25. A method for determining a query to identify digital content with a client computer, the method enabling operations, comprising:

receiving a first equivalency between:

a primary-language query term in a primary language; and

a user-specified secondary-language query term in a secondary language;

receiving a second equivalency between the primary-language query term and an alternate secondary-language query term in the secondary language;

determining whether to apply a unique identifier to either of the user-specified secondary-language query term or the alternate secondary-language query term with a translations generator in communication with the client computer, wherein the unique identifier refines the meaning of a query term and indicates a structured query term;

designating a primary translation as one of the user-specified secondary-language query term and the alternate secondary-language query term based on the primary-language query term;

receiving a search query in the secondary language; and

determining the primary-language query term with a translating machine in communication with the client computer and the translations generator, wherein the determination is based at least in part on the search query, the user-specified secondary-language query term, and the alternate secondary-language query term.

26. The method of claim 25 , wherein the primary language is English and the secondary language is one of a plurality of languages other than English.

27. The method of claim 25 , wherein the alternate secondary language query term comprises at least one of the following; a dictionary entry, a user-defined Boolean expression, and a synonym.

28. The method of claim 25 , further comprising providing the primary-language query term to a search engine for searching at least one of the following:

structured text metadata, which comprises a controlled vocabulary of keywords associated with content; and

corresponding unstructured free-text metadata, which comprises categories of text that need not conform to a controlled vocabulary.

29. The method of claim 25 , further comprising removing a noise word from at least one of the following; the user-specified secondary-language query term and the alternate secondary-language query term.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 23, 2023
From: BRANDED ENTERTAINMENT NETWORK, INC.
To: BEN GROUP, INC.
Reel/Frame 062456/0176 →
CHANGE OF NAME Recorded Jun 13, 2017
From: CORBIS CORPORATION
To: BRANDED ENTERTAINMENT NETWORK, INC.
Reel/Frame 043265/0819 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 17, 2007
From: SUMMERLIN, JOEL; FUNNELL, JARETT; UHLIG, HEIKE; YERIGAN, WAYNE
To: CORBIS CORPORATION
Reel/Frame 019173/0732 →
Continuity (2)
Provisional Application 60886649 · Jan 25, 2007
Related Publication 20080275691A1 · Nov 6, 2008