IP Library Granted Patent US 7,523,102
Granted Patent B2
US 7,523,102 · App. 11/149,453 · Granted Apr 21, 2009

Content search in complex language, such as Japanese

Assignee: Getty Images, Inc.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,523,102
App. No.
11/149,453
Granted
Apr 21, 2009
Kind
B2
Abstract

A search facility provides searching capabilities in languages such as Japanese. The facility may use a vocabulary knowledge base organized by concepts. For example, each concept may be associated with at least one keyword (as well as any synonyms or variant forms) by applying one or more rules that relate to identifying common main forms, script variants, alternative grammatical forms, phonetic variants, proper noun variants, numerical variants, scientific name, cultural relevance, etc. The contents of the vocabulary knowledge base are then used in executing search queries. A user may enter a search query in which keywords (or synonyms associated with those key words) may be identified, along with various stopwords that facilitate segmentation of the search query and other actions. Execution of the search query may result in a list of assets or similar indications being returned, which relate to concepts identified within the search query.

Claims (41)

1. A system for searching for content identified using text and symbols of a complex language associated with multiple written forms, the system comprising:

a computer-readable storage medium having stored thereon an asset repository associated with multiple searchable assets;

a computer-readable storage medium having stored thereon a vocabulary knowledge base for storing vocabulary information associated with the complex language, wherein the vocabulary knowledge base stores information related to multiple semantic concepts that are usable to identify assets within the asset repository, wherein the vocabulary knowledge base is generated or updated by a repeatable method comprising:

assigning an identifier to a semantic concept;

identifying a main written form for the semantic concept, wherein the main written form is based on at least one of the multiple written forms;

for at least one of the multiple written forms associated with the complex language, associating at least one synonymous written form with the semantic concept, wherein the synonymous written form is at least partially distinct from the main written form; and

storing the identifier, the main written form, and the at least one synonymous written form in a data storage component associated with the system; and

a computing system having a processor to execute a search engine for receiving and executing queries for the searchable assets, wherein the execution is based, at least in part, on the contents of the vocabulary knowledge base.

2. The system of claim 1 , further comprising, a classification component for classifying the searchable assets in the asset repository to facilitate matching between the searchable assets and the vocabulary information.

3. The system of claim 1 wherein the asset repository stores data associated with images.

4. The system of claim 1 wherein the asset repository stores data associated with documents.

5. The system of claim 1 wherein the vocabulary knowledge base is a subcomponent of a meta-data repository, and wherein the meta-data repository further comprises a concept-to-asset repository, for storing information associated with relationships between concepts from the vocabulary database and assets from the asset repository.

6. The system of claim 1 wherein the vocabulary knowledge base is structured hierarchically, such that concepts located farther down in the hierarchy provide more specificity than concepts located higher up in the hierarchy.

7. The system of claim 1 wherein the identified main written form is tagged as a keyword for searching.

8. The system of claim 1 wherein the language is Japanese, and wherein the multiple written forms include, kanji script, katakana script and hiragana script, okurigana variant, and romaji written form.

9. A computer-implemented method for executing a search query, the method comprising:

receiving a search query including a textual expression, wherein the textual expression is written in a language that at least occasionally lacks discrete boundaries between words or autonomous language units;

referencing a structured vocabulary knowledge base to determine whether the textual expression comprises a keyword or synonym associated with the structured vocabulary knowledge base, wherein the structured vocabulary knowledge base is for storing vocabulary information associated with a language having multiple orthographic forms or scripts, and wherien the structured vocabulary knowledge base is generated prior to receiving the search query by a repeatable method comprising:

assigning an identifier to a semantic concept that is usable as a keyword or key phrase;

identifying a main written form for the semantic concept, wherein the main written form is based on at least one of the multiple written forms;

for at least one of the multiple written forms associated with the complex language, associating at least one synonymous written form with the semantic concept, wherein the synonymous written form is at least partially distinct from the main written form; and

storing the identifier, the main written form, and the at least one synonymous written form in a data storage component associated with the vocabulary knowledge base;

if the textual expression does not comprise a keyword, key phrase or synonym associated with the structured vocabulary knowledge base, performing segmentation on the textual expression, wherein the segmentation includes systematically splitting the textual expression into two or more segments, and identifying at least one keyword from the vocabulary knowledge based on the textual expression and the two or more segments;

performing the search query using the at least one identified keyword; and

providing for display results of the search query.

10. The method of claim 9 wherein the systematic splitting includes identifying predetermined stopwords that are not keywords, wherein at least some of the predetermined stopwords are prepositions.

11. The method of claim 9 wherein the systematic splitting includes splitting the textual expression into the longest segments possible in an effort to preserve the intended meaning of the textual expression.

12. The method of claim 9 wherein the systematic splitting includes an application of both linguistic rules and contrived rules.

13. The method of claim 9 wherein the systematic splitting includes identifying predetermined stopwords that are not keywords, wherein at least some of the predetermined stopwords represent Boolean values used in executing the search query.

14. The method of claim 9 wherein the language is Japanese, and wherein determining whether the textual expression in its entirety comprises a keyword or synonym associated with the structured vocabulary knowledge base includes checking for matches in kanji groups with trailing hiragana.

15. The method of claim 9 wherein the language is Japanese, and wherein determining whether the textual expression in its entirety comprises a keyword or synonym associated with the structured vocabulary knowledge base includes checking for matches in character groups of only kanji, only hiragana, only katakana, and only non-kanji non-hiragana non-katakana characters.

16. The method of claim 9 wherein the systematic splitting includes segmenting a textual expression WXYZ as: WXY, XYZ, WX, XY, YZ, W, X, Y, Z, wherein W, X, Y, and Z each represent a character or autonomous grouping of characters associated with the language.

17. The method of claim 9 wherein the search query is received via a network connection.

18. A method in a first computer system for retrieving a media content unit from a second computer system having a plurality of media content units that have been classified according to keyword terms of a structured vocabulary, comprising:

sending a request for a media content unit, the request specifying a search term;

receiving an indication of at least one media content unit that corresponds to the specified search term, wherein the search term is located within the structured vocabulary and is used to determine at least one media content unit that corresponds to the search term, and wherein orthographic variations of the search term are automatically provided to assist in determining the at least one media content unit that corresponds to the search term, wherein the structured vocabulary is generated at the second computer system prior to the first computer sending the request by a repeatable method comprising:

assigning an identifier to a semantic concept that is usable as a keyword or key phrase;

identifying a main written form for the semantic concept, wherein the main written form is based on at least one of the multiple written forms;

for at least one of the multiple written forms associated with the complex language, associating at least one synonymous written form with the semantic concept, wherein the synonymous written form is at least partially distinct from the main written form; and

storing the identifier, the main written form, and the at least one synonymous written form in a data storage component associated with the vocabulary knowledge base; and displaying the at least one media content unit on a display.

19. The method of claim 18 wherein the keyword terms of the structured vocabulary are ordered such that the relationship of each term to each other term is inherent in the ordering.

Assignments (12)
NOTES SECURITY AGREEMENT Recorded May 6, 2025
From: GETTY IMAGES, INC.
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION, AS NOTES COLLATERAL AGENT
Reel/Frame 071183/0081 →
SECURITY AGREEMENT Recorded Feb 21, 2019
From: GETTY IMAGES, INC.; GETTY IMAGES (US), INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 048390/0709 →
RELEASE OF SECURITY INTEREST Recorded Feb 20, 2019
From: WILMINGTON TRUST, NATIONAL ASSOCIATION
To: GETTY IMAGES, INC.; GETTY IMAGES (US), INC.; GETTY IMAGES (SEATTLE), INC.
Reel/Frame 048384/0420 →
RELEASE OF SECURITY INTEREST Recorded Feb 20, 2019
From: JPMORGAN CHASE BANK, N.A.
To: GETTY IMAGES, INC.; GETTY IMAGES (US), INC.; GETTY IMAGES (SEATTLE), INC.
Reel/Frame 048384/0341 →
NOTICE OF SUCCESSION OF AGENCY Recorded Mar 23, 2018
From: BARCLAYS BANK PLC, AS PRIOR AGENT
To: JPMORGAN CHASE BANK, N.A., AS SUCCESSOR AGENT
Reel/Frame 045680/0532 →
NOTICE AND CONFIRMATION OF GRANT OF SECURITY INTEREST IN PATENTS Recorded Dec 14, 2015
From: GETTY IMAGES, INC.
To: WILMINGTON TRUST, NATIONAL ASSOCIATION
Reel/Frame 037289/0736 →
NOTICE AND CONFIRMATION OF GRANT OF SECURITY INTEREST IN PATENTS Recorded Oct 25, 2012
From: GETTY IMAGES, INC.
To: BARCLAYS BANK PLC
Reel/Frame 029190/0917 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS Recorded Oct 18, 2012
From: GENERAL ELECTRIC CAPITAL CORPORATION
To: GETTY IMAGES, INC.
Reel/Frame 029155/0882 →
SECURITY AGREEMENT Recorded Nov 23, 2010
From: GETTY IMAGES INC.
To: GENERAL ELECTRIC CAPITAL CORPORATION, AS COLLATERAL AGENT
Reel/Frame 025412/0499 →
RELEASE OF SECURITY INTEREST Recorded Nov 22, 2010
From: GENERAL ELECTRIC CAPITAL CORPORATION, AS COLLATERAL AGENT
To: GETTY IMAGES INC; GETTY IMAGES (US) INC.
Reel/Frame 025389/0825 →
SECURITY AGREEMENT Recorded Jul 8, 2008
From: GETTY IMAGES, INC.
To: GENERAL ELECTRIC CAPITAL CORPORATION
Reel/Frame 021204/0871 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 25, 2005
From: BJARNESTAM, ANNA; MACGUFFIE, MONIKA; CAROTHERS, DAVID
To: GETTY IMAGES, INC.
Reel/Frame 017153/0588 →
Continuity (4)
Provisional Application 6061074100 · Sep 17, 2004
Provisional Application 6058275900 · Jun 25, 2004
Provisional Application 6057913000 · Jun 12, 2004
Related Publication 20060031207A1 · Feb 9, 2006