IP Library Granted Patent US 10,268,674
Granted Patent B2
US 10,268,674 · App. 15/483,877 · Granted Apr 23, 2019

Linguistic intelligence using language validator

Inventors: Bhuvaneswari Tamilchelvan (Singapore, SG); Lakshmi Narasimhan M C (Bangalore, IN)
Assignee: Dell Products L.P.
G06F17/275G06F17/2725
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,268,674
App. No.
15/483,877
Granted
Apr 23, 2019
Kind
B2
Abstract

A language validation system includes a language validation module (LVM) and a linguistic intelligence engine (LIE). The LVM may perform language validation on text in a source document. The LVM may identify a language of origin for each word or sentence in the document. The LVM may determine whether a word/sentence not of the expected language constitutes a true defect. The LVM may provide input to an LIE database as the LVM identifies new words/sentences. The LIE may remain dormant for an initial interval until it obtains a sufficient vocabulary. Eventually, the LIE may be sufficiently knowledgeable to perform automated translations requests to translate original documents in the original language into one or more source documents in one or more source language. The LIE may adapt or otherwise implement a particular style in accordance with a style guide or analogous information.

Claims (47)

1. A language validation method, comprising:

automatically extracting, by an information handling system, text content from a plurality of source files into a single textual string (STS) document, wherein the text content in each of the plurality of source files comprises a translation, from an original language to a source language, of a corresponding original file;

parsing, by the information handling system, the STS document to identify each of a plurality of textual components in the STS document, wherein, said parsing comprises:

performing a first parsing of the STS document, wherein the first parsing comprises, parsing individual words in the STS into a first list document comprising a plurality of words; and

performing a second parsing of the STS document, wherein the second parsing comprises parsing whole sentences in the STS into a second list document comprising a plurality of sentences;

performing language validation operations for each of the plurality of textual components in the first and second list documents, wherein the plurality of textual components in the first list document correspond to the plurality of words and wherein the plurality of textual components in the second list document correspond to the plurality of sentences, and wherein the language validation operations include:

determining, by a language categorization utility, a language of origin of the textual component;

determining whether the language of origin of the textual component matches the source language;

responsive to the language of origin failing to match the source language, performing potential defect operations comprising:

passing the textual component to a defect manager comprising one or more localization rules;

determining, in accordance with the one or more localization rules, whether the textual component comprises a source file defect;

responsive to determining that the textual component comprises a source file defect, logging defect information corresponding to the source file defect to a validation output file; and

responsive to determining that the textual component does not comprise a source file defect, adding the textual component to a linguistic intelligence database, wherein textual components in the linguistic intelligence database are indexed according an original language equivalent of the textual component; and

wherein one or more of the source files is generated from an original file by performing source file generation operations, wherein the source file generation operations include:

identifying the plurality of textual components in the original file; and

for each textual component in the plurality of textual components in the original file:

indexing the linguistic intelligence database to identify a source language equivalent; and

replacing the textual component with the source language equivalent.

2. The method of claim 1 , wherein the defect information includes: the textual component, a page number identifying a location of the textual component, and a globally unique identifier of the source file defect; and a publication version is retrieved from the document.

3. The method of claim 1 , wherein determining in accordance with the localization rules comprises determining instances in which:

the language of origin is the original language; and

the source language lacks a native equivalent of the textual component.

4. The method of claim 1 , further comprising:

performing said source file generation operations is subject to determining that the linguistic intelligence database exceeds a threshold coverage of the source language.

5. An information handling system, comprising:

a processor;

a storage medium, accessible to the processor, including processor-executable program instructions that, when executed by the processor, cause the processor to perform operations comprising:

extracting text content from a plurality of source files into a single textual string (STS) document, wherein the text content in each of the plurality of source files comprises a translation, from an original language to a source language, of a corresponding original file;

parsing, by the information handling system, the STS document to identify each of a plurality of textual components in the STS document, wherein, said parsing comprises:

performing a first parsing of the STS document, wherein the first parsing comprises, parsing individual words in the STS into a first list document comprising a plurality of words; and

performing a second parsing of the STS document, wherein the second parsing comprises parsing whole sentences in the STS into a second list document comprising a plurality of sentences;

performing language validation operations for each of the plurality of textual components in the first and second list documents, wherein the plurality of textual components in the first list document correspond to the plurality of words and wherein the plurality of textual components in the second list document correspond to the plurality of sentences, and wherein the language validation operations include:

determining, by a language categorization utility, a language of origin of the textual component;

determining whether the language of origin of the textual component matches the source language;

responsive to the language of origin failing to match the source language, performing potential defect operations comprising:

passing the textual component to a defect manager comprising one or more localization rules;

determining, in accordance with the one or more localization rules, whether the textual component comprises a source file defect;

responsive to determining that the textual component comprises a source file defect, logging defect information corresponding to the source file defect to a validation output file; and

responsive to determining that the textual component does not comprise a source file defect, adding the textual component to a linguistic intelligence database, wherein textual components in the linguistic intelligence database are indexed according an original language equivalent of the textual component; and

wherein one or more of the source files is generated from an original file by performing source file generation operations, wherein the source file generation operations include:

identifying the plurality of textual components in the original file; and

for each textual component in the plurality of textual components in the original file:

indexing the linguistic intelligence database to identify a source language equivalent; and

replacing the textual component with the source language equivalent.

6. The information handling system of claim 5 , wherein the defect information includes: the textual component, a page number identifying a location of the textual component, and a globally unique identifier of the source file defect; and a publication version is retrieved from the document.

7. The information handling system of claim 5 , wherein determining in accordance with the one or more localization rules comprises determining instances wherein the language of origin is an acceptable substitute for the source language.

8. The information handling system of claim 5 , wherein performing said source file generation operations is subject to determining that the linguistic intelligence database exceeds a threshold coverage of the source language.

Assignments (10)
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (053546/0001) Recorded Jun 23, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL MARKETING L.P. (ON BEHALF OF ITSELF AND AS SUCCESSOR-IN-INTEREST TO CREDANT TECHNOLOGIES, INC.); DELL INTERNATIONAL L.L.C.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; DELL MARKETING CORPORATION (SUCCESSOR-IN-INTEREST TO FORCE10 NETWORKS, INC. AND WYSE TECHNOLOGY L.L.C.); EMC IP HOLDING COMPANY LLC
Reel/Frame 071642/0001 →
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (050724/0466) Recorded Jun 23, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC; DELL MARKETING CORPORATION (SUCCESSOR-IN-INTEREST TO WYSE TECHNOLOGY L.L.C.)
Reel/Frame 060753/0486 →
RELEASE OF SECURITY INTEREST IN PATENTS PREVIOUSLY RECORDED AT REEL/FRAME (042769/0001) Recorded Apr 26, 2022
From: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS NOTES COLLATERAL AGENT
To: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC (ON BEHALF OF ITSELF AND AS SUCCESSOR-IN-INTEREST TO MOZY, INC.); DELL MARKETING CORPORATION (SUCCESSOR-IN-INTEREST TO WYSE TECHNOLOGY L.L.C.)
Reel/Frame 059803/0802 →
RELEASE OF SECURITY INTEREST AT REEL 042768 FRAME 0585 Recorded Nov 2, 2021
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
To: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC; MOZY, INC.; WYSE TECHNOLOGY L.L.C.
Reel/Frame 058297/0536 →
SECURITY AGREEMENT Recorded Apr 22, 2020
From: CREDANT TECHNOLOGIES INC.; DELL INTERNATIONAL L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; FORCE10 NETWORKS, INC.; WYSE TECHNOLOGY L.L.C.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A.
Reel/Frame 053546/0001 →
PATENT SECURITY AGREEMENT (NOTES) Recorded Oct 15, 2019
From: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC; WYSE TECHNOLOGY L.L.C.
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS COLLATERAL AGENT
Reel/Frame 050724/0466 →
SECURITY AGREEMENT Recorded Mar 21, 2019
From: CREDANT TECHNOLOGIES, INC.; DELL INTERNATIONAL L.L.C.; DELL MARKETING L.P.; DELL PRODUCTS L.P.; DELL USA L.P.; EMC CORPORATION; FORCE10 NETWORKS, INC.; WYSE TECHNOLOGY L.L.C.; EMC IP HOLDING COMPANY LLC
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A.
Reel/Frame 049452/0223 →
PATENT SECURITY INTEREST (NOTES) Recorded Jun 12, 2017
From: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC; MOZY, INC.; WYSE TECHNOLOGY L.L.C.
To: THE BANK OF NEW YORK MELLON TRUST COMPANY, N.A., AS COLLATERAL AGENT
Reel/Frame 042769/0001 →
PATENT SECURITY INTEREST (CREDIT) Recorded Jun 12, 2017
From: DELL PRODUCTS L.P.; EMC CORPORATION; EMC IP HOLDING COMPANY LLC; MOZY, INC.; WYSE TECHNOLOGY L.L.C.
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
Reel/Frame 042768/0585 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 10, 2017
From: TAMILCHELVAN, BHUVANESWARI; M C, LAKSHMI NARASIMHAN
To: DELL PRODUCTS L.P.
Reel/Frame 041949/0885 →
Continuity (1)
Related Publication 20180293231A1 · Oct 11, 2018