IP Library › Granted Patent US 12,411,802
Granted Patent B2
US 12,411,802 · App. 17/533,491 · Granted Sep 9, 2025

Re-ordering files by keyword for migration

Inventors: Tohru Hasegawa (Tokyo, JP); Hiroshi Itagaki (Yokohama, JP); Tsuyoshi Miyamura (Yokohama, JP); Atsushi Abe (Ebina, JP); Shinsuke Mitsuma (Machida, JP); Noriko Yamamoto (Tokyo, JP)
Assignee: International Business Machines Corporation
G06F16/119G06F7/08G06F40/279G10L13/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,411,802
App. No.
17/533,491
Granted
Sep 9, 2025
Kind
B2
Abstract

A computer implemented method includes identifying a set of target files to be migrated from a primary storage to a secondary storage, extracting text data from the set of target files, identifying a set of keywords corresponding to the extracted text data from the set of target files, determining a number of keyword appearances in each file of the set of target files, assigning an order of migration corresponding to the set of target files such that the target files are written to the secondary storage in order of decreasing number of keywords, and migrating the files to the secondary storage according to the assigned order of migration. The method may additionally include writing the files with the greatest number of keywords closest to the default position of the tape reader. A computer program product and computer system corresponding to the method are also disclosed herein.

Claims (39)

1. A computer implemented method comprising:

identifying a set of target files to be migrated from a primary storage to a secondary storage;

applying a natural language processing program to the set of target files to extract text data from the set of target files;

identifying a set of keywords corresponding to the extracted text data from the set of target files;

determining a number of keyword appearances in each file of the set of target files;

assigning an order of migration corresponding to the set of target files, wherein the set of target files are ordered in descending order and grouped into a first group and a second group; and

migrating the files to the secondary storage according to the assigned order of migration by selecting from the first group or the second group based on a wrap number of a tape on which files are written being even or odd.

2. The computer implemented method of claim 1 , wherein determining a number of keyword appearances in each file of the set of target files includes determining a keyword prevalence with respect to the target files, wherein the keyword prevalence indicates a prevalence of the identified set of keywords with respect to the target file in its entirety.

3. The computer implemented method of claim 1 , wherein extracting text data from the set of target files includes leveraging text to speech techniques to extract text data from one or more portions of audio data from the set of target files.

4. The computer implemented method of claim 1 , wherein extracting text data from the set of target files includes leveraging content detection techniques to extract text data from one or more portions of non-audio, non-textual data from the set of target files.

5. The computer implemented method of claim 1 , wherein assigning an order of migration corresponding to the set of target files includes adjusting or altering an existing order of migration.

6. The computer implemented method of claim 1 , wherein migrating the files to the secondary storage according to the assigned order of migration occurs when a file that is written in the primary storage and that has not been accessed for a certain period of time has used up one whole tape.

7. A computer program product comprising:

one or more computer readable storage media and program instructions stored on the one or more computer readable storage media, the program instructions comprising instructions to:

identify a set of target files to be migrated from a primary storage to a secondary storage;

apply a natural language processing program to the set of target files to extract text data from the set of target files;

identify a set of keywords corresponding to the extracted text data from the set of target files;

determine a number of keyword appearances in each file of the set of target files;

assign an order of migration corresponding to the set of target files, wherein the set of target files are ordered in descending order and grouped into a first group and a second group; and

migrate the files to the secondary storage according to the assigned order of migration by selecting from the first group or the second group based on a wrap number of a tape on which files are written being even or odd.

8. The computer program product of claim 7 , wherein the program instructions to determine a number of keyword appearances in each file of the set of target files comprise instructions to determine a keyword prevalence with respect to the target files, wherein the keyword prevalence indicates a prevalence of the identified set of keywords with respect to the target file in its entirety.

9. The computer program product of claim 7 , wherein the program instructions to extract text data from the set of target files comprise instructions to leverage text to speech techniques to extract text data from one or more portions of audio data from the set of target files.

10. The computer program product of claim 7 , wherein the program instructions to extract text data from the set of target files comprise instructions to leverage content detection techniques to extract text data from one or more portions of non-audio, non-textual data from the set of target files.

11. The computer program product of claim 7 , wherein the program instructions to assign an order of migration corresponding to the set of target files comprise instructions to adjust or alter an existing order of migration.

12. The computer program product of claim 7 , wherein migrating the files to the secondary storage according to the assigned order of migration occurs when a file that is written in the primary storage and that has not been accessed for a certain period of time has used up one whole tape.

13. A computer system comprising:

one or more computer processors;

one or more computer-readable storage media;

program instructions stored on the computer-readable storage media for execution by at least one of the one or more processors, the program instructions comprising instructions to:

identify a set of target files to be migrated from a primary storage to a secondary storage;

apply a natural language processing program to the set of target files to extract text data from the set of target files;

identify a set of keywords corresponding to the extracted text data from the set of target files;

determine a number of keyword appearances in each file of the set of target files;

assign an order of migration corresponding to the set of target files, wherein the set of target files are ordered in descending order and grouped into a first group and a second group; and

migrate the files to the secondary storage according to the assigned order of migration by selecting from the first group or the second group based on a wrap number of a tape on which files are written being even or odd.

14. The computer system of claim 13 , wherein the program instructions to determine a number of keyword appearances in each file of the set of target files comprise instructions to determine a keyword prevalence with respect to the target files, wherein the keyword prevalence indicates a prevalence of the identified set of keywords with respect to the target file in its entirety.

15. The computer system of claim 13 , wherein the program instructions to extract text data from the set of target files comprise instructions to leverage text to speech techniques to extract text data from one or more portions of audio data from the set of target files.

16. The computer system of claim 13 , wherein the program instructions to extract text data from the set of target files comprise instructions to leverage content detection techniques to extract text data from one or more portions of non-audio, non-textual data from the set of target files.

17. The computer system of claim 13 , wherein the program instructions to assign an order of migration corresponding to the set of target files comprise instructions to adjust or alter an existing order of migration.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 23, 2021
From: HASEGAWA, TOHRU; ITAGAKI, HIROSHI; MIYAMURA, TSUYOSHI; ABE, ATSUSHI; MITSUMA, SHINSUKE; YAMAMOTO, NORIKO
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 058194/0405 →
Continuity (1)
Related Publication 20230161731A1 · May 25, 2023
References Cited (16)
US 10649697B2 · Hasegawa · 2020 [cited by applicant]
US 11100048B2 · Dain · 2021 [cited by applicant]
US 20060136525A1 · Akelbein · 2006 [cited by examiner]
US 20080307527A1 · Kaczmarski · 2008 [cited by examiner]
US 20120005193A1 · Nemoto · 2012 [cited by examiner]
US 20120101995A1 · Agetsuma · 2012 [cited by examiner]
US 20120323934A1 · Amir · 2012 [cited by examiner]
US 20150161161A1 · Iwanaga · 2015 [cited by examiner]
US 20150309929A1 · Ishii · 2015 [cited by examiner]
US 20150310107A1 · Alhakimi · 2015 [cited by examiner]
US 20160117259A1 · Hasegawa · 2016 [cited by examiner]
US 20180004760A1 · Bataller · 2018 [cited by examiner]
US 20190361622A1 · Hasegawa · 2019 [cited by applicant]
US 20210149590A1 · Miyamura · 2021 [cited by applicant]
US 20210165834A1 · Hasegawa · 2021 [cited by applicant]
JP 6242326 · 2017 [cited by applicant]