Content hiding software identification and/or extraction system and method
An exemplary system and method facilitate the identify and/or extract content hiding software, e.g., in a software curation environment (e.g., Apple's App Store). In some embodiments, the exemplary system and method may be applied to U.S.-based platforms as well as international platforms in Russia, India, China, among others.
1. A method to identify content hiding software, the method comprising:
pre-identifying content hiding software in curation platforms, wherein pre-identifying comprises:
scanning titles and/or subtitles in a curation platform using a first set of keywords;
returning title, subtitle, and bundle identifiers of a plurality of potential content hiding software;
classifying, via one or more trained machine learning classifiers, a software from the plurality of potential content hiding software as content hiding software or non-content hiding software, wherein the one or more classifiers have been trained using a feature set comprising a second set of keywords associated with a title, subtitle, and/or bundle identifiers of software, wherein the first set of key words is smaller than the second set of keywords; and
storing the classified content hiding software in a database as pre-identified content hiding software;
acquiring, from a user device or a remote computing device operatively coupled to the user device, configuration data of the user device;
parsing the configuration data into a list of installed software and respective bundle identifiers; and
identifying installed software on the user device as a content hiding software by (i) comparing the list of installed software to pre-identified content hiding software in the database and (ii) extracting hidden information from artifacts of stored data of the installed software on the user device.
2. The method of claim 1 , wherein the extracting of the hidden information comprises:
iterating, by a processor, without user input or intervention, through an installed software directory for a classified content hiding software and extract relevant artifacts.
3. The method of claim 1 , wherein the one or more classifiers include at least one of Gaussian Naive Bayes (GNB), Support Vector Machine (SVM) and Decision Tree (DT).
4. The method of claim 1 , wherein the one or more classifiers are trained using the second set of keywords that includes at least one of: ‘private’, ‘sensitive’, ‘censor’, ‘protect’, ‘decoy’, ‘privacy’, ‘secret’, ‘hide’, ‘vault’, ‘secure’, ‘password protected’, ‘browser’, ‘private browser’.
5. The method of claim 1 , wherein the curation platform is for mobile computing device or non-mobile computing devices.
6. A system comprising:
a processor; and
a memory operatively coupled to the processor, the memory having instructions stored thereon, wherein execution of the instructions by the processor, causes the processor to:
scan for titles and subtitles in a curation platform using a first set of keywords;
return title, subtitle, and bundle identifiers of a plurality of potential content hiding software;
classify, via one or more classifiers, a software from the plurality of potential content hiding software as content hiding software or non-content hiding software, wherein the one or more classifiers have been trained using a feature set comprising a second set of keywords associated with a title, subtitle, and bundle identifiers of software in published fields for software in the curation platform, wherein the first set of keywords is smaller than the second set of keywords;
store the classified content hiding software in a database as pre-identified content hiding software;
acquire, from a user device or a remote computing device operatively coupled to the user device, configuration data of the user device;
parse the configuration data into a list of installed software and respective bundle identifiers; and
identify installed software on the user device as a content hiding software by (i) comparing the list of installed software to pre-identified content hiding software in the database and (ii) extracting hidden information from artifacts of stored data of the installed software on the user device.
7. The system of claim 6 , wherein the instructions to extract the hidden information comprises:
instructions to iterate, without user input or intervention, through an installed software directory for a classified content hiding software and extract relevant artifacts.
8. The system of claim 6 , wherein the one or more classifiers include at least one of Gaussian Naive Bayes (GNB), Support Vector Machine (SVM) and Decision Tree (DT).
9. The system of claim 6 , wherein the one or more classifiers are trained using the second set of keywords that includes at least one of: ‘private’, ‘sensitive’, ‘censor’, ‘protect’, ‘decoy’, ‘privacy’, ‘secret’, ‘hide’, ‘vault’, ‘secure’, ‘safe’, ‘photos’, ‘videos’, ‘notes’, ‘passwords’, ‘contacts’, ‘password-protected’, ‘password protected’, ‘browser’, ‘private browser’.
10. The system of claim 6 , wherein the curation platform is for mobile computing device or non-mobile computing devices.
11. A non-transitory computer-readable medium having instructions stored thereon, wherein execution of the instructions by a processor, causes the processor to:
scan for titles and subtitles in a curation platform using a first set of keywords;
return title, subtitle, and bundle identifiers of a plurality of potential content hiding software;
classify, via one or more classifiers, a software from the plurality of potential content hiding software as content hiding software or non-content hiding software, wherein the one or more classifiers have been trained using a feature set comprising a second set of keywords associated title, subtitle, and bundle identifiers of software in published fields for software in the curation platform, wherein the first set of keywords is smaller than the second set of keywords;
store the classified content hiding software in a database as pre-identified content hiding software;
acquire, from a user device or a remote computing device operatively coupled to the user device, configuration data of the user device;
parse the configuration data into a list of installed software and respective bundle identifiers; and
identify installed software on the user device as a content hiding software by (i) comparing the list of installed software to pre-identified content hiding software in the database and (ii) extracting hidden information from artifacts of stored data of the installed software on the user device.
12. The computer-readable medium of claim 11 , wherein the instructions to extract the hidden information comprises:
instructions to iterate, without user input or intervention, through an installed software directory for a classified content hiding software and extract relevant artifacts.
13. The computer-readable medium of claim 11 , wherein the one or more classifiers include at least one of Gaussian Naive Bayes (GNB), Support Vector Machine (SVM) and Decision Tree (DT).
14. The computer-readable medium of claim 11 , wherein the one or more classifiers are trained using the second set of keywords that includes at least one of: ‘private’, ‘sensitive’, ‘censor’, ‘protect’, ‘decoy’, ‘privacy’, ‘secret’, ‘hide’, ‘vault’, ‘secure’, ‘safe’, ‘photos’, ‘videos’, ‘notes’, ‘passwords’, ‘contacts’, ‘password-protected’, ‘password protected’, ‘browser’, ‘private browser’.
15. The method of claim 1 , wherein the stored data comprises plist files, json files, and solite database files for one type of software application, and wherein the extracted hidden information comprises decoy data, spying data, dual mode data, supports encryption data, and/or password storage data for a second type of software application.
16. The system of claim 6 , wherein the stored data comprises plist files, json files, and sqlite database files for one type of software application, and wherein the extracted hidden information comprises decoy data, spying data, dual mode data, supports encryption data, and/or password storage data for a second type of software application.
17. The computer-readable medium of claim 11 , wherein the stored data comprises plist files, json files, and sqlite database files for one type of software application, and wherein the extracted hidden information comprises decoy data, spying data, dual mode data, supports encryption data, and/or password storage data for a second type of software application.