IP Library Granted Patent US 10,460,041
Granted Patent B2
US 10,460,041 · App. 15/403,047 · Granted Oct 29, 2019

Efficient string search

Inventors: Thomas E. Raffill (Sunnyvale, CA); Shunhui Zhu (San Jose, CA); Roman Yanovsky (Los Altos, CA); Boris Yanovsky (Saratoga, CA); John Gmuender (San Jose, CA)
Assignee: SONICWALL INC.
G06F17/2863G06F16/90344G06F17/2705G06F17/28
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,460,041
App. No.
15/403,047
Granted
Oct 29, 2019
Kind
B2
Abstract

Some embodiments of an efficient string search have been presented. In one embodiment, a string of bytes representing content written in a non-delimited language is received, wherein the content has been classified into a predetermined category. In a single pass through the string of bytes, a set of N-grams is searched for simultaneously. Statistical information on occurrences of the N-grams, if any, in the string of bytes is collected. In some embodiments, a model is generated based on the statistical information, where the model is usable by a content filter to classify content.

Claims (45)

1. A method for filtering received content, the method comprising:

generating, by a hardware processor executing instructions out of a memory, a model that corresponds to a digital content file that is pre-classified under a first classification, the generated model including a number of conditions associated with the first classification that are related to statistical data regarding a set of keywords found in the digital content file;

storing the model in a model repository in the memory;

receiving a string of bytes over a data communication interface that is communicatively coupled to a communication network;

performing a string search on the received string of bytes, wherein the string search generates statistical information indicating whether the received string of bytes includes any keywords from the set of keywords associated with at least the model corresponding to the first classification;

comparing the generated statistical information with the model corresponding to the first classification;

identifying that at least a portion of the received string of bytes include a number of the keywords from the set of keywords associated with the model corresponding to the first classification;

identifying whether the number of conditions associated with the first classification are satisfied based on the number of keywords identified in the received string of bytes that are from the set of keywords associated with the model; and

processing the portion of the received string of bytes based on whether the number of conditions are identified as being satisfied, wherein the portion is allowed to be provided to a user device when the number of keywords does not satisfy the conditions associated with the first classification, and wherein the portion is not allowed to be provided to the user device when the number of keywords does satisfy the conditions associated with the first classification.

2. The method of claim 1 , wherein the portion of the received string of bytes includes text in a non-delimited language.

3. The method of claim 2 , wherein the non-delimited language is selected from the group consisting of Chinese, Japanese, and Thai.

4. The method of claim 1 , wherein each keyword corresponds to one or more N-grams.

5. The method of claim 4 , further comprising identifying a plurality of states associated with each of the one or more N-grams.

6. The method of claim 1 , wherein the number of keywords associated with the first classification are part of one or more other models for classifying the digital content.

7. The method of claim 1 , further comprising generating one or more other models for classifying the digital content.

8. A non-transitory computer-readable storage medium having embodied thereon a program executable by a hardware processor for performing a method for filtering received content, the method comprising:

generating, by the hardware processor, a model that corresponds to a digital content file that is pre-classified under a first classification, the generated model including a number of conditions associated with the first classification that are related to statistical data regarding a set of keywords found in the digital content file;

storing the model in a model repository in memory;

receiving a string of bytes over a data communication interface that is communicatively coupled to a communication network;

performing a string search on the received string of bytes, wherein the string search generates statistical information indicating whether the received string of bytes includes any keywords from the set of keywords associated with the model corresponding to the first classification;

comparing the generated statistical information with the model corresponding to the first classification;

identifying that at least a portion of the received string of bytes include a number of the keywords from the set of keywords associated with the model corresponding to the first classification;

identifying whether the number of conditions associated with the first classification are satisfied based on the number of keywords identified in the received string of bytes that are from the set of keywords associated with the model; and

processing the portion of the received string of bytes based on whether the number of conditions are identified as being satisfied, wherein the portion is allowed to be provided to a user device when the number of keywords does not satisfy the conditions associated with the first classification, and wherein the portion is not allowed to be provided to the user device when the number of keywords does satisfy the conditions associated with the first classification.

9. The non-transitory computer-readable storage medium of claim 8 , wherein the portion of the received string of bytes includes text in a non-delimited language.

10. The non-transitory computer-readable storage medium of claim 9 , wherein the non-delimited language is selected from the group consisting of Chinese, Japanese, and Thai.

11. The non-transitory computer-readable storage medium of claim 8 , wherein each keyword corresponds to one or more N-grams.

12. The non-transitory computer-readable storage medium of claim 11 , wherein the program further comprises instructions executable to identify a plurality of states associated with each of the one or more N-grams.

13. The non-transitory computer-readable storage medium of claim 8 , wherein the number of keywords associated with the first classification are part of one or more other models for classifying the digital content.

14. The non-transitory computer-readable storage medium of claim 8 , wherein the program further comprises instructions executable to generate one or more other models for classifying the digital content.

15. An apparatus for performing a method for filtering received content, the apparatus comprising:

a memory;

a model repository that stores a model that corresponds to a digital content file that is pre-classified under a first classification, the model including a number of conditions associated with the first classification that are related to statistical data regarding a set of keywords found in the digital content file;

a communication interface that receives a string of bytes, the communication interface communicatively coupled to a communication network; and

a hardware processor that executes instructions stored in the memory, wherein execution of the instructions by the processor:

performs a string search on the received string of bytes, wherein the string search generates statistical information indicating whether the received string of bytes includes any keywords from the set of keywords associated with the model corresponding to the first classification;

compares the generated statistical information with the model corresponding to the first classification;

identifies that at least a portion of the received string of bytes include a number of the keywords from the set of keywords associated with the model corresponding to the first classification;

identifies whether the number of conditions associated with the first classification are satisfied based on the number of keywords identified in the received string of bytes as being from the set of keywords associated with the model; and

identifies that the portion of the received string of bytes corresponds to content that is allowed to be provided to a user device based on identifying that the number of keywords does not satisfy the number of conditions associated with the first classification, wherein the portion of the received string of bytes is identified as content that is not allowed to be provided to the user device based on identifying that the number of keywords does satisfy the conditions associated with the first classification.

16. The apparatus of claim 15 , wherein the portion of the received string of bytes includes text in a non-delimited language.

17. The apparatus of claim 16 , wherein the non-delimited language is selected from the group consisting of Chinese, Japanese, and Thai.

18. The apparatus of claim 15 , wherein each keyword corresponds to one or more N-grams.

19. The apparatus of claim 15 , wherein the number of keywords associated with the first classification are part of one or more other models for classifying the digital content.

20. The apparatus of claim 15 , wherein the processor executes further instructions to generate one or more other models for classifying the digital content.

Assignments (11)
FIRST LIEN IP SUPPLEMENT Recorded Jun 30, 2025
From: SONICWALL US HOLDINGS INC.
To: UBS AG, STAMFORD BRANCH, AS COLLATERAL AGENT
Reel/Frame 071777/0641 →
RELEASE OF SECOND LIEN SECURITY INTEREST IN PATENTS RECORDED AT RF 046321/0393 Recorded Jun 16, 2025
From: UBS AG, STAMFORD BRANCH, AS COLLATERAL AGENT
To: SONICWALL US HOLDINGS INC.
Reel/Frame 071625/0887 →
SECOND LIEN PATENT SECURITY AGREEMENT Recorded Jun 7, 2018
From: SONICWALL US HOLDINGS INC.
To: UBS AG, STAMFORD BRANCH, AS COLLATERAL AGENT
Reel/Frame 046321/0393 →
FIRST LIEN PATENT SECURITY AGREEMENT Recorded Jun 7, 2018
From: SONICWALL US HOLDINGS INC.
To: UBS AG, STAMFORD BRANCH, AS COLLATERAL AGENT
Reel/Frame 046321/0414 →
CHANGE OF NAME Recorded Nov 28, 2017
From: SONICWALL, INC.
To: SONICWALL L.L.C.
Reel/Frame 044528/0853 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 28, 2017
From: RAFFILL, THOMAS E; ZHU, SHUNHUI; YANOVSKY, ROMAN; YANOVSKY, BORIS; GMUENDER, JOHN
To: SONICWALL, INC.
Reel/Frame 044241/0161 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 28, 2017
From: QUEST SOFTWARE INC.
To: SONICWALL US HOLDINGS INC.
Reel/Frame 044822/0223 →
CHANGE OF NAME Recorded Nov 28, 2017
From: DELL SOFTWARE INC.
To: QUEST SOFTWARE INC.
Reel/Frame 044528/0863 →
MERGER Recorded Nov 28, 2017
From: SONICWALL, INC.
To: PSM MERGER SUB (DELAWARE), INC.
Reel/Frame 044241/0192 →
CHANGE OF NAME Recorded Nov 28, 2017
From: PSM MERGER SUB (DELAWARE), INC.
To: SONICWALL, INC.
Reel/Frame 044241/0226 →
MERGER Recorded Nov 28, 2017
From: SONICWALL L.L.C.
To: DELL SOFTWARE INC.
Reel/Frame 044241/0315 →
Continuity (5)
Continuation 14326230 · Jul 8, 2014
Continuation 13973859 · Aug 22, 2013
Continuation 13335743 · Dec 22, 2011
Continuation 11881556 · Jul 27, 2007
Related Publication 20170147565A1 · May 25, 2017