IP Library › Granted Patent US 12,693,994
Granted Patent B2
US 12,693,994 · App. 18/369,351 · Granted Jul 28, 2026

Method for reducing false-positives for identification of digital content

Inventors: William Johnston Buchanan (Edinburgh Lothian, GB); Owen Chin Wai Lo (Edinburgh Lothian, GB); Philip Penrose (Keith Aberdeen, GB); Richard Macfarlane (Edinburgh Lothian, GB); Ian Stevenson (Edinburgh Lothian, GB); Bruce Ramsay (Edinburgh Lothian, GB)
Assignee: CYACOMB LIMITED
G06F16/152G06F21/10G06F21/56G06F21/564G06F21/6227G06F21/6272G06F21/64
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,693,994
App. No.
18/369,351
Filed
Sep 18, 2023
Granted
Jul 28, 2026
Kind
B2
Examiner
ZHU, ZHIMEI
Art Unit
2495
USPC
726/26
Abstract

Many areas of investigation require searching through data that may be of interest. In a first method step, a digital content element is provided. The digital content element may have any suitable format or data structure of interest to a searching entity. The digital content element may be a particular data file that is of interest to a searching entity. In a second step, the digital content element is compared with a first set of data provided by a combination of a second set of data and a third set of data. The first set of data is a collection of known digital content elements that are of interest to a searching entity, for example contraband digital content elements or digital content elements owned by or represented by the searching entity. In a third method step, the digital content element is identified as known if the digital content element is detected within the first set of data.

Claims (51)

1 . A method for identifying at least one digital content element, the digital content element forming a part of a set of digital content, the method comprising:

providing the digital content element;

comparing the digital content element with a first set of data provided by a combination of a second set of data and a third set of data, wherein the second set of data represents known data of interest to a searching entity and wherein the third set of data represents known non-identifying digital content elements including one or more non-unique digital content elements relating to a file structure or meta data;

if the digital content element is detected within the first set of data, then identifying the digital content element as digital content of interest; and

adjusting the second set of data or the third set of data to include one or more additional data elements based on determining whether the one or more additional data elements are not already included in either of the second set of data and the third set of data.

2 . The method according to claim 1 , wherein the step of comparing comprises:

comparing the digital content element with the second set of data; and

if the digital content element is detected within the second set of data, then comparing the digital content element with the third set of data, and

wherein the step of identifying comprises:

if the digital content element is detected within the second set of data and if the digital content element is not detected within the third set of data, then identifying the digital content element as digital content of interest.

3 . The method according to claim 1 , wherein the step of comparing comprises:

comparing the digital content element with the third set of data; and

if the digital content element is not detected within the third set of data, then comparing the digital content element with the second set of data, and wherein the step of identifying comprises:

if the digital content element is detected within the second set of data and if the digital content is not detected within the third set of data, then identifying the digital content as digital content of interest.

4 . The method according to claim 1 , wherein the step of comparing comprises:

creating the first set of data by subtracting the third set of data from the second set of data; and

comparing the digital content element with the first set of data, and wherein the step of identifying comprises:

if the digital content element is detected within the first set of data, then identifying the digital content element as digital content of interest.

5 . The method according to claim 1 , wherein creating the first set of data comprises:

comparing each element in the second set of data with each element in the third set of data; and

if an element of the second set of data is not detected within the third set of data, then adding the element to the first set of data.

6 . The method according to claim 1 , wherein the set of digital content comprises:

at least one data file, and wherein the digital content element is a fragment of the data file.

7 . The method according to claim 1 , wherein the digital content element is defined in the structure of the set of digital content.

8 . The method according to claim 1 , wherein the digital content element is a block.

9 . The method according to claim 8 , wherein the block corresponds to a network packet or a payload portion of a network packet.

10 . The method according to claim 8 , wherein the block corresponds to one of: a memory block; a disk storage; a disk storage sector; or a block comprising at least one data file.

11 . The method according to claim 8 , wherein the block has a fixed size.

12 . The method according to claim 1 , wherein the digital content element has been encoded by the way of one of: a hashing function; or a locality-sensitive hashing function.

13 . The method according to claim 1 , wherein at least one of the second set of data or the third set of data have been encoded by way of a hashing function.

14 . The method according to claim 1 , wherein the second set of data and the third set of data is one of: a cuckoo filter; or a bloom filter.

15 . The method according to claim 1 , further comprising: receiving a fourth set of data identified as known, wherein the fourth set of data comprises misidentified digital content elements;

comparing the fourth set of data with the second set of data; and

if a misidentified digital content element is detected within the second set of data, adding the misidentified digital content element to the third set of data.

16 . The method according to claim 1 , wherein at least one of the second set of data and the third set of data comprises a plurality of respective subsets of data.

17 . The method of claim 1 , wherein the one or more additional data elements comprise at least one set of population data comprising a plurality of population data elements, and adjusting the second set of data or the third set of data comprises:

comparing each population data element with the third set of data;

if a population data element is not detected within the third set of data, then compare the population data element with the second set of data;

if a population data element is not detected within the second set of data, then adding the population data element of the second set of data;

and if a population data element is detected within the second set of data, then adding the population data element to the third set of data.

18 . The method according to claim 17 , wherein the step of providing comprises:

comparing the at least one set of population data with a population database, the population database comprising at least one known sets of population data;

if a set of population data is not detected in the population database, then adding the set of population data to the population database;

and if a set of population data is detected in the population database, then ignoring the set of population data.

19 . The method according to claim 18 , wherein the population database comprises at least one representation of at least one known sets of population data, and wherein the step of comparing comprises:

providing a representation of each of the at least one set of population data; and

comparing the representation of each of the at least one set of population data with each of the at least one representation of the at least one known sets of population data.

20 . The method of claim 1 , wherein the one or more additional data elements comprise at least one set of population data comprising a plurality of population data elements, and adjusting the second set of data or the third set of data comprises:

comparing each population data element with the second set of data;

if a population data element is not detected within the second set of data, then adding the population data element to the second set of data; and

if a population data element is detected within the second set of data, then adding the population data element to the third set of data.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 6, 2026
From: BUCHANAN, WILLIAM JOHNSTON; LO, OWEN CHIN WAI; PENROSE, PHILIP; MACFARLANE, RICHARD; STEVENSON, IAN; RAMSAY, BRUCE
To: CYAN FORENSICS LIMITED
Reel/Frame 074277/0244 →
CHANGE OF NAME Recorded Apr 6, 2026
From: CYAN FORENSICS LIMITED
To: CYACOMB LIMITED
Reel/Frame 074285/0407 →
Priority Claims (1)
GB 1705334 · Apr 3, 2017 · national
Continuity (2)
Continuation 16500736 · Mar 12, 2018
Related Publication 20240004964A1 · Jan 4, 2024
References Cited (48)
US 7640589B1 · Mashevsky · 2009 [cited by examiner]
US 7984304B1 · Waldspurger · 2011 [cited by examiner]
US 8375450B1 · Oliver · 2013 [cited by examiner]
US 8561180B1 · Nachenberg · 2013 [cited by examiner]
US 8621625B1 · Bogorad · 2013 [cited by examiner]
US 8826444B1 · Kalle · 2014 [cited by examiner]
US 8990944B1 · Singh · 2015 [cited by examiner]
US 9858413B1 · Zuo · 2018 [cited by examiner]
US 10032023B1 · Kuo · 2018 [cited by examiner]
US 10193902B1 · Caspi · 2019 [cited by examiner]
US 20020082837A1 · Pitman et al. · 2002 [cited by applicant]
US 20020116195A1 · Pitman et al. · 2002 [cited by applicant]
US 20060206935A1 · Choi · 2006 [cited by examiner]
US 20070083757A1 · Nakano · 2007 [cited by examiner]
US 20070226799A1 · Gopalan · 2007 [cited by examiner]
US 20070226801A1 · Gopalan · 2007 [cited by examiner]
US 20070226802A1 · Gopalan · 2007 [cited by examiner]
US 20070240222A1 · Tuvell et al. · 2007 [cited by applicant]
US 20080040807A1 · Lu et al. · 2008 [cited by applicant]
US 20080104186A1 · Wieneke · 2008 [cited by examiner]
US 20080168558A1 · Kratzer · 2008 [cited by examiner]
US 20080256647A1 · Kim et al. · 2008 [cited by applicant]
US 20100030722A1 · Goodson · 2010 [cited by examiner]
US 20100082811A1 · Van Der Merwe · 2010 [cited by examiner]
US 20110083176A1 · Martynenko · 2011 [cited by examiner]
US 20110126286A1 · Nazarov · 2011 [cited by examiner]
US 20110173698A1 · Polyakov · 2011 [cited by examiner]
US 20120310994A1 · Wionzek · 2012 [cited by examiner]
US 20120324579A1 · Jarrett · 2012 [cited by examiner]
US 20130139265A1 · Romanenko · 2013 [cited by examiner]
US 20130179995A1 · Basile et al. · 2013 [cited by applicant]
US 20140223566A1 · Zaitsev · 2014 [cited by examiner]
US 20140280155A1 · Elliot · 2014 [cited by examiner]
US 20160164901A1 · Mainieri · 2016 [cited by examiner]
US 20170180394A1 · Crofton · 2017 [cited by examiner]
US 20180063146A1 · Nakata · 2018 [cited by examiner]
US 20210294878A1 · Buchanan et al. · 2021 [cited by applicant]
GB 2483246A · 2012 [cited by applicant]
JP 2003167970A · 2003 [cited by applicant]
WO 2007087141A1 · 2007 [cited by applicant]
WO WO2007117582A2 · 2007 [cited by examiner]
Neil Youngman, “Fine-tuning SpamAssassin”, Aug. 2004, obtained online from <https://linuxgazette.net/105/youngman.html>, retrieved on May 3, 2025. (Year: 2004). [cited by examiner]
“Appending an id to a list if not already present in the list”, obtained online from <https://stackoverflow.com/questions/17370984/appending-an-id-to-a-list-if-not-already-present-in-the-list>, retrieved on Nov. 10, 202… [cited by examiner]
V. Roussev, “Hashing and Data Fingerprinting in Digital Forensics,” in IEEE Security & Privacy, vol. 7, No. 2, pp. 49-55, Mar.-Apr. 2009 (Year: 2009). [cited by examiner]
Deepanshu Bhalla, “SAS SQL: Comparing Two Tables”, 2015, Obtained online from <https://web.archive.org/web/20151208233324/https://www.listendata.com/2015/04/sas-sqi-comparing-two-tables.html>, retreived on Apr. 27, 2023… [cited by applicant]
Great Britain Search Report issued in counterpart GB Application No. GB1705334.9 dated Oct. 2, 2017 (two (2) pages). [cited by applicant]
International Search Report and Written Opinion for PCT/GB2018/050617; International Filing Date Mar. 12, 2018; Date of Mailing: Aug. 4, 2018; 10 pages. [cited by applicant]
Frei, Stefan, “IP Address & Range Calculator”; obtained on line from: https://techzoom.net/lab/ip-address-calculator/; 3 pages; retrieved on Dec. 28, 2021. [cited by applicant]