IP Library › Granted Patent US 12,591,617
Granted Patent B1
US 12,591,617 · App. 19/012,240 · Granted Mar 31, 2026

Neural network retraining based on image classification feedback

Inventors: Huaijin Wang (Redwood City, CA); Alexander Jordan Bildner (San Francisco, CA); Amey Patel (Mountain View, CA); Andrew Ellison (Newark, CA); John Goddard (Palo Alto, CA); Ryan Wong (Palo Alto, CA); Darin Tay (Palo Alto, CA); Reza Bosagh Zadeh (Palo Alto, CA)
Assignee: Matroid, Inc.
G06F16/73G06F16/75G06F40/279G06F40/40G06V10/764G06V10/7747G06V10/82G06V2201/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,591,617
App. No.
19/012,240
Granted
Mar 31, 2026
Kind
B1
Abstract

Described herein are systems and methods that highlight target objects in media content. The detection system trains a neural network to identify objects within an image. The detection system receives a text input requesting a search within the image and applies a search language model to the text input, which identifies a target object associated with the requested search. The detection system applies the neural network to the image to identify instances of the target object. The detection system modifies a user interface to include the image and modifies the image to highlight the identified instances of the target object. The detection system receives feedback that modifies the highlighted instances of the target object within the user interface and retrains the neural network based on the modified highlighted instances.

Claims (81)

1 . A method comprising:

training, by a detection system, a neural network to identify objects within a first image;

receiving, at the detection system, a text input requesting a search within the first image;

applying, by the detection system, an LLM to the received text input to identify a target object associated with the requested search;

applying, by the detection system, the neural network to the first image to identify instances of the target object;

modifying, by the detection system, a first portion of a user interface to include the first image;

modifying, by the detection system, the first image within the first portion of the user interface to highlight the identified instances of the target object;

receiving, by the detection system, text indicative of one or more modifications to the highlighted instances of the target object;

applying, by the detection system, a second LLM to determine modification criteria for the one or more modifications;

generating, by the detection system, a second image based on application of the modification criteria to the first image; and

retraining, by the detection system, the neural network based on the modified highlighted instances of the target object.

2 . The method of claim 1 , wherein receiving text indicative of one or more modifications from the user comprises one or more of:

receiving an indication of approval by the user of the modified image; and

receiving an indication of disapproval of one or more highlighted instances of the target object.

3 . The method of claim 1 , further comprising:

receiving an interaction with a slider bar presented at the user interface, wherein the interaction is associated with a confidence level for the neural network; and

in response to receiving the interaction:

applying the neural network to the first image to identify a new set of instances of the target object based on the confidence level; and

modifying the first image within the first portion of the user interface to highlight the new set of identified instances of the target object.

4 . The method of claim 1 , further comprising:

retraining the LLM based on the text indicative of one or more modifications from the user.

5 . The method of claim 1 , wherein modifying the first image within the first portion of the user interface to highlight the identified instances of the target object comprises:

overlaying portions of the first image corresponding to instances of the target object with circles or boxes.

6 . The method of claim 1 , wherein modifying the first image within the first portion of the user interface to highlight the identified instances of the target object comprises:

overlaying portions of the first image corresponding to instances of the target object with outlines of the instances of the target object.

7 . The method of claim 1 , wherein modifying the first image within the first portion of the user interface to highlight the identified instances of the target object comprises:

overlaying portions of the first image corresponding to instances of the target object with labels corresponding to the instances of the target object.

8 . The method of claim 7 , wherein the text indicative of one or more modifications comprises a textual edit of a label overlaid on the first image, the method further comprising:

tuning the LLM based on the textual edit.

9 . A non-transitory computer-readable storage medium storing instructions that, when executed, cause a processor to perform operations comprising:

training, by a detection system, a neural network to identify objects within a first image;

receiving, at the detection system, a text input requesting a search within the first image;

applying, by the detection system, an LLM to the received text input to identify a target object associated with the requested search;

applying, by the detection system, the neural network to the first image to identify instances of the target object;

modifying, by the detection system, a first portion of a user interface to include the first image;

modifying, by the detection system, the first image within the first portion of the user interface to highlight the identified instances of the target object;

receiving, by the detection system, text indicative of one or more modifications to the highlighted instances of the target object;

applying, by the detection system, a second LLM to determine modification criteria for the one or more modifications;

generating, by the detection system, a second image based on application of the modification criteria to the first image; and

retraining, by the detection system, the neural network based on the modified highlighted instances of the target object.

10 . The non-transitory computer-readable storage medium of claim 9 , wherein the operation of receiving text indicative of one or more modifications comprises one or more of:

receiving an indication of approval by the user of the modified image; and

receiving an indication of disapproval of one or more highlighted instances of the target object.

11 . The non-transitory computer-readable storage medium of claim 9 , the operations further comprising:

receiving an interaction with a slider bar presented at the user interface, wherein the interaction is associated with a confidence level for the neural network; and

in response to receiving the interaction:

applying the neural network to the first image to identify a new set of instances of the target object based on the confidence level; and

modifying the first image within the first portion of the user interface to highlight the new set of identified instances of the target object.

12 . The non-transitory computer-readable storage medium of claim 9 , the operations further comprising:

retraining the LLM based on the text indicative of one or more modifications from the user.

13 . The non-transitory computer-readable storage medium of claim 9 , wherein the operation of modifying the first image within the first portion of the user interface to highlight the identified instances of the target object comprises:

overlaying portions of the first image corresponding to instances of the target object with circles or boxes.

14 . The non-transitory computer-readable storage medium of claim 9 , wherein the operation of modifying the first image within the first portion of the user interface to highlight the identified instances of the target object comprises:

overlaying portions of the first image corresponding to instances of the target object with outlines of the instances of the target object.

15 . The non-transitory computer-readable storage medium of claim 9 , wherein the operation of modifying the first image within the first portion of the user interface to highlight the identified instances of the target object comprises:

overlaying portions of the first image corresponding to instances of the target object with labels corresponding to the instances of the target object.

16 . The non-transitory computer-readable storage medium of claim 15 , wherein the text indicative of one or more modifications comprises a textual edit of a label overlaid on the first image, the operations further comprising:

tuning the LLM based on the textual edit.

17 . A system comprising:

a processor; and

a non-transitory computer-readable storage medium storing instructions that, when executed, cause the processor to perform operations comprising:

training, by a detection system, a neural network to identify objects within a first image;

receiving, at the detection system, a text input requesting a search within the first image;

applying, by the detection system, an LLM to the received text input to identify a target object associated with the requested search;

applying, by the detection system, the neural network to the first image to identify instances of the target object;

modifying, by the detection system, a first portion of a user interface to include the first image;

modifying, by the detection system, the first image within the first portion of the user interface to highlight the identified instances of the target object;

receiving, by the detection system, text indicative of one or more modifications to the highlighted instances of the target object;

applying, by the detection system, a second LLM to determine modification criteria for the one or more modifications;

generating, by the detection system, a second image based on application of the modification criteria to the first image; and

retraining, by the detection system, the neural network based on the modified highlighted instances of the target object.

18 . The system of claim 17 , wherein the operation of receiving text indicative of one or more modifications from the user comprises one or more of:

receiving an indication of approval by the user of the modified image; and

receiving an indication of disapproval of one or more highlighted instances of the target object.

19 . The system of claim 17 , the operations further comprising:

receiving an interaction with a slider bar presented at the user interface, wherein the interaction is associated with a confidence level for the neural network; and

in response to receiving the interaction:

applying the neural network to the first image to identify a new set of instances of the target object based on the confidence level; and

modifying the first image within the first portion of the user interface to highlight the new set of identified instances of the target object.

20 . The system of claim 17 , the operations further comprising:

retraining the LLM based on the text indicative of one or more modifications from the user.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 17, 2025
From: WANG, HUAIJIN; BILDNER, ALEXANDER JORDAN; PATEL, AMEY; ELLISON, ANDREW; GODDARD, JOHN; WONG, RYAN; TAY, DARIN; ZADEH, REZA BOSAGH
To: MATROID, INC.
Reel/Frame 072930/0430 →
References Cited (7)
US 12198224B2 · Yuan · 2025 [cited by examiner]
US 20220261579A1 · Jindal · 2022 [cited by examiner]
US 20230206525A1 · Harikumar · 2023 [cited by examiner]
US 20240264718A1 · Benedetto · 2024 [cited by examiner]
US 20240282130A1 · Green · 2024 [cited by examiner]
US 20250094484A1 · Zheng · 2025 [cited by examiner]
Anastasios D. Doulamis et al., “Retrainable Neural Networks for Image Analysis and Classification”, IEEE, pp. 3558-3563 (Year: 1997). [cited by examiner]