IP Library › Granted Patent US 12,361,734
Granted Patent B2
US 12,361,734 · App. 18/162,077 · Granted Jul 15, 2025

Method for detecting image by semantic segmentation

Inventors: Hsiang-Chen Wang (Chiayi, TW); Kuan-Lin Chen (Chiayi, TW); Yu-Ming Tsao (Chiayi County, TW); Jen-Feng Hsu (Tainan, TW)
Assignee: National Chung Cheng University
G06V20/70G06V10/26G06V10/751G06V10/764G06V10/7715G06V10/82G06V2201/03
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,734
App. No.
18/162,077
Granted
Jul 15, 2025
Kind
B2
Abstract

The present application discloses a method for detecting image by using semantic segmentation. To input an image with data augmentation, and then encode and decode using a neural network. At least one semantically divided, and finally the at least one semantically divided is compared with the sample to classify as a target or a non-target. In this way, the CNN is used to detect whether the image is the SCC image or not, and locate the section, thereby assisting the doctor in interpreting the esophagus image.

Claims (23)

1. A method for detecting image by using semantic segmentation, comprising steps of:

an image extraction unit of a host extracting a first image;

said host performing data augmentation on said first image using a data augmentation function to generate a second image;

said host generating one or more semantic segmentation block according to a residual learning model of a neural network and an encoding-decoding method, and said encoding-decoding method comprising steps of:

generating a plurality of first pooling images from said second image using maximum pooling along a first contracting path, said maximum pooling performing dimension reduction on said second image for extracting a plurality of feature values, and the resolution of said second image is halved after pooling;

generating a plurality of second pooling images from said plurality of first pooling images using maximum pooling along a second contracting path, said maximum pooling performing dimension reduction on said plurality of first pooling images for extracting said plurality of feature values, and the resolution of said plurality of first pooling images is halved after pooling;

generating a plurality of third pooling images from said plurality of second pooling images using maximum pooling along a third contracting path, said maximum pooling performing dimension reduction on said plurality of second pooling images for extracting said plurality of feature values, and the resolution of said plurality of second pooling images is halved after pooling;

generating a plurality of fourth pooling images from said plurality of third pooling images using maximum pooling along a fourth contracting path, said maximum pooling performing dimension reduction on said plurality of third pooling images for extracting said plurality of feature values, and the resolution of said plurality of third pooling images is halved after pooling;

performing two or more layers of convolution calculations on said plurality of fourth pooling images along a first expansive path using a plurality of kernels after upsampling and concatenating said plurality of third pooling images for generating a plurality of first output images, said upsampling locating said plurality of feature values and doubling the resolution of said plurality of fourth pooling images, and said two or more layers of convolution calculations reducing the increased channel number after concatenation;

performing two or more layers of convolution calculations on said plurality of first output images along a second expansive path using a plurality of kernels after upsampling and concatenating said plurality of second pooling images for generating a plurality of second output images, and said upsampling locating said plurality of feature values and doubling the resolution of said plurality of first output images;

performing two or more layers of convolution calculations on said plurality of second output images along a third expansive path using a plurality of kernels after upsampling and concatenating said plurality of first pooling images for generating a plurality of third output images, and said upsampling locating said plurality of feature values and doubling the resolution of said plurality of second output images; and

performing two or more layers of convolution calculations on said plurality of third output images along a fourth expansive path using a plurality of kernels after upsampling and concatenating said plurality of second pooling images for generating a fourth output images, said fourth output image including said one or more semantic segmentation block, said upsampling locating said plurality of feature values and doubling the resolution of said plurality of third output images, and the resolution of said fourth output image equal to the resolution of said second image;

said host comparing said one or more semantic segmentation block with a sample image and producing a comparison result if the comparison matches; and

said host classifying said one or more semantic segmentation block as a target-object image according to said comparison result.

2. The method for detecting image of claim 1 , wherein said maximum pooling includes a plurality of kernels with 2×2 kernel size.

3. The method for detecting image of claim 1 , wherein said upsampling includes a plurality of deconvolution kernels with 2×2 kernel size.

4. The method for detecting image of claim 1 , wherein said data augmentation function is the function ImageDataGenerator in a Keras library.

5. The method for detecting image of claim 4 , wherein in said function ImageDataGenerator, a rotation_range is set to 60; a shear_range is set to 0.5; a fill_mode is set to ‘nearest’; and a validation_split is set to 0.1.

6. The method for detecting image of claim 1 , wherein said neural network is U-NET.

7. The method for detecting image of claim 1 , wherein said step of an image extraction unit of a host extracting a first image, said image extraction unit extracts said first image and adjusts said first image to a default size.

8. The method for detecting image of claim 1 , wherein said step of an image extraction unit of a host extracting a first image, said image extraction unit extracts said first image and examples of said first image include a white light image or a narrow-band image.

9. The method for detecting image of claim 1 , wherein said step of said host comparing said one or more semantic segmentation block with a sample image and producing a comparison result if the comparison matches, said host compares a plurality of feature values corresponding to each of said one or more semantic segmentation block with a plurality of feature values of said sample image; and if said feature values match, a comparison result is produced.

10. The method for detecting image of claim 1 , wherein said step of said host classifying said one or more semantic segmentation block as a target-object image according to said comparison result, when said host detects matches between said plurality of feature values corresponding to each of said one or more semantic segmentation block and said plurality of feature values of said sample image, said host classifies said one or more semantic segmentation block as said target-object image; otherwise, said host classifies said one or more semantic segmentation block as a non-target-object image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 1, 2023
From: WANG, HSIANG-CHEN; CHEN, KUAN-LIN; TSAO, YU-MING; HSU, JEN-FENG
To: NATIONAL CHUNG CHENG UNIVERSITY
Reel/Frame 062560/0056 →
Priority Claims (1)
TW 111108094 · Mar 4, 2022 · national
Continuity (1)
Related Publication 20230282010A1 · Sep 7, 2023
References Cited (19)
US 10445881B2 · Spizhevoy · 2019 [cited by examiner]
US 10565729B2 · Vajda · 2020 [cited by examiner]
US 10922393B2 · Spizhevoy · 2021 [cited by examiner]
US 11107232B2 · Li · 2021 [cited by examiner]
US 11176709B2 · Pillai · 2021 [cited by examiner]
US 11257252B2 · Wen · 2022 [cited by examiner]
US 11386583B2 · Wen · 2022 [cited by examiner]
US 11461998B2 · Liu · 2022 [cited by examiner]
US 12020437B2 · Cha · 2024 [cited by examiner]
US 12167031B1 · Galvin · 2024 [cited by examiner]
US 20200085382A1 · Taerum · 2020 [cited by examiner]
US 20200125852A1 · Carreira · 2020 [cited by examiner]
US 20200218948A1 · Mao · 2020 [cited by examiner]
US 20210271847A1 · Courtiol · 2021 [cited by examiner]
US 20210279503A1 · Qi · 2021 [cited by examiner]
US 20210406582A1 · Wang · 2021 [cited by examiner]
US 20220180199A1 · Xu · 2022 [cited by examiner]
US 20220366682A1 · Cha · 2022 [cited by examiner]
US 20230281818A1 · Wang · 2023 [cited by examiner]