IP Library Granted Patent US 12,555,356
Granted Patent B2
US 12,555,356 · App. 18/580,609 · Granted Feb 17, 2026

Computer vision system, computer vision method, computer vision program, and learning method

Inventors: Takayoshi Yamashita (Kasugai, JP); Hironobu Fujiyoshi (Kasugai, JP); Tsubasa Hirakawa (Kasugai, JP); Mitsuru Nakazawa (Tokyo, JP); Yeongnam Chae (Tokyo, JP); Bjorn Stenger (Tokyo, JP)
Assignees: RAKUTEN GROUP, INC.; CHUBU UNIVERSITY EDUCATIONAL FOUNDATION
G06V10/771G06N3/02G06V10/7715G06V20/42
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,555,356
App. No.
18/580,609
Granted
Feb 17, 2026
Kind
B2
Abstract

A computer vision system, with at least one processor configured to: acquire, from a sports match video, a plurality of pieces of consecutive image data indicating a portion of the sports match video, the plurality of pieces of consecutive image data including a plurality of pieces of first consecutive image data that are consecutive; and execute an estimation, by using a machine learning model, of whether the portion is of a predetermined scene type.

Claims (62)

1 . A computer vision system, comprising at least one processor configured to:

acquire, from a sports match video, a plurality of pieces of consecutive image data indicating a portion of the sports match video, the plurality of pieces of consecutive image data including a plurality of pieces of first consecutive image data that are consecutive; and

execute an estimation, by using a machine learning model, of whether the portion is of a predetermined scene type,

wherein, in the estimation, the at least one processor is configured to:

acquire, from the plurality of pieces of first consecutive image data, a plurality of first features each corresponding to one piece of the plurality of pieces of first consecutive image data and each indicating a feature of the one piece of the plurality of pieces of first consecutive image data;

acquire a plurality of second features from the plurality of first features by calculating a plurality of first salience degrees each corresponding to one of the plurality of first features and each indicating saliency of the one of the plurality of first features, and weighting each of the plurality of first features by corresponding one of the plurality of first salience degrees; and

acquire a result of the estimation based on the plurality of second features.

2 . The computer vision system according to claim 1 , wherein each of the plurality of first salience degrees is calculated based on a similarity between the plurality of first features.

3 . The computer vision system according to claim 1 , wherein the at least one processor is configured to:

acquire a plurality of pieces of frame image data indicating the portion, a number of pieces of the frame image data to be acquired being different from a number of pieces of the consecutive image data; and

acquire from the plurality of pieces of frame image data a same number of pieces of the consecutive image data as the number of pieces of the consecutive image data input to the machine learning model.

4 . The computer vision system according to claim 1 ,

wherein the plurality of pieces of consecutive image data further include a plurality of pieces of second consecutive image data that are consecutive after the plurality of pieces of first consecutive image data,

wherein, in the estimation, the at least one processor is further configured to:

acquire, from the plurality of pieces of second consecutive image data, a plurality of third features each corresponding to one piece of the plurality of pieces of second consecutive image data and each indicating a feature of the one piece of the plurality of pieces of second consecutive image data; and

acquire a plurality of fourth features from the plurality of third features by calculating a plurality of second salience degrees each corresponding to one of the plurality of third features and each indicating saliency of the one of the plurality of third features, and weighting each of the plurality of third features by corresponding one of the plurality of second salience degrees, and

acquire the result of the estimation based on the plurality of second features and the plurality of fourth features.

5 . The computer vision system according to claim 4 , wherein a number of pieces of the first consecutive image data is equal to a number of pieces of the second consecutive image data.

6 . The computer vision system according to claim 4 ,

wherein the at least one processor is configured to estimate which of a plurality of scene types including a first scene type and a second scene type the portion is of,

wherein the at least one processor is configured to:

acquire from the sports match video a plurality of pieces of first frame image data indicating the portion, a number of pieces of the first frame image data corresponding to the first scene type, and a plurality of pieces of second frame image data indicating the portion, a number of pieces of the second frame image data corresponding to the second scene type;

acquire from the plurality of pieces of first frame image data a same number of pieces of the consecutive image data relating to the first scene type as a number of pieces of the consecutive image data input to the machine learning model; and

acquire from the plurality of pieces of second frame image data the same number of pieces of the consecutive image data relating to the second scene type as the number of pieces of the consecutive image data input to the machine learning model, and

wherein, in the estimation, the at least one processor is configured to:

acquire first determination data as to whether the portion is of the first scene type based on the consecutive image data relating to the first scene type;

acquire second determination data as to whether the portion is of the second scene type based on the consecutive image data relating to the second scene type; and

acquire the result of the estimation as to which of the plurality of scene types the portion is of based on the first determination data and the second determination data.

7 . The computer vision system according to claim 4 , wherein the machine learning model is generated by:

acquiring a plurality of pieces of training consecutive image data including a plurality of pieces of first training consecutive image data that are consecutive and a plurality of pieces of second training consecutive image data that are consecutive after the plurality of pieces of first training consecutive image data, and label data associated with the plurality of pieces of training consecutive image data, the label data indicating the scene type relating to the plurality of pieces of training consecutive image data;

inputting the plurality of pieces of training consecutive image data to the machine learning model and acquiring a result of an estimation of the scene type relating to the plurality of pieces of training consecutive image data; and

training the machine learning model based on the result of the estimation and the label data.

8 . The computer vision system according to claim 7 ,

wherein the plurality of pieces of first training consecutive image data correspond to before an event characterizing the scene type relating to the plurality of pieces of training consecutive image data, and

wherein the plurality of pieces of second training consecutive image data correspond to after the event.

9 . A computer vision method, comprising:

acquiring, from a sports match video, a plurality of pieces of consecutive image data indicating a portion of the sports match video, the plurality of pieces of consecutive image data including a plurality of pieces of first consecutive image data that are consecutive; and

executing an estimation, by using a machine learning model, of whether the portion is of a predetermined scene type,

wherein the estimation comprises:

acquiring, from the plurality of pieces of first consecutive image data, a plurality of first features each corresponding to one piece of the plurality of pieces of first consecutive image data and each indicating a feature of the one piece of the plurality of pieces of first consecutive image data;

acquiring a plurality of second features from the plurality of first features by calculating a plurality of first salience degrees each corresponding to one of the plurality of first features and each indicating saliency of the one of the plurality of first features, and weighting each of the plurality of first features by corresponding one of the plurality of first salience degrees; and

acquiring a result of the estimation based on the plurality of second features.

10 . A non-transitory computer-readable information storage medium for storing a program for causing a computer to:

acquire, from a sports match video, a plurality of pieces of consecutive image data indicating a portion of the sports match video, the plurality of pieces of consecutive image data including a plurality of pieces of first consecutive image data that are consecutive; and

execute an estimation, by using a machine learning model, of whether the portion is of a predetermined scene type,

wherein the estimation further causing a computer to:

acquire, from the plurality of pieces of first consecutive image data, a plurality of first features each corresponding to one piece of the plurality of pieces of first consecutive image data and each indicating a feature of the one piece of the plurality of pieces of first consecutive image data;

acquire a plurality of second features from the plurality of first features by calculating a plurality of first salience degrees each corresponding to one of the plurality of first features and each indicating saliency of the one of the plurality of first features, and weighting each of the plurality of first features by corresponding one of the plurality of first salience degrees; and

acquire a result of the estimation based on the plurality of second features.

11 . A learning method for training a machine learning model configured to estimate whether a portion of a sports match video is of a predetermined scene type, based on a plurality of pieces of consecutive image data indicating the portion, the plurality of pieces of consecutive image data including a plurality of pieces of first consecutive image data that are consecutive and a plurality of pieces of second consecutive image data that are consecutive after the plurality of pieces of first consecutive image data, the learning method comprising:

acquiring a plurality of pieces of training consecutive image data including a plurality of pieces of first training consecutive image data that are consecutive and a plurality of pieces of second training consecutive image data that are consecutive after the plurality of pieces of first training consecutive image data, and label data associated with the plurality of pieces of training consecutive image data, the label data indicating the scene type relating to the plurality of pieces of training consecutive image data;

inputting the plurality of pieces of training consecutive image data to the machine learning model and acquiring a result of an estimation of the scene type relating to the plurality of pieces of training consecutive image data; and

training the machine learning model based on the result of the estimation and the label data,

wherein the estimation comprises:

acquiring from the plurality of pieces of first consecutive image data a plurality of first features each corresponding to one piece of the plurality of pieces of first consecutive image data and each indicating a feature of the one piece of the plurality of pieces of first consecutive image data;

acquiring a plurality of second features from the plurality of first features by calculating a plurality of first salience degrees each corresponding to one of the plurality of first features and each indicating saliency of the one of the plurality of first features, and weighting each of the plurality of first features by corresponding one of the plurality of first salience degrees;

acquiring from the plurality of pieces of second consecutive image data a plurality of third features each corresponding to one piece of the plurality of pieces of second consecutive image data and each indicating a feature of the one piece of the plurality of pieces of second consecutive image data;

acquiring a plurality of fourth features from the plurality of third features by calculating a plurality of second salience degrees each corresponding to one of the plurality of third features and each indicating saliency of the one of the plurality of third features, and weighting each of the plurality of third features by corresponding one of the plurality of second salience degrees; and

acquiring a result of the estimation based on the plurality of second features and the plurality of fourth features.

12 . The learning method according to claim 11 ,

wherein the plurality of pieces of first training consecutive image data correspond to before an event characterizing the scene type relating to the plurality of pieces of training consecutive image data, and

wherein the plurality of pieces of second training consecutive image data correspond to after the event.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 19, 2024
From: YAMASHITA, TAKAYOSHI; FUJIYOSHI, HIRONOBU; HIRAKAWA, TSUBASA; NAKAZAWA, MITSURU; CHAE, YEONGNAM; STENGER, BJORN
To: RAKUTEN GROUP, INC.; CHUBU UNIVERSITY EDUCATIONAL FOUNDATION
Reel/Frame 066173/0877 →
Continuity (1)
Related Publication 20240338930A1 · Oct 10, 2024
References Cited (11)
US 11715486B2 · Sainath · 2023 [cited by examiner]
US 20090092313A1 · Negi · 2009 [cited by examiner]
US 20170262996A1 · Jain · 2017 [cited by examiner]
US 20180330183A1 · Tsunoda · 2018 [cited by examiner]
US 20200302185A1 · Hussein · 2020 [cited by examiner]
US 20210124987A1 · Gan · 2021 [cited by examiner]
Baccouche, Moez, Franck Mamalet, Christian Wolf, Christophe Garcia, and Atilla Baskurt. “Action classification in soccer videos with long short-term memory recurrent neural networks.” In International Conference on Arti… [cited by examiner]
Rahmad, Nur Azmina, et al. “A survey of video based action recognition in sports.” Indonesian Journal of Electrical Engineering and Computer Science 11.3 (2018): 987-993. [cited by examiner]
Wu, Fei, et al. “A survey on video action recognition in sports: Datasets, methods and applications.” IEEE Transactions on Multimedia 25 (2022): 7943-7966. [cited by examiner]
Jeff Donahue et al., “Long-term recurrent convolutional networks for visual recognition and description,” 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA, Jun. 2015, pp. 2625-2634… [cited by applicant]
Anthony Cioppa et al., “A Context-Aware Loss Function for Action Spotting in Soccer Videos,” 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, Jun. 2020, pp. 13123-13133, doi:… [cited by applicant]