IP Library Granted Patent US 12670896
Granted Patent B2
US 12670896 · App. 18/091,393 · Granted Jun 30, 2026

Information processing method, non-transitory recording medium, information processing apparatus, and information processing system

Inventors: Masaki Nose (Kanagawa, JP); Akihiro Kato (Kanagawa, JP); Hiroyuki Nagano (Tokyo, JP); Yuuto Gotoh (Kanagawa, JP)
Assignee: RICOH COMPANY, LTD.
G10L13/08G10L15/063G10L15/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12670896
App. No.
18/091,393
Granted
Jun 30, 2026
Kind
B2
Abstract

An information processing method includes obtaining speech data based on a distance between a sound collection device and a speaker, obtaining text data input in a service for exchanging messages, and outputting first learning data that is based on the speech data and second learning data that includes the text data.

Claims (46)

1 . An information processing method, comprising:

generating first learning data from first audio data acquired based on a distance between a microphone and a speaker;

generating second learning data using text data that is independent from the first audio data;

storing third learning data corresponding to a speech recognition usage environment by pairing second audio data that is different from the first audio data with annotated text data corresponding to the second audio data;

performing pre-training of a speech recognition model using the first learning data and the second learning data to generate a first speech recognition model, the pre-training being performed independently of the speech recognition usage environment;

performing training by modifying parameters of the first speech recognition model based on the third learning data corresponding to the speech recognition usage environment to generate a second speech recognition model; and

performing speech recognition on sampled audio data from the speech recognition usage environment using the second speech recognition model and generating recognized text data.

2 . The information processing method according to claim 1 , further comprising:

obtaining the first audio data using a first microphone which is less than a predetermined distance from a mouth of the speaker;

obtaining the text data from a received text message; and

obtaining the second audio data using a second microphone which is greater than a predetermined distance from the mouth of the speaker.

3 . The information processing method of claim 2 , wherein:

the obtaining the first audio data utilizes a system that implements a remote conference via the Internet.

4 . The information processing method of claim 3 , wherein:

the obtaining the second audio data utilizes a system that implements a face-to-face conference.

5 . A non-transitory recording medium storing a plurality of instructions which, when executed by one or more processors, causes the processors to perform a method, the method comprising:

generating first learning data from first audio data acquired based on a distance between a microphone and a speaker;

generating second learning data using text data that is independent of the first audio data;

storing third learning data corresponding to a speech recognition usage environment by pairing second audio data that is different from the first audio data with annotated text data corresponding to the second audio data;

performing pre-training of a speech recognition model using the first learning data and the second learning data to generate a first speech recognition model, the pre-training being performed independently of the speech recognition usage environment;

performing by modifying parameters of the first speech recognition model based on the third learning data corresponding to the speech recognition usage environment to generate a second speech recognition model; and

performing speech recognition on sampled audio data from the speech recognition usage environment using the second speech recognition model and generating recognized text data.

6 . The non-transitory recording medium according to claim 5 , wherein the method further comprises:

obtaining the first audio data using a first microphone which is less than a predetermined distance from a mouth of the speaker;

obtaining the text data from a received text message; and

obtaining the second audio data using a second microphone which is greater than a predetermined distance from the mouth of the speaker.

7 . The non-transitory recording medium according to claim 6 , wherein:

the obtaining the first audio data utilizes a system that implements a remote conference via the Internet.

8 . The non-transitory recording medium according to claim 7 , wherein:

the obtaining the second audio data utilizes a system that implements a face-to-face conference.

9 . An information processing system, comprising:

circuitry configured to:

generate first learning data from first audio data acquired based on a distance between a microphone and a speaker;

generate second learning data using text data that is independent of the first audio data;

store third learning data corresponding to a speech recognition usage environment by pairing second audio data that is different from the first audio data with annotated text data corresponding to the second audio data;

perform pre-training of a speech recognition model using the first learning data and the second learning data to generate a first speech recognition model, the pre-training being performed independently of the speech recognition usage environment;

perform training by modifying parameters of the first speech recognition model based on the third learning data corresponding to the speech recognition usage environment to generate a second speech recognition model; and

perform speech recognition on sampled audio data from the speech recognition usage environment using the second speech recognition model and generate recognized text data.

10 . The information processing system according to claim 9 , wherein the circuitry is further configured to:

obtain the first audio data using a first microphone which is less than a predetermined distance from a mouth of the speaker;

obtain the text data from a received text message; and

obtain the second audio data using a second microphone which is greater than a predetermined distance from the mouth of the speaker.

11 . The information processing system according to claim 10 , wherein:

the obtaining the first audio data utilizes a remote conference system using the Internet.

12 . The information processing system according to claim 11 , wherein:

the obtaining the second audio data utilizes a system that implements a face-to-face conference.