IP Library Granted Patent US 10,984,801
Granted Patent B2
US 10,984,801 · App. 16/609,553 · Granted Apr 20, 2021

ASR training and adaptation

Inventors: Volodya Grancharov (Solna, SE); Erlendur Karlsson (Uppsala, SE); Sigurdur Sverrisson (Kungsängen, SE); Maxim Teslenko (Sollentuna, SE); Konstantinos Vandikas (Solna, SE); Aneta Vulgarakis Feljan (Stockholm, SE)
Assignee: Telefonaktiebolaget LM Ericsson (publ)
G10L17/04G10L17/00G10L17/06
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,984,801
App. No.
16/609,553
Granted
Apr 20, 2021
Kind
B2
Abstract

AM and LM parameters to be used for adapting an ASR model are derived for each audio segment of an audio stream comprising multiple audio programs. A set of identifiers, including a speaker identifier, a speaker domain identifier and a program domain identifier, is obtained for each audio segment. The set of identifiers are used to select most suitable AM and LM parameters for the particular audio segment. The embodiments enable provision of maximum constraints on the AMs and LMs and enable adaptation of the ASR model on the fly for audio streams of multiple audio programs, such as broadcast audio. This means that the embodiments enable selecting AM and LM parameters that are most suitable in terms of ASR performance for each audio segment.

Claims (34)

1. An audio processing method for automatic speech recognition, ASR, said method comprising:

for each audio segment of multiple audio segments in an audio stream comprising audio data of multiple audio programs, each audio segment comprising speech of a single speaker:

obtaining a speaker identifier of a speaker of said audio segment;

determining a speaker domain identifier for said audio segment based on a program domain identifier associated with said speaker identifier; and

associating said speaker identifier, said speaker domain identifier and a program domain identifier with said audio segment to enable generation of ASR adaptation parameters based on said speaker identifier, said speaker domain identifier and said program domain identifier, wherein said program domain identifier is associated with an audio program of said multiple audio programs, and said audio segment comprises audio data of said audio program.

2. The method according to claim 1 , wherein obtaining said speaker identifier comprises:

performing a speaker recognition process on said audio segment to determine a speaker model of said speaker;

retrieving said speaker identifier from a database based on said speaker model if said database comprises said speaker identifier; and

otherwise assigning a speaker identifier of said speaker and storing said speaker identifier and said speaker model in said database.

3. The method according to claim 1 , wherein determining said speaker domain identifier comprises:

retrieving, based on said speaker identifier, said program domain identifier associated with said speaker from a database storing program domain identifiers for different speakers with a respective speaker identifier if said database comprises a program domain identifier for said speaker identifier; and

assigning said program domain identifier associated with said speaker as speaker domain identifier; and

otherwise assigning a default speaker domain identifier to said audio segment.

4. A device configured for audio processing for automatic speech recognition, ASR, wherein

said device is configured to obtain, for each audio segment of multiple audio segments in an audio stream comprising audio data of multiple audio programs, each audio segment comprising speech of a single speaker, a speaker identifier of a speaker of said audio segment;

said device is configured to determine, for each audio segment of said multiple audio segments, a speaker domain identifier for said audio segment based on a program domain identifier associated with said speaker identifier; and

said device is configured to associate, for each audio segment of said multiple audio segments, said speaker identifier, said speaker domain identifier and a program domain identifier with said audio segment to enable generation of ASR adaptation parameters based on said speaker identifier, said speaker domain identifier and said program domain identifier, wherein said program domain identifier is associated with an audio program of said multiple audio programs, and said audio segment comprises audio data of said audio program.

5. The device according to claim 4 , wherein

said device is configured to perform, for each audio segment of said multiple audio segments, a speaker recognition process on said audio segment to determine a speaker model of said speaker; and

said device is configured to determine, for each audio segment of said multiple audio segments, said speaker identifier based on said speaker model.

6. The device according to claim 5 , wherein

said device is configured to retrieve, for each audio segment of said multiple audio segments, said speaker identifier from a database based on said speaker model if said database comprises said speaker identifier; and

said device is configured to assign, for each audio segment of said multiple audio segments and if said database does not comprise said speaker identifier, a speaker identifier of said speaker and store said speaker identifier and said speaker model in said database.

7. The device according to claim 4 , wherein

said device is configured to retrieve, for each audio segment of said multiple audio segments and based on said speaker identifier, said program domain identifier associated with said speaker from a database storing program domain identifiers for different speakers with a respective speaker identifier if said database comprises a program domain identifier for said speaker identifier;

said device is configured to assign, for each audio segment of said multiple audio segments, said program domain identifier associated with said speaker as speaker domain identifier; and

said device is configured to assign, for each audio segment of said multiple audio segments and if said database does not comprise said program domain identifier for said speaker identifier, a default speaker domain identifier to said audio segment.

8. The device according to claim 4 , wherein said device is configured to determine said program domain identifier associated with said audio program based on a media description of said audio program.

9. The device according to claim 4 , further comprising:

a processor; and

a memory comprising instructions executable by said processor, wherein said processor is operative to

obtain said speaker identifier;

determine said speaker domain identifier; and

associate said speaker identifier, said speaker domain identifier and said program domain identifier with said audio segment.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 19, 2019
From: GRANCHAROV, VOLODYA; KARLSSON, ERLENDUR; TESLENKO, MAXIM; SVERRISSON, SIGURDUR; VANDIKAS, KONSTANTINOS; VULGARAKIS FELJAN, ANETA
To: TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
Reel/Frame 051051/0881 →
Continuity (1)
Related Publication 20200066281A1 · Feb 27, 2020