IP Library › Granted Patent US 11,727,939
Granted Patent B2
US 11,727,939 · App. 17/568,931 · Granted Aug 15, 2023

Voice-controlled management of user profiles

Inventors: Volodya Grancharov (Solna, SE); Tomer Amiaz (Tel Aviv, IL); Hadar Gecht (Ramat-HaSharon, IL); Harald Pobloth (Täby, SE)
Assignee: TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
G10L17/02G06F17/18G10L15/08G10L21/0308
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,727,939
App. No.
17/568,931
Granted
Aug 15, 2023
Kind
B2
Abstract

A network node in a communication network receives, from a user equipment, a cluster of audio segments. The network node calculates a first confidence measure representing a first probability that a first speaker model represents a speaker of the cluster of audio segments. The network node also calculates a second confidence measure representing a second probability that a second speaker model represents the speaker of the cluster of audio segments. In response to the first confidence measure and the second confidence measure both representing probabilities that are higher than a target probability, the network node updates a first user profile associated with the first speaker model and a second user profile associated with the second speaker model based on a user preference assigned to the cluster of audio segments.

Claims (61)

1. A method of managing user profiles, implemented by a network node in a communication network, the method comprising:

receiving, from a user equipment, a cluster of audio segments;

calculating a first confidence measure representing a first probability that a first speaker model represents a speaker of the cluster of audio segments;

calculating a second confidence measure representing a second probability that a second speaker model represents the speaker of the cluster of audio segments; and

responsive to the first confidence measure and the second confidence measure both representing probabilities that are higher than a target probability, updating a first user profile associated with the first speaker model and a second user profile associated with the second speaker model based on a user preference assigned to the cluster of audio segments.

2. The method of claim 1 , further comprising:

receiving, from the user equipment, a further cluster of audio segments;

calculating a further first confidence measure representing a further first probability that the first speaker model represents a speaker of the further cluster of audio segments;

calculating a further second confidence measure representing a further second probability that the second speaker model represents the speaker of the further cluster of audio segments; and

responsive to the further first probability being higher than the target probability, updating the user profile associated with the first speaker model based on a further user preference assigned to the further cluster of audio segments.

3. The method of claim 2 , further comprising sending, to the user equipment, an identity of the speaker of the further cluster of audio segments in response to the further first probability being higher than the target probability.

4. The method of claim 1 , wherein receiving the cluster of audio segments comprises receiving the cluster of audio segments in an audio stream from the user equipment, the method further comprising performing speaker diarisation on the audio stream to form the cluster of audio segments such that the cluster of audio segments comprises speech of a single speaker.

5. The method of claim 4 , wherein performing speaker diarisation comprises:

detecting speech active segments from the audio stream;

detecting speaker change points in the speech active segments to form the audio segments of the single speaker; and

clustering the audio segments of the single speaker to form the cluster of audio segments.

6. The method of claim 1 , further comprising updating the first speaker model and the second speaker model based on the cluster of audio segments in response to the first confidence measure and the second confidence measure both representing a probability that is higher than the target probability.

7. The method of claim 1 , further comprising assigning the user preference to the cluster of audio segments.

8. The method of claim 7 , further comprising performing automatic speech recognition on the cluster of audio segments to identify the user preference.

9. The method of claim 1 , further comprising identifying which of the first speaker model and the second speaker model is associated with a highest probability between the first confidence measure and the second confidence measure.

10. The method of claim 1 , further comprising:

calculating a further first confidence measure representing a further first probability that the first speaker model represents a speaker of a further cluster of audio segments;

calculating a further second confidence measure representing a further second probability that the second speaker model represents the speaker of the further cluster of audio segments;

responsive to the further first confidence measure and the further second confidence measure both representing probabilities that are not higher than the target probability:

creating a new speaker model;

updating a default user profile based on a user preference associated with the further cluster of audio segments; and

associating the updated default user profile with the new speaker model.

11. A network node comprising:

processing circuitry;

memory containing instructions executable by the processing circuitry whereby the network node is configured to:

receive, from a user equipment, a cluster of audio segments;

calculate a first confidence measure representing a first probability that a first speaker model represents a speaker of the cluster of audio segments;

calculate a second confidence measure representing a second probability that a second speaker model represents the speaker of the cluster of audio segments; and

responsive to the first confidence measure and the second confidence measure both representing probabilities that are higher than a target probability, update a first user profile associated with the first speaker model and a second user profile associated with the second speaker model based on a user preference assigned to the cluster of audio segments.

12. The network node of claim 11 , wherein the network node is further configured to:

receive, from the user equipment, a further cluster of audio segments;

calculate a further first confidence measure representing a further first probability that the first speaker model represents a speaker of the further cluster of audio segments;

calculate a further second confidence measure representing a further second probability that the second speaker model represents the speaker of the further cluster of audio segments; and

responsive to the further first probability being higher than the target probability, update the user profile associated with the first speaker model based on a further user preference assigned to the further cluster of audio segments.

13. The network node of claim 12 , wherein the network node is further configured to send, to the user equipment, an identity of the speaker of the further cluster of audio segments in response to the further first probability being higher than the target probability.

14. The network node of claim 11 , wherein to receive the cluster of audio segments the network node is configured to receive an audio stream comprising the cluster of audio segments from the user equipment and the network node is further configured to perform speaker diarisation on the audio stream to form the cluster of audio segments such that the cluster of audio segments comprises speech of a single speaker.

15. The network node of claim 14 , wherein the network node is further configured to:

detect speech active segments from the audio stream;

detect speaker change points in the speech active segments to form audio segments of the single speaker; and

cluster the audio segments of the single speaker to form the cluster of audio segments.

16. The network node of claim 11 , wherein the network node is further configured to update the first speaker model and the second speaker model based on the cluster of audio segments in response to the first confidence measure and the second confidence measure both representing a probability that is higher than the target probability.

17. The network node of claim 11 , wherein the network node is further configured to assign the user preference to the cluster of audio segments.

18. The network node of claim 17 , wherein the network node is further configured to perform automatic speech recognition on the cluster of audio segments to identify the user preference.

19. The network node of claim 11 , wherein the network node is further configured to identify which of the first speaker model and the second speaker model is associated with a highest probability between the first confidence measure and the second confidence measure.

20. The network node of claim 19 , wherein the network node is further configured to:

calculate a further first confidence measure representing a further first probability that the first speaker model represents a speaker of a further cluster of audio segments;

calculate a further second confidence measure representing a further second probability that the second speaker model represents the speaker of the further cluster of audio segments;

responsive to the further first confidence measure and the further second confidence measure both representing probabilities that are not higher than the target probability:

create a new speaker model;

update a default user profile based on a user preference associated with the further cluster of audio segments; and

associate the updated default user profile with the new speaker model.

21. A non-transitory computer readable medium storing a computer program for controlling a programmable network node in a communication network, the computer program comprising software instructions that, when executed by processing circuitry of the programmable network node, cause the programmable network node to:

receive, from a user equipment, a cluster of audio segments;

calculate a first confidence measure representing a first probability that a first speaker model represents a speaker of the cluster of audio segments;

calculate a second confidence measure representing a second probability that a second speaker model represents the speaker of the cluster of audio segments; and

responsive to the first confidence measure and the second confidence measure both representing probabilities that are higher than a target probability, update a first user profile associated with the first speaker model and a second user profile associated with the second speaker model based on a user preference assigned to the cluster of audio segments.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 5, 2022
From: AMIAZ, TOMER; GECHT, HADAR; GRANCHAROV, VOLODYA; POBLOTH, HARALD
To: TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
Reel/Frame 058556/0698 →
Continuity (2)
Continuation 16644531
Related Publication 20220130395A1 · Apr 28, 2022
Cited By (1)
US 12,272,349