IP Library › Granted Patent US 12,537,020
Granted Patent B2
US 12,537,020 · App. 18/090,296 · Granted Jan 27, 2026

Role separation method, meeting summary recording method, role display method and apparatus, electronic device, and computer storage medium

Inventors: Siqi Zheng (Hangzhou, CN); Xianliang Wang (Hangzhou, CN); Hongbin Suo (Hangzhou, CN)
Assignee: Alibaba Group Holding Limited
G10L25/78G06F21/32G10L17/02H04R1/406
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,537,020
App. No.
18/090,296
Granted
Jan 27, 2026
Kind
B2
Abstract

A role separation method, a meeting summary recording method, a role display method and apparatus, an electronic device, and a computer storage medium, relating to the field of speech processing. The role separation method comprises: obtaining sound source angle data corresponding to a speech data frame, acquired by a speech acquisition device, of a role to be separated (S 102 ); on the basis of the sound source angle data, performing identity recognition on the role to be separated to obtain a first identity recognition result of the role to be separated (S 104 ); and separating the role on the basis of the first identity recognition result of the role to be separated (S 106 ). The role is separated in real time, thus making user experience smooth.

Claims (78)

1 . A method comprising:

acquiring sound source angle data corresponding to a speech data frame of a to-be-separated role collected by a speech collection device, the speech collection device including a microphone array, the acquiring the sound source angle data corresponding to the speech data frame of the to-be-separated role collected by the speech collection device including:

acquiring a covariance matrix of the speech data frame received by at least some microphones in the microphone array;

performing eigenvalue decomposition on the covariance matrix to obtain multiple eigenvalues;

selecting a first quantity of largest eigenvalues from the multiple eigenvalues, and forming a speech signal sub-space based on eigenvectors corresponding to the selected eigenvalues; and

determining the sound source angle data based on the speech signal sub-space;

performing identification on the to-be-separated role based on the sound source angle data to obtain a first identification result of the to-be-separated role; and

separating a role based on the first identification result of the to-be-separated role.

2 . The method according to claim 1 , wherein after the acquiring the sound source angle data corresponding to the speech data frame of the to-be-separated role collected by the speech collection device, the method further comprises:

performing voice activity detection on the speech data frame of the to-be-separated role to obtain a speech data frame with a voice activity;

filtering and smoothing the speech data frame with the voice activity based on an energy spectrum of the speech data frame of the to-be-separated role to obtain a filtered and smoothed speech data frame; and

updating the sound source angle data based on the filtered and smoothed speech data frame to obtain updated sound source angle data.

3 . The method according to claim 2 , wherein the filtering and smoothing the speech data frame with the voice activity based on the energy spectrum of the speech data frame of the to-be-separated role to obtain the filtered and smoothed speech data frame comprises:

filtering and smoothing the speech data frame with the voice activity through a median filter based on a spectral flatness of the energy spectrum of the speech data frame of the to-be-separated role to obtain the filtered and smoothed speech data frame.

4 . The method according to claim 1 , wherein the performing the identification on the to-be-separated role based on the sound source angle data to obtain the first identification result of the to-be-separated role comprises:

performing sequence clustering on the sound source angle data to obtain a sequence clustering result of the sound source angle data; and

determining that a role identifier corresponding to the sequence clustering result of the sound source angle data is the first identification result of the to-be-separated role.

5 . The method according to claim 4 , wherein the performing the sequence clustering on the sound source angle data to obtain the sequence clustering result of the sound source angle data comprises:

determining a distance between the sound source angle data and a sound source angle sequence clustering center; and

determining the sequence clustering result of the sound source angle data based on the distance between the sound source angle data and the sound source angle sequence clustering center.

6 . The method according to claim 1 , wherein after the obtaining the first identification result of the to-be-separated role, the method further comprises:

performing voiceprint identification on a speech data frame of the to-be-separated role within a preset time period to obtain a second identification result of the to-be-separated role; and

in response to determining that the first identification result is different from the second identification result, using the second identification result to correct the first identification result to obtain a final identification result of the to-be-separated role.

7 . The method according to claim 1 , wherein the first quantity is equivalent to an estimated quantity of sound sources.

8 . The method according to claim 6 , wherein after the obtaining the final identification result of the to-be-separated role, the method further comprises:

acquiring face image data of the to-be-separated role collected by an image collection device;

performing face recognition on the face image data to obtain a third identification result of the to-be-separated role; and

in response to determining that the third identification result is different from the second identification result, using the third identification result to correct the second identification result to obtain the final identification result of the to-be-separated role.

9 . An apparatus comprising:

one or more processors; and

one or more memories storing thereon computer-readable instructions that, executable by the one or more processors, cause the one or more processors to perform acts comprising:

acquiring sound source angle data corresponding to a speech data frame of a role collected by a speech collection device, the speech collection device including a microphone array, the acquiring the sound source angle data corresponding to the speech data frame of the role collected by the speech collection device including:

acquiring a covariance matrix of the speech data frame received by at least some microphones in the microphone array;

performing eigenvalue decomposition on the covariance matrix to obtain multiple eigenvalues;

selecting a first quantity of largest eigenvalues from the multiple eigenvalues, and forming a speech signal sub-space based on eigenvectors corresponding to the selected eigenvalues; and

determining the sound source angle data based on the speech signal sub-space;

performing identification on the role based on the sound source angle data to obtain a first identification result of the role; and

displaying identity data of the role on an interactive interface of the speech collection device based on the first identification result of the role.

10 . The apparatus according to claim 9 , wherein after the acquiring the sound source angle data corresponding to the speech data frame of the role collected by the speech collection device, the acts further comprise:

switching on a lamp of the speech collection device in a sound source direction indicated by the sound source angle data.

11 . The apparatus according to claim 9 , wherein the acts further comprise:

displaying a speaking action image or a speech waveform image of the role on the interactive interface of the speech collection device.

12 . The apparatus according to claim 9 , wherein the acts further comprise:

performing voiceprint identification on the speech data frame within a preset time period to obtain a second identification result; and

in response to determining that the first identification result is different from the second identification result, using the second identification result to correct the first identification result to obtain a final identification result of the role.

13 . The apparatus according to claim 12 , wherein after the obtaining the final identification result of the role, the acts further comprise:

acquiring face image data of the role collected by an image collection device;

performing face recognition on the face image data to obtain a third identification result of the role; and

in response to determining that the third identification result is different from the second identification result, using the third identification result to correct the second identification result to obtain the final identification result of the role.

14 . The apparatus according to claim 12 , wherein the acts further comprise displaying identity data of the role on the interactive interface of the speech collection device based on the final identification result of the role.

15 . One or more memories storing thereon computer-readable instructions that, executable by one or more processors, cause the one or more processors to perform acts comprising:

acquiring sound source angle data corresponding to a speech data frame of a to-be-separated role collected by a speech collection device;

performing identification on the to-be-separated role based on the sound source angle data to obtain a first identification result of the to-be-separated role;

performing voiceprint identification on the speech data frame of the to-be-separated role within a preset time period to obtain a second identification result of the to-be-separated role;

in response to determining that the first identification result is different from the second identification result, using the second identification result to correct the first identification result;

acquiring face image data of the to-be-separated role collected by an image collection device;

performing face recognition on the face image data to obtain a third identification result of the to-be-separated role;

in response to determining that the third identification result is different from the second identification result, using the third identification result to correct the second identification result to obtain a final identification result of the to-be-separated role; and

separating a role based on the final identification result of the to-be-separated role.

16 . The one or more memories according to claim 15 , wherein after the acquiring the sound source angle data corresponding to the speech data frame of the to-be-separated role collected by the speech collection device, the acts further comprise:

performing voice activity detection on the speech data frame of the to-be-separated role to obtain a speech data frame with a voice activity;

filtering and smoothing the speech data frame with the voice activity based on an energy spectrum of the speech data frame of the to-be-separated role to obtain a filtered and smoothed speech data frame; and

updating the sound source angle data based on the filtered and smoothed speech data frame to obtain updated sound source angle data.

17 . The one or more memories according to claim 16 , wherein the filtering and smoothing the speech data frame with the voice activity based on the energy spectrum of the speech data frame of the to-be-separated role to obtain the filtered and smoothed speech data frame comprises:

filtering and smoothing the speech data frame with the voice activity through a median filter based on a spectral flatness of the energy spectrum of the speech data frame of the to-be-separated role to obtain the filtered and smoothed speech data frame.

18 . The one or more memories according to claim 15 , wherein the performing the identification on the to-be-separated role based on the sound source angle data to obtain the first identification result of the to-be-separated role comprises:

performing sequence clustering on the sound source angle data to obtain a sequence clustering result of the sound source angle data; and

determining that a role identifier corresponding to the sequence clustering result of the sound source angle data is the first identification result of the to-be-separated role.

19 . The one or more memories according to claim 18 , wherein the performing the sequence clustering on the sound source angle data to obtain the sequence clustering result of the sound source angle data comprises:

determining a distance between the sound source angle data and a sound source angle sequence clustering center; and

determining the sequence clustering result of the sound source angle data based on the distance between the sound source angle data and the sound source angle sequence clustering center.

20 . The one or more memories according to claim 15 , wherein:

the speech collection device comprises a microphone array; and

the acquiring the sound source angle data corresponding to the speech data frame of the to-be-separated role collected by the speech collection device comprises:

acquiring a covariance matrix of the speech data frame received by at least some microphones in the microphone array;

performing eigenvalue decomposition on the covariance matrix to obtain multiple eigenvalues;

selecting a first quantity of largest eigenvalues from the multiple eigenvalues, and forming a speech signal sub-space based on eigenvectors corresponding to the selected eigenvalues; and

determining the sound source angle data based on the speech signal sub-space.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 7, 2023
From: ZHENG, SIQI; WANG, XIANLIANG; SUO, HONGBIN
To: ALIBABA GROUP HOLDING LIMITED
Reel/Frame 063885/0477 →
Priority Claims (1)
CN 202010596049.3 · Jun 28, 2020 · national
Continuity (2)
Continuation PCTCN2021101956 · Jun 24, 2021
Related Publication 20230162757A1 · May 25, 2023
References Cited (26)
US 9807497B2 · Nakadai · 2017 [cited by applicant]
US 11527242B2 · Wu · 2022 [cited by examiner]
US 20150025888A1 · Sharp · 2015 [cited by examiner]
US 20150082404A1 · Goldstein · 2015 [cited by applicant]
US 20160021242A1 · Hodge · 2016 [cited by examiner]
US 20190115030A1 · Lesso · 2019 [cited by examiner]
US 20190272844A1 · Sharma · 2019 [cited by examiner]
US 20200077218A1 · Nakadai · 2020 [cited by examiner]
CN 203351200U · 2013 [cited by examiner]
CN 105023572A · 2015 [cited by examiner]
CN 110062200A · 2019 [cited by applicant]
CN 110298252A · 2019 [cited by applicant]
CN 111048095A · 2020 [cited by applicant]
CN 111199741A · 2020 [cited by applicant]
CN 111260313A · 2020 [cited by applicant]
EP 2942975A1 · 2015 [cited by examiner]
JP 2016133304A · 2016 [cited by examiner]
JP 6589042B1 · 2019 [cited by examiner]
KR 20050035562A · 2005 [cited by examiner]
KR 20110038447B2 · 2011 [cited by examiner]
English Translation of Chinese First Office Action for corresponding Chinese Application No. 202010596049.3 dated Jun. 22, 2023, 12 pages. [cited by applicant]
English Translation of PCT Search Report dated Sep. 24, 2021 for corresponding PCT Application No. PCT/CN2021/101956 2 pages. [cited by applicant]
English Translation of PCT Written Opinion dated Sep. 24, 2021 for corresponding PCT Application No. PCT/CN2021/101956, 4 pages. [cited by applicant]
English Translation of Chinese First Search Report for corresponding Chinese Application No. 202010596049.3 dated Jun. 22, 2023, 2 pages. [cited by applicant]
English Translation of Chinese Second Office Action for corresponding Chinese Application No. 202010596049.3 dated Feb. 22, 2024, 7 pages. [cited by applicant]
Written Opinion for Singapore Application No. 11202261617P, dated Oct. 7, 2025, 6 pages. [cited by applicant]