IP Library › Granted Patent US 12,315,527
Granted Patent B2
US 12,315,527 · App. 17/428,015 · Granted May 27, 2025

Method and system for speech recognition

Inventors: Shiliang Zhang (Hangzhou, CN); Ming Lei (Hangzhou, CN)
Assignee: ALIBABA GROUP HOLDING LIMITED
G10L21/0216G10L15/063H04R3/005H04R3/04G10L2021/02166
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,315,527
App. No.
17/428,015
Filed
Aug 3, 2021
Granted
May 27, 2025
Kind
B2
Art Unit
2653
USPC
704/200
Abstract

Embodiments of the disclosure provide a method and system for speech recognition. The method comprises dividing space into a plurality of regions based on preset DOA angles to allocate a signal source to the plurality of regions, wherein signals in the plurality of regions are enhanced and recognized, the result of which are fused to obtain a recognition result of the signal source.

Claims (41)

1. A method comprising:

allocating a signal source based on different directions of arrival (DOAs) by dividing a physical space into a plurality of non-overlapping regions to allocate the signal source into the plurality of regions, the allocation performed prior to beamforming, the plurality of non-overlapping regions based on preset DOA angles comprising at least two of: an angle of 30 degrees, an angle of 60 degrees, an angle of 90 degrees, an angle of 120 degrees, and an angle of 150 degrees;

enhancing signals of the signal source for each of the regions to obtain enhanced signals corresponding to the regions, the enhancing performed independently for each region;

performing speech recognition on the enhanced signals corresponding to the regions to obtain recognition results corresponding to the regions;

providing the recognition results corresponding to the regions to respective acoustic models, each acoustic model trained for its corresponding region based on enhanced signal samples from that region; and

fusing outputs of the acoustic models to obtain a recognition result, wherein fusing analyzes outputs from all preset regions regardless of estimated signal source direction.

2. The method of claim 1 , the enhancing signals of the signal source corresponding to the regions comprising performing delay-and-sum (DAS) beamforming on the signals of the signal source corresponding to the regions to obtain the enhanced signals.

3. The method of claim 1 , the enhancing signals of the signal source corresponding to the regions comprising performing Minimum Variance Distortionless Response (MVDR) beamforming on the signals of the signal source corresponding to the regions to obtain the enhanced signals.

4. The method of claim 1 , further comprising:

dividing, prior to the allocating the signal source, a space into regions according to the regions;

performing speech enhancement on speech signals in the different regions to obtain different enhanced signal samples; and

using the obtained samples to perform training to obtain the acoustic models corresponding to the regions.

5. The method of claim 1 , the fusing output results is performed by using a Recognizer Output Voting Error Reduction (ROVER) based fusion system.

6. A system comprising:

a processor; and

a storage medium for tangibly storing thereon program logic for execution by the processor, the stored program logic comprising:

logic, executed by the processor, for allocating a signal source based on different directions of arrival (DOAs) by dividing a physical space into a plurality of non-overlapping regions to allocate the signal source into the plurality of regions, the allocation performed prior to beamforming, the plurality of non-overlapping regions based on preset DOA angles comprising at least two of: an angle of 30 degrees, an angle of 60 degrees, an angle of 90 degrees, an angle of 120 degrees, and an angle of 150 degrees,

logic, executed by the processor, for enhancing signals of the signal source for each of the regions to obtain enhanced signals corresponding to the regions, the enhancing performed independently for each region, respectively,

logic, executed by the processor, for performing speech recognition on the enhanced signals corresponding to the regions to obtain recognition results corresponding to the regions, each acoustic model trained for its corresponding region, logic, executed by the processor, for providing the recognition results corresponding to the regions to respective acoustic models based on enhanced signal samples from that region, and

logic, executed by the processor, for fusing output results from the acoustic models to obtain a recognition result, wherein the fusing analyzes outputs from all preset regions regardless of estimated signal source direction.

7. The system of claim 6 , the logic for allocating a signal source based on regions comprising: logic, executed by the processor, for dividing a space into a plurality of regions to allocate the signal source into the plurality of regions formed based on DOA angles.

8. The system of claim 6 , the logic for enhancing signals of the signal source corresponding to the regions comprising: logic, executed by the processor, for performing delay-and-sum (DAS) beamforming on the signals of the signal source corresponding to the regions to obtain the enhanced signals.

9. The system of claim 6 , the logic for enhancing signals of the signal source corresponding to the regions comprising: logic, executed by the processor, for performing Minimum Variance Distortionless Response (MVDR) beamforming on the signals of the signal source corresponding to the regions to obtain the enhanced signals.

10. The system of claim 6 , the stored program logic further comprising:

logic, executed by the processor, prior to the allocating the signal source, for dividing a space into regions according to the regions,

logic, executed by the processor, for performing speech enhancement on speech signals in the different regions to obtain different enhanced signal samples, and

logic, executed by the processor, for using the obtained samples to perform training to obtain the acoustic models corresponding to the regions.

11. The system of claim 6 , the fusing output results is performed by using a Recognizer Output Voting Error Reduction (ROVER) based fusion system.

12. A non-transitory computer-readable storage medium for tangibly storing computer program instructions capable of being executed by a computer processor, the computer program instructions defining:

allocating a signal source based on different directions of arrival (DOAs) by dividing a physical space into a plurality of generally non-overlapping regions to allocate the signal source into the plurality of regions, the allocation performed prior to beamforming, the plurality of non-overlapping regions based on preset DOA angles comprising at least two of: an angle of 30 degrees, an angle of 60 degrees, an angle of 90 degrees, an angle of 120 degrees, and an angle of 150 degrees;

enhancing signals of the signal source for each of the regions to obtain enhanced signals corresponding to the regions, the enhancing performed independently for each region;

performing speech recognition on the enhanced signals corresponding to the regions to obtain recognition results corresponding to the regions;

providing the recognition results corresponding to the regions to respective acoustic models, each acoustic model trained for its corresponding region based on enhanced signal samples from that region; and

fusing output results from the acoustic models to obtain a recognition result, wherein the fusing analyzes outputs from all preset regions regardless of estimated signal source direction.

13. The computer-readable storage medium of claim 12 , the allocating a signal source based on regions comprising: dividing a space into a plurality of regions to allocate the signal source into the plurality of regions.

14. The computer-readable storage medium of claim 12 , the enhancing signals of the signal source corresponding to the regions comprising: performing delay-and-sum (DAS) beamforming on the signals of the signal source corresponding to the regions to obtain the enhanced signals.

15. The computer-readable storage medium of claim 12 , the enhancing signals of the signal source corresponding to the regions comprising performing Minimum Variance Distortionless Response (MVDR) beamforming on the signals of the signal source corresponding to the regions to obtain the enhanced signals.

16. The computer-readable storage medium of claim 12 , further comprising:

dividing, prior to the allocating the signal source, a space into regions according to the regions;

performing speech enhancement on speech signals in the different regions to obtain different enhanced signal samples; and

using the obtained samples to perform training to obtain the acoustic models corresponding to the regions.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 19, 2021
From: ZHANG, SHILIANG; LEI, MING
To: ALIBABA GROUP HOLDING LIMITED
Reel/Frame 057223/0769 →
Priority Claims (1)
CN 201910111593.1 · Feb 12, 2019 · national
Continuity (1)
Related Publication 20220028404A1 · Jan 27, 2022
References Cited (74)
US 5586191A · Elko et al. · 1996 [cited by applicant]
US 5854999A · Hirayama · 1998 [cited by applicant]
US 6138094A · Miet et al. · 2000 [cited by applicant]
US 6574597B1 · Mohri et al. · 2003 [cited by applicant]
US 6633842B1 · Gong · 2003 [cited by applicant]
US 8762145B2 · Ouchi et al. · 2014 [cited by applicant]
US RE45379E · Rowe · 2015 [cited by applicant]
US 8976978B2 · Kitazawa et al. · 2015 [cited by applicant]
US 9076450B1 · Sadek · 2015 [cited by examiner]
US 9286897B2 · Bisani et al. · 2016 [cited by applicant]
US 9443516B2 · Katuri et al. · 2016 [cited by applicant]
US 9576582B2 · Ljolje et al. · 2017 [cited by applicant]
US 9653070B2 · Chang et al. · 2017 [cited by applicant]
US 10349172B1 · Huang · 2019 [cited by examiner]
US 10622004B1 · Zhang · 2020 [cited by examiner]
US 10943583B1 · Gandhe · 2021 [cited by examiner]
US 10971158B1 · Patangay · 2021 [cited by examiner]
US 11574628B1 · Kumatani · 2023 [cited by examiner]
US 20020042712A1 · Yajima et al. · 2002 [cited by applicant]
US 20020120443A1 · Epstein · 2002 [cited by examiner]
US 20040024599A1 · Deisher · 2004 [cited by applicant]
US 20080089531A1 · Koga et al. · 2008 [cited by applicant]
US 20090018828A1 · Nakadai · 2009 [cited by examiner]
US 20090018833A1 · Kozat · 2009 [cited by examiner]
US 20090030552A1 · Nakadai · 2009 [cited by examiner]
US 20100217590A1 · Nemer · 2010 [cited by examiner]
US 20110293107A1 · Kitazawa et al. · 2011 [cited by applicant]
US 20130332165A1 · Beckley et al. · 2013 [cited by applicant]
US 20140112487A1 · Laska · 2014 [cited by examiner]
US 20150095026A1 · Bisani · 2015 [cited by examiner]
US 20150161999A1 · Kalluri · 2015 [cited by examiner]
US 20160005394A1 · Hiroe · 2016 [cited by examiner]
US 20160034811A1 · Paulik · 2016 [cited by examiner]
US 20160171977A1 · Siohan et al. · 2016 [cited by applicant]
US 20160217789A1 · Lee · 2016 [cited by examiner]
US 20160275954A1 · Park et al. · 2016 [cited by applicant]
US 20160322055A1 · Sainath · 2016 [cited by examiner]
US 20170105074A1 · Jensen · 2017 [cited by examiner]
US 20170278513A1 · Li · 2017 [cited by examiner]
US 20180233129A1 · Bakish et al. · 2018 [cited by applicant]
US 20180240471A1 · Markovich Golan · 2018 [cited by examiner]
US 20180270565A1 · Ganeshkumar · 2018 [cited by examiner]
US 20180330745A1 · Ebenezer · 2018 [cited by examiner]
US 20190073999A1 · Prémont · 2019 [cited by examiner]
US 20190115039A1 · Du · 2019 [cited by examiner]
US 20190341050A1 · Diamant · 2019 [cited by examiner]
US 20190341053A1 · Zhang · 2019 [cited by examiner]
US 20200075033A1 · Hijazi · 2020 [cited by examiner]
US 20200175961A1 · Thomson · 2020 [cited by examiner]
US 20200335088A1 · Gao · 2020 [cited by examiner]
US 20200342846A1 · Cai · 2020 [cited by examiner]
US 20200342887A1 · Xu et al. · 2020 [cited by applicant]
US 20210005184A1 · Rao · 2021 [cited by examiner]
US 20210312914A1 · Hedayatnia · 2021 [cited by examiner]
CN 101194182A · 2008 [cited by examiner]
CN 102271299A · 2011 [cited by applicant]
CN 105161092A · 2015 [cited by applicant]
CN 105765650A · 2016 [cited by applicant]
CN 107742522A · 2018 [cited by applicant]
CN 108877827A · 2018 [cited by applicant]
CN 108922553A · 2018 [cited by applicant]
CN 109272989A · 2019 [cited by applicant]
CN 110047478A · 2019 [cited by examiner]
CN 108702458B · 2021 [cited by examiner]
EP 2710400B1 · 2021 [cited by examiner]
JP 2004198656A · 2004 [cited by examiner]
KR 101658001B1 · 2016 [cited by applicant]
WO WO2018171223A1 · 2018 [cited by examiner]
WO WO2020034095A1 · 2020 [cited by examiner]
Rogozan, Alexandrina, and Paul Deléglise. “Adaptive fusion of acoustic and visual sources for automatic speech recognition.” Speech Communication 26.1-2 (1998): 149-161. (Year: 1998). [cited by examiner]
Stefanakis, Nikolaos, Despoina Pavlidi, and Athanasios Mouchtaris. “Perpendicular cross-spectra fusion for sound source localization with a planar microphone array.” IEEE/ACM Transactions on Audio, Speech, and Language … [cited by examiner]
Alexandridis, Anastasios, and Athanasios Mouchtaris. “Multiple sound source location estimation in wireless acoustic sensor networks using DOA estimates: The data-association problem.” IEEE/ACM Transactions on Audio, Sp… [cited by examiner]
Vincent, Emmanuel, et al. “An analysis of environment, microphone and data simulation mismatches in robust speech recognition.” Computer Speech & Language 46 (2017): 535-557. (Year: 2017). [cited by examiner]
International Search Report to corresponding International Application No. PCT/CN2020/074178, mailed Apr. 21, 2020 (2 pages). [cited by applicant]