IP Library › Granted Patent US 12,260,853
Granted Patent B2
US 12,260,853 · App. 17/654,270 · Granted Mar 25, 2025

Speech processing method and apparatus

Inventors: Yinping Zhang (Beijing, CN); Chenyu Zhang (Beijing, CN); Lili Guo (Beijing, CN)
Assignee: LENOVO (BEIJING) LIMITED
G10L15/1815G10L15/05G10L15/063G10L15/16G10L15/22G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,260,853
App. No.
17/654,270
Granted
Mar 25, 2025
Kind
B2
Abstract

A speech processing method includes obtaining first speech information from a user, determining one or more similar speech segments in the first speech information and deleting one or more similar frames each of the one or more similar speech segments to obtain second speech information, and analyzing the second speech information to determine a user intent corresponding to the first speech information. A duration of the first speech information exceeds a preset analysis duration threshold, and a duration of the second speech information does not exceed the preset analysis duration threshold.

Claims (54)

1. A speech processing method, comprising:

obtaining first speech information from a user, wherein a duration of the first speech information exceeds a preset analysis duration threshold;

determining one or more similar speech segments in the first speech information and deleting one or more similar frames in the one or more similar speech segments to obtain second speech information, wherein a duration of the second speech information does not exceed the preset analysis duration threshold, and deleting the one or more similar frames in the one or more similar speech segments to obtain the second speech information comprises:

determining, according to the duration of the first speech information and the preset analysis duration threshold, a deletion ratio; and

deleting, according to the deletion ratio, the one or more similar frames in the one or more similar speech segments to obtain the second speech information; and

analyzing the second speech information to determine a user intent corresponding to the first speech information.

2. The method according to claim 1 , wherein before obtaining the first speech information from the user, the method further comprises:

obtaining third speech information from the user, wherein a duration of the third speech information does not exceed the preset analysis duration threshold; and

in response to the user intent not being determinable according to the third speech information, obtaining the first speech information from the user,

wherein an end time point of the first speech information is the same as an end time point of the third speech information.

3. The method according to claim 1 , wherein after obtaining the first speech information from the user, the method further comprises:

determining a plurality of signal frames corresponding to the first speech information;

deleting one or more signal frames not containing speech to obtain one or more frames containing speech; and

perform splicing processing on the one or more frames containing speech to obtain fourth speech information.

4. The method according to claim 3 , wherein after obtaining the fourth speech information, the method further comprises:

in response to a duration of the fourth speech information not exceeding the preset analysis duration threshold, analyzing the fourth speech information to determine the user intent corresponding to the first speech information.

5. The method of claim 4 , further comprising:

in response to the duration of the fourth speech information exceeding the preset analysis duration threshold, determining one or more similar speech segments in the fourth speech information, and deleting one or more similar frames in the similar speech segments to obtain second speech information.

6. The method according to claim 1 , wherein determining one or more similar speech segments in the first speech information comprises:

determining one or more frames containing speech corresponding to the first speech information, performing a frame-to-frame similarity analysis to the one or more frames containing speech, and determining one or more similar speech segments according to the one or more frames containing speech that satisfy a similarity standard.

7. The method according to claim 1 ,

wherein deleting, according to the deletion ratio, the one or more similar frames in the one or more similar speech segments to obtain the second speech information includes deleting, according to the deletion ratio, the one or more similar frames in the one or more similar speech segments to obtain fifth speech information,

the method further comprising:

performing smoothing on the fifth speech information to obtain the second speech information.

8. The method according to claim 1 , wherein analyzing the second speech information to determine the user intent corresponding to the first speech information comprises:

analyzing the second speech information through a wake-up word model to determine whether the first speech information is used to wake up an electronic apparatus,

wherein the wake-up word model is obtained by training a wake-up word speech sample set through a neural network, and a duration of each wake-up word speech in the wake-up word sample set does not exceed the preset analysis duration threshold.

9. A speech processing apparatus, comprising:

an acquisition module configured to obtain first speech information from a user, wherein a duration of the first speech information exceeds a preset analysis duration threshold;

a determination module configured to determine one or more similar speech segments in the first speech information and delete one or more similar frames in the one or more similar speech segments to obtain second speech information, wherein a duration of the second speech information does not exceed the preset analysis duration threshold, and the determination module comprises:

a determination sub-module configured to determine a corresponding deletion ratio according to the duration of the first speech information and the preset analysis duration threshold; and

a deletion sub-module configured to delete, according to the deletion ratio, the one or more similar frames in the one or more similar speech segments to obtain the second speech information; and

an analysis module configured to analyze the second speech information to determine a user intent corresponding to the first speech information.

10. The speech processing apparatus according to claim 9 , wherein the acquisition module is further configured to obtain third speech information from the user, wherein a duration of the third speech information does not exceed the preset analysis duration threshold.

11. The speech processing apparatus according to claim 10 , wherein the acquisition module is further configured to, in response to a user intent not being determinable according to the third speech information, obtain the first speech information from the user, wherein an end time point of the first speech information is the same as an end time point of the third speech information.

12. The speech processing apparatus according to claim 9 , wherein the determination module is further configured to determine a plurality of signal frames corresponding to the first speech information.

13. The speech processing apparatus according to claim 12 , further comprising:

a deletion module configured to delete one or more signal frames not containing speech from the plurality of signal frames to obtain one or more frames containing speech; and

a splicing module configured to perform splicing on the one or more frames containing speech to obtain fourth speech information.

14. The speech processing apparatus according to claim 13 , wherein the analysis module is further configured to analyze the fourth speech information in response to a duration of the fourth speech information not exceeding the preset analysis duration threshold, to determine the user intent corresponding to the first speech information.

15. The speech processing apparatus according to claim 13 , wherein the determination module is further configured to, in response to a duration of the fourth speech information exceeding the preset analysis duration threshold, determine similar speech segments in the fourth speech information and delete similar frames in the similar speech segments to obtain the second speech information.

16. The speech processing apparatus according to claim 9 , wherein the determination module is further configured to determine one or more frames containing speech corresponding to the first speech information, perform frame-to-frame similarity analysis on the frames containing speech, and determine one or more similar speech segments according to the one or more frames containing speech that satisfy a preset similarity degree.

17. The speech processing apparatus according to claim 9 ,

wherein the deletion sub-module is configured to delete, according to the deletion ratio, the one or more similar frames in the one or more similar speech segments to obtain fifth speech information,

the speech processing apparatus further comprising:

a smoothing module configured to perform smoothing processing on the fifth speech information to obtain the second speech information.

18. An electronic device, comprising:

a memory for storing program; and

a processing for executing the program stored in the memory to:

obtain first speech information from a user, wherein a duration of the first speech information exceeds a preset analysis duration threshold;

determine one or more similar speech segments in the first speech information and delete one or more similar frames in the one or more similar speech segments to obtain second speech information, wherein a duration of the second speech information does not exceed the preset analysis duration threshold, and deleting the one or more similar frames in the one or more similar speech segments to obtain the second speech information comprises:

determining, according to the duration of the first speech information and the preset analysis duration threshold, a deletion ratio; and

deleting, according to the deletion ratio, the one or more similar frames in the one or more similar speech segments to obtain the second speech information; and

analyze the second speech information to determine a user intent corresponding to the first speech information.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 10, 2022
From: ZHANG, YINPING; ZHANG, CHENYU; GUO, LILI
To: LENOVO (BEIJING) LIMITED
Reel/Frame 059222/0093 →
Priority Claims (1)
CN 202110645953.3 · Jun 10, 2021 · national
Continuity (1)
Related Publication 20220399012A1 · Dec 15, 2022
References Cited (13)
US 10482879B2 · Tang · 2019 [cited by examiner]
US 11342003B1 · Siagian · 2022 [cited by examiner]
US 20040193406A1 · Yamato · 2004 [cited by examiner]
US 20070265839A1 · Sasaki · 2007 [cited by examiner]
US 20090103896A1 · Harrington · 2009 [cited by examiner]
US 20100228548A1 · Liu · 2010 [cited by examiner]
US 20130054236A1 · Garcia Martinez · 2013 [cited by examiner]
US 20140195227A1 · Rudzicz · 2014 [cited by examiner]
US 20150058013A1 · Pakhomov · 2015 [cited by examiner]
US 20180039888A1 · Ge · 2018 [cited by examiner]
US 20200349927A1 · Stoimenov · 2020 [cited by examiner]
US 20220076023A1 · Shin · 2022 [cited by examiner]
US 20230178080A1 · Min · 2023 [cited by examiner]