IP Library Granted Patent US 11,244,686
Granted Patent B2
US 11,244,686 · App. 16/355,164 · Granted Feb 8, 2022

Method and apparatus for processing speech

Inventor: Ya Wu (Beijing, CN)
Assignees: Baidu Online Network Technology (Beijing) Co., Ltd.; ShangHai Xiaodu Technology Co. Ltd.
G10L15/32G10L15/08G10L15/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,244,686
App. No.
16/355,164
Granted
Feb 8, 2022
Kind
B2
Abstract

Embodiments of a method and apparatus for processing a speech are provided. The method can include: acquiring, in response to determining at least one speech interaction device in a target speech interaction device set receiving an input speech, a speech feature of the input speech received by a speech interaction device of the at least one speech interaction device; and selecting, based on the speech feature of the input speech received by the speech interaction device in the at least one speech interaction device, a first speech interaction device from the at least one speech interaction device to process the input speech. Some embodiments realize the selection of a targeted speech interaction device.

Claims (48)

1. A method for processing speech, the method comprising:

acquiring, in response to determining at least one speech interaction device in a target speech interaction device set receiving an input speech, a speech feature of the input speech received by a speech interaction device of the at least one speech interaction device; and

selecting, based on the speech feature of the input speech received by the speech interaction device in the at least one speech interaction device, a first speech interaction device from the at least one speech interaction device to process the input speech,

wherein the speech feature comprises sound pressure;

wherein selecting, based on the speech feature of the input speech received by the speech interaction device in the at least one speech interaction device, the first speech interaction device from the at least one speech interaction device to process the input speech, comprises:

selecting, according to the sound pressure of the input speech received by the speech interaction device in the at least one speech interaction device in descending order, a preset first number of the at least one speech interaction device from the at least one speech interaction device to process the input speech; and

wherein the method is performed by at least one hardware processor.

2. The method according to claim 1 , wherein the speech feature further comprises loudness; and

the selecting, based on the speech feature of the input speech received by the speech interaction device in the at least one speech interaction device, a first speech interaction device from the at least one speech interaction device to process the input speech, further comprises:

selecting, according to the loudness of the input speech received by the speech interaction device in the at least one speech interaction device in descending order, a preset second number of the at least one speech interaction device from the at least one speech interaction device to process the input speech.

3. The method according to claim 1 , wherein the selecting, based on the speech feature of the input speech received by the speech interaction device in the at least one speech interaction device, a first speech interaction device from the at least one speech interaction device to process the input speech, comprises:

selecting, in response to determining that the input speech comprises a preset wake-up word, based on the speech feature of the input speech received by the speech interaction device in the at least one speech interaction device, the first speech interaction device from the at least one speech interaction device for being woken up so that the woken first speech interaction device processes the input speech.

4. The method according to claim 1 , wherein before the selecting a first speech interaction device from the at least one speech interaction device to process the input speech, the method further comprises:

analyzing the input speech to obtain an analysis result; and

the selecting a first speech interaction device from the at least one speech interaction device to process the input speech, comprises:

selecting the first speech interaction device from the at least one speech interaction device, and sending the analysis result to the selected first speech interaction device, so that the selected first speech interaction device performs an operation indicated by the analysis result.

5. An apparatus for processing speech, the apparatus comprising:

at least one processor; and

a memory storing instructions, the instructions when executed by the at least one processor, cause the at least one processor to perform operations, the operations comprising:

acquiring, in response to determining at least one speech interaction device in a target speech interaction device set receiving an input speech, a speech feature of the input speech received by a speech interaction device of the at least one speech interaction device; and

selecting, based on the speech feature of the input speech received by the speech interaction device in the at least one speech interaction device, a first speech interaction device from the at least one speech interaction device to process the input speech;

wherein the speech feature comprises sound pressure; and

wherein selecting, based on the speech feature of the input speech received by the speech interaction device in the at least one speech interaction device, the first speech interaction device from the at least one speech interaction device to process the input speech, comprises:

selecting, according to the sound pressure of the input speech received by the speech interaction device in the at least one speech interaction device in descending order, a preset first number of the at least one speech interaction device from the at least one speech interaction device to process the input speech.

6. The apparatus according to claim 5 , wherein the speech feature further comprises loudness; and

the selecting, based on the speech feature of the input speech received by the speech interaction device in the at least one speech interaction device, a first speech interaction device from the at least one speech interaction device to process the input speech, further comprises:

selecting, according to the loudness of the input speech received by the speech interaction device in the at least one speech interaction device in descending order, a preset second number of the at least one speech interaction device from the at least one speech interaction device to process the input speech.

7. The apparatus according to claim 5 , wherein the selecting, based on the speech feature of the input speech received by the speech interaction device in the at least one speech interaction device, a first speech interaction device from the at least one speech interaction device to process the input speech, comprises:

selecting, in response to determining that the input speech comprises a preset wake-up word, based on the speech feature of the input speech received by the speech interaction device in the at least one speech interaction device, the first speech interaction device from the at least one speech interaction device for being woken up so that the woken first speech interaction device processes the input speech.

8. The apparatus according to claim 5 , wherein before the selecting a first speech interaction device from the at least one speech interaction device to process the input speech, the operations further comprise:

analyzing the input speech to obtain an analysis result; and

the selecting a first speech interaction device from the at least one speech interaction device to process the input speech, comprises:

selecting the first speech interaction device from the at least one speech interaction device, and sending the analysis result to the selected first speech interaction device, so that the selected first speech interaction device performs an operation indicated by the analysis result.

9. A non-transitory computer-readable storage medium storing a computer program, the computer program, when executed by one or more processors, causes the one or more processors to perform operations, the operations comprising:

acquiring, in response to determining at least one speech interaction device in a target speech interaction device set receiving an input speech, a speech feature of the input speech received by a speech interaction device of the at least one speech interaction device; and

selecting, based on the speech feature of the input speech received by the speech interaction device in the at least one speech interaction device, a first speech interaction device from the at least one speech interaction device to process the input speech;

wherein the speech feature comprises sound pressure; and

wherein selecting, based on the speech feature of the input speech received by the speech interaction device in the at least one speech interaction device, the first speech interaction device from the at least one speech interaction device to process the input speech, comprises:

selecting, according to the sound pressure of the input speech received by the speech interaction device in the at least one speech interaction device in descending order, a preset first number of the at least one speech interaction device from the at least one speech interaction device to process the input speech.

10. The non-transitory computer-readable storage medium according to claim 9 , wherein the speech feature further comprises loudness; and

wherein selecting, based on the speech feature of the input speech received by the speech interaction device in the at least one speech interaction device, the first speech interaction device from the at least one speech interaction device to process the input speech, further comprises:

selecting, according to the loudness of the input speech received by the speech interaction device in the at least one speech interaction device in descending order, a preset second number of the at least one speech interaction device from the at least one speech interaction device to process the input speech.

11. The non-transitory computer-readable storage medium according to claim 9 , wherein selecting, based on the speech feature of the input speech received by the speech interaction device in the at least one speech interaction device, the first speech interaction device from the at least one speech interaction device to process the input speech, comprises:

selecting, in response to determining that the input speech comprises a preset wake-up word, based on the speech feature of the input speech received by the speech interaction device in the at least one speech interaction device, the first speech interaction device from the at least one speech interaction device for being woken up so that the woken first speech interaction device processes the input speech.

12. The non-transitory computer-readable storage medium according to claim 9 , wherein before selecting the first speech interaction device from the at least one speech interaction device to process the input speech, the method further comprises:

analyzing the input speech to obtain an analysis result; and

wherein selecting the first speech interaction device from the at least one speech interaction device to process the input speech, comprises:

selecting the first speech interaction device from the at least one speech interaction device, and sending the analysis result to the selected first speech interaction device such that the selected first speech interaction device performs an operation indicated by the analysis result.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 30, 2021
From: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.; SHANGHAI XIAODU TECHNOLOGY CO. LTD.
Reel/Frame 056811/0772 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 15, 2019
From: WU, YA
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
Reel/Frame 048616/0678 →
Priority Claims (1)
CN 201810718087.4 · Jun 29, 2018 · national
Continuity (1)
Related Publication 20200005793A1 · Jan 2, 2020