IP Library Granted Patent US 11,158,102
Granted Patent B2
US 11,158,102 · App. 16/668,963 · Granted Oct 26, 2021

Method and apparatus for processing information

Inventors: Xiao Liu (Beijing, CN); Fuqiang Lyu (Beijing, CN); Jianxiang Wang (Beijing, CN); Jianchao Ji (Beijing, CN)
Assignee: Beijing Baidu Netcom Science and Technology Co., Ltd.
G06T13/205G06F3/011G06F3/017G06F3/16G06K9/00302G06T13/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,158,102
App. No.
16/668,963
Granted
Oct 26, 2021
Kind
B2
Abstract

Embodiments of the present disclosure provide a method and apparatus for processing information. A method may include: generating voice response information based on voice information sent by a user; generating a phoneme sequence based on the voice response information; generating mouth movement information based on the phoneme sequence, the mouth movement information being used for controlling a mouth movement of a displayed three-dimensional human image when playing the voice response information; and playing the voice response information, and controlling the mouth movement of the three-dimensional human image based on the mouth movement information.

Claims (74)

1. A method for processing information, comprising:

generating voice response information to voice information sent by a user based on the voice information sent by the user;

generating a phoneme sequence based on the voice response information;

using each of first phonemes in the phoneme sequence as a target phoneme, matching the target phoneme with second phonemes in a pre-stored corresponding relationship table to generate a mouth movement parameter sequence including mouth movement parameters and corresponding to the phoneme sequence, and using the mouth movement parameter sequence as mouth movement information, the pre-stored corresponding relationship table recording the second phonemes and the mouth movement parameters, the mouth movement information being used for controlling a mouth movement of a displayed three-dimensional human image when playing the voice response information; and

playing the voice response information, and controlling the mouth movement of the three-dimensional human image based on the mouth movement information.

2. The method according to claim 1 , wherein the method further comprises:

acquiring a video of the user captured when the user is sending the voice information;

performing, for a video frame of the video, facial expression identification on a face image in the video frame, to obtain an expression identification result; and

playing the video, and presenting, in a played current video frame, an expression identification result corresponding to a face image in the current video frame.

3. The method according to claim 2 , wherein before the playing the video, the method further comprises:

receiving a request for face image decoration sent by the user, wherein the request for face image decoration comprises decorative image selection information;

selecting a target decorative image from a preset decorative image set based on the decorative image selection information; and

adding the target decorative image to the video frame of the video.

4. The method according to claim 3 , wherein the adding the target decorative image to the video frame of the video comprises:

selecting a video frame from the video at intervals of a first preset frame number, to obtain at least one video frame; and

performing, for a video frame of the at least one video frame, face key point detection of a face image in the video frame, to obtain positions of face key points; and adding the target decorative image to the video frame and a second preset frame number of video frames after the video frame based on the positions of the face key points in the video frame.

5. The method according to claim 1 , wherein the method further comprises:

generating gesture change information based on the phoneme sequence, the gesture change information being used for controlling a gesture change of the displayed three-dimensional human image when playing the voice response information; and

the playing the voice response information, and controlling the mouth movement of the three-dimensional human image based on the mouth movement information comprises:

playing the voice response information, and controlling the mouth movement and the gesture change of the three-dimensional human image based on the mouth movement information and the gesture change information.

6. The method according to claim 1 , wherein the method further comprises:

generating to-be-displayed information based on the voice information, and displaying the to-be-displayed information.

7. The method according to claim 1 , wherein the method further comprises:

determining a target service type based on the voice information; and

determining target expression information based on the target service type, and controlling an expression of the three-dimensional human image based on the target expression information.

8. An apparatus for processing information, comprising:

at least one processor; and

a memory storing instructions, wherein the instructions when executed by the at least one processor, cause the at least one processor to perform operations, the operations comprising:

generating voice response information to voice information sent by a user based on the voice information sent by the user;

generating a phoneme sequence based on the voice response information;

using each of first phonemes in the phoneme sequence as a target phoneme, matching the target phoneme with second phonemes in a pre-stored corresponding relationship table to generate a mouth movement parameter sequence including mouth movement parameters and corresponding to the phoneme sequence, and using the mouth movement parameter sequence as mouth movement information, the pre-stored corresponding relationship table recording the second phonemes and the mouth movement parameters, the mouth movement information being used for controlling a mouth movement of a displayed three-dimensional human image when playing the voice response information; and

playing the voice response information, and control the mouth movement of the three-dimensional human image based on the mouth movement information.

9. The apparatus according to claim 8 , wherein the operations further comprise:

acquiring a video of the user captured when the user is sending the voice information;

performing, for a video frame of the video, facial expression identification on a face image in the video frame, to obtain an expression identification result; and

playing the video, and presenting, in a played current video frame, an expression identification result corresponding to a face image in the current video frame.

10. The apparatus according to claim 9 , wherein before the playing the video, the operations further comprise:

receiving a request for face image decoration sent by the user, wherein the request for face image decoration comprises decorative image selection information;

selecting a target decorative image from a preset decorative image set based on the decorative image selection information; and

adding the target decorative image to the video frame of the video.

11. The apparatus according to claim 10 , wherein the adding the target decorative image to the video frame of the video comprises:

selecting a video frame from the video at intervals of a first preset frame number, to obtain at least one video frame; and

performing, for a video frame of the at least one video frame, face key point detection of a face image in the video frame, to obtain positions of face key points; and adding the target decorative image to the video frame and a second preset frame number of video frames after the video frame based on the positions of the face key points in the video frame.

12. The apparatus according to claim 8 , wherein the operations further comprise:

generating gesture change information based on the phoneme sequence, the gesture change information being used for controlling a gesture change of the displayed three-dimensional human image when playing the voice response information; and

the playing the voice response information, and controlling the mouth movement of the three-dimensional human image based on the mouth movement information comprises:

playing the voice response information, and controlling the mouth movement and the gesture change of the three-dimensional human image based on the mouth movement information and the gesture change information.

13. The apparatus according to claim 8 , wherein the operations further comprise:

generating to-be-displayed information based on the voice information, and displaying the to-be-displayed information.

14. The apparatus according to claim 8 , wherein the operations further comprise:

determining a target service type based on the voice information; and

determining target expression information based on the target service type, and controlling an expression of the three-dimensional human image based on the target expression information.

15. A non-transitory computer readable medium, storing a computer program thereon, wherein the program, when executed by a processor, causes the processor to perform operations, the operations comprising:

generating voice response information to voice information sent by a user based on the voice information sent by the user;

generating a phoneme sequence based on the voice response information;

using each of first phonemes in the phoneme sequence as a target phoneme, matching the target phoneme with second phonemes in a pre-stored corresponding relationship table to generate a mouth movement parameter sequence including mouth movement parameters and corresponding to the phoneme sequence, and using the mouth movement parameter sequence as mouth movement information, the pre-stored corresponding relationship table recording the second phonemes and the mouth movement parameters, the mouth movement information being used for controlling a mouth movement of a displayed three-dimensional human image when playing the voice response information; and

playing the voice response information, and control the mouth movement of the three-dimensional human image based on the mouth movement information.

16. The non-transitory computer readable medium according to claim 15 , wherein the operations further comprise:

acquiring a video of the user captured when the user is sending the voice information;

performing, for a video frame of the video, facial expression identification on a face image in the video frame, to obtain an expression identification result; and

playing the video, and presenting, in a played current video frame, an expression identification result corresponding to a face image in the current video frame.

17. The non-transitory computer readable medium according to claim 16 , wherein before the playing the video, the operations further comprise:

receiving a request for face image decoration sent by the user, wherein the request for face image decoration comprises decorative image selection information;

selecting a target decorative image from a preset decorative image set based on the decorative image selection information; and

adding the target decorative image to the video frame of the video.

18. The non-transitory computer readable medium according to claim 17 , wherein the adding the target decorative image to the video frame of the video comprises:

selecting a video frame from the video at intervals of a first preset frame number, to obtain at least one video frame; and

performing, for a video frame of the at least one video frame, face key point detection of a face image in the video frame, to obtain positions of face key points; and adding the target decorative image to the video frame and a second preset frame number of video frames after the video frame based on the positions of the face key points in the video frame.

19. The non-transitory computer readable medium according to claim 15 , wherein the operations further comprise:

generating gesture change information based on the phoneme sequence, the gesture change information being used for controlling a gesture change of the displayed three-dimensional human image when playing the voice response information; and

the playing the voice response information, and controlling the mouth movement of the three-dimensional human image based on the mouth movement information comprises:

playing the voice response information, and controlling the mouth movement and the gesture change of the three-dimensional human image based on the mouth movement information and the gesture change information.

20. The non-transitory computer readable medium according to claim 15 , wherein the operations further comprise:

generating to-be-displayed information based on the voice information, and displaying the to-be-displayed information.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 30, 2019
From: LIU, XIAO; LYU, FUQIANG; WANG, JIANXIANG; JI, JIANCHAO
To: BEIJING BAIDU NETCOM SCIENCE AND TECHNOLOGY CO., LTD.
Reel/Frame 050885/0714 →
Priority Claims (1)
CN 201910058552.0 · Jan 22, 2019 · national
Continuity (1)
Related Publication 20200234478A1 · Jul 23, 2020
Cited By (4)
US 12,361,621 US 12,591,947 US 12,602,849 US 12,682,436