IP Library Granted Patent US 11,527,242
Granted Patent B2
US 11,527,242 · App. 16/610,254 · Granted Dec 13, 2022

Lip-language identification method and apparatus, and augmented reality (AR) device and storage medium which identifies an object based on an azimuth angle associated with the AR field of view

Inventors: Naifu Wu (Beijing, CN); Xitong Ma (Beijing, CN); Lixin Kou (Beijing, CN); Sha Feng (Beijing, CN)
Assignee: Beijing BOE Technology Development Co., Ltd.
G10L15/22G06V20/20G06V40/171G10L15/1815G10L15/25
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,527,242
App. No.
16/610,254
Granted
Dec 13, 2022
Kind
B2
Abstract

A lip-language identification method and an apparatus thereof, an augmented reality device and a storage medium. The lip-language identification method includes: acquiring a sequence of face images for an object to be identified; performing lip-language identification based on a sequence of face images so as to determine semantic information of speech content of the object to be identified corresponding to lip actions in a face image; and outputting the semantic information.

Claims (46)

1. A lip-language identification method based on an augmented reality device, the augmented reality device comprising a camera device and an infrared sensor, the method comprising:

acquiring, by the augmented reality device, a sequence of face images for an object to be identified;

sending, by the augmented reality device, the sequence of face images to a server;

performing, by the server, lip-language identification based on the sequence of face images, so as to determine semantic information of speech content of the object to be identified corresponding to lip actions in the face images; and

receiving, by the augmented reality device, the semantic information sent by the server and outputting the semantic information,

wherein acquiring the sequence of face images for the object to be identified, comprises:

acquiring a sequence of images including the object to be identified;

positioning the object to be identified and acquiring azimuth of the object to be identified; and

determining a position of a face region of the object to be identified in each frame of image in the sequence of images according to the positioned azimuth of the object to be identified; and generating the sequence of face images by cropping an image of the face region of the object to be identified from each frame of the images; and

wherein positioning the azimuth of the object to be identified, comprises:

positioning the azimuth of the object to be identified according to a voice signal emitted when the object to be identified is speaking, and

positioning the azimuth of the object to be identified by sensing the object to be identified through the infrared sensor;

wherein the azimuth of the object to be identified is an angle between the position of the object to be identified and a central axis of the field of view range of the camera device;

wherein the semantic information is semantic text information and/or semantic audio information;

wherein outputting the semantic information comprises:

displaying, by the augmented reality device, the semantic text information within a visual field of a user wearing the augmented reality device, in response to receiving a display mode instruction; and

playing, by the augmented reality device, the semantic audio information, in response to receiving an audio mode instruction.

2. The lip-language identification method according to claim 1 , further comprising saving the sequence of face images, after acquiring the sequence of face images for the object to be identified.

3. The lip-language identification method according to claim 2 , wherein sending the sequence of face images to the server comprises:

sending the saved sequence of face images to the server upon receiving a sending instruction.

4. A lip-language identification apparatus, comprising:

a processor; and

a machine-readable storage medium, storing instructions that are executed by the processor for performing the lip-language identification method according to claim 1 .

5. A storage medium that stores non-transitorily computer readable instructions that, when executed by a computer, the computer executes instructions for the lip-language identification method according to claim 1 .

6. A lip-language identification apparatus, comprising:

a face image sequence acquiring unit, configured to acquire a sequence of face images for an object to be identified;

a sending unit, configured to send the sequence of face images to a server, wherein the server determines semantic information corresponding to lip actions in the face images by performing lip-language identification; and

a receiving unit, configured to receive semantic information from the server,

an output unit, configured to output semantic information;

wherein the face image sequence acquiring unit comprises:

an image sequence acquiring subunit, configured to acquire a sequence of images for the object to be identified;

a positioning subunit, configured to position an azimuth of the object to be identified; and

a face image sequence generation subunit, configured to determine a position of a face region of the object to be identified in each frame of image in the sequence of images according to the positioned azimuth of the object to be identified; and crop an image of the face region of the object to be identified from the each frame image so as to generate the sequence of face images; and

wherein the positioning subunit is further configured to position the azimuth of the object to be identified according to a voice signal emitted when the object to be identified is speaking, and

position the azimuth of the object to be identified by sensing the object to be identified through an infrared sensor;

wherein the azimuth of the object to be identified is an angle between the position of the object to be identified and a central axis of the field of view range of a camera device;

wherein the output unit comprises:

an output mode instruction generation subunit, configured to generate a display mode instruction, wherein the output mode instruction includes a display mode instruction and an audio mode instruction;

wherein the semantic information is semantic text information and/or semantic audio information, and the output unit further comprises:

a display subunit, configured to display the semantic text information within a visual field of a user wearing an augmented reality device upon receiving the display mode instruction; and

a play subunit, configured to play the semantic audio information upon receiving the audio mode instruction.

7. An augmented reality device, comprising the lip-language identification apparatus according to claim 6 .

8. The augmented reality device according to claim 7 , further comprising a camera device, a display device or a play device;

wherein the camera device is configured to capture an image of the object to be identified;

the display device is configured to display semantic information; and

the play device is configured to play the semantic information.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 3, 2022
From: BOE TECHNOLOGY GROUP CO., LTD.
To: BEIJING BOE TECHNOLOGY DEVELOPMENT CO., LTD.
Reel/Frame 061644/0638 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 1, 2019
From: WU, NAIFU; MA, XITONG; KOU, LIXIN; FENG, SHA
To: BOE TECHNOLOGY GROUP CO., LTD.
Reel/Frame 050891/0739 →
Priority Claims (1)
CN 201810384886.2 · Apr 26, 2018 · national
Continuity (1)
Related Publication 20200058302A1 · Feb 20, 2020
Cited By (1)
US 12,537,020