IP Library Granted Patent US 11,587,362
Granted Patent B2
US 11,587,362 · App. 17/124,420 · Granted Feb 21, 2023

Techniques for determining sign language gesture partially shown in image(s)

Inventors: Jampierre Vieira Rocha (Indaiatuba, BR); Jeniffer Lensk (Votorantim, BR); Marcelo da Costa Ferreira (Campinas, BR)
Assignee: Lenovo (Singapore) Pte. Ltd.
G06V40/28G06F3/017G06F40/10G06F40/56G06N3/02G06V40/107
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,587,362
App. No.
17/124,420
Granted
Feb 21, 2023
Kind
B2
Abstract

In one aspect, a device may include a processor and storage accessible to the processor. The storage may include instructions executable by the processor to receive at least one image that indicates a first gesture being made by a person using a hand-based sign language, with at least part of the first gesture extending out of the image frame of the image. The instructions may then be executable to provide the image to a gesture classifier and to receive plural candidate first text words for the first gesture from the gesture classifier. The instructions may then be executable to use at least a second text word correlated to a second gesture to select one of the candidate first text words, combine the second text word with the selected first text word to establish a text string, and provide the text string to an apparatus different from the device.

Claims (59)

1. A first device, comprising:

at least one processor; and

storage accessible to the at least one processor and comprising instructions executable by the at least one processor to:

present, on a display, a settings graphical user interface (GUI), the settings GUI comprising plural settings, the settings GUI comprising a first setting that is selectable to configure the first device to classify ambiguous gestures using one or more images, the settings GUI comprising a second setting that is selectable to configure the first device to make inferences regarding ambiguous gestures using context;

identify user selection of the first setting;

identify user selection of the second setting;

responsive to identification of user selection of the first setting, configure the first device to classify ambiguous gestures using one or more images;

responsive to identification of user selection of the second setting, configure the first device to make inferences regarding ambiguous gestures using context;

receive one or more first images from a camera, the one or more first images indicating a first gesture that is ambiguous and that is being made by a person using a hand-based sign language, at least part of the ambiguous first gesture extending out of at least one respective image frame of the one or more first images;

classify the ambiguous first gesture using the one or more first images and a sign language gesture classifier established at least in part by an artificial neural network;

receive, from the sign language gesture classifier, plural candidate first text words for the ambiguous first gesture;

use context related to a second text word correlated to a second gesture different from the ambiguous first gesture to infer one of the candidate first text words as an intended word for the ambiguous first gesture;

combine the second text word with the intended word to establish a text string;

provide the text string to an application executing at one or more of: the first device, a second device; and

use the application to execute an operation based on the text string.

2. The first device of claim 1 , wherein at least part of the ambiguous first gesture extends out of each respective image frame of the one or more first images.

3. The first device of claim 1 , wherein natural language understanding is executed to infer one of the candidate first text words as the intended word using the second text word.

4. The first device of claim 1 , wherein the sign language gesture classifier is configured for receiving as input images of respective gestures and providing as output one or more respective text words corresponding to respective gestures from the input.

5. The first device of claim 4 , wherein the sign language gesture classifier uses a database of image frames corresponding to respective gestures to provide the output.

6. The first device of claim 1 , wherein the instructions are further executable to:

use at least the second text word and a third text word correlated to a third gesture different from the first and second gestures to infer one of the candidate first text words as the intended word for the ambiguous first gesture, wherein the second gesture as indicated in one or more second images from the camera was gestured before the ambiguous first gesture and wherein the third gesture as indicated in one or more third images from the camera was gestured after the ambiguous first gesture; and

combine the second and third text words with the intended word for the ambiguous first gesture to establish the text string, the text string comprising the second text word placed before the intended word and comprising the third text word placed after the intended word.

7. The first device of claim 1 , further comprising the camera.

8. The first device of claim 1 , wherein the instructions are further executable to:

provide the text string to the application executing at the first device.

9. The first device of claim 1 , wherein the instructions are further executable to:

provide the text string to the application executing at the second device.

10. The first device of claim 1 , further comprising the display.

11. The first device of claim 1 , wherein the first setting is different from the second setting, the first and second settings being listed vertically on the settings GUI.

12. A method, comprising:

presenting, on a display accessible to a first device, a settings graphical user interface (GUI), the settings GUI comprising plural settings, the settings GUI comprising a first setting that is selectable to configure the first device to classify ambiguous gestures using one or more images, the settings GUI comprising a second setting that is selectable to configure the first device to make inferences regarding ambiguous gestures using context;

identifying user selection of the first setting;

identifying user selection of the second setting;

responsive to identifying user selection of the first setting, configuring the first device to classify ambiguous gestures using one or more images;

responsive to identifying user selection of the second setting, configuring the first device to make inferences regarding ambiguous gestures using context;

providing, at the first device, at least a first image showing a first gesture into a gesture classifier to receive, as inferred output from the gesture classifier, a first text word corresponding to the first gesture;

providing, at the first device, at least a second image partially but not fully showing a second gesture that is ambiguous into the gesture classifier to receive, as inferred output from the gesture classifier, a second text word corresponding to the second gesture, the second gesture being different from the first gesture, the second text word being different from the first text word;

providing, to an application, a text string indicating the first text word and the second text word; and

using the application to execute an operation based on the text string.

13. The method of claim 12 , wherein the gesture classifier extrapolates additional portions of the second gesture extending out of the second image partially but not fully showing the second gesture, and wherein the gesture classifier uses the extrapolation to output the second text word.

14. The method of claim 12 , wherein the first setting is different from the second setting, the first and second settings being listed vertically on the settings GUI.

15. At least one computer readable storage medium (CRSM) that is not a transitory signal, the computer readable storage medium comprising instructions executable by at least one processor to:

present, on a display of a first device, a settings graphical user interface (GUI), the settings GUI comprising plural settings, the settings GUI comprising a first setting that is selectable to configure the first device to classify ambiguous gestures using one or more images, the settings GUI comprising a second setting that is selectable to configure the first device to make inferences regarding ambiguous gestures using context;

identify user selection of the first setting;

identify user selection of the second setting;

responsive to identification of user selection of the first setting, configure the first device to classify ambiguous gestures using one or more images;

responsive to identification of user selection of the second setting, configure the first device to make inferences regarding ambiguous gestures using context;

receive one or more first images from a camera, the one or more first images indicating a first gesture that is ambiguous and that is being made by a person using a hand-based sign language, at least part of the ambiguous first gesture extending out of at least one respective image frame of the one or more first images;

classify the ambiguous first gesture using the one or more first images and a sign language gesture classifier established at least in part by an artificial neural network;

receive, from the sign language gesture classifier, plural candidate first text words for the ambiguous first gesture;

use context related to a second text word correlated to a second gesture different from the ambiguous first gesture to infer one of the candidate first text words as an intended word for the ambiguous first gesture;

combine the second text word with the intended word to establish a text string;

provide the text string to an application executing at one or more of: the first device, a second device; and

use the application to execute an operation based on the text string.

16. The at least one CRSM of claim 15 , wherein at least part of the ambiguous first gesture extends out of each respective image frame of the one or more first images.

17. The at least one CRSM of claim 15 , wherein natural language understanding is executed to infer one of the candidate first text words as the intended word using the second text word.

18. The at least one CRSM of claim 15 , wherein the sign language gesture classifier is configured for receiving as input images of respective gestures and providing as output one or more respective text words corresponding to respective gestures from the input.

19. The at least one CRSM of claim 18 , wherein the sign language gesture classifier uses a database of image frames corresponding to respective gestures to provide the output.

20. The at least one CRSM of claim 15 , wherein the first setting is different from the second setting, the first and second settings being listed vertically on the settings GUI.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 20, 2025
From: LENOVO PC INTERNATIONAL LIMITED
To: LENOVO SWITZERLAND INTERNATIONAL GMBH
Reel/Frame 070269/0092 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 19, 2025
From: LENOVO (SINGAPORE) PTE LTD.
To: LENOVO PC INTERNATIONAL LIMITED
Reel/Frame 070266/0906 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 22, 2020
From: ROCHA, JAMPIERRE VIEIRA; LENSK, JENIFFER; FERREIRA, MARCELO DA COSTA
To: LENOVO (SINGAPORE) PTE. LTD.
Reel/Frame 054729/0522 →
Continuity (1)
Related Publication 20220188538A1 · Jun 16, 2022