IP Library Granted Patent US 12,190,866
Granted Patent B2
US 12,190,866 · App. 17/707,029 · Granted Jan 7, 2025

Voice based manual image review

Inventor: Lee D Roche (Apex, NC)
Assignee: Conduent Business Services, LLC
G10L15/08G06F3/0481G10L15/22G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,190,866
App. No.
17/707,029
Granted
Jan 7, 2025
Kind
B2
Abstract

Methods and systems for manual-based image review can involve associating an image with a group of keyword utterances, the image displayable in a display screen of a computing device, the group of keyword utterances including different keyword utterances. A prompt for the user to utter a keyword utterance can be displayed in a first area of the image in the display screen and another prompt for the user to utter another keyword utterance can be displayed in another area of the image in the display screen. Audio of the keyword utterances displayed in the display screen can be captured and processed by natural language processing (NLP), when uttered by the user. The utterances can be displayed respectively as text in the first area and the other area of the image in response to processing by NLP of the audio. Thus, instead of users typing in the results and changing their focus between screen and keyboard, for example, the user can speak to the results, which increases the throughput of the results.

Claims (38)

1. A method for manual-based image review, comprising:

associating an image with a plurality of keyword utterances, the image displayable in a display screen of a computing device, the plurality of keyword utterances including different keyword utterances, wherein each area of the display screen displays at least one keyword utterance among the plurality of keyword utterances;

displaying for a user, a prompt for the user to utter the at least one keyword utterance among the plurality of keyword utterances in a first area of the image displayed in the display screen and another prompt for the user to utter at least one other keyword utterance among the plurality of keyword utterances in another area of the image displayed in the display screen;

capturing and processing by natural language processing (NLP), audio of the at least one keyword utterance and the at least one other keyword utterance when uttered by the user through an audio device operable to operate with the NLP, for display of the at least one keyword utterance and the at least one other keyword utterance as text in the respective first area and the another area of the image, in response to processing by NLP of the audio, wherein the NLP is operable to understand the at least one keyword utterance and post NLP results resulting from the capturing and the processing by the NLP;

wherein the NLP understands the at least one keyword utterance and the at least one other keyword utterance uttered by the user posts NLP results to a back office after the processing of the audio; and

the NLP comprises at least one of: NLP based on a hybrid sequence-to-sequence approach, and NLP processing review and override based on confidence analysis.

2. The method of claim 1 wherein the capturing and the processing by NLP of the at least one other keyword utterance further comprises: detecting the at least one other keyword utterance when uttered by the user.

3. The method of claim 1 wherein:

the capturing and the processing by NLP of the at least one other keyword utterance further comprises: detecting the at least one other keyword utterance when uttered by the user; and

the NLP understands the at least one keyword utterance and the at least one other keyword utterance uttered by the user posts NLP results to a back office after the processing of the audio.

4. The method of claim 1 wherein the image comprises an image of a vehicle.

5. The method of claim 4 wherein the plurality of keyword utterances associated with the image includes at least one of: a license plate number of the vehicle, a state of the vehicle, and a type of the vehicle.

6. The method of claim 1 wherein the image comprises an image of a vehicle and wherein the plurality of keyword utterances associated with the image includes at least one of: a license plate number of the vehicle, a state of the vehicle, and a type of the vehicle.

7. The method of claim 1 wherein the image comprises an image of a vehicle and wherein the plurality of keyword utterances associated with the image includes a license plate number of the vehicle, a state of the vehicle, and a type of the vehicle.

8. A system for manual-based image review, comprising:

a display screen for displaying an image associated with a plurality of keyword utterances, wherein the plurality of keyword utterances includes different keyword utterances;

a user interface for displaying for a user, a prompt for the user to utter at least one keyword utterance among the plurality of keyword utterances in a first area of the image displayed in the display screen and another prompt for the user to utter at least one other keyword utterance among the plurality of keyword utterances in another area of the image displayed in the display screen; and

an audio device for capturing and processing by natural language processing (NLP), audio of the at least one keyword utterance and the at least one other keyword utterance when uttered by the user through an audio device operable to operate with the NLP, for display of the at least one keyword utterance and the at least one other keyword utterance as text in the respective first area and the another area of the image, in response to processing by NLP of the audio, wherein the NLP is operable to understand the at least one keyword utterance and post NLP results resulting from the capturing and the processing by the NLP;

wherein the NLP understands the at least one keyword utterance and the at least one other keyword utterance uttered by the user posts NLP results to a back office after the processing of the audio; and

the NLP comprises at least one of: NLP based on a hybrid sequence-to-sequence approach, and NLP processing review and override based on confidence analysis.

9. The system of claim 8 wherein the capturing and the processing by NLP of the at least one other keyword utterance involves detecting the at least one other keyword utterance when uttered by the user.

10. The system of claim 8 wherein:

the capturing and the processing by NLP of the at least one other keyword utterance further comprises: detecting the at least one other keyword utterance when uttered by the user; and

the NLP understands the at least one keyword utterance and the at least one other keyword utterance uttered by the user posts NLP results to a back office after the processing of the audio.

11. The system of claim 8 wherein the image comprises an image of a vehicle.

12. The system of claim 11 wherein the plurality of keyword utterances associated with the image includes at least one of: a license plate number of the vehicle, a state of the vehicle, and a type of the vehicle.

13. The system of claim 11 wherein the image comprises an image of a vehicle and wherein the plurality of keyword utterances associated with the image includes at least one of: a license plate number of the vehicle, a state of the vehicle, and a type of the vehicle.

14. The system of claim 11 wherein the image comprises an image of a vehicle and wherein the plurality of keyword utterances associated with the image includes a license plate number of the vehicle, a state of the vehicle, and a type of the vehicle.

15. A computer program product for facilitating manual-based image review of images, the computer program product comprising one or more non-transitory computer readable storage media and program instructions collectively stored on the one or more non-transitory computer readable storage media, the program instructions comprising program instructions to:

associate an image with a plurality of keyword utterances, the image displayable in a display screen of a computing device, the plurality of keyword utterances including different keyword utterances;

display for a user, a prompt for the user to utter at least one keyword utterance among the plurality of keyword utterances in a first area of the image displayed in the display screen and another prompt for the user to utter at least one other keyword utterance among the plurality of keyword utterances in another area of the image displayed in the display screen; and

capture and process by natural language processing (NLP), audio of the at least one keyword utterance and the at least one other keyword utterance when uttered by the user through an audio device operable to operate with the NLP, for display of the at least one keyword utterance and the at least one other keyword utterance as text in the respective first area and the another area of the image, in response to processing by NLP of the audio, wherein the NLP is operable to understand the at least one keyword utterance and post NLP results resulting from the capturing and the processing by the NLP;

wherein the NLP understands the at least one keyword utterance and the at least one other keyword utterance uttered by the user posts NLP results to a back office after the processing of the audio; and

the NLP comprises at least one of: NLP based on a hybrid sequence-to-sequence approach, and NLP processing review and override based on confidence analysis.

16. The computer program product of claim 15 wherein the program instructions further comprise program instructions to detect the at least one other keyword utterance when uttered by the user.

17. The computer program product of claim 15 wherein:

the image comprises an image of a vehicle; and

the plurality of keyword utterances associated with the image includes at least one of: a license plate number of the vehicle, a state of the vehicle, and a type of the vehicle.

Assignments (3)
SECURITY AGREEMENT Recorded Oct 20, 2025
From: CONDUENT BUSINESS SERVICES, LLC
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 073114/0679 →
SECURITY AGREEMENT Recorded Aug 26, 2025
From: CONDUENT BUSINESS SERVICES, LLC
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 072556/0233 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 29, 2022
From: ROCHE, LEE D
To: CONDUENT BUSINESS SERVICES, LLC
Reel/Frame 059424/0832 →
Continuity (1)
Related Publication 20230317061A1 · Oct 5, 2023
References Cited (21)
US 9082038B2 · Kozitsky et al. · 2015 [cited by applicant]
US 9501707B2 · Bulan et al. · 2016 [cited by applicant]
US 9965677B2 · Bulan et al. · 2018 [cited by applicant]
US 10026004B2 · Mizes et al. · 2018 [cited by applicant]
US 10909845B2 · Bernal et al. · 2021 [cited by applicant]
US 10929661B1 · Manyam · 2021 [cited by examiner]
US 11093487B2 · Erpenbach et al. · 2021 [cited by applicant]
US 20050084134A1 · Toda · 2005 [cited by examiner]
US 20110202338A1 · Inghelbrecht · 2011 [cited by examiner]
US 20170136631A1 · Li · 2017 [cited by examiner]
US 20170262723A1 · Kozitsky et al. · 2017 [cited by applicant]
US 20170358295A1 · Roux et al. · 2017 [cited by applicant]
US 20180350229A1 · Yigit · 2018 [cited by examiner]
US 20190228276A1 · Lei · 2019 [cited by examiner]
US 20210097306A1 · Crary et al. · 2021 [cited by applicant]
US 20210232762A1 · Munro et al. · 2021 [cited by applicant]
US 20220375235A1 · Tatematsu · 2022 [cited by examiner]
EP 2887333B1 · 2018 [cited by applicant]
Satadal Saha, “A Review on Automatic License Plate Recognition System”; Students' Technical Article Competition: PRAYAS-2018, Apr. 29, 2018. [cited by applicant]
Orhan Bulan, et al., “Segmentation- and Annotation-Free License Plate Recognition With Deep Localization and Failure Identification”; IEEE Transactions on Intelligent Transportation Systems. 2017. [cited by applicant]
Wikipedia, “Natural Language Processing”, page last edited Jan. 30, 2022. [cited by applicant]