IP Library Granted Patent US 7,295,975
Granted Patent B1
US 7,295,975 · App. 10/970,215 · Granted Nov 13, 2007

Systems and methods for extracting meaning from multimodal inputs using finite-state devices

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,295,975
App. No.
10/970,215
Granted
Nov 13, 2007
Kind
B1
Abstract

Multimodal utterances contain a number of different modes. These modes can include speech, gestures, and pen, haptic, and gaze inputs, and the like. This invention use recognition results from one or more of these modes to provide compensation to the recognition process of one or more other ones of these modes. In various exemplary embodiments, a multimodal recognition system inputs one or more recognition lattices from one or more of these modes, and generates one or more models to be used by one or more mode recognizers to recognize the one or more other modes. In one exemplary embodiment, a gesture recognizer inputs a gesture input and outputs a gesture recognition lattice to a multimodal parser. The multimodal parser generates a language model and outputs it to an automatic speech recognition system, which uses the received language model to recognize the speech input that corresponds to the recognized gesture input.

Claims (45)

1. Apparatus for recognizing an utterance comprising a plurality of modes, said apparatus comprising:

means for recognizing a first mode in said plurality of modes;

means for outputting a recognition result for said first mode in said plurality of modes;

means for generating a recognition model for use in recognizing a second mode in said plurality of modes, said recognition model a function of said recognition result associated with said first mode in said plurality of modes; and

means for recognizing said second mode in said plurality of modes using said recognition model.

2. The apparatus of claim 1 , wherein said recognition result comprises a first mode recognition lattice.

3. The apparatus of claim 2 , wherein said means for generating a recognition model comprises:

means for inputting said first mode recognition lattice;

means for inputting a first finite-state transducer, said first finite-state transducer relating said first mode to said second mode;

means for outputting a second finite-state transducer, said second finite-state transducer a function of said first mode recognition lattice and said first finite-state transducer.

4. The apparatus of claim 1 , wherein said first mode in said plurality of modes comprises a gesture mode and said second mode in said plurality of modes comprises a speech mode.

5. The apparatus of claim 4 , wherein said means for generating a recognition model comprises:

means for receiving a gesture recognition lattice and a first finite-state transducer, said first finite-state transducer relating a first portion of an utterance comprising said first mode to a second portion of an utterance comprising said second mode; and

means for generating a second finite-state transducer,

wherein said second finite-state transducer comprises a gesture/speech recognition model finite state transducer based on said gesture recognition lattice and said first finite-state transducer.

6. The apparatus of claim 5 further comprising:

means for receiving said second finite-state transducer; and

means for generating a second mode recognition model as a function of said second finite-state transducer.

7. A recognition system for receiving and recognizing an utterance comprising a plurality of modes, the recognition system comprising:

a first mode recognition subsystem adapted to generate a first recognition result associated with a first mode in said plurality of modes;

a multimodal recognition subsystem adapted to generate a recognition model based on said first recognition result; and

a second mode recognition subsystem adapted to generate a second recognition result associated with a second mode in said plurality of modes as a function of said first recognition model.

8. The multimodal recognition subsystem of claim 7 , wherein said first recognition result comprises a first mode recognition lattice.

9. The multimodal recognition subsystem of claim 8 , wherein said multimodal recognition subsystem is adapted to generate a first finite-state transducer, said finite-state transducer a function of said first mode recognition lattice and a second finite-state transducer that relates said first mode to said second mode.

10. The multimodal recognition subsystem of claim 9 , wherein said recognition model comprises a projection of said second finite-state transducer, said projection a function of said second finite-state transducer.

11. The multimodal recognition system of claim 7 , wherein said first mode recognition subsystem comprises a gesture recognition subsystem and said second mode recognition subsystem comprises a speech recognition subsystem.

12. The multimodal recognition subsystem of claim 11 , wherein said first recognition result comprises a gesture recognition lattice.

13. The multimodal recognition subsystem of claim 12 , wherein said recognition model comprises a speech recognition model, said speech recognition model comprising:

a projection of a gesture/speech recognition model finite-state transducer, said gesture/speech recognition model finite state transducer generated being a function of said gesture recognition lattice and said first finite-state transducer.

14. The multimodal recognition subsystem of claim 13 , wherein said speech recognition subsystem is adapted to generate a speech recognition of a speech mode as a function of said speech recognition model.

15. The multimodal recognition system of claim 14 , wherein said speech recognition subsystem comprises:

a speech processing subsystem adapted to generate a feature vector lattice as a function of a speech signal;

a phonetic recognition subsystem adapted to generate a phone lattice as a function of said feature vector lattice and an acoustic model;

a word recognition subsystem adapted to generate a word lattice as a function of said phone lattice and a lexicon lattice; and

a spoken mode recognition subsystem adapted to generate said speech recognition result as a function of said word lattice and said speech recognition model.

16. The multimodal recognition subsystem of claim 13 , wherein said speech recognition model is a speech recognition lattice.

17. The multimodal recognition system of claim 16 wherein said speech recognition lattice is one of a grammar model lattice and a language model lattice.

18. The multimodal recognition system of claim 12 , wherein said gesture recognition subsystem comprises

a gesture feature recognition subsystem adapted to generate a gesture feature lattice as a function of a gesture mode; and

a gesture recognition subsystem adapted to generate said gesture recognition lattice as a function of said gesture feature lattice.

19. The multimodal recognition system of claim 7 , further comprising a plurality of mode input devices, at least two of said plurality of mode input devices adapted to receive different modes.

20. The multimodal recognition system of claim 19 , wherein said plurality of mode input devices comprise at least two of a gesture input device, a speech input device, a pen input device, a computer vision device, a haptic input device, a gaze input device, and a body motion input device.

21. The multimodal recognition system of claim 19 , wherein at least two of said plurality of input devices are combined into a single multimodal input device.

22. The multimodal recognition system of claim 7 , wherein said first mode subsystem comprises at least one of a gesture recognition subsystem, a speech recognition subsystem, a pen input recognition subsystem, a computer vision recognition system, a haptic recognition subsystem, a gaze recognition subsystem, and a body motion recognition system.

23. The multimodal recognition system of claim 7 , wherein said second mode subsystem comprises at least one of a gesture recognition subsystem, a speech recognition subsystem, a pen input recognition subsystem, a computer vision recognition system, a haptic recognition subsystem, a gaze recognition subsystem, and a body motion recognition system.

Assignments (18)
RELEASE OF SECURITY INTEREST Recorded Sep 4, 2025
From: RUNWAY GROWTH FINANCE CORP., AS AGENT
To: INTERACTIONS CORPORATION; INTERACTIONS LLC
Reel/Frame 072802/0931 →
CORRECTIVE ASSIGNMENT TO CORRECT THE THE APPLICATION NUMBER PREVIOUSLY RECORDED AT REEL: 060445 FRAME: 0733. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Feb 1, 2023
From: INTERACTIONS LLC; INTERACTIONS CORPORATION
To: RUNWAY GROWTH FINANCE CORP.
Reel/Frame 062919/0063 →
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL/FRAME: 036100/0925 Recorded Jul 1, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060559/0576 →
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL/FRAME: 049388/0082 Recorded Jun 30, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060558/0474 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jun 27, 2022
From: INTERACTIONS LLC; INTERACTIONS CORPORATION
To: RUNWAY GROWTH FINANCE CORP.
Reel/Frame 060445/0733 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY Recorded May 23, 2022
From: ORIX GROWTH CAPITAL, LLC
To: INTERACTIONS CORPORATION; INTERACTIONS LLC
Reel/Frame 061749/0825 →
RELEASE OF SECURITY INTEREST Recorded May 18, 2020
From: BEARCUB ACQUISITIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 052693/0866 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jun 5, 2019
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 049388/0082 →
ASSIGNMENT OF IP SECURITY AGREEMENT Recorded Nov 17, 2017
From: ARES VENTURE FINANCE, L.P.
To: BEARCUB ACQUISITIONS LLC
Reel/Frame 044481/0034 →
CORRECTIVE ASSIGNMENT TO CORRECT THE CHANGE PATENT 7146987 TO 7149687 PREVIOUSLY RECORDED ON REEL 036009 FRAME 0349. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Nov 17, 2015
From: INTERACTIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 037134/0712 →
FIRST AMENDMENT TO INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jul 13, 2015
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 036100/0925 →
SECURITY INTEREST Recorded Jun 23, 2015
From: INTERACTIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 036009/0349 →
SECURITY INTEREST Recorded Dec 19, 2014
From: INTERACTIONS LLC
To: ORIX VENTURES, LLC
Reel/Frame 034677/0768 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2014
From: AT&T ALEX HOLDINGS, LLC
To: INTERACTIONS LLC
Reel/Frame 034642/0640 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 10, 2014
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: AT&T ALEX HOLDINGS, LLC
Reel/Frame 034467/0822 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 8, 2014
From: BANGALORE, SRINIVAS; JOHNSTON, MICHAEL J.
To: AT&T CORP.
Reel/Frame 034429/0452 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 8, 2014
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 034429/0474 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 8, 2014
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 034429/0467 →