METHOD AND APPARATUS FOR IMPROVED ENTITY EXTRACTION FROM AUDIO CALLS
In a method and apparatus for improved entity extraction in an audio of a conversation or a call, the method includes generating, at a server, from speech data of a conversation between at least two persons, text data and associated preliminary entity prediction data, using an automated speech recognition (ASR) engine comprising one or more neural networks trained via multi-task training. The method further includes identifying, using the text data and associated preliminary entity prediction data, at least one named entity in said speech data.
1 . A computing apparatus for improved entity detection in a conversation, the apparatus comprising:
a processor; and
a memory storing instructions that, when executed by the processor, configure the apparatus to:
generate, at a server, from speech data of a conversation between at least two persons, text data and associated preliminary entity prediction data, using an automated speech recognition (ASR) engine comprising one or more neural networks trained via multi-task training; and
identify, at the server, using the text data and associated preliminary entity prediction data, at least one named entity in said speech data.
2 . The computing apparatus of claim 1 , wherein the instructions further configure the apparatus to generate a call summary comprising the at least one named entity.
3 . The computing apparatus of claim 2 , wherein the instructions further configure the apparatus to send the call summary or the at least one named entity for display on a device remote to the server.
4 . The computing apparatus of claim 3 , wherein the call summary or the at least one named entity is sent for display while the call is active.
5 . A method for improved entity detection in a conversation, the method comprising:
generating, at a server, from speech data of a conversation between at least two persons, text data and associated preliminary entity prediction data, using an automated speech recognition (ASR) engine comprising one or more neural networks trained via multi-task training; and
identifying, at the server, using the text data and associated preliminary entity prediction data, at least one named entity in said speech data.
6 . The method of claim 5 , further comprising generating a call summary comprising the at least one named entity.
7 . The method of claim 6 , further comprising, sending the call summary or the at least one named entity for display on a device remote to the server.
8 . The method of claim 7 , wherein the call summary or the at least one named entity is sent for display while the call is active.
9 . A non-transitory computer-readable storage medium, the computer-readable storage medium including instructions that when executed by a computer, cause the computer to:
generate, at a server, from speech data of a conversation between at least two persons, text data and associated preliminary entity prediction data, using an automated speech recognition (ASR) engine comprising one or more neural networks trained via multi-task training; and
identify, at the server, using the text data and associated preliminary entity prediction data, at least one named entity in said speech data.
10 . The computer-readable storage medium of claim 9 , wherein the instructions further configure the computer to generate a call summary comprising the at least one named entity.
11 . The computer-readable storage medium of claim 10 , wherein the instructions further configure the computer to send the call summary or the at least one named entity for display on a device remote to the server.
12 . The computer-readable storage medium of claim 11 , wherein the call summary or the at least one named entity is sent for display while the call is active.