IP Library › Granted Patent US 12,658,184
Granted Patent B2
US 12,658,184 · App. 16/041,405 · Granted Jun 16, 2026

Visualization interface for voice input

Inventors: Ferhan Ture (Washington, DC); Md Iftekhar Tanveer (Rochester, NY)
Assignee: Comcast Cable Communication LLC
G10L15/1822G06F16/322G06F16/358G10L15/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,658,184
App. No.
16/041,405
Filed
Jul 20, 2018
Granted
Jun 16, 2026
Kind
B2
Examiner
HE, JIALONG
Art Unit
2659
USPC
704/257
Abstract

Voice inputs may be analyzed based on syntactic properties and respective dependency tree structure (e.g., parse tree) may be generated. Dependency tree structures may be grouped/clustered based on an associated root word/phrase. A visual interface may display the dependency tree structure in a manner that maps each voice input to a proper response (e.g., an action, an operation, a command, etc.). Mapping voice inputs to responses may be used to train a neural network.

Claims (66)

1 . A method comprising:

determining, based on a first vectorization of a first root word and a first plurality of dependent words of a received natural language voice input, a first intent associated with the received natural language voice input;

determining, by an artificial neural network, based on the first intent, a second intent associated with a stored voice input, wherein the stored voice input is associated with a content-based response;

causing output, by the artificial neural network, via an interactive content-based response user interface, and based on the first intent and the second intent, of a cluster comprising the received natural language voice input and the content-based response; and

causing output, based on a user interaction configured to expand the cluster, of an expanded cluster.

2 . The method of claim 1 , wherein the content-based response comprises a deep learning response.

3 . The method of claim 1 , wherein determining the first intent comprises mapping at least a portion of the received natural language voice input to stored text.

4 . The method of claim 1 , further comprising determining, based on the stored voice input one or more of: a command or an action.

5 . The method of claim 4 , wherein determining the command or the action comprises mapping the stored voice input to stored text.

6 . The method of claim 1 , wherein the cluster comprises one or more selectable subclusters, and wherein the one or more selectable subclusters are associated with one or more additional content-based responses.

7 . The method of claim 6 , further comprising:

receiving, via the one or more selectable subclusters, one or more user inputs; and

displaying, based on the one or more user inputs, one or more visualization interfaces associated with the one or more additional content-based responses.

8 . The method of claim 1 , wherein the stored voice input is associated with statistical information, wherein the statistical information is based on a quantity of voice inputs received by a voice enabled device.

9 . The method of claim 1 , further comprising:

comparing the first intent and the second intent;

determining, based on the comparison of the first intent and the second intent, a similarity between the first intent and the second intent; and

causing output, based on the similarity, a visualization of the first intent and the second intent.

10 . The method of claim 1 , further comprising receiving a user input configured to expand the cluster.

11 . A system comprising:

a computing device associated with an artificial neural network configured to:

determine, based on a first vectorization of a first root word and a first plurality of dependent words of a received natural language voice input, a first intent associated with the received natural language voice input;

determine, by an artificial neural network, based on the first intent, a second intent associated with a stored voice input, wherein the stored voice input is associated with a content-based response;

cause output, by the artificial neural network, via an interactive content-based response user interface, and based on the first intent and the second intent, of a cluster comprising the received natural language voice input and the content-based response; and

cause output, based on a user interaction configured to expand the cluster, an expanded cluster; and

a user device configured to:

send the natural language voice input.

12 . The system of claim 11 , wherein the content-based response comprises a deep learning response.

13 . The system of claim 11 , wherein the computing device is configured to determine the first intent by mapping at least a portion of the received natural language voice input to stored text.

14 . The system of claim 11 , wherein the computing device is further configured to determine, based on the stored voice input one or more of: a command or an action.

15 . The system of claim 14 , wherein the computing device is configured to determine the command or the action by mapping the stored voice input to stored text.

16 . The system of claim 11 , wherein the cluster comprises one or more selectable subclusters, and wherein the one or more selectable subclusters are associated with one or more additional content-based responses.

17 . The system of claim 16 , wherein the computing device is further configured to:

receive, via the one or more selectable subclusters, one or more user inputs; and

display, based on the one or more user inputs, one or more visualization interfaces associated with the one or more additional content-based responses.

18 . The system of claim 11 , wherein the stored voice input is associated with statistical information, wherein the statistical information is based on a quantity of voice inputs received by a voice enabled device.

19 . An apparatus comprising:

one or more processors; and

memory storing processor-executable instructions that, when executed by the one or more processors, cause the one or more processors to:

determine, based on a first vectorization of a first root word and a first plurality of dependent words of a received natural language voice input, a first intent associated with the received natural language voice input;

determine, by an artificial neural network, based on the first intent, a second intent associated with a stored voice input, wherein the stored voice input is associated with a content-based response;

cause output, by the artificial neural network, via an interactive content-based response user interface, and based on the first intent and the second intent, of a cluster comprising the received natural language voice input and the content-based response; and

cause output, based on a user interaction configured to expand the cluster, an expanded interactive cluster.

20 . The apparatus of claim 19 , wherein the content-based response comprises a deep learning response.

21 . The apparatus of claim 19 , wherein the processor-executable instructions that, when executed by the one or more processors, cause the one or more processors to determine the first intent further cause the one or more processors to map at least a portion of the received natural language voice input to stored text.

22 . The apparatus of claim 19 , wherein the processor-executable instructions, when executed by the one or more processors, further cause the one or more processors to determine, based on the stored voice input one or more of: a command or an action.

23 . The apparatus of claim 22 , wherein the processor-executable instructions that, when executed by the one or more processors, cause the one or more processors to determine the command or the action further cause the one or more processors to map the stored voice input to stored text.

24 . The apparatus of claim 19 , wherein the cluster comprises one or more selectable subclusters, and wherein the one or more selectable subclusters are associated with one or more additional content-based responses.

25 . The apparatus of claim 24 , wherein the processor-executable instructions, when executed by the one or more processors, further cause the one or more processors to:

receive, via the one or more selectable subclusters, one or more user inputs; and

display, based on the one or more user inputs, one or more visualization interfaces associated with the one or more additional content-based responses.

26 . The apparatus of claim 19 , wherein the stored voice input is associated with statistical information, wherein the statistical information is based on a quantity of voice inputs received by a voice enabled device.

27 . One or more non-transitory computer-readable media storing processor-executable instructions thereon, that, when executed by at least one processor, cause the at least one processor to:

determine, based on a first vectorization of a first root word and a first plurality of dependent words of a received natural language voice input, a first intent associated with the received natural language voice input;

determine, by an artificial neural network, based on the first intent, a second intent associated with a stored voice input, wherein the stored voice input is associated with a content-based response;

cause output, by the artificial neural network, via an interactive content-based response user interface, and based on the first intent and the second intent, of a cluster comprising the received natural language voice input and the content-based response; and

cause output, based on a user interaction configured to expand the cluster, an expanded cluster.

28 . The one or more non-transitory computer-readable media of claim 27 , wherein the content-based response comprises a deep learning response.

29 . The one or more non-transitory computer-readable media of claim 27 , wherein the processor-executable instructions that, when executed by the at least one processor, cause the at least one processor to determine the first intent further cause the at least one processor to map at least a portion of the received natural language voice input to stored text.

30 . The one or more non-transitory computer-readable media of claim 27 , wherein the processor-executable instructions, when executed by the at least one processor, further cause the at least one processor to determine, based on the stored voice input one or more of: a command or an action.

31 . The one or more non-transitory computer-readable media of claim 30 , wherein the processor-executable instructions that, when executed by the at least one processor, cause the at least one processor to determine the command or the action further cause the at least one processor to map the stored voice input to stored text.

32 . The one or more non-transitory computer-readable media of claim 27 , wherein the cluster comprises one or more selectable subclusters, and wherein the one or more selectable subclusters are associated with one or more additional content-based responses.

33 . The one or more non-transitory computer-readable media of claim 32 , wherein the processor-executable instructions, when executed by the at least one processor, further cause the at least one processor to:

receive, via the one or more selectable subclusters, one or more user inputs; and

display, based on the one or more user inputs, one or more visualization interfaces associated with the one or more additional content-based responses.

34 . The one or more non-transitory computer-readable media of claim 27 , wherein the stored voice input is associated with statistical information, wherein the statistical information is based on a quantity of voice inputs received by a voice enabled device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 31, 2020
From: TURE, FERHAN; TANVEER, MD IFTEKHAR
To: COMCAST CABLE COMMUNICATIONS, LLC
Reel/Frame 053645/0913 →
Continuity (1)
Related Publication 20200027446A1 · Jan 23, 2020
References Cited (21)
US 5832428A · Chow et al. · 1998 [cited by applicant]
US 8694488B1 · Garg · 2014 [cited by examiner]
US 20040111253A1 · Luo et al. · 2004 [cited by applicant]
US 20060168515A1 · Dorsett, Jr. · 2006 [cited by examiner]
US 20110238409A1 · Larcheveque · 2011 [cited by examiner]
US 20110238410A1 · Larcheveque · 2011 [cited by examiner]
US 20140052444A1 · Roberge · 2014 [cited by examiner]
US 20160259851A1 · Hopkins · 2016 [cited by examiner]
US 20170069310A1 · Hakkani-Tur · 2017 [cited by examiner]
US 20180182381A1 · Singh · 2018 [cited by examiner]
US 20180300310A1 · Shinn · 2018 [cited by examiner]
US 20200020318A1 · Canada · 2020 [cited by examiner]
US 20200050940A1 · Li · 2020 [cited by examiner]
Socher, Richard, et al. “Parsing with compositional vector grammars.” Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (vol. 1: Long Papers). 2013. (Year: 2013). [cited by examiner]
Tanveer, Md Iftekhar, and Ferhan Türe. “SyntaViz: Visualizing Voice Queries through a Syntax-Driven Hierarchical Ontology.” Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System … [cited by examiner]
Derrick, Donald, and Daniel Archambault. “TreeForm: Explaining and exploring grammar through syntax trees.” Literary and linguistic computing 25.1 (2010): 53-66. (Year: 2010). [cited by examiner]
Almeida-Martínez, Francisco J., and Jaime Urquiza-Fuentes. “Syntax trees visualization in language processing courses.” 2009 Ninth IEEE International Conference on Advanced Learning Technologies. IEEE, 2009. (Year: 2009… [cited by examiner]
Bird et al. “Natural Language Processing with Python”, published by O'reilly, 2009. (Year: 2009). [cited by examiner]
Petrov, “Announcing SyntaxNet: The World's Most Accurate Parser Goes Open Source”, [online] https://research.google/blog/announcing-syntaxnet-the-worlds-most-accurate-parser-goes-open-source/, published in 2016. (Year: … [cited by examiner]
Tanveer et al: “SyntaViz: Visualizing Voice Queries through a Syntax-Driven Hierarchical Ontology”, Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, (2018), … [cited by applicant]
European Search Report and Written Opinion were mailed on Nov. 27, 2019 by the European Patent Office for EP Application No. 19187405.6, filed on Jul. 20, 2018 and published as EP 3598436 on Jan. 22, 2020 (Applicant—Com… [cited by applicant]