IP Library › Granted Patent US 12,300,247
Granted Patent B2
US 12,300,247 · App. 17/601,428 · Granted May 13, 2025

Voice-based social network

Inventor: Pramod Kumar Verma (Calabasas, CA)
G10L15/30G10L15/22G10L15/26G06Q50/01G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,300,247
App. No.
17/601,428
Granted
May 13, 2025
Kind
B2
Abstract

This invention presents a novel voice-based social network, where users can compose, explore, and share voice posts. Each voice post is composed of audio, text with dictation or transcription from speech, and other optional elements such as picture, video, contact, etc. During the composition step, the user speaks to the microphone, and the system generates text using the text-to-speech method. Users optionally attach a picture or video and category. Each voice post is visualized as a text on the top of the picture as an overlay. Text is highlighted with a synced part-of the speech. Users can explore posts using search interfaces using keywords and categories. Users can also comment using voice posts. This system also provides advanced interfaces such as recommendation interface where users can see related posts, connection interface where users can connect each other, message interface where users can communicate with each other via voice messages.

Claims (87)

1. A voice-based social networking system, leveraging voice-only posts, where each post is augmented with voice-over and automatically generated or inferred related elements, including images, hashtags, and text transcriptions from speech using Artificial Intelligence/Machine Learning (AI/ML), operating on one or more interconnected computing devices via network interfaces, comprising:

(a) a user interface system module executing on client devices, the user interface configured to send and receive data from an Application Programming Interface (API) module and an object server, further comprising:

i. a signup interface;

ii. a recent feed interface for displaying recent voice posts;

iii. an explore interface for performing search queries;

iv. a search results interface for displaying search outcomes;

v. a post visualization interface for viewing individual voice posts;

vi. a comment interface for displaying voice comments;

vii. a profile interface for viewing a user's post feed;

viii. an add-post interface for creating new voice posts;

ix. a messaging interface;

x. a notification interface for displaying notifications; and

xi. a settings interface;

(b) an Application Programming Interface (API) server module, configured to send and receive data from a social-graph and database server module and a data processor module, and to provide endpoints to client devices;

(c) an object server module, configured to send and receive object data, including video, audio, files, and images, from the user interface module and data processor;

(d) a social-graph and database server module, configured to store and retrieve processed data from the Application Programming Interface (API) module and data processor module, maintaining social graph information and other relevant data in the database; and

(e) a data processor module, configured to exchange data with the social-graph and database module, the Application Programming Interface (API) server module, and the object server module.

2. The voice-based social networking system of claim 1 , wherein the user interface system further comprises:

(a) a connection interface for managing friends, followers, following, interests, networks, and location; and

(b) a connection history interface for analyzing past interactions and connections.

3. The voice-based social networking system of claim 1 , wherein the search interface further comprises:

(a) an action interface for the page, enabling filtering, sorting, group actions, changing views, and playing voice or video posts;

(b) a filter interface for refining search results to display voice posts;

(c) a sort interface for organizing search results based on voice posts; and

(d) a group interface for executing group actions on search results, including the option to play all voice posts.

4. The voice-based social networking system of claim 1 , wherein the signup interface further comprises:

(a) a login interface that navigates to a home or tab interface;

(b) a guest interface for guest login;

(c) a signup interface for creating a new account using basic user information, including username, full name, password, network, location, about, contact details, and address;

(d) a network interface for selecting the user's public/private network and location settings; and

(e) a reset password interface, along with links to account related information.

5. The voice-based social networking system of claim 1 , wherein the recent interface further comprises:

(a) recently added voice posts, tags, and categories, with or without voice wave visualization; and

(b) a query interface for making searches using keywords and navigating to the search interface.

6. The voice-based social networking system of claim 1 , wherein the search interface further comprises:

(a) a query interface for searching with keywords and navigating to the search interface;

(b) a search results interface displaying voice posts based on the query, with navigation to the page interface;

(c) a page interface displaying page information, including layout, visualization of voice waveforms, and animations of augmented voice with transcribed or dictated post titles and voice comments; and

(d) an action interface for the page, providing actions including voice comment, send voice message, share, save, like, rate, replay/play voice/video, report, widget actions, and form actions.

7. The voice-based social networking system of claim 1 , wherein the page interface further comprises:

(a) a page layout featuring visualization of the voice waveform, animation of augmented voice with dictated or transcribed post titles, and voice comments; and

(b) an action interface providing actions on the page, including voice comment, send voice message, share, save, like, rate, replay/play voice/video, report, widget actions, and form actions.

8. The voice-based social networking system of claim 1 , wherein the profile interface further comprises:

(a) a posts interface displaying the user's voice posts based on queries, with navigation to the page interface;

(b) a button for adding a new post, navigating to the add-post interface to create a new voice post; and

(c) an action interface providing actions for the user's voice posts, including share, edit, delete, and view statistics.

9. The voice-based social networking system of claim 1 , wherein the add-post interface further comprises:

(a) a basic input interface for entering a title/description, cover picture, voice, category/tag, and settings, including language, privacy, and location;

(b) a speech-to-text interface for dictating or transcribing user speech, with navigation from the add-post interface, and saving the voice input;

(c) an input text interface for editing or correcting errors in the text generated by the speech-to-text interface; and

(d) an action interface providing actions for the user's voice post, including share, edit, delete, and view statistics.

10. The voice-based social networking system of claim 1 , wherein the message interface further comprises:

(a) a list of voice messages, with navigation to a thread interface containing voice message threads;

(b) a thread interface displaying communication threads for a given voice post;

(c) a compose interface within the thread, allowing users to compose messages using a speech-to-text interface with an option to correct errors via an input text interface; and

(d) a link to the related post on the page interface through the thread interface.

11. The voice-based social networking system of claim 1 , wherein the settings interface further comprises:

(a) an option to update user information, including username, full name, password, link or website, age, and profile photo;

(b) a navigation option to advanced settings interfaces;

(c) a navigation option to the notification interface;

(d) a navigation option to the language settings interface;

(e) a navigation option to the logout button;

(f) a navigation option to the delete or deactivate account button;

(g) an option to change or customize the voice wave animation; and

(h) an advanced settings interface providing navigation to legal settings and additional advanced options.

12. The voice-based social networking system of claim 1 , wherein the user interface further comprises a statistics interface for displaying visitor information for posts added to the search index.

13. The voice-based social networking system of claim 1 , wherein the user interface further comprises:

(a) an advertisement interface for placing ads in the search index, with an option to set payment information and create ad posts for available search categories; and

(b) a payment and store interface for in-app purchases.

14. The voice-based social networking system of claim 1 , wherein the comment interface further comprises:

(a) a search interface for searching comments;

(b) an action interface for interacting with comments;

(c) a visualization of comments in the form of voice posts, with navigation to the post page; and

(d) a user interface for viewing additional comments.

15. The voice-based social networking system of claim 1 , wherein the add-post interface further comprises:

(a) a file input interface;

(b) a video input interface;

(c) a 3D input interface;

(d) a map input interface;

(e) a form input interface;

(f) a Uniform Resource Locator (URL) input interface; and

(g) a contact input interface.

16. The voice-based social networking system of claim 1 , wherein the page interface further comprises a recommendation interface for providing recommended voice posts for a given voice post.

17. The voice-based social networking system of claim 1 , wherein the Artificial Intelligence/Machine Learning (AI/ML) module is configured to support multi-language support for posting a voice post in any language of choice.

18. The voice-based social networking system of claim 1 , wherein the data server is further configured to execute an image fitting algorithm.

19. The voice-based social networking system of claim 1 , wherein the data server is further configured to infer missing information, including tags (hashtags), and categories.

20. The voice-based social networking system of claim 1 , wherein the system is configured to detect spam by error correction and identifying discrepancies between user-edited transcribed text and the original audio.

Continuity (1)
Related Publication 20220208196A1 · Jun 30, 2022
References Cited (37)
US 8345934B2 · Obrador · 2013 [cited by examiner]
US 9336512B2 · Outerbridge · 2016 [cited by examiner]
US 10410108B2 · Shaji · 2019 [cited by examiner]
US 10637811B2 · Outerbridge · 2020 [cited by examiner]
US 20110206191A1 · Tengler · 2011 [cited by examiner]
US 20120014560A1 · Obrador · 2012 [cited by examiner]
US 20120209902A1 · Outerbridge · 2012 [cited by examiner]
US 20140039871A1 · Crawford · 2014 [cited by examiner]
US 20150052209A1 · Vorotyntsev · 2015 [cited by examiner]
US 20150149321A1 · Salameh · 2015 [cited by examiner]
US 20150332067A1 · Gorod · 2015 [cited by examiner]
US 20150363001A1 · Malzbender · 2015 [cited by examiner]
US 20170019363A1 · Outerbridge · 2017 [cited by examiner]
US 20180039879A1 · Shaji · 2018 [cited by examiner]
Maribeth Back, Jonathan Cohen, Rich Gold, Steve Harrison, and Scott Minneman. 2001. Listen Reader: An Electronically Augmented Paper-based Book. In Proceedings of the SIGCHI Conference on Human Factors in Computing Syst… [cited by applicant]
Joëlle Bitton, Stefan Agamanolis, and Matthew Karau. 2004. RAW: Conveying Minimally-mediated Impressions of Everyday Life with an Audio-photographic Tool. In Proceedings of the SIGCHI Conference on Human Factors in Comp… [cited by applicant]
Robin N. Brewer, Leah Findlater, Joseph ‘Jofish’ Kaye, Walter Lasecki, Cosmin Munteanu, and Astrid Weber. 2018. Accessible Voice Interfaces. In Companion of the 2018 ACM Conference on Computer Supported Cooperative Work… [cited by applicant]
Ruy Cervantes and Nithya Sambasivan. 2008. VoiceList: User-driven Telephone-based Audio Content. In Proceedings of the 10th International Conference on Human Computer Interaction with Mobile Devices and Services (Mobile… [cited by applicant]
David Frohlich and Ella Tallyn. 1999. Audiophotography: Practice and Prospects. In CHI '99 Extended Abstracts on Human Factors in Computing Systems (CHI EA '99). ACM, New York, NY, USA, 296-297. DOI: http://dx.doi.org/1… [cited by applicant]
Xuedong Huang, James Baker, and Raj Reddy. 2014. A Historical Perspective of Speech Recognition. Commun. ACM 57, 1 (Jan. 2014), 94-103. DOI: http://dx.doi.org/10.1145/2500887. [cited by applicant]
Scott R. Klemmer, Jamey Graham, Gregory J. Wolff, and James A. Landay. 2003. Books with Voices: Paper Transcripts as a Physical Interface to Oral Histories. In Proceedings of the SIGCHI Conference on Human Factors in Co… [cited by applicant]
Gilad Mishne, David Carmel, Ron Hoory, Alexey Roytman, and Aya Soffer. 2005. Automatic Analysis of Call-center Conversations. In Proceedings of the 14th ACM International Conference on Information and Knowledge Manageme… [cited by applicant]
Cosmin Munteanu and Gerald Penn. 2017. Speech-based Interaction: Myths, Challenges, and Opportunities. In Proceedings of the 2017 CHI Conference Extended Abstracts on Human Factors in Computing Systems (CHI EA '17). ACM… [cited by applicant]
Neil Patel, Sheetal Agarwal, Nitendra Rajput, Amit Nanavati, Paresh Dave, and Tapan S. Parikh. 2009. A Comparative Study of Speech and Dialed Input Voice Interfaces in Rural India. In Proceedings of the SIGCHI Conferenc… [cited by applicant]
Neil Patel, Deepti Chittamuru, Anupam Jain, Paresh Dave, and Tapan S. Parikh. 2010. Avaaj Otalo: A Field Study of an Interactive Voice Forum for Small Farmers in Rural India. In Proceedings of the SIGCHI Conference on H… [cited by applicant]
Jennifer Pearson, Simon Robinson, and Matt Jones. 2015. PaperChains: Dynamic Sketch+Voice Annotations. In Proceedings of the 18th ACM Conference on Computer Supported Cooperative Work & Social Computing (CSCW '15). … [cited by applicant]
Madeline Plauché and Udhyakumar Nallasamy. 2007. Speech Interfaces for Equitable Access to Information Technology. Inf. Technol. Int. Dev. 4, 1 (Oct. 2007), 69-86. DOI: http://dx.doi.org/10.1162/itid.2007.4.1.69. [cited by applicant]
Anand Rajaraman and Jeffrey David Ullman. 2011. Mining of Massive Datasets. Cambridge University Press, New York, NY, USA. [cited by applicant]
Lisa J. Stifelman, Barry Arons, Chris Schmandt, and Eric A. Hulteen. 1993. VoiceNotes: A Speech Interface for a Hand-held Voice Notetaker. In Proceedings of the Interact '93 and CHI '93 Conference on Human Factors in Co… [cited by applicant]
Altman, N. S. 1992. An Introduction to Kernel and Nearest- Neighbor Nonparametric Regression. The American Statis-tician 46(3): 175-185. ISSN 00031305. URL http://www. jstor.org/stable/2685209. [cited by applicant]
Anantha, R.; Chappidi, S.; and Dawoodi, A. W. 2020. Learn-ing to Rank Intents in Voice Assistants. URL https://arxiv.org/pdf/2005.00119.pdf. [cited by applicant]
Capes, T.; Coles, P.; Conkie, A.; Golipour, L.; Hadjitarkhani, A.; Hu, Q.; Huddleston, N.; Hunt, M.; Li, J.; Neeracher, M.; Prahallad, K.; Raitio, T.; Rasipuram, R.; Townsend, G.; Williamson, B.; Winarsky, D.; Wu, Z.; a… [cited by applicant]
Chen, X. C.; Sagar, A.; Kao, J. T.; Li, T. Y.; Klein, C.; Pul-man, S.; Garg, A.; and Williams, J. D. 2019. Active Learning for Domain Classification in a Commercial Spoken Personal Assistant. URL https://arxiv.org/pdf/1… [cited by applicant]
Gysel, C. V.; Tsagkias, M.; Pusateri, E.; and Oparin, I. 2020. Predicting Entity Popularity to Improve Spoken En-tity Recognition by Virtual Assistants. URL https://arxiv. org/pdf/2005.12816.pdf. [cited by applicant]
Lipton, Z. C. 2015. A Critical Review of Recurrent Neural Networks for Sequence Learning. CoRR abs/1506.00019. URL http://arxiv.org/abs/1506.00019. [cited by applicant]
McAllaster, G. M. K. A.-H. 2019. Bandwidth Embed-dings for Mixed-Bandwidth Speech Recognition. URL https://arxiv.org/pdf/1909.02667.pdf. [cited by applicant]
Rajaraman, A.; and Ullman, J. D. 2011. ing, 1-17. Cambridge University Press. CBO9781139058452.002. Data Min—doi:10.1017/. [cited by applicant]