IP Library Granted Patent US 10,311,858
Granted Patent B1
US 10,311,858 · App. 15/385,493 · Granted Jun 4, 2019

Method and system for building an integrated user profile

Inventors: Bernard Mont-Reynaud (Sunnyvale, CA); Jun Huang (Fremont, CA); Kiran Garaga Lokeswarappa (Mountain View, CA); Joel Gedalius (Baltimore, MD)
Assignee: SoundHound, Inc.
G10L15/02G06F17/274G06F17/2705G06N20/00G06Q30/0276G10L15/063G10L15/1815G10L25/90H04L67/306G10L2015/025
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,311,858
App. No.
15/385,493
Granted
Jun 4, 2019
Kind
B1
Abstract

A system and method are provided for adding user characterization information to a user profile by analyzing user's speech. User properties such as age, gender, accent, and English proficiency may be inferred by extracting and deriving features from user speech, without the user having to configure such information manually. A feature extraction module that receives audio signals as input extracts acoustic, phonetic, textual, linguistic, and semantic features. The module may be a system component independent of any particular vertical application or may be embedded in an application that accepts voice input and performs natural language understanding. A profile generation module receives the features extracted by the feature extraction module and uses classifiers to determine user property values based on the extracted and derived features and store these values in a user profile. The resulting profile variables may be globally available to other applications.

Claims (78)

1. A non-transitory computer-readable medium storing instructions, which when executed by a processor of a car navigation system, cause the processor to:

extract an acoustic feature from user speech;

derive a second acoustic feature from the extracted acoustic feature;

use an accent classifier on the derived second acoustic feature to determine a value of an accent property;

assign the value to the accent property in a user profile; and

generate, by the car navigation system, voice instruction addressed to the user, the voice instruction bearing an accent that corresponds to the value of the accent property.

2. The non-transitory computer-readable medium of claim 1 , further comprising instructions that, when executed by a processor of a car navigation system, cause the car navigation system to:

extract a linguistic feature from user speech;

determine a value of a region property in the user profile based on the extracted linguistic feature; and

generate, by the car navigation system, voice instructions addressed to the user, the voice instructions comprising words from a vocabulary that corresponds to the value of the region property.

3. The non-transitory computer-readable medium of claim 1 , wherein the classifier is trained using machine learning.

4. The non-transitory computer-readable medium of claim 1 , wherein the classifier uses pitch statistics.

5. The non-transitory computer-readable medium of claim 1 , wherein the extracted acoustic feature comprises at least one of: a cepstrum and a spectrogram.

6. The non-transitory computer-readable medium of claim 1 , further comprising instructions that, when executed by a processor of a car navigation system, cause the car navigation system to:

compare the value of the accent property to location context information to determine compatibility,

wherein the assigning of the value to the accent property in the user profile is performed responsive to determining that the value of the accent property is compatible with the location context information.

7. A method of advertisement selection in an online ad system, the method comprising:

extracting an acoustic feature from user speech;

deriving a second acoustic feature from the extracted acoustic feature;

using an accent classifier on the derived second acoustic feature to determine a value of an accent property;

assigning the value to the accent property in a user profile; and

selecting, by the online ad system, an ad for display, the ad being aimed at users of a specific demographic group indicated by the value of the accent property.

8. The method of claim 7 , further comprising:

extracting a linguistic feature from user speech;

determining a value of a region property in the user profile based on the extracted linguistic feature; and

selecting, by the online ad system, an ad for display, the ad comprising words from a vocabulary that corresponds to the value of the region property.

9. The method of claim 7 , wherein the classifier is trained using machine learning.

10. The method of claim 7 , wherein the classifier uses pitch statistics.

11. The method of claim 7 , wherein the extracted acoustic feature comprises at least one of: a cepstrum and a spectrogram.

12. The method of claim 7 , further comprising:

comparing the value of the accent property to location context information to determine compatibility,

wherein the assigning of the value to the accent property in the user profile is responsive to determining that the value of the accent property is compatible with the location context information.

13. The method of claim 7 , wherein the ad is a music ad.

14. A method of advertisement selection in an online ad system, the method comprising:

extracting an acoustic feature from user speech;

deriving a second acoustic feature from the extracted acoustic feature;

using an age classifier on the derived second acoustic feature to determine a value of an age range property;

assigning the value to the age range property in a user profile; and

selecting, by the online ad system, an ad for display, the ad being aimed at users having an age within the age range assigned to the age range property.

15. The method of claim 14 , further comprising:

extracting a linguistic feature from user speech;

determining the value of the age range property in the user profile based on the extracted linguistic feature; and

selecting, by the online ad system, an ad for display, the ad comprising words from a vocabulary that corresponds to the value of the age range property.

16. The method of claim 14 , wherein the classifier is trained using machine learning.

17. The method of claim 14 , wherein the classifier uses pitch statistics.

18. The method of claim 14 , wherein the extracted acoustic feature comprises at least one of: a cepstrum and a spectrogram.

19. The method of claim 14 , wherein the ad is a music ad.

20. A non-transitory computer-readable medium storing instructions, which when executed by a processor of an online add system, cause the processor to:

extract an acoustic feature from user speech;

derive a second acoustic feature from the extracted acoustic feature;

use an age classifier on the derived second acoustic feature to determine a value of an age range property;

assign the value to the age range property in a user profile; and

select, by the online ad system, an ad for display, the ad being aimed at users having an age within the age range assigned to the age range property.

21. The car navigation system of claim 20 , further comprising instructions that, when executed by the one or more processors, cause the car navigation system implement actions comprising:

extracting a linguistic feature from user speech;

determining a value of a region property in the user profile based on the extracted linguistic feature; and

generating, by the car navigation system, voice instructions addressed to the user, the voice instructions comprising words from a vocabulary that corresponds to the value of the region property.

22. The car navigation system of claim 20 , wherein the classifier is trained using machine learning.

23. The car navigation system of claim 20 , wherein the classifier uses pitch statistics.

24. The car navigation system of claim 20 , wherein the extracted acoustic feature comprises at least one of: a cepstrum and a spectrogram.

25. The car navigation system of claim 20 , further comprising instructions that, when executed by the one or more processors, cause the car navigation system implement actions comprising:

comparing the value of the accent property to location context information to determine compatibility, wherein the assigning of the value to the accent property in the user profile is performed responsive to determining that the value of the accent property is compatible with the location context information.

26. A non-transitory computer-readable medium storing instructions for advertisement selection, which when executed by a processor of an online ad system, cause the processor to implement a method comprising:

extracting an acoustic feature from user speech;

deriving a second acoustic feature from the extracted acoustic feature;

using an accent classifier on the derived second acoustic feature to determine a value of an accent property;

assigning the value to the accent property in a user profile; and

selecting, by the online ad system, an ad for display, the ad being aimed at users of a specific demographic group indicated by the value of the accent property.

27. The non-transitory computer-readable medium of claim 26 , wherein the method implemented by the processor further comprises:

extracting a linguistic feature from user speech;

determining a value of a region property in the user profile based on the extracted linguistic feature; and

selecting, by the online ad system, an ad for display, the ad comprising words from a vocabulary that corresponds to the value of the region property.

28. The non-transitory computer-readable medium of claim 26 , wherein the classifier is trained using machine learning.

29. The non-transitory computer-readable medium of claim 26 , wherein the classifier uses pitch statistics.

30. The non-transitory computer-readable medium of claim 26 , wherein the extracted acoustic feature comprises at least one of: a cepstrum and a spectrogram.

31. The non-transitory computer-readable medium of claim 26 , the method further comprising:

comparing the value of the accent property to location context information to determine compatibility, wherein the assigning of the value to the accent property in the user profile is responsive to determining that the value of the accent property is compatible with the location context information.

32. The non-transitory computer-readable medium of claim 26 , wherein the ad is a music ad.

Assignments (10)
SECURITY INTEREST Recorded Aug 9, 2024
From: SOUNDHOUND, INC.
To: MONROE CAPITAL MANAGEMENT ADVISORS, LLC, AS COLLATERAL AGENT
Reel/Frame 068526/0413 →
RELEASE OF SECURITY INTEREST Recorded Jun 11, 2024
From: ACP POST OAK CREDIT II LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
Reel/Frame 067698/0845 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2023
From: SOUNDHOUND AI IP HOLDING, LLC
To: SOUNDHOUND AI IP, LLC
Reel/Frame 064205/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2023
From: SOUNDHOUND, INC.
To: SOUNDHOUND AI IP HOLDING, LLC
Reel/Frame 064083/0484 →
RELEASE OF SECURITY INTEREST Recorded Apr 21, 2023
From: FIRST-CITIZENS BANK & TRUST COMPANY, AS AGENT
To: SOUNDHOUND, INC.
Reel/Frame 063411/0396 →
RELEASE OF SECURITY INTEREST Recorded Apr 19, 2023
From: OCEAN II PLO LLC, AS ADMINISTRATIVE AGENT AND COLLATERAL AGENT
To: SOUNDHOUND, INC.
Reel/Frame 063380/0625 →
SECURITY INTEREST Recorded Apr 17, 2023
From: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
To: ACP POST OAK CREDIT II LLC
Reel/Frame 063349/0355 →
CORRECTIVE ASSIGNMENT TO CORRECT THE COVER SHEET PREVIOUSLY RECORDED AT REEL: 056627 FRAME: 0772. ASSIGNOR(S) HEREBY CONFIRMS THE SECURITY INTEREST. Recorded Apr 12, 2023
From: SOUNDHOUND, INC.
To: OCEAN II PLO LLC, AS ADMINISTRATIVE AGENT AND COLLATERAL AGENT
Reel/Frame 063336/0146 →
SECURITY INTEREST Recorded Jun 18, 2021
From: OCEAN II PLO LLC, AS ADMINISTRATIVE AGENT AND COLLATERAL AGENT
To: SOUNDHOUND, INC.
Reel/Frame 056627/0772 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2016
From: MONT-REYNAUD, BERNARD; HUANG, JUN; LOKESWARAPPA, KIRAN GARAGA; GEDALIUS, JOEL
To: SOUNDHOUND, INC.
Reel/Frame 040696/0708 →
Continuity (2)
Continuation 14704833 · May 5, 2015
Provisional Application 61992172 · May 12, 2014