IP Library › Granted Patent US 12,230,288
Granted Patent B2
US 12,230,288 · App. 17/828,116 · Granted Feb 18, 2025

Systems and methods for automated customized voice filtering

Inventors: Jin Zhang (San Mateo, CA); Celeste Bean (San Mateo, CA); Sepideh Karimi (San Mateo, CA); Sudha Krishnamurthy (San Mateo, CA)
Assignees: SONY INTERACTIVE ENTERTAINMENT LLC; SONY INTERACTIVE ENTERTAINMENT INC.
G10L21/013G10L15/187G10L15/22G10L25/51G10L25/90
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,230,288
App. No.
17/828,116
Granted
Feb 18, 2025
Kind
B2
Abstract

Systems and methods for audio processing are described. An audio processing system receives audio content that includes a voice sample. The audio processing system analyzes the voice sample to identify a sound type in the voice sample. The sound type corresponds to pronunciation of at least one specified character in the voice sample. The audio processing system generates a filtered voice sample at least in part by filtering the voice sample to modify the sound type. The audio processing system outputs the filtered voice sample.

Claims (36)

1. An apparatus for audio processing, the apparatus comprising:

at least one memory storing instructions; and

at least one processor that executes the instructions, wherein execution of the instructions by the at least one processor causes the at least one processor to:

receive audio content that includes a voice sample of a voice of a user saying at least one word, the at least one word including a plurality of characters;

analyze the voice sample to identify a sound type in the voice sample, wherein the sound type corresponds to a pronunciation by the user in the voice sample of at least one specified character of the plurality of characters;

generate a filtered voice sample using a personalized filter at least in part by filtering the voice sample to modify the sound type, wherein the personalized filter is customized to the voice of the user based on at least one additional voice sample of the voice of the user; and

output the filtered voice sample.

2. The apparatus of claim 1 , wherein the sound type includes a sibilance.

3. The apparatus of claim 1 , wherein the sound type corresponds to at least a voice type in the pronunciation by the user in the voice sample of the at least one specified character.

4. The apparatus of claim 3 , wherein the voice type corresponds to at least one of a gender or a sex.

5. The apparatus of claim 3 , wherein the voice type corresponds to at least one of an age, an accent, a dialect, or an ethnic background.

6. The apparatus of claim 1 , wherein the sound type corresponds to at least a voice frequency in the pronunciation by the user in the voice sample of the at least one specified character.

7. The apparatus of claim 1 , wherein the sound type corresponds to at least a relative position between a microphone and a person in the pronunciation by the user in the voice sample of the at least one specified character during recording of the audio content using the microphone, wherein the voice sample is spoken by the person.

8. The apparatus of claim 1 , wherein the sound type corresponds to a speech dysfluency, wherein filtering the voice sample includes correcting the speech dysfluency.

9. The apparatus of claim 1 , wherein filtering the voice sample includes applying a filter to the voice sample, wherein the filter includes a de-esser.

10. The apparatus of claim 1 , wherein filtering the voice sample includes applying a filter to the voice sample, wherein the filter includes a compressor that filters a specified frequency range, wherein the audio content includes audio in the specified frequency range, wherein the specified frequency range corresponds to the sound type.

11. The apparatus of claim 1 , wherein filtering the voice sample includes applying a filter to the voice sample, wherein the filter is customized to a voice type corresponding to the voice sample.

12. The apparatus of claim 11 , wherein the filter includes a trained machine learning model that is customized to the voice type, the trained machine learning model having been trained using training data that includes one or more additional voice samples associated with the voice type.

13. The apparatus of claim 12 , wherein the execution of the instructions by the at least one processor causes the at least one processor to:

request at least one of the one or more additional voice samples from a user device, wherein the user device is associated with the voice type.

14. The apparatus of claim 12 , wherein the execution of the instructions by the at least one processor causes the at least one processor to:

update the trained machine learning model using additional training data, wherein the additional training data includes at least the voice sample and the filtered voice sample.

15. The apparatus of claim 1 , wherein outputting the filtered voice sample includes causing the filtered voice sample to be output using an audio output device.

16. The apparatus of claim 1 , wherein outputting the filtered voice sample includes transmitting the filtered voice sample to a recipient device over a communication interface.

17. The apparatus of claim 1 , wherein the apparatus includes a digital signal processor (DSP).

18. A method of audio processing, the method comprising:

receiving audio content that includes a voice sample of a voice of a user saying at least one word, the at least one word including a plurality of characters;

analyzing the voice sample to identify a sound type in the voice sample, wherein the sound type corresponds to a pronunciation by the user in the voice sample of at least one specified character of the plurality of characters;

generating a filtered voice sample using a personalized filter at least in part by filtering the voice sample to modify the sound type, wherein the personalized filter is customized to the voice of the user based on at least one additional voice sample of the voice of the user; and

outputting the filtered voice sample.

19. The method of claim 18 , wherein filtering the voice sample includes applying a filter to the voice sample, wherein the filter is customized to a voice type corresponding to the voice sample.

20. A non-transitory computer readable storage medium having embodied thereon a program, wherein the program is executable by a processor to perform a method of audio processing, the method comprising:

receiving audio content that includes a voice sample of a voice of a user saying at least one word, the at least one word including a plurality of characters;

analyzing the voice sample to identify a sound type in the voice sample, wherein the sound type corresponds to a pronunciation by the user in the voice sample of at least one specified character of the plurality of characters;

generating a filtered voice sample using a personalized filter at least in part by filtering the voice sample to modify the sound type, wherein the personalized filter is customized to the voice of the user based on at least one additional voice sample of the voice of the user; and

outputting the filtered voice sample.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 31, 2022
From: BEAN, CELESTE; ZHANG, JIN; KARIMI, SEPIDEH; KRISHNAMURTHY, SUDHA
To: SONY INTERACTIVE ENTERTAINMENT LLC; SONY INTERACTIVE ENTERTAINMENT INC.
Reel/Frame 060052/0073 →
Continuity (1)
Related Publication 20230410824A1 · Dec 21, 2023
References Cited (19)
US 20070033042A1 · Marcheret et al. · 2007 [cited by applicant]
US 20080111887A1 · Cooper · 2008 [cited by examiner]
US 20090094630A1 · Brown · 2009 [cited by examiner]
US 20110066434A1 · Li · 2011 [cited by examiner]
US 20110300806A1 · Lindahl et al. · 2011 [cited by applicant]
US 20140314261A1 · Selig et al. · 2014 [cited by applicant]
US 20170223471A1 · Foo et al. · 2017 [cited by applicant]
US 20180242098A1 · Pratt et al. · 2018 [cited by applicant]
US 20190259360A1 · Yoelin · 2019 [cited by applicant]
US 20200077190A1 · Solis · 2020 [cited by examiner]
US 20200382883A1 · Woodruff · 2020 [cited by examiner]
US 20200382892A1 · Mehta et al. · 2020 [cited by applicant]
US 20210092534A1 · Fichtl · 2021 [cited by applicant]
US 20230388705A1 · Tinklenberg et al. · 2023 [cited by applicant]
WO WO2023235084 · 2023 [cited by applicant]
WO WO2023235088 · 2023 [cited by applicant]
Application No. PCT/US2023/020519, International Search Report and Written Opinion dated Aug. 1, 2023. [cited by applicant]
Application No. PCT/US2023/020720, International Search Report and Written Opinion dated Aug. 1, 2023. [cited by applicant]
U.S. Appl. No. 17/828,668, Non-Final Office Action dated Aug. 1, 2024. [cited by applicant]