IP Library › Granted Patent US 12,283,283
Granted Patent B2
US 12,283,283 · App. 18/324,365 · Granted Apr 22, 2025

Machine learning-based audio codec switching

Inventors: Hsin Fu Henry Chiang (Bellevue, WA); Yasmin Karimli (Kirkland, WA); Ming Shan Kwok (Seattle, WA)
Assignee: T-Mobile USA, Inc.
G10L19/22G10L19/0208G10L25/51H04W76/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,283,283
App. No.
18/324,365
Granted
Apr 22, 2025
Kind
B2
Abstract

Described herein are techniques, devices, and systems for selectively using a music-capable audio codec on-demand during a communication session. A user equipment (UE) may adaptively transition between using a first audio codec that provides a first audio bandwidth and a second audio codec (e.g., the EVS-FB codec) that provides a second audio bandwidth that is greater than the first audio bandwidth. The transition to the second audio codec may occur in response to determining that sound in the environment of the UE includes frequencies outside of a range of frequencies associated with a human voice, such as by determining that music is being played in the environment of the UE, which allows for selectively using a music-capable audio codec when it would be beneficial to do so.

Claims (62)

1. A computer-implemented method comprising:

establishing, by a user equipment (UE), and via a serving base station, a voice call using a first Enhanced Voice Services (EVS) audio codec that provides a first audio bandwidth;

determining, by the UE, that music is being played in an environment of the UE;

sending, by the UE, and based at least in part on the determining that the music is being played in the environment, a message to the serving base station for transitioning from using the first EVS audio codec to using a second EVS audio codec that provides a second audio bandwidth greater than the first audio bandwidth; and

continuing, by the UE, and via the serving base station, the voice call using the second EVS audio codec.

2. The computer-implemented method of claim 1 , wherein the second EVS audio codec is an EVS Full Band (EVS-FB) codec.

3. The computer-implemented method of claim 1 , further comprising, in response to the determining that the music is being played in the environment:

outputting, via the UE, a user prompt indicating that the UE detected background music and associated with transitioning from using the first EVS audio codec to using the second EVS audio codec; and

receiving user input via the UE to transition from using the first EVS audio codec to using a second EVS audio codec,

wherein the sending of the message occurs in response to the receiving of the user input.

4. The computer-implemented method of claim 1 , further comprising, after the continuing the voice call using the second EVS audio codec:

determining, by the UE, that the music is no longer being played in the environment;

sending, by the UE, and based at least in part on the determining that the music is no longer being played in the environment, a second message to the serving base station for transitioning from using the second EVS audio codec to using the first EVS audio codec; and

continuing, by the UE, and via the serving base station, the voice call using the first EVS audio codec.

5. The computer-implemented method of claim 4 , further comprising, in response to the determining that the music is no longer being played in the environment:

outputting, via the UE, a second user prompt indicating that the UE has ceased detecting background music and associated with transitioning from using the second EVS audio codec to using the first EVS audio codec; and

receiving second user input via the UE to transition from using the second EVS audio codec to using the first EVS audio codec,

wherein the sending of the second message occurs in response to the receiving of the second user input.

6. The computer-implemented method of claim 1 , further comprising:

determining, by the UE, and during the voice call, that a value indicative of a radio frequency (RF) condition is equal to or greater than a threshold value,

wherein the sending of the message is further based on the determining that the value indicative of the RF condition is equal to or greater than the threshold value.

7. A user equipment (UE) comprising:

a processor; and

memory storing computer-executable instructions that, when executed by the processor, cause the UE to:

establish, via a serving base station, a communication session using a first audio codec that provides a first audio bandwidth;

determine that sound in an environment of the UE includes frequencies outside of a range of frequencies associated with a human voice;

send, based at least in part on determining that the sound includes the frequencies outside of the range of frequencies associated with the human voice, a message to the serving base station for transitioning from using the first audio codec to using a second audio codec that provides a second audio bandwidth greater than the first audio bandwidth; and

continue, via the serving base station, the communication session using the second audio codec.

8. The UE of claim 7 , wherein the second audio codec is an Enhanced Voice Services Full Band (EVS-FB) codec.

9. The UE of claim 7 , wherein the computer-executable instructions, when executed by the processor, further cause the UE to, in response to determining that the sound includes the frequencies outside of the range of frequencies associated with the human voice:

output a user prompt associated with transitioning from using the first audio codec to using the second audio codec; and

receive user input to transition from using the first audio codec to using a second audio codec,

wherein sending the message occurs in response to receiving the user input.

10. The UE of claim 9 , wherein:

the user prompt requests a selection of a bit rate among multiple available bit rates to use with the second audio codec;

receiving the user input comprises receiving the selection of the bit rate as a selected bit rate; and

continuing the communication session using the second audio codec comprises using the second audio codec at the selected bit rate.

11. The UE of claim 7 , wherein sending the message occurs without user intervention in response to determining that the sound includes the frequencies outside of the range of frequencies associated with the human voice.

12. The UE of claim 7 , wherein the message includes a capability indicator indicating that the UE supports the second audio codec, and wherein continuing the communication session using the second audio codec is based at least in part on a second UE involved in the communication session also supporting the second audio codec.

13. The UE of claim 7 , wherein determining that the sound includes the frequencies outside of the range of frequencies associated with the human voice comprises determining that music is being played in the environment.

14. The UE of claim 13 , wherein the determining that the music is being played in the environment comprises:

providing audio data based on the sound as input to a trained machine learning model; and

generating, as output from the trained machine learning model, a probability that a source of the sound is not the human voice.

15. A computer-implemented method comprising:

establishing, by a user equipment (UE), and via a serving base station, a communication session using a first audio codec that provides a first audio bandwidth;

determining, by the UE, that sound in an environment of the UE includes frequencies outside of a range of frequencies associated with a human voice;

sending, by the UE, and based at least in part on determining that the sound includes the frequencies outside of the range of frequencies associated with the human voice, a message to the serving base station for transitioning from using the first audio codec to using a second audio codec that provides a second audio bandwidth greater than the first audio bandwidth; and

continuing, by the UE, and via the serving base station, the communication session using the second audio codec.

16. The computer-implemented method of claim 15 , wherein:

the first audio codec is at least one of:

an Enhanced Voice Services Wideband (EVS-WB) codec; or

an Enhanced Voice Services Super Wideband (EVS-SWB) codec; and

the second audio codec is an Enhanced Voice Services Full Band (EVS-FB) codec.

17. The computer-implemented method of claim 15 , further comprising:

determining, by the UE, and during the communication session, that a value indicative of a radio frequency (RF) condition is equal to or greater than a threshold value,

wherein the sending of the message is further based on the determining that the value indicative of the RF condition is equal to or greater than the threshold value.

18. The computer-implemented method of claim 15 , wherein the communication session is established using the first audio codec as a default codec.

19. The computer-implemented method of claim 15 , wherein the determining that the sound includes the frequencies outside of the range of frequencies associated with the human voice comprises determining that music is being played in the environment.

20. The computer-implemented method of claim 19 , further comprising, after the continuing the communication session using the second audio codec:

determining, by the UE, that the music is no longer being played in the environment;

sending, by the UE, and based at least in part on the determining that the music is no longer being played in the environment, a second message to the serving base station for transitioning from using the second audio codec to using the first audio codec; and

continuing, by the UE, and via the serving base station, the communication session using the first audio codec.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 26, 2023
From: CHIANG, HSIN FU HENRY; KARIMLI, YASMIN; KWOK, MING SHAN
To: T-MOBILE USA, INC.
Reel/Frame 065018/0689 →
Continuity (2)
Continuation 17115463 · Dec 8, 2020
Related Publication 20230298605A1 · Sep 21, 2023
References Cited (21)
US 9591048B2 · Torgersrud · 2017 [cited by examiner]
US 9860766B1 · Pawar · 2018 [cited by examiner]
US 10231282B1 · Lee · 2019 [cited by examiner]
US 10904798B2 · Choi · 2021 [cited by examiner]
US 11252612B2 · Chang · 2022 [cited by examiner]
US 20140067405A1 · Patel et al. · 2014 [cited by applicant]
US 20170092288A1 · Dewasurendra et al. · 2017 [cited by applicant]
US 20170289319A1 · Kwok · 2017 [cited by examiner]
US 20170303114A1 · Johansson · 2017 [cited by examiner]
US 20170346954A1 · Wang et al. · 2017 [cited by applicant]
US 20180206087A1 · Guan et al. · 2018 [cited by applicant]
US 20180302515A1 · Li · 2018 [cited by examiner]
US 20180376004A1 · Kodali · 2018 [cited by examiner]
US 20200236154A1 · Ragot et al. · 2020 [cited by applicant]
US 20220180883A1 · Chiang et al. · 2022 [cited by applicant]
US 20220182495A1 · Kwok et al. · 2022 [cited by applicant]
US 20220394133A1 · Kwok · 2022 [cited by applicant]
Dietz, et al., “Overview of the EVS Codec Architecture”, 2015 IEEE International Conference on Accoustics, Speech, and Signal Processing ( ICASSP), doi: 10,1109/ICASSP.2015.7179063, Apr. 2015, pp. 5698-5702. [cited by applicant]
Eksler, et al., “Audio Bandwidth Detection in the EVS Codec”, 2015 IEEE Global Conference on Signal ind Information Processing, (GlobalSIP), 2015, 10.1109/GlobalSIP.2015.7418243, Dec. 2015, pp. 488-492. [cited by applicant]
Kang, et al., “Improvement of Speech/Music Classification for 3GPP EVS Based on LSTM”, Symmetry 2018, 10,605, doi:10.3390/sym10110605, Nov. 2018, 8 pgs. [cited by applicant]
Office Action for U.S. Appl. No. 17/115,463, mailed on Jul. 21, 2022, Chang, “Machine Learning-Based Audio Codec Switching”, 22 pages. [cited by applicant]