IP Library Granted Patent US 12,437,763
Granted Patent B2
US 12,437,763 · App. 17/623,372 · Granted Oct 7, 2025

Methods to employ concatenation in ASR service usage based on speech-to-text cost

Inventors: Ankur Anil Aher (Maharashtra, IN); Jeffry Copps Robert Jose (Tamil Nadu, IN)
Assignee: Adeia Guides Inc.
G10L15/26G10L15/04G10L15/22G10L15/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,437,763
App. No.
17/623,372
Granted
Oct 7, 2025
Kind
B2
Abstract

Systems and methods for processing audio streams are disclosed herein. In a disclosed method, N number of audio streams is received from an independent source, each audio stream includes speech content. A set of the N audio streams is concatenated to generate a concatenated audio stream. Based on the N received audio streams or the set of audio streams, N−1 or one less that the total number of audio streams in the set of audio stream separators, respectively, are generated. An audio stream separator is inserted between every two adjacent audio streams of the concatenated audio stream to generate a single audio stream payload. The single audio stream payload is transmitted for transcription of the audio stream speech content to text content and in response to transmitting the payload, a text file is received including text content corresponding to the audio streams delineated by the audio stream separators.

Claims (40)

1. A method of processing audio streams, the method comprising:

receiving a plurality of audio streams during a time window, each audio stream received from an independent source, including speech content, and having an audio stream duration;

selecting a subset of the plurality of audio streams based on the respective audio stream durations of the plurality of audio streams, the subset comprising N number of audio streams, wherein a combined duration of the N number of audio streams is below a threshold duration, and wherein the threshold duration is based on a transaction cost for speech-to-text conversion;

generating N−1 number of audio stream separators;

concatenating the N number of audio streams to generate a concatenated audio stream;

inserting an audio stream separator between every two adjacent audio streams of the concatenated audio stream to generate a single audio stream payload, each audio stream separator delineating a beginning of a next audio stream and an end of a preceding audio stream in the concatenated audio stream;

transmitting the single audio stream payload from a buffer for transcription of the speech content of the set of audio streams to text content, a transcription of each of the speech content of the set of audio streams; and

in response to transmitting the single audio stream payload for transcription, receiving a text file including the text content delineated with the audio stream separators.

2. The method of claim 1 , further comprising concatenating the N number of audio streams in the buffer.

3. The method of claim 2 , further comprising inserting the audio stream separator between every two adjacent audio streams of the concatenated audio stream in the buffer.

4. The method of claim 1 , wherein a size of the buffer is based on a multiple of a maximum speech content duration among the durations of the speech content of the N number of audio streams.

5. The method of claim 4 , wherein the maximum speech content duration corresponds to a minimum base transcription price.

6. The method of claim 4 , wherein minimum transcription charges for transcription of the audio streams are based on the maximum speech content duration.

7. The method of claim 1 , further comprising concatenating the N number of audio streams in a storage.

8. The method of claim 7 , further comprising inserting the audio stream separator between every two adjacent audio streams of the concatenated audio stream in the storage.

9. The method of claim 1 , wherein a size of the buffer is based on a duration of the time window.

10. The method of claim 1 , further comprising performing the transmitting the single audio stream payload from a buffer for transcription no later than an end of a wait time, the wait time starting from a request to onboard the buffer with one or more of the audio streams of the audio stream payload and ending at a subsequent request to re-onboard the buffer with a subsequent set of the plurality of audio streams.

11. The method of claim 1 , wherein performing the transmitting regardless of a density of the buffer to prevent noticeable delay in receiving the text file.

12. A system for processing audio streams, the system comprising: input/output (I/O) circuitry configured to:

receive a plurality of audio streams during a time window, each audio stream received from an independent source, including speech content, and having an audio stream duration; and

control circuitry configured to:

select a subset of the plurality of audio streams based on the respective audio stream durations of the plurality of audio streams, the subset comprising N number of audio streams, wherein a combined duration of the N number of audio streams is below a threshold duration, and wherein the threshold duration is based on a transaction cost for speech-to-text conversion;

generate N−1 number of audio stream separators;

concatenate the N number of audio streams to generate a concatenated audio stream;

insert an audio stream separator between every two adjacent audio streams of the concatenated audio stream to generate a single audio stream payload, each audio stream separator delineating a beginning of a next audio stream and an end of a preceding audio stream in the concatenated audio stream;

transmit the single audio stream payload from a buffer for transcription of the speech content of the set of audio streams to text content, a transcription of each of the speech content of the set of audio streams; and

in response to transmitting the single audio stream payload for transcription, receive a text file including the text content delineated with the audio stream separators.

13. The system of claim 12 , wherein:

a size of the buffer is based on a multiple of a maximum speech content duration among the durations of the speech content of the N number of audio streams.

14. The system of claim 13 , wherein the maximum speech content duration corresponds to a minimum base transcription price.

15. The system of claim 13 , wherein minimum transcription charges for transcription of the audio streams are based on the maximum speech content duration.

16. The system of claim 12 , wherein the control circuitry is further configured to:

concatenating the N number of audio streams in the buffer; and

inserting the audio stream separator between every two adjacent audio streams of the concatenated audio stream in the buffer.

17. The system of claim 12 , wherein a size of the buffer is based on a duration of the time window.

18. The system of claim 12 , wherein the control circuitry is further configured to:

concatenate the N number of audio streams in a storage; and

insert the audio stream separator between every two adjacent audio streams of the concatenated audio stream in the storage.

19. The system of claim 12 , wherein the control circuitry is further configured to perform the transmitting the single audio stream payload from a buffer for transcription no later than an end of a wait time, the wait time starting from a request to onboard the buffer with one or more of the audio streams of the audio stream payload and ending at a subsequent request to re-onboard the buffer with a subsequent set of the plurality of audio streams.

20. The system of claim 12 , wherein the control circuitry is further configured to perform the transmitting regardless of a density of the buffer to prevent noticeable delay in receiving the text file.

Assignments (3)
CHANGE OF NAME Recorded Oct 3, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069106/0238 →
SECURITY INTEREST Recorded May 3, 2023
From: ADEIA GUIDES INC.; ADEIA IMAGING LLC; ADEIA MEDIA HOLDINGS LLC; ADEIA MEDIA SOLUTIONS INC.; ADEIA SEMICONDUCTOR ADVANCED TECHNOLOGIES INC.; ADEIA SEMICONDUCTOR BONDING TECHNOLOGIES INC.; ADEIA SEMICONDUCTOR INC.; ADEIA SEMICONDUCTOR SOLUTIONS LLC; ADEIA SEMICONDUCTOR TECHNOLOGIES LLC; ADEIA SOLUTIONS LLC
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 063529/0272 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 28, 2021
From: AHER, ANKUR ANIL; ROBERT JOSE, JEFFRY COPPS
To: ROVI GUIDES, INC.
Reel/Frame 058597/0648 →