IP Library Granted Patent US 12,488,794
Granted Patent B2
US 12,488,794 · App. 17/452,499 · Granted Dec 2, 2025

Apparatus and method for analysis of audio recordings

Inventors: Fabio Cappello (London, GB); Danjeli Schembri (London, GB); Oliver Hume (London, GB)
Assignee: Sony Interactive Entertainment Inc.
G10L15/187G06F40/279G10L15/10G11B27/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,488,794
App. No.
17/452,499
Granted
Dec 2, 2025
Kind
B2
Abstract

A data processing apparatus includes storage circuitry to store audio data for a plurality of respective dialogue recordings for a content and to store text data indicative of a sequence of respective words within the audio data for each of the plurality of respective dialogue recordings, analysis circuitry to compare the text data for a current dialogue recording with predetermined text data for the content and to output comparison data for the current dialogue recording, the comparison data indicative of one or more differences between the text data for the current dialogue recording and the predetermined text data, selection circuitry to select one or more candidate dialogue recordings from the plurality of respective dialogue recordings for the content in dependence upon the comparison data, and recording circuitry to modify at least a portion of the audio data for the current dialogue recording in dependence upon the audio data for one or more of the candidate dialogue recordings to obtain modified audio data and to store the modified audio data for the current dialogue recording.

Claims (58)

1 . A data processing apparatus, comprising:

storage circuitry to store a plurality of respective audio recordings for a content and to store text data indicative of respective sequences of words within each of the plurality of respective audio recordings;

analysis circuitry to:

compare, by the analysis circuitry, a sequence of words indicated by the text data for a current audio recording with predetermined text data for the content to detect whether the sequence of words is present in the predetermined text data,

if the sequence of words is present in the predetermined text data, determine that the current audio recording comprises a correct recital of the predetermined text data and that modification of the current audio recording is not required, and

if the sequence of words is not present in the predetermined text data, determine that the current audio recording does not comprise a correct recital of the predetermined text data and output comparison data for the current audio recording, the comparison data indicative of a plurality of respective words that are missing from the current audio recording in comparison to the predetermined text data;

selection circuitry to:

select one or more candidate audio recordings from the plurality of respective audio recordings for the content in dependence upon the comparison data so that each of the one or more candidate audio recordings includes at least one of the plurality of respective words that are missing from the current audio recording;

assign one or more priority ratings to each of the one or more candidate audio recordings in dependence upon one or more properties associated with at least one of the candidate audio recording and the text data for the candidate audio recording, wherein the one or more priority ratings include a first priority rating indicative of how many of the plurality of respective words that are missing from the current audio recording are present in each of the one or more candidate audio recordings;

rank the one or more candidate audio recordings based on the one or more priority ratings; and

select a candidate audio recording from the one or more candidate audio recordings according to the ranking; and

recording circuitry to:

modify at least a first portion of the current audio recording to include at least a second portion of the selected candidate audio recording obtain a modified audio recording; and

store the modified audio recording for the current audio recording.

2 . The data processing apparatus according to claim 1 , wherein the analysis circuitry is configured to detect a sequence of respective words in the predetermined text data having a highest degree of match with the text data for the current audio recording in dependence upon the comparison of the text data for the current audio recording with the predetermined text data.

3 . The data processing apparatus according to claim 2 , wherein the comparison data is indicative of one or more of the respective words present in the sequence of respective words in the predetermined text data having the highest degree of match that are not present in the sequence of words in the text data for the current audio recording.

4 . The data processing apparatus according to claim 1 , wherein the analysis circuitry is configured to compare each of the plurality of respective audio recordings for the content with the predetermined text data for the content and to assign a match score to each of the plurality of respective audio recordings, wherein the match score for a given audio recording is indicative of a degree of match between the sequence of words in the text data for the given audio recording and the sequence of respective words in the predetermined text data, and wherein the analysis circuitry is configured to select an audio recording having a highest match score as the current audio recording.

5 . The data processing apparatus according to claim 1 , wherein the selection circuitry is configured to assign a second priority rating to each of the one or more candidate audio recording in dependence upon a number of respective words included in the text data for the candidate audio recording that match a respective word included in the text data for the current audio recording.

6 . The data processing apparatus according to claim 1 , wherein the selection circuitry is configured to:

detect, in a sequence of words included in the predetermined text data, a respective word that is adjacent to a respective word indicated by the comparison data;

detect, in a sequence of words included in the text data for a candidate audio recording of the one or more candidate audio recordings, a respective word that is adjacent to a respective word indicated by the comparison data; and

assign a third priority rating to the candidate audio recording in dependence upon whether the detected respective words match each other.

7 . The data processing apparatus according to claim 1 , wherein the selection circuitry is configured to calculate a number of words per unit time for both the candidate audio recording and the current audio recording and to assign a fourth priority rating to the candidate audio recording in dependence upon a magnitude of a difference between the number of words per unit time for the candidate audio recording and the current audio recording.

8 . The data processing apparatus according to claim 1 , wherein the selection circuitry is configured to detect, for both the candidate audio recording and the current audio recording, an amplitude of an audio signal in the audio recording and to assign a fifth priority rating to the candidate audio recording in dependence upon a magnitude of a difference between the detected amplitudes.

9 . The data processing apparatus according to claim 1 , wherein the selection circuitry is configured to detect, for both the candidate audio recording and the current audio recording, a pitch of a voice in the audio recording and to assign a sixth priority rating to the candidate audio recording in dependence upon a magnitude of a difference between the detected pitches.

10 . The data processing apparatus according to claim 1 , wherein the selection circuitry is configured to:

calculate a confidence score for each of the plurality of candidate audio recordings;

wherein;

the confidence score for a candidate audio recording is dependent upon one or more of the priority ratings assigned to that candidate audio recording; and

selecting the candidate audio recording from the one or more candidate audio recordings according to the ranking comprises selecting the candidate audio recording having the highest confidence score.

11 . The data processing apparatus according to claim 1 , wherein the second portion of the selected candidate audio recording includes a respective word of the plurality of respective words that are missing from the current audio recording.

12 . The data processing apparatus according to claim 1 , wherein the analysis circuitry is configured to generate the text data for each of the plurality of respective audio recordings from each respective audio recording.

13 . The data processing apparatus according to claim 1 , wherein the plurality of respective audio recordings each correspond to a same voice actor.

14 . The data processing apparatus according to claim 1 , wherein the predetermined text data for the content includes at least one sequence of respective words, and wherein the sequence of respective words is updatable in response to a user input.

15 . A data processing method comprising:

storing a plurality of respective audio recordings for a content;

storing text data indicative of respective sequences of words within each of the plurality of respective audio recordings;

comparing a sequence of words indicated by the text data for a current audio recording with predetermined text data for the content;

outputting, based on the comparing, comparison data for the current audio recording, the comparison data indicative of a plurality of respective words that are missing from the current audio recording in comparison to the predetermined text data;

selecting one or more candidate audio recordings from the plurality of respective audio recordings for the content in dependence upon the comparison data so that each of the one or more candidate audio recordings includes at least one of the plurality of respective words that are missing from the current audio recording;

assigning one or more priority ratings to each of the one or more candidate audio recordings in dependence upon one or more properties associated with at least one of the candidate audio recording and the text data for the candidate audio recording, wherein the one or more priority ratings include a first priority rating indicative of how many of the plurality of respective words that are missing from the current audio recording are present in each of the one or more candidate audio recordings;

ranking the one or more candidate audio recordings according to the one or more priority ratings;

selecting a candidate audio recording from the one or more candidate audio recordings according to the ranking;

modifying at least a first portion of the current audio recording to include at least a second portion of the selected candidate audio recording to obtain a modified audio recording; and

storing the modified audio recording for the current audio recording.

16 . A non-transitory, computer readable storage medium containing computer software which, when executed by a computer, causes the computer to perform a data processing method comprising the steps of:

storing for a plurality of respective audio recordings for a content;

storing text data indicative of respective sequences of words within each of the plurality of respective audio recordings;

comparing a sequence of words indicated by the text data for a current audio recording with predetermined text data for the content to detect whether the sequence of words is present in the predetermined text data;

if the sequence of words is present in the predetermined text data, determining that the current audio recording comprises a correct recital of the predetermined text data and that modification of the current audio recording is not required;

if the sequence of respective words is not present in the predetermined text data:

determining that the current audio recording does not comprise a correct recital of the predetermined text data and outputting comparison data for the current audio recording, the comparison data indicative of differences between the text data for a plurality of respective words that are missing from the current audio recording in comparison to the predetermined text data;

selecting one or more candidate audio recordings from the plurality of respective audio recordings for the content in dependence upon the comparison data so that each of the one or more candidate audio recordings includes at least one of the plurality of respective words that are missing from the current audio recording;

assigning one or more priority ratings to each of the one or more candidate audio recordings in dependence upon one or more properties associated with at least one of the candidate audio recording and the text data for the candidate audio recording, wherein the one or more priority ratings include a first priority rating indicative of how many of the plurality of respective words that are missing from the current audio recording are present in each of the one or more candidate audio recordings;

ranking the one or more candidate audio recordings according to the one or more priority ratings;

selecting a candidate audio recording from the one or more candidate audio recordings according to the ranking;

modifying at least a first portion of the current audio recording to include at least a second portion of the selected candidate audio recording to obtain a modified audio recording; and

storing the modified audio recording for the current audio recording.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 28, 2022
From: SONY INTERACTIVE ENTERTAINMENT EUROPE LIMITED
To: SONY INTERACTIVE ENTERTAINMENT INC.
Reel/Frame 059761/0698 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 28, 2022
From: CAPPELLO, FABIO
To: SONY INTERACTIVE ENTERTAINMENT EUROPE LIMITED
Reel/Frame 059820/0812 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 27, 2021
From: SCHEMBRI, DANJELI; HUME, OLIVER
To: SONY INTERACTIVE ENTERTAINMENT INC.
Reel/Frame 057935/0121 →
Priority Claims (1)
GB 2017768 · Nov 11, 2020 · national
Continuity (1)
Related Publication 20220148584A1 · May 12, 2022
References Cited (18)
US 5758323A · Case · 1998 [cited by examiner]
US 10455297B1 · Mahyar · 2019 [cited by examiner]
US 10878802B2 · Yamamoto · 2020 [cited by examiner]
US 11836181B2 · Saggi · 2023 [cited by examiner]
US 20060095262A1 · Danieli · 2006 [cited by applicant]
US 20080154601A1 · Stifelman · 2008 [cited by examiner]
US 20130124984A1 · Kuspa · 2013 [cited by examiner]
US 20130151251A1 · Herz · 2013 [cited by applicant]
US 20140164371A1 · Tesch · 2014 [cited by examiner]
US 20140201631A1 · Pornprasitsakul · 2014 [cited by examiner]
US 20180233162A1 · Venkataramani · 2018 [cited by examiner]
US 20190295531A1 · Rao · 2019 [cited by examiner]
US 20190311745A1 · Shenkler · 2019 [cited by examiner]
US 20200066293A1 · Sanchez · 2020 [cited by examiner]
CN 105788588A · 2016 [cited by examiner]
CN 105788588B · 2020 [cited by applicant]
Extended European Search Report for corresponding EP Application No. 21204783.1, 27 pages, dated Mar. 11, 2022. [cited by applicant]
Combined Search Report and Examination Report for corresponding GB Application No. 2017768.9, 7 pages, dated May 11, 2021. [cited by applicant]