IP Library Granted Patent US 12,651,607
Granted Patent B2
US 12,651,607 · App. 18/691,047 · Granted Jun 9, 2026

Information processing apparatus and method that use learning models to alter watermark information based on mixed audio signals

Inventor: Naoya Takahashi (Tokyo, JP)
Assignee: SONY GROUP CORPORATION
G10L25/30G10L21/0272G10L25/03
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,651,607
App. No.
18/691,047
Granted
Jun 9, 2026
Kind
B2
Abstract

An information processing apparatus includes a decoder that extracts, using a certain learning model, additional information included in a mixed sound signal where a plurality of sound source signals is mixed together. A change in the sound source signals and the mixed sound signal caused by addition of the additional information is imperceptible.

Claims (42)

1 . An information processing apparatus comprising:

a decoder that extracts, using a certain learning model, additional information included in a mixed sound signal where a plurality of sound source signals is mixed together,

wherein a change in the sound source signals and the mixed sound signal caused by addition of the additional information is imperceptible;

wherein the learning model is a model obtained by performing learning for minimizing a loss function based on an error function between sound source signals before and after the additional information is included in the sound source signals, an error function between signals obtained by including the additional information in the sound source signals and signals corresponding to the sound source signals obtained through sound source separation, an error function between pieces of the additional information before and after the additional information is included in the sound source signals, and an error function between the additional information before the additional information is included in the sound source signals and the additional information included in the sound source signals obtained through the sound source separation.

2 . The information processing apparatus according to claim 1 ,

wherein the additional information is information added to the mixed sound signal where the plurality of sound source signals is mixed together.

3 . The information processing apparatus according to claim 1 , wherein the additional information is digital watermark information.

4 . The information processing apparatus according to claim 1 ,

wherein the decoder extracts the additional information from separation signals obtained by performing sound source separation on the mixed sound signal.

5 . The information processing apparatus according to claim 4 , further comprising

a sound source separation unit that performs the sound source separation.

6 . The information processing apparatus according to claim 1 ,

wherein at least one of the sound source signals constituting the mixed sound signal includes the additional information.

7 . The information processing apparatus according to claim 6 ,

wherein each of the sound source signals constituting the mixed sound signal includes the additional information.

8 . An information processing apparatus comprising a concealer that includes, using a certain learning model, additional information in a plurality of sound source signals or a mixed sound signal where the plurality of sound source signals is mixed together, wherein a change in the sound source signals and the mixed sound signal caused by addition of the additional information is imperceptible;

wherein the learning model is a model obtained by performing learning for minimizing a loss function based on an error function between sound source signals before and after the additional information is included in the sound source signals, an error function between signals obtained by including the additional information in the sound source signals and signals corresponding to the sound source signals obtained through sound source separation, an error function between pieces of the additional information before and after the additional information is included in the sound source signals, and an error function between the additional information before the additional information is included in the sound source signals and the additional information included in the sound source signals obtained through the sound source separation.

9 . The information processing apparatus according to claim 8 ,

wherein the concealer adds the additional information to the mixed sound signal where the plurality of sound source signals is mixed together.

10 . The information processing apparatus according to claim 8 ,

wherein the learning model to be used when the additional information is included is selectable.

11 . The information processing apparatus according to claim 8 , wherein the additional information is digital watermark information.

12 . The information processing apparatus according to claim 8 ,

wherein the concealer adds the additional information to at least one of the plurality of sound source signals.

13 . The information processing apparatus according to claim 12 ,

wherein the concealer adds the additional information to all the plurality of sound source signals.

14 . An information processing method comprising:

extracting, using a decoder and a certain learning model, additional information included in a mixed sound signal where a plurality of sound source signals is mixed together,

wherein a change in the sound source signals and the mixed sound signal caused by addition of the additional information is imperceptible;

wherein the learning model is a model obtained by performing learning for minimizing a loss function based on an error function between sound source signals before and after the additional information is included in the sound source signals, an error function between signals obtained by including the additional information in the sound source signals and signals corresponding to the sound source signals obtained through sound source separation, an error function between pieces of the additional information before and after the additional information is included in the sound source signals, and an error function between the additional information before the additional information is included in the sound source signals and the additional information included in the sound source signals obtained through the sound source separation.

15 . A non-transitory computer-readable medium storing a program causing a computer to perform an information processing method comprising:

extracting, using a decoder and a certain learning model, additional information included in a mixed sound signal where a plurality of sound source signals is mixed together,

wherein a change in the sound source signals and the mixed sound signal caused by addition of the additional information is imperceptible;

wherein the learning model is a model obtained by performing learning for minimizing a loss function based on an error function between sound source signals before and after the additional information is included in the sound source signals, an error function between signals obtained by including the additional information in the sound source signals and signals corresponding to the sound source signals obtained through sound source separation, an error function between pieces of the additional information before and after the additional information is included in the sound source signals, and an error function between the additional information before the additional information is included in the sound source signals and the additional information included in the sound source signals obtained through the sound source separation.

16 . An information processing method comprising:

including, using a concealer and a certain learning model, additional information in a plurality of sound source signals or a mixed sound signal where the plurality of sound source signals is mixed together,

wherein a change in the sound source signals and the mixed sound signal caused by addition of the additional information is imperceptible;

wherein the learning model is a model obtained by performing learning for minimizing a loss function based on an error function between sound source signals before and after the additional information is included in the sound source signals, an error function between signals obtained by including the additional information in the sound source signals and signals corresponding to the sound source signals obtained through sound source separation, an error function between pieces of the additional information before and after the additional information is included in the sound source signals, and an error function between the additional information before the additional information is included in the sound source signals and the additional information included in the sound source signals obtained through the sound source separation.

17 . A non-transitory computer-readable medium storing a program causing a computer to perform an information processing method comprising:

including, using a concealer and a certain learning model, additional information in a plurality of sound source signals or a mixed sound signal where the plurality of sound source signals is mixed together,

wherein a change in the sound source signals and the mixed sound signal caused by addition of the additional information is imperceptible;

wherein the learning model is a model obtained by performing learning for minimizing a loss function based on an error function between sound source signals before and after the additional information is included in the sound source signals, an error function between signals obtained by including the additional information in the sound source signals and signals corresponding to the sound source signals obtained through sound source separation, an error function between pieces of the additional information before and after the additional information is included in the sound source signals, and an error function between the additional information before the additional information is included in the sound source signals and the additional information included in the sound source signals obtained through the sound source separation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 12, 2024
From: TAKAHASHI, NAOYA
To: SONY GROUP CORPORATION
Reel/Frame 066723/0557 →
Priority Claims (1)
JP 2021-157816 · Sep 28, 2021 · national
Continuity (1)
Related Publication 20240379117A1 · Nov 14, 2024
References Cited (18)
US 10236006B1 · Gurijala · 2019 [cited by examiner]
US 20120203362A1 · Parvaix · 2012 [cited by examiner]
US 20160140968A1 · Paulus · 2016 [cited by examiner]
US 20190198036A1 · Osako · 2019 [cited by applicant]
US 20200014817A1 · Deshmukh · 2020 [cited by examiner]
US 20200134447A1 · Bennett · 2020 [cited by examiner]
US 20210223780A1 · Bramley · 2021 [cited by examiner]
US 20210256978A1 · Zeyu · 2021 [cited by applicant]
US 20240371390A1 · Takahashi · 2024 [cited by examiner]
US 20240379117A1 · Takahashi · 2024 [cited by examiner]
JP 2004500728A · 2004 [cited by applicant]
JP 2018186464A · 2018 [cited by applicant]
JP 2021157816 · 2021 [cited by applicant]
WO 2018047643A1 · 2018 [cited by applicant]
WO WO2021009319A1 · 2021 [cited by applicant]
International Search Report and Written Opinion mailed on Apr. 19, 2022, received for PCT Application PCT/JP2022/006048, filed on Feb. 16, 2022, 9 pages including English Translation. [cited by applicant]
Felix Kreuk et al: “Hide and Speak: Towards Deep Neural Networks for Speech Steganography”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Jul. 27, 2020 (Jul. 27, 2020) , XP… [cited by applicant]
Naoya Takahashi et al: “Source Mixing and Separation Robust Audio Steganography”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Oct. 11, 2021 (Oct. 11, 2021), XP091075534. [cited by applicant]