IP Library Granted Patent US 10,783,875
Granted Patent B2
US 10,783,875 · App. 16/027,111 · Granted Sep 22, 2020

Unsupervised non-parallel speech domain adaptation using a multi-discriminator adversarial network

Inventors: Ehsan Hosseini-Asl (Palo Alto, CA); Caiming Xiong (Palo Alto, CA); Yingbo Zhou (San Jose, CA); Richard Socher (Menlo Park, CA)
Assignee: salesforce.com, inc.
G10L15/07G06N3/08G10L15/16G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,783,875
App. No.
16/027,111
Granted
Sep 22, 2020
Kind
B2
Abstract

A system for domain adaptation includes a domain adaptation model configured to adapt a representation of a signal in a first domain to a second domain to generate an adapted presentation and a plurality of discriminators corresponding to a plurality of bands of values of a domain variable. Each of the plurality of discriminators is configured to discriminate between the adapted representation and representations of one or more other signals in the second domain.

Claims (55)

1. A system comprising:

a domain adaptation model configured to adapt a representation of a signal in a first domain to a second domain to generate an adapted presentation; and

a plurality of discriminators corresponding to a plurality of bands,

wherein each of the plurality of bands corresponds to a domain variable range of a domain variable of the first and second domains,

wherein the plurality of bands is determined based on a variation of a characteristic feature associated with the domain variable between the first domain and second domain, and

wherein each of the plurality of discriminators is configured to discriminate between the adapted representation and representations of one or more other signals in the second domain.

2. The system of claim 1 , wherein bandwidths of the plurality of bands are determined based on the corresponding characteristic feature variations.

3. The system of claim 1 , wherein a first discriminator of the plurality of discriminators corresponds to a first band of the plurality of bands having a first width of the domain variable, and

wherein a second discriminator of the plurality of discriminators corresponds to a second band of the plurality of bands having a second width of the domain variable different from the first range.

4. The system of claim 1 , wherein the first domain is a first speech domain and the second domain is a second speech domain.

5. The system of claim 4 , wherein the domain variable includes an audio frequency.

6. The system of claim 5 , wherein the characteristic feature includes a frequency amplitude variation rate for a fixed time window.

7. The system of claim 1 , further comprising:

a second domain adaptation model configured to adapt a second representation of a second signal in the second domain to the first domain; and

a plurality of second discriminators corresponding to a plurality of second bands, each of the plurality of second discriminators being configured to discriminate between the adapted second representation and representations of one or more other signals in the first domain.

8. A non-transitory machine-readable medium comprising a plurality of machine-readable instructions which, when executed by one or more processors, are adapted to cause the one or more processors to perform a method comprising:

providing a domain adaptation model configured to adapt a representation of a signal in a first domain to a second domain to generate an adapted presentation; and

providing a plurality of discriminators corresponding to a plurality of bands,

wherein each of the plurality of bands corresponds to a domain variable range of a domain variable of the first and second domains,

wherein the plurality of bands is determined based on a variation of a characteristic feature associated with the domain variable between the first domain and second domain, and

wherein each of the plurality of discriminators is configured to discriminate between the adapted representation and representations of one or more other signals in the second domain.

9. The non-transitory machine-readable medium of claim 8 , wherein wherein bandwidths of the plurality of bands are determined based on the corresponding characteristic feature variations.

10. The non-transitory machine-readable medium of claim 8 , wherein a first band of the plurality of bands has a first domain variable range; and

wherein a second band of the plurality of bands has a second domain variable range different from the first domain variable range.

11. The non-transitory machine-readable medium of claim 8 , where a first band and a second band of the plurality of bands overlap.

12. The non-transitory machine-readable medium of claim 8 , wherein the first domain is a first speech domain and the second domain is a second speech domain.

13. The non-transitory machine-readable medium of claim 12 , wherein the domain variable is an audio frequency.

14. The non-transitory machine-readable medium of claim 8 , wherein the method further comprises:

providing a second domain adaptation model configured to adapt a second representation of a second signal in the second domain to the first domain; and

providing a plurality of second discriminators corresponding to a plurality of second bands, each of the plurality of second discriminators being configured to discriminate between the adapted second representation and representations of one or more other signals in the first domain.

15. A method for training parameters of a first domain adaptation model using multiple independent discriminators, comprising:

providing a plurality of first discriminator models corresponding to a plurality of first bands, each of the plurality of bands corresponding to a domain variable range of a domain variable of a source domain and a target domain,

wherein the plurality of bands is determined based on a variation of a characteristic feature associated with the domain variable between the first domain and second domain;

evaluating the plurality of first discriminator models based on:

one or more first training representations adapted from the source domain to the target domain by the first domain adaptation model, and

one or more second training representations in the target domain, yielding a first multi-discriminator objective;

evaluating a learning objective based on the first multi-discriminator objective; and

updating the parameters of the first domain adaptation model based on the learning objective.

16. The method of claim 15 , further comprising:

evaluating a plurality of second discriminator models corresponding to a plurality of second bands of values of the domain variable based on:

one or more third training representations adapted from the target domain to the source domain by a second domain adaptation model, and

one or more fourth training representations in the source domain, yielding a second multi-discriminator objective;

wherein the evaluating the learning objective includes:

evaluating the learning objective based on the first multi-discriminator objective and second multi-discriminator objective.

17. The method of claim 16 , further comprising:

evaluating a cycle consistency objective based on:

one or more fifth training representations adapted from the source domain to the target domain by the first domain adaptation model and from the target domain to the source domain by the second domain adaptation model; and

one or more sixth training representations adapted from the target domain to the source domain by the second domain adaptation model and from the source domain to the target domain by the first domain adaptation model;

wherein the evaluating the learning objective includes:

evaluating the learning objective based on the first multi-discriminator objective, second multi-discriminator objective, and cycle consistency objective.

18. The method of claim 15 , wherein the source domain is a first speech domain and the target domain is a second speech domain.

19. The method of claim 16 , wherein the domain variable is an audio frequency.

20. The method of claim 15 , further comprising:

wherein a first discriminator of the plurality of discriminators corresponds to a first band of the plurality of bands having a first range of the domain variable, and

wherein a second discriminator of the plurality of discriminators corresponds to a second band of the plurality of bands having a second range of the domain variable different from the first range.

Assignments (2)
CHANGE OF NAME Recorded Dec 18, 2024
From: SALESFORCE.COM, INC.
To: SALESFORCE, INC.
Reel/Frame 069717/0353 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 26, 2019
From: HOSSEINI-ASL, EHSAN; XIONG, CAIMING; ZHOU, YINGBO; SOCHER, RICHARD
To: SALESFORCE.COM, INC.
Reel/Frame 051366/0733 →
Continuity (3)
Provisional Application 62647459 · Mar 23, 2018
Provisional Application 62644313 · Mar 16, 2018
Related Publication 20190295530A1 · Sep 26, 2019
Cited By (5)
US 12,265,909 US 12,299,982 US 12,493,799 US 12,530,560 US 12,681,769