DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant's arguments filed 05/18/26 have been fully considered but they are not persuasive.
Applicant argues that Liao et al. do not teach the first adversarial network comprises: a first generator configured to perform a task; and a first discriminator configured to recognize the first attribute, the second adversarial network comprises: a second generator configured to recognize the first attribute; and a second discriminator, configured to perform the task; providing the information elimination model, a first adversarial network, and a second adversarial network" and "two input layers, including one input layer from each of the first adversarial network and the second adversarial network, are generated based on an output layer of the information elimination model and the input feature (Amendment, pages 10 - 12).
The examiner disagrees, since Liao et al. disclose “acoustic features and a vocoder can be used for voice conversion: first, an automatic speech recognition (Automatic Speech Recognition, ASR) technology is used to analyze a source speech to obtain phonetic posteriorgram (Phonetic Posteriorgram, PPG) features of the source speech, and then the PPG features are input to a conversion model to obtain acoustic speech parameters to be sent to the vocoder, and finally the vocoder converts the source speech into a target speech according to the speech acoustic features…the purpose of voice conversion can be achieved by continuously and cyclically generating and discriminating the target speech through two speech generators and two discriminators based on a CycleGAN mechanism… The multiple discriminators can perform discrimination on output results of different stacking numbers of decoders, that is, the output results of different layers of decoders can be obtained respectively, and the output result is used as an input of each discriminator and is used as an input of the next layer of decoder… The model convergence condition can be predetermined according to an actual need. For example, the model convergence condition can be that the value of the first loss function is less than or equal to a preset first loss threshold, and the value of the second loss function is less than or equal to a preset second loss threshold.” (paragraphs 66, 70, 134, and 143).
Applicant argues that Liao et al. do not teach the loss function is associated with a disentangling loss of the two input layers of the first adversarial network and the second adversarial network, a first loss of the first generator, a second loss of the first discriminator, a third loss of the second generator, and a fourth loss of the second discriminator (Amendment, pages 12, 13).
The examiner disagrees, since Liao et al. disclose “the purpose of voice conversion can be achieved by continuously and cyclically generating and discriminating the target speech through two speech generators and two discriminators based on a CycleGAN mechanism… The core principle of GAN is as follows: The generator generates data for the discriminator to determine whether the generated data is true or false. Until the discriminator cannot distinguish between the generated data and the original data, the training is completed. Therefore, the loss function in the embodiment of the present invention may include two parts: the first loss function corresponding to the generator and the second loss function corresponding to the discriminator.”(paragraphs 70, 129).
Applicant’s arguments, see pages 7 - 9, filed 05/18/26, with respect to claims 1 - 16 have been fully considered and are persuasive. The rejection of claims 1 – 16 under 35 U.S.C 101 has been withdrawn.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1, 3, 5-7, 10 – 16 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Liao et al. (US PAP 2024/0412749).
As per claims 1, 15, 16, Liao et al. teach a computer-implemented method for training an information elimination model, the information elimination model configured to eliminate, from an input feature, first information that allows a first attribute to be recognizable, the computer-implemented method comprising:
providing the information elimination model, a first adversarial network, and a second adversarial network (“generating and discriminating the target speech through two speech generators and two discriminators based on a CycleGAN mechanism.”; paragraphs 70, 168); and
minimizing a loss function to train the information elimination model, wherein two input layers, including one input layer from each of the first adversarial network and the second adversarial network, are generated based on an output layer of the information elimination model and the input feature (“it can be determined that the value of the first loss function does not meet the model convergence condition; on the contrary, it can be determined that the value of the first loss function meets the model convergence condition; if the value of the second loss function is greater than the second loss threshold, it can be determined that the value of the second loss function does not meet the model convergence condition; on the contrary, it can be determined that the value of the second loss function meets the model convergence condition.”; paragraphs 27 - 35, 141 – 143),
the first adversarial network comprises: a first generator configured to perform a task; and a first discriminator configured to recognize the first (“generating and discriminating the target speech through two speech generators and two discriminators based on a CycleGAN mechanism…first, an automatic speech recognition (Automatic Speech Recognition, ASR) technology is used to analyze a source speech to obtain phonetic posteriorgram (Phonetic Posteriorgram, PPG) features of the source speech, and then the PPG features are input to a conversion model to obtain acoustic speech parameters to be sent to the vocoder, and finally the vocoder converts the source speech into a target speech according to the speech acoustic features.”; paragraphs 66, 70, 168);
the second adversarial network comprises: a second generator configured to recognize the first attribute; and a second discriminator, configured to perform the task (“generating and discriminating the target speech through two speech generators and two discriminators based on a CycleGAN mechanism…first, an automatic speech recognition (Automatic Speech Recognition, ASR) technology is used to analyze a source speech to obtain phonetic posteriorgram (Phonetic Posteriorgram, PPG) second
features of the source speech, and then the PPG features are input to a conversion model to obtain acoustic speech parameters to be sent to the vocoder, and finally the vocoder converts the source speech into a target speech according to the speech acoustic features.”; paragraphs 66, 70, 168); and
the loss function is associated with a disentangling loss of the two input layers of the first adversarial network and the second adversarial network, a first loss of the first generator, a second loss of the first discriminator, a third loss of the second generator, and a fourth loss of the second discriminator (paragraphs 97, 128 - 131).
Claim 15 further recites receiving an original signal; generating a feature not including first information based on the original signal and an information elimination model, the first information allowing a first attribute to be recognizable; and performing the task based on the feature and the machine learning model (“first, an automatic speech recognition (Automatic Speech Recognition, ASR) technology is used to analyze a source speech to obtain phonetic posteriorgram (Phonetic Posteriorgram, PPG) features of the source speech”; paragraphs 66, 70, 168).
As per claim 3, Liao et al. further disclose the output layer of the information elimination model comprises a first control signal corresponding to the first adversarial network and a second control signal corresponding to the second adversarial network (paragraphs 5, 70, 134, 168).
As per claim 5, Liao et al. further disclose the first control signal generated by the information elimination model after training is configured to eliminate the first information from the input feature (paragraphs 5, 70, 134, 168).
As per claim 6, Liao et al. further disclose providing a model configured to perform the task and taking the model as the first generator and the second discriminator; and providing a recognition model for the first attribute and taking the recognition model as the first discriminator and the second generator (“an automatic speech recognition (Automatic Speech Recognition, ASR) technology is used to analyze a source speech to obtain phonetic posteriorgram (Phonetic Posteriorgram, PPG) features of the source speech, and then the PPG features are input to a conversion model to obtain acoustic speech parameters to be sent to the vocoder, and finally the vocoder converts the source speech into a target speech according to the speech acoustic features.”; paragraphs 66 – 68).
As per claim 7, Liao et al. further disclose minimizing the loss function to train the information elimination model comprises: keeping a plurality of parameters of the first adversarial network and the second adversarial network unchanged when minimizing the loss function (“updating network parameters of the discriminator network based on the value of the second loss function; [0034] an iterative training step: performing iterative training based on the speech generation network and the discriminator network after the parameters are updated, until the value of the first loss function and the value of the second loss function both meet a model convergence condition; and [0035] a generator determining step: using the trained speech generation network as a speech” paragraphs 32 – 40).
As per claim 10, 11, Liao et al. further disclose the task comprises a speech recognition-related task; the task comprises an automatic speech recognition (automatic speech recognition (“Automatic Speech Recognition, ASR) technology is used to analyze a source speech to obtain phonetic posteriorgram (Phonetic Posteriorgram, PPG) features of the source speech”; paragraph 66).
As per claim 12, Liao et al. further disclose a gradient reversal layer is included before each of the first discriminator and the second discriminator (paragraphs 139 – 143, 157 - 168).
As per claim 13, Liao et al. further disclose each of the second loss and the third loss comprises a recall loss(paragraphs 139 – 143, 157 - 168).
As per claim 14, Liao et al. further disclose each of the first loss and the fourth loss comprises a connectionist temporal classification loss(paragraphs 139 – 143, 157 - 168).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable Liao et al. (US PAP 2024/0412749) in view of CHANGIZ REZAEI et al. (US PAP 2023/0140142).
As per claim 4, Liao et al. do not specifically teach the information elimination model generates the first control signal and the second control signal using a Gumbel-Softmax function.
CHANGIZ REZAEI et al. disclose that the generator of the generative adversarial network generates a plurality of generated neural network architectures responsive to the received search space. The discriminator of the generative adversarial network selects an optimal neural network architecture from among the plurality of generated neural network architectures… Both SNAS and GDAS methods employ the use of Gumbel Softmax when operating on the search space. The search algorithm is more specific in describing the architectures it selects (Abstract, paragraph 9).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use Gumbel Softmax as taught by CHANGIZ REZAEI et al. in Liao et al., because that would help improve the existing models (paragraph 54).
Claims 8, 9 are rejected under 35 U.S.C. 103 as being unpatentable over Liao et al. (US PAP 2024/0412749) in view of Wang et al. (US PAP 2023/0363679).
As per claims 8, 9, Liao et al. do not specifically teach the first attribute comprises an attribute related to vulnerable populations; the first attribute comprises an attribute of dementia.
Wang et al. disclose a novel adversarial training loss is also introduced to obtain identity-independent and stroke-discriminative features… The “cookie theft” task requires the subject to retrieve, think, organize, and express the information, which will end up evaluating the subject's speech ability from both motor and cognitive aspects. Such an analysis has also been successful in identifying subjects with Alzheimer's-related dementia, aphasia, and some other cognitive-communication impairments, and would be suitable for stroke screening purposes (paragraphs 25, 70).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to analyze cognitive disorder as taught by Wang et al. in Liao et al., because that would help identify subjects with Alzheimer's-related dementia, aphasia, and other cognitive-communication impairments. (paragraph 113).
Allowable Subject Matter
Claim 2 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter:
As to claim 2, the prior art made of record does not teach or suggest providing a third adversarial network, wherein an input layer of the third adversarial network is generated based on the output layer of the information elimination model and the input feature, the first adversarial network further comprises a third discriminator configured to recognize the second attribute, the second adversarial network further comprises a fourth discriminator configured to recognize the second attribute, the third adversarial network comprises: a third generator configured to recognize the second attribute;
a fifth discriminator configured to recognize the first attribute; a sixth discriminator configured to perform the task, the disentangling loss is further associated with the input layer of the third adversarial network, the loss function is further associated with a fifth loss of the third generator, a sixth loss of the fifth discriminator, a seventh loss of the sixth discriminator, an eighth loss of the third discriminator, and a nineth loss of the fourth discriminator.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to LEONARD SAINT-CYR whose telephone number is (571)272-4247. The examiner can normally be reached Monday- Friday.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached at (571)272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/LEONARD SAINT-CYR/ Primary Examiner, Art Unit 2658