DETAILED ACTION
Introduction
1. This office action is in response to Applicant's submission filed on 01/23/2025. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims 1-15 are currently pending and examined below.
Drawings
2. The drawings filed on 01/23/2025 have been accepted and considered by the Examiner.
Information Disclosure Statement
3. The Information Statement (IDS) filed on 01/23/2025 has been accepted/considered and is in compliance with the provisions of 37 CFR 1.97.
Priority
4. The Applicants priority to Korean Patent Application # 10-2024-0011694, filed on January 25, 2024, has been accepted and considered in this office action.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
5. Claim 15 is rejected under 35 U.S.C. 101 as it explicitly claims a computer program. Computer programs claimed as such or as computer listings per se, i.e., the descriptions or expressions of the programs, are not physical "things." They are neither computer components nor statutory processes, as they are not "acts" being performed. Such claimed computer programs do not define any structural and functional interrelationships between the computer program and other claimed elements of a computer which permit the computer program's functionality to be realized. In contrast, a claimed non-transitory computer-readable medium encoded with a computer program is a computer element which defines structural and functional interrelationships between the computer program and the rest of the computer which permit the computer program's functionality to be realized and is thus statutory. See Lowry, 32 F.3d at 1583-84, 32 USPQ2d at 1035. Additionally, the medium recited in this claim should be explicitly qualified as non-transitory.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(2) The claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
6. Claims 1-5, 8-12 and 15 are rejected under 35 U.S.C. 102 (a) (2) as being anticipated by Wu (U.S. Patent Application Publication # 2026/0141898 A1).
With regards to claim 1, Wu teaches a method performed by an electronic device using artificial intelligence, comprising receiving a first speech signal (Para 7, teaches input speech data);
And outputting a first text corresponding to the first speech signal from a pre-trained first artificial intelligence algorithm module using the first speech signal as input, wherein the pre-trained first artificial intelligence algorithm module is pre-trained based on a first loss (Para 38, teaches that the system is configured to compute a transducer loss corresponding to the first set of layers which predict the blank token. Para 64, teaches using a pretrained language model, such as a BERT or RoBERTa model);
and wherein the first loss is determined based on a similarity between at least one speech embedding output from the first artificial intelligence algorithm module using a second speech signal as input and at least one text embedding output from a second artificial intelligence algorithm module using a second text as input (Para 38, teaches that the system is configured to compute a transducer loss corresponding to the first set of layers which predict the blank token. The objective function of the transducer model is to minimize the negative log probability over all possible alignments between the acoustic features and label sequences. The system is also configured to compute a language model loss, accounting for cross-entropy, corresponding to the second set of layers which predict the vocabulary token. The language model loss is combined with the transducer loss by subtracting the language model loss from the transducer loss. In some instances, the weighting factor is determined and applied to the language model loss).
With regards to claim 2, Wu teaches the method of claim 1 wherein the first artificial intelligence algorithm module includes a CTC (connectionist temporal classification) model, and wherein the second artificial intelligence algorithm module includes a BERT (bidirectional encoder representations from transformers) model (Para 20, teaches modifying factorized neural transducers to obtain even greater accuracy when performing ASR tasks including implementing CTC criterion into the training process to provide for faster, more efficient training processes, with improved prediction functionality. Para 64, teaches using a pretrained language model, such as a BERT or RoBERTa model).
With regards to claim 3, Wu teaches the method of claim 1 wherein the first loss is determined based on a CTC-BERT score, which is determined based on an average value of the similarity between the at least one speech embedding and the at least one text embedding (Para 48 and equation 4, teach the scoring according to Bayes' theorem, wherein the acoustic and language model scores are combined by weighted sum in the log probability domain. So, by converting the encoder output to a log probability by adding CTC criterion, the encoder output can be added with a weighted log probability of the vocabulary predictor output).
With regards to claim 4, Wu teaches the method of claim 1 wherein the pre-trained first artificial intelligence algorithm module is further pre-trained based on the first loss and a second loss, and wherein the second loss is determined based on a reference token sequence output from the first artificial intelligence algorithm module using the second speech signal as input (Para 38, teaches that the system is configured to compute a transducer loss corresponding to the first set of layers which predict the blank token. The objective function of the transducer model is to minimize the negative log probability over all possible alignments between the acoustic features and label sequences. The system is also configured to compute a language model loss, accounting for cross-entropy, corresponding to the second set of layers which predict the vocabulary token. The language model loss is combined with the transducer loss by subtracting the language model loss from the transducer loss. In some instances, the weighting factor is determined and applied to the language model loss).
With regards to claim 5, Wu teaches the method of claim 1 wherein the second artificial intelligence algorithm module is pre-trained and has a fixed model parameter (Para 64, teaches using a pretrained language model, such as a BERT or RoBERTa model which is pre-trained using a large number of training pairs. During training of the entire M-FNT model, the pre-trained language model is frozen i.e., not modified during the training process).
With regards to claims 8-12, these are device claims for the corresponding method claims 1-5. These two sets of claims are related as method and device of using the same, with each claimed device element's function corresponding to the claimed method step. Accordingly, claims 8-12 are similarly rejected under the same rationale as applied above with respect to method claims 1-5.
With regards to claim 15, this is a program claim for the corresponding method claim 1. These two claims are related as method and program of using the same, with each claimed program element's function corresponding to the claimed method step. Accordingly, claim 15 is similarly rejected under the same rationale as applied above with respect to method claim 1.
Allowable Subject Matter
7. Claims 6-7 and 13-14 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The prior art of record, alone or in combination, does not currently suggest or teach the invention as outlined in these claims. More detailed reasons for allowance will be outlined as and when the Application proceeds to allowability.
Conclusion
8. The following prior art, made of record but not relied upon, is considered pertinent to applicant's disclosure: Radford (U.S. Patent # 12079587 B1), Srinivasan (U.S. Patent # 12614552 B1). These references are also included in the PTO-892 form attached with this office action.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. If you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). In case you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NEERAJ SHARMA whose contact information is given below. The examiner can normally be reached on Monday to Friday 8 am to 5 pm. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre Louis-Desir can be reached on 571-272-7799 (Direct Phone). The fax number for the organization where this application or proceeding is assigned is 571-273-8300.
/NEERAJ SHARMA/
Primary Examiner, Art Unit 2659
571-270-5487 (Direct Phone)
571-270-6487 (Direct Fax)
neeraj.sharma@uspto.gov (Direct Email)