DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Specification
The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed.
The following title is suggested: --Method and Apparatus for Domain Adaptation of an Artificial Intelligence Speech Recognition Model using Loss Ratio-Driven Data Filtering--.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 17 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
In Claim 17, Line 15, "the ratio" lacks antecedent basis and it is unclear what term is being referenced by this limitation. For claim interpretation in the interest of compact prosecution, "the ratio" will be construed as --a ratio-- similar to claims 1 and 9.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-17 are rejected under 35 U.S.C. 101 for being directed towards a patent ineligible mental process under the broadest reasonable interpretation.
As a first matter, however, under step 1 of the 2019 Patent Subject Matter Eligibility Guidelines (2019 PEG), Claim 17 is rejected under 35 U.S.C. 101 as failing to fall within one of the four statutory categories of invention. Specifically, claim 17 is directed to a “program stored in a medium.” The instant specification does not provide a special definition or disavowal of claim scope that excludes signals per se from the ordinary and customary meaning of “medium” under the broadest reasonable interpretation. Accordingly, claim 17 is directed towards a signal per se under the BRI that does not fall within the four statutory categories of invention and as such has been rejected under 35 U.S.C. 101. Moreover, it is worth noting that the “medium” is not “computer-readable” raising the issue of whether the medium would impart program functionality when executed by a processor. Thus, Applicant is recommended to replace “medium” in claim 17 with –non-transitory computer-readable medium-- to resolve these issues.
Continuing on to Step 2A prong 1 of the 2019 PEG, independent claims 1, 9, and 17 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
In regards to the process of these independent claims, the claimed functionality could be practiced as a mental process in the following manner:
receiving a speech signal (a human could listen to and mentally understand spoken words/statements); and
generating a text corresponding to the speech signal by using the speech signal as input in a pre-trained first artificial intelligence algorithm model (a human can read and copy down displayed text resulting from repeating the heard speech as an input to a pre-trained AI algorithm model; it should be noted that the Examiner is aware that Applicant may take issue in the inclusion of the claimed AI model being part of the abstract idea grouping of Step 2A prong 1, however, as claimed under the BRI, use of the AI model despite its training particulars included in proceeding wherein clauses is completely passive and reads on a human process. Although claimed active use of the particularly trained AI model in speech-to-text would reasonably constitute a technical improvement in the filed of AI-based speech recognition, the claim as drafted with passive use of the AI model as only using speech as an input reads on a mental process under the BRI. Applicant is directed to MPEP 2106.04(d)(1)- "Second, if the specification sets forth an improvement in technology, the claim must be evaluated to ensure that the claim itself reflects the disclosed improvement. That is, the claim includes the components or steps of the invention that provide the improvement described in the specification. The claim itself does not need to explicitly recite the improvement described in the specification (e.g., "thereby increasing the bandwidth of the channel")."
This judicial exception is not integrated into a practical application. Outside of the identified abstract idea, the claimed invention only includes generic computer components (e.g., processor, modem, and memory/medium) which amount to no more than mere instructions to implement an otherwise abstract idea using generic computer components wherein the computer components are used as a tool to carry out an otherwise abstract idea and not improved as a tool.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The above identified additional generic computer components are no more than mere instructions to apply the exception using generic computer components that are well-known, routine, and conventional as is evidenced by Bancorp Services v. Sun Life (Fed. Cir. 2012) and Alice Corp. v. CLS Bank (2014). See also buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) for a computer sending and receiving information over a network.
Accordingly, at least independent claims 1, 9, and 17 are not patent eligible under 35 U.S.C. 101 because they are directed to a judicial exception under the BRI. As noted above, a simple fix is available to Applicant in actively indicating that the particularly trained AI model related to a technical improvement under Step 2A prong 2 is actively involved in generating the text by processing the received speech instead of the passive involvement under the BRI raising this judicial exception rejection under 35 U.S.C. 101.
The remaining dependent claims fail to add patent eligible subject matter to their respective parent claims because they only serve to narrow details of a training procedure of a AI model that is passively involved in text generation so as to read on a human mental process under the BRI. Thus, due to the passive involvement of the AI model being inherited from the independent claims and not resolved in the dependent claims by only narrowing the model structure/training procedures of a passively involved model, these dependent claims are also directed towards patent ineligible subject matter under 35 U.S.C. 101.
Potentially Allowable Subject Matter
Claims 1-17 would be potentially allowable over the prior art of record if amended to overcome the preceding rejections under 35 U.S.C. 101 and 112(b).
The following is a statement of reasons for the indication of potentially allowable subject matter:
With respect to independent Claims 1, 9, and 17, the prior art of record taken individually or as a combination fails to explicitly teach or fairly suggest a respective method, electronic device, and processor-executable medium for performing automatic speech-to-text (STT) recognition by inputting speech into a particularly trained AI model. In particular, this AI model is trained according to a sequence of wherein clauses recited in the independent claims: wherein the first artificial intelligence algorithm model including a first model parameter outputs a first predicted text by using a synthetic speech and a reference text as input, extracts a first loss based on the first predicted text and the reference text, and performs a first pre-training based on the first loss, wherein the first artificial intelligence algorithm model that performed the first pre-training includes a second model parameter, wherein the first artificial intelligence algorithm model that performed the first pre-training outputs a second predicted text by using the synthetic speech and the reference text as input, and extracts a second loss based on the second predicted text and the reference text, wherein the electronic device determines a loss rate which is a ratio of the first loss and the second loss, determines an adaptation parameter based on the second model parameter and the second loss if the loss rate is below a threshold value, and determines a third model parameter based on the adaptation parameter and the second model parameter, wherein the first artificial intelligence algorithm model is repeatedly pre-trained such that the first artificial intelligence algorithm model that performed a second pre-training is configured to include the third model parameter.
Most pertinent Prior Art:
Similar to the claimed invention, Yue, et al. ("Exploring Machine Speech Chain for Domain Adaptation," 2022) deals with domain adaptation for end-to-end automatic speech recognition (Abstract). Yue, relies upon a process first using a text-to-speech (TTS) model to generate speech for a target domain text corresponding to the claimed synthetic speech for a reference text followed by using the TTS data as input to an ASR model for fine-tuning/adjusting model parameters (Section 1, Page 6757; Section 3, Page 6759). Yue, however, fails to teach the particularly claimed loss rate as a ratio of the first loss and second loss and how such a loss rate is considered in determining the claimed adaptation parameter.
Fazel, et al. ("SynthASR: Unlocking Synthetic Data for Speech Recognition," 2021) teaches adapting an ASR model to a new task using "synthetic speech." The adaptation process involves the use of a loss function that takes into consideration model parameters from a previous synthetic speech training iteration and a current iteration (Section 3.5, Page 3, Fig. 1). Like, Yue, however, Fazel also fails to teach the loss ratio-driven data filtering approach set forth in the independent claims.
Patent-based prior art in the form of Chen, et al. (U.S. PG Publication: 2021/0350786 A1) teaches the use of synthetic speech 306 for unpaired training text data for ASR model training based upon a loss function relying upon a difference between synthetic and non-synthetic recognition results (Fig. 3, Elements 304 and 306 (Fig. 3a; Paragraphs 0047 and 0073). Chen, however, also fails to teach the loss ratio-driven data filtering approach set forth in the independent claims.
Thus, the prior art of record fails to explicitly teach or fairly suggest the potentially allowable subject matter set forth in independent claims 1, 9, and 17.
The remaining dependent claims further limit and inherit the subject matter of respective independent claims containing potentially allowable subject matter, and thus, are also potentially allowable over the prior art of record by virtue of their dependency.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: see the preceding discussions of Yue, et al., Fazel, et al., and Chen, et al. provided in the Potentially Allowable Subject Matter section.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JAMES S WOZNIAK whose telephone number is (571)272-7632. The examiner can normally be reached 7-3, off alternate Fridays.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant may use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Flanders can be reached at (571)272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
JAMES S. WOZNIAK
Primary Examiner
Art Unit 2655
/JAMES S WOZNIAK/Primary Examiner, Art Unit 2655