DETAILED ACTION
Notice of Pre-AIA or AIA Status
1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments/Amendments
2. With respect to 101 Abstract idea, Applicant argues on pages 1-3 of the Remarks that
“Amended claim 1 now recites two technical features that reflect a specific technical improvement to speech recognition technology. First, claim 1 recites "convert the first text data on an assumption of a speech error of the first text data, thereby to generate converted text data." This feature specifies how the conversion is performed - the text data is converted based on an assumption of a speech error, which is a specific data augmentation technique for generating training data that cannot practically be performed in the human mind. Second, claim 1 recites "perform learning of a speech recognition unit… so that the speech recognition unit recognizes a speech error in the speech data and generates text data in which the speech error is corrected." This feature specifies the technical improvement achieved by the learning - the speech recognition unit is trained to recognize and correct speech errors.
The specification confirms this technical improvement. As described in non-limiting paragraph [0041], "in a case where the converted text data are generated on the assumption of the speech error of the first text data, the speech recognizer 50 is allowed to recognize the speech error in the speech data and to generate the text data. Therefore, the speech recognizer 50 is also allowed to generate the text data in which the speech error is automatically corrected." As-Filed Specification, paragraph [0041]. The amended claims directly reflect this improvement.
In Ex parte Desjardins, Appeal 2024-000567 (ARP, Sept. 26, 2025), the USPTO Appeals Review Panel held that claims reciting a specific training strategy for a machine learning model that reflects a technical improvement described in the specification integrate the abstract idea into a practical application. The ARP stated that "the eligibility determination should turn on whether the claims are directed to an improvement to computer functionality versus being directed to an abstract idea.” Similarly here, amended claim 1 recites a specific training approach that helps to improve the speech recognition unit's ability to recognize and correct speech errors.
In summary, the claimed process as a whole - converting text data on an assumption of a speech error, generating speech data from the converted text, and performing learning using the text data and converted speech data so that the speech recognition unit recognizes and corrects speech errors - constitutes a specific technical improvement to speech recognition technology that helps to integrate any recited abstract idea into a practical application.
For at least the foregoing reasons, claim 1, as amended, is directed to statutory patent eligible subject matter. Accordingly, reconsideration and withdrawal of rejection of claim 1 are respectfully requested.
Claims 2-10 depend from claim 1, recite additional features, and are eligible for at least the same reasons as claim 1 and/or for the additional features cited therein. Accordingly, reconsideration and withdrawal of rejections of claims 2-10 are respectfully requested.
Claims 12 and 13 are independent claims reciting features analogous to claim 1. Therefore, claims 12 and 13 are eligible for reasons analogous to claim 1. Accordingly, reconsideration and withdrawal of rejections of claims 12 and 13 are respectfully requested.”
First, Examiner respectfully notes that the limitation recites in the independent claims as drafted covers a mental process. More specifically, the underlying abstract idea revolved around what happens once a human receives a misrecognized transcription (e.g., innovation). The human corrects the misrecognized transcription by converting the misrecognized transcription into corrected transcriptions (e.g., ivation, innoinnovation and innoashow), pronounces the corrected transcription and recognizes speech. The claim recites performing learning of a speech recognition unit by using the misrecognized transcription and the pronounce of the corrected transcription as input. The speech recognition unit trained with the training data are the misrecognized transcription and the pronunciation of the corrected transcription. However, claim does not recite any technical details on how the speech recognition unit is trained. Rather, the limitation only defines the input (e.g., training data) in training the speech recognition unit and does not include any details of how the training/learning is performed with the training data. The training step is recited at a high level of generality and thus is insignificant extra-solution activity. See MPEP 2106.05(g) (whether the limitation is significant”). Additionally, the recitation of “speech recognition unit” generally links the abstract idea to a particular field of use. See MPEP 2106.05(h).
Second, Examiner respectfully notes that Desjardins involved a particular approach to training a machine learning model leading to an improvement in the field of artificial intelligent model training. The present process differs in that the information processing system is not improved as a tool but uses the information processing system as a tool to implement an otherwise abstract mental process practically performed by a human. Thus, the claimed invented differs from Desjardins. See MPEP 2106.05(f) for a discussion of mere instructions to implement an otherwise abstract idea.
Applicant’s arguments are not persuasive, and thus for these reasons, Examiner respectfully disagrees. Consequently, the 101 abstract idea rejection is maintained.
Claim Rejections - 35 USC § 101
3. 35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
4. Claims 1-10 and 12-13 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e., a law of nature, a natural phenomenon, or an abstract idea) without significantly more.
Claim 1 recites
“[Claim 1] (Currently Amended) An information processing system comprising:
at least one memory that is configured to store instructions; and
at least one processor that is configured to execute the instructions to
acquire first text data;
convert the first text data on an assumption of a speech error of the first text data, thereby to generate converted text data;
generate converted speech data corresponding to the converted text data; and
perform learning of a speech recognition unit that generates, from speech data, text data corresponding to the speech data, by using the first text data and the converted speech data as inputs, so that the speech recognition unit recognizes a speech error in the speech data and generates text data in which the speech error is corrected.”
The independent claims 1, 12 and 13 recite substantially the same concept but do so in the context of a system, a method and a non-transitory recording medium.
The limitations recited in the independent claims as drafted cover mental processes. More specifically, the underlying abstract idea revolved around what happens once a human receives a misrecognized transcription (e.g., innovation). The human corrects the misrecognized transcription by converting the misrecognized transcription into corrected transcriptions (e.g., ivation, innoinnovation and innoashow), pronounces the corrected transcription and recognizes the speech. (Step 2A, Prong One:YES.)
Step 2A, Prong Two: This part of the eligibility analysis evaluates whether the claim as a whole integrates the recited judicial exception into a practical application of the exception or whether the claim is “directed to” the judicial exception. This evaluation is performed by (1) identifying whether there are any additional elements recited in the claim beyond the judicial exception, and (2) evaluating those additionally and in combination to determine whether the claim as a whole integrates the exception into a practical application.
The claim recites “perform learning of a speech recognition unit that generates, from speech data, text data corresponding to the speech data, by using the first text data and the converted speech data as inputs, so that the speech recognition unit recognizes a speech error in the speech data and generates text data in which the speech error is corrected.”
The limitation “perform learning of a speech recognition unit that generates, from speech data, text data corresponding to the speech data, by using the first text data and the converted speech data as inputs, so that the speech recognition unit recognizes a speech error in the speech data and generates text data in which the speech error is corrected” recites a mental process of recognizing speech. Claim also recites performing learning of a speech recognition unit. Claim does not recite any technical details on how the speech recognition unit is trained. Rather, the limitation only defines the input (e.g., training data) in training the speech recognition unit and does not include any details of how the training/learning is performed with the training data. The training step is recited at a high level of generality and thus is insignificant extra-solution activity. See MPEP 2106.05(g) (whether the limitation is significant”). Additionally, the recitation of “speech recognition unit” generally links the abstract idea to a particular field of use. See MPEP 2106.05(h).
Claims recite the additional limitations of a memory, a processor, a computer, and a non-transitory recording medium. The additional element(s) or combination of elements such as a memory, a processor, a computer, and a non-transitory recording medium in the claim(s) other than the abstract idea per se amount(s) to no more than (i) mere instructions to implement the idea on a computer, and/or (ii) recitation of generic computer structure that serves to perform generic computer functions that are well-understood, routine, and conventional activities previously known to the pertinent industry. Viewed as a whole, these additional claim element(s) do not provide meaningful limitation(s) to transform the abstract idea into a patent eligible application of the abstract idea such that the claim(s) amounts to significantly more than the abstract idea itself. Therefore, the claim(s) are rejected under 35 U.S.C. 101 as being directed to non-statutory subject matter. The paragraph [0015] discloses “[0015] The processor 11 reads a computer program. For example, the processor 11 is configured to read a computer program stored by at least one of the RAM 12, the ROM 13 and the storage apparatus 14. Alternatively, the processor 11 may read a computer program stored in a computer-readable recording medium, by using a not-illustrated recording medium reading apparatus. The processor 11 may acquire (i.e., may read) a computer program from a not- illustrated apparatus disposed outside the information processing system 10, through a network interface. The processor 11 controls the RAM12, the storage apparatus 14, the input apparatus 15, and the output apparatus 16 by executing the read computer program. Especially in the present example embodiment, when the processor 11 executes the read computer program, a functional block for performing learning/training of a speech recognizer, is realized or implemented in the processor 11. That is, the processor 11 may function as a controller for executing each control of the information processing system 10.” As filed in the specification, the computer is listed as a general-purpose computer and are mainly used as an application thereof. Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea.
Even when viewed in combination, these additional elements do not integrate the recited judicial exception into a practical application (Step 2A, Prong Two: NO), and the claim is directed to the judicial exception. (Step 2A: YES).
Step 2B: This part of the eligibility analysis evaluates whether the claim as a whole amounts to significantly more than the recited exception, i.e., whether any additional element, or combination of additional elements, adds an inventive concept to the claim. See MPEP 2106.05.
The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, the additional element of using a computer is noted as a general computer. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The claims are not patent eligible.
The dependent claims do not remedy the issues noted above. More specifically, claim 2 recites a mental process of “generate first speech data corresponding to the first text data;” For example, a human could pronounce a text. The limitation “perform the learning of the speech recognition unit by using the first text data, the converted speech data, and the first speech data, as inputs” recites performing the learning of the speech recognition unit at high level of generality. There is no technical detail about the performing is accomplished. Claim 3 recites a mental process of generating a plurality of texts based on conversation rule. Claim 4 recites a mental process of converting a text to a plurality of texts. The limitation “perform learning of the text data conversion unit by using the second text data” recites performing at high level of generality. There is no technical detail about the performing is accomplished. Claim 5 recites a mental process of determining whether two words is a speech error of the other. Claim 6 recites a mental process of presenting the second text to the user and acquiring the third text based on the feedback from the user. The limitation of “perform the learning of the text data conversion unit by using the second text data and the third text data” recites performing at high level of generality. There is no technical detail about the performing is accomplished. Claim 7 recites mental process of acquiring. Claim 8 recites a mental process of correcting speech error. Claim 9 recites using score to indicate a possibility that the speech data includes the speech error. Claim 10 recites a mental process of determining degree of tension in the meeting and determining whether or not to correct the speech error.
For at least the supra provided reasons, claims 1-10, 12-13 are rejected under 35 U.S.C. 101 as being directed to non-statutory subject matter.
Allowable Subject Matter
5. Claims 1-10, 12-13 are allowed in view of the prior art of record. The claims stand rejected under 101 Abstract idea, and for the application to pass to allowance this rejection need to be overcome. Any amendments to overcome the 101 rejection that results in any change in scope require further search and/or consideration in order to determine it allowability.
The following is a statement of reasons for the indication of allowable subject matter: the prior art(s) taken alone or in combination fail(s) to teach the following element(s) in combination with the other recited elements in the claim(s).
“convert the first text data on an assumption of a speech error of the first text data, thereby to generate converted text data;
generate converted speech data corresponding to the converted text data; and
perform learning of a speech recognition unit that generates, from speech data, text data corresponding to the speech data, by using the first text data and the converted speech data as inputs, so that the speech recognition unit recognizes a speech error in the speech data and generates text data in which the speech error is corrected.”
Claims 12 and 13 recite similar features as Claim 1.
The closest prior arts found as follows.
a. Zhou et al. (US 2019/0130896 A1). In this reference, Zhou et al. discloses a method and a system for training a deep end-to-end speech recognition model, wherein the deep end-to-end speech recognition model is used to generate text data from speech data. Zhou et al. synthesizes sample speech variations on original speech samples labelled with text transcriptions, and modifying a particular original speech sample to independently vary tempo and pitch of the original speech sample while retaining the labelled text transcription of the original speech sample, thereby producing multiple sample speech variations having multiple degrees of variation from the original speech sample (Zhou et al. [0009] Further sample speech variations can include synthesizing sample speech variations by further modifying the particular original speech sample to vary its volume, independently of varying the tempo and the pitch, and by applying temporal alignment offsets to the particular original speech sample, producing additional sample speech variations from the particular original speech sample and having the labelled text transcription of the original speech sample. Another disclosed variation can include a shift of the alignment between the original speech sample and the sample speech variation with temporal alignment offset of zero milliseconds to ten milliseconds. Some implementations of the disclosed method also include synthesizing sample speech variations by applying pseudo-random noise to the particular original speech sample, producing additional sample speech variations. In some implementations, the pseudo-random noise is generated from recordings of sound and combined with the original speech sample as random background noise.) Zhou et al. performs learning of a speech recognition unit by using the original speech sample, the labelled text transcription of the original speech sample and the sample speech variations from the particular original speech sample. Zhou et al. generates the sample speech variations from the particular original speech sample by modifying the original speech sample. Zhou et al. does not generate the sample speech variations from the converted text, wherein the converted text is generated from the first text (i.e., the labelled text transcription). Thus, Zhou et al. fails to teach and/or suggest the allowable subject matter noted above.
b. Lee et al. (US 2018/0137855 A1). In this reference, Lee et al. discloses a method and a system for training the natural language processing model (Lee et al. [0013] The second word may be generated by changing a portion of characters in the first word to other characters, or adding another character to the first word, [0014] The natural language processing method may also include: receiving a voice signal; extracting features from the received voice signal; recognizing a phoneme sequence from the extracted features through an acoustic model; and generating the sentence data by recognizing words from the phoneme sequence through a language model, [0015] In accordance with another embodiment, there may be provided a training method including: generating a changed word by applying noise to a word in sentence data; converting the changed word and another word to which noise may be not applied to corresponding word vectors; converting characters in the changed word and characters in the other word to corresponding character vectors; and generating a sentence vector based on the word vectors and the character vectors.) The training data for training the natural language processing model includes the converted text from the first text. However, Lee et al. does not generate speech corresponding to the converted text for training the natural language processing. Thus, Lee et al. fails to teach and/or suggest the allowable subject matter noted above.
b. Laird-McConnell et al. (US 11,349,679 B1). In this reference, Laird-McConnell et al. disclose a method for using training data include speech samples performed by different speakers reading a same passage in training an automatic speech recognition (Laird-McConnell et al. col. 5 lines 44-56 In some embodiments, training data used in training an automatic speech recognition model include speech samples performed by different speakers reading a same passage. In some embodiments, these different speakers include participants from different countries and have different language backgrounds. These speech samples are then fed into a deep neural network to train the speech recognition model. In some embodiments, training data used in training a natural language model or a text language model include different phrases that are tagged as a decision and/or a task. In some embodiments, a learning decision tree is implemented to determine a similarity between current communication with each agenda item.) Laird-McConnell et al. performs learning of speech recognition unit by using speech samples performed by different speakers reading the same passage. Laird-McConnell et al. does not convert the passage to another passage and generates the speech samples from another passage. Thus, Laird-McConnell et al. fail to teach and/or suggest the allowable subject matter noted above.
Conclusion
6. The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. See PTO-892.
a. Chen et al. (US 2021/0350786 A1.) In this reference, Chen et al. disclose a method for training a generative adversarial network (GAN)-based text-to-speech (TTS) model and a speech recognition model.
b. Williams, III et al. (US 2021/0335503 A1.) In this reference, Williams, III et al. disclose a method for updating speech and speaker recognition algorithm based on the received transcript feedback information.
c. MCQUISTON et al. (US 2020/0258525 A1). In this reference, MCQUISTON et al. disclose a method for updating the speech recognition algorithm based on the edited transcription.
7. THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
8. Any inquiry concerning this communication or earlier communications from the examiner should be directed to THUYKHANH LE whose telephone number is (571)272-6429. The examiner can normally be reached Mon-Fri: 9am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew C. Flanders can be reached on 571-272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/THUYKHANH LE/Primary Examiner, Art Unit 2655