Prosecution Insights
Last updated: October 04, 2026
Application No. 18/897,849

MODEL TRAINING METHOD AND APPARATUS, ELECTRONIC DEVICE AND COMPUTER READABLE MEDIUM

Final Rejection §101§102§103§112
Filed
Sep 26, 2024
Priority
Mar 22, 2024 — CN 202410334093.5
Examiner
LELAND III, EDWIN S
Art Unit
2654
Tech Center
2600 — Communications
Assignee
Mashang Consumer Finance Co. Ltd.
OA Round
2 (Final)
75%
Grant Probability
Favorable
3-4
OA Rounds
5m
Est. Remaining
75%
With Interview

Examiner Intelligence

Grants 75% — above average
75%
Career Allowance Rate
353 granted / 470 resolved
+13.1% vs TC avg
Minimal -0% lift
Without
With
+-0.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 5m
Avg Prosecution
14 currently pending
Career history
481
Total Applications
across all art units

Statute-Specific Performance

§101
17.8%
-22.2% vs TC avg
§103
44.1%
+4.1% vs TC avg
§102
16.0%
-24.0% vs TC avg
§112
14.4%
-25.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 470 resolved cases

Office Action

§101 §102 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Receipt is acknowledged of papers submitted under 35 U.S.C. 119(a)-(d), which papers have been placed of record in the file. Status of Claims Claims 15-17 are cancelled and claims 21-23 are newly added leaving claims 1-14 and 18-23 pending in this application. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 9, and 21-23 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Specifically, the term “in real time”, while denoting something that is done very quickly, has no commonly understood or defined upper bound on when an action transitions from being in real time, to not being in real time. The claims are therefore indefinite. Claims 22 and 23 inherit this deficiency from claim 9. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-14 and 18-23 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claims recite updating parameters based on total model loss and repeating until the total model loss converges, which is a mathematical concept. This judicial exception is not integrated into a practical application because the only additional elements are generic computer components performing generic computing tasks. The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the only additional elements are generic computer components performing generic computing tasks. As per claim 1, the following limitations are recited: A - performing feature extraction from a speech sample to obtain a speech feature, wherein the speech sample is a conversation speech between a customer and an agent in a customer service system; B - inputting the speech feature into an encoding network in a model, wherein the encoding network comprises cascaded encoding layers, and the encoding layer comprises a first encoding layer and a second encoding layer; C - decoding a first encoding feature to obtain an additional loss, wherein the first encoding feature is an encoding feature output by the first encoding layer; D - obtaining an encoding loss based on a second encoding feature output by the second encoding layer and an encoding label; E - obtaining a total encoding loss based on the additional loss, the encoding loss, and a preset first loss weight; F - inputting the second encoding feature output by the second encoding layer into a decoding network for decoding processing to obtain a total decoding loss; G - obtaining a total model loss based on the total encoding loss, the total decoding loss, and a preset second loss weight; H - updating parameters in the encoding network and the decoding network based on the total model loss, and training the model according to the updated parameters, until the total model loss converges, obtaining a trained model. Limitations A-H are all directed to a mathematical concept. The Subject Matter Eligibility analysis is as follows: Step 1: Is the claim to a process, machine, manufacture or composition of matter? YES Step 2A, prong 1: Does the claim recite an abstract idea, law of nature or natural phenomenon? YES Step 2A, prong 2: Does the claim recite additional elements that integrate the Judicial Exception into a practical application? NO Step 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception? NO Therefore the claim is not subject matter eligible. As per claim 10, the following limitations are recited: I - at least one processor J - a memory communicatively connected with the at least one processor, wherein, the memory stores one or more computer programs executed by the at least one processor, the one or more computer programs are executed by the at least one processor to enable the at least one processor to A - performing feature extraction from a speech sample to obtain a speech feature, wherein the speech sample is a conversation speech between a customer and an agent in a customer service system; B - inputting the speech feature into an encoding network in a model, wherein the encoding network comprises cascaded encoding layers, and the encoding layer comprises a first encoding layer and a second encoding layer; C - decoding a first encoding feature to obtain an additional loss, wherein the first encoding feature is an encoding feature output by the first encoding layer; D - obtaining an encoding loss based on a second encoding feature output by the second encoding layer and an encoding label; E - obtaining a total encoding loss based on the additional loss, the encoding loss, and a preset first loss weight; F - inputting the second encoding feature output by the second encoding layer into a decoding network for decoding processing to obtain a total decoding loss; G - obtaining a total model loss based on the total encoding loss, the total decoding loss, and a preset second loss weight; H - updating parameters in the encoding network and the decoding network based on the total model loss, and training the model according to the updated parameters, until the total model loss converges, obtaining a trained model. Limitations A-H are all directed to a mathematical concept and limitations I & J are directed to generic computing components performing generic computing tasks. The Subject Matter Eligibility analysis is as follows: Step 1: Is the claim to a process, machine, manufacture or composition of matter? YES Step 2A, prong 1: Does the claim recite an abstract idea, law of nature or natural phenomenon? YES Step 2A, prong 2: Does the claim recite additional elements that integrate the Judicial Exception into a practical application? NO Step 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception? NO Therefore the claim is not subject matter eligible. As per claim 18, the following limitations are recited: K - A non-transitory computer-readable storage medium storing a computer program thereon, wherein the computer program, when executed by a processor, implements the following steps A - performing feature extraction from a speech sample to obtain a speech feature, wherein the speech sample is a conversation speech between a customer and an agent in a customer service system; B - inputting the speech feature into an encoding network in a model, wherein the encoding network comprises cascaded encoding layers, and the encoding layer comprises a first encoding layer and a second encoding layer; C - decoding a first encoding feature to obtain an additional loss, wherein the first encoding feature is an encoding feature output by the first encoding layer; D - obtaining an encoding loss based on a second encoding feature output by the second encoding layer and an encoding label; E - obtaining a total encoding loss based on the additional loss, the encoding loss, and a preset first loss weight; F - inputting the second encoding feature output by the second encoding layer into a decoding network for decoding processing to obtain a total decoding loss; G - obtaining a total model loss based on the total encoding loss, the total decoding loss, and a preset second loss weight; H - updating parameters in the encoding network and the decoding network based on the total model loss, and training the model according to the updated parameters, until the total model loss converges, obtaining a trained model. Limitations A-H are all directed to a mathematical concept and limitation K is directed to generic computing components performing generic computing tasks. The Subject Matter Eligibility analysis is as follows: Step 1: Is the claim to a process, machine, manufacture or composition of matter? YES Step 2A, prong 1: Does the claim recite an abstract idea, law of nature or natural phenomenon? YES Step 2A, prong 2: Does the claim recite additional elements that integrate the Judicial Exception into a practical application? NO Step 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception? NO Therefore the claim is not subject matter eligible. As per claims 2, 11 and 19, the following limitations are added: L - decoding the first encoding feature by using an additional decoding network to obtain an additional decoding feature; M - obtaining the additional loss based on the additional decoding feature and a preset additional decoding label. Limitations L & M are directed to a mathematical concept. The Subject Matter Eligibility analysis therefore remains unchanged. As per claims 3, 12 and 20, the following limitations are added: N - the first encoding feature comprises the first feature at one-third of the encoding network, and/or the first encoding feature at two-thirds of the encoding network. Limitation N is directed to a mathematical concept. The Subject Matter Eligibility analysis therefore remains unchanged. As per claims 4 and 13, the following limitations are added: O - obtaining a first loss based on a decoding feature output by a first decoding layer and a first decoding label, wherein the decoding network comprises the first decoding layer and a second decoding layer; P - obtaining a decoding loss based on the decoding feature output by the 2nd decoding layer and a decoding label; Q - obtaining the total decoding loss based on the first loss, the decoding loss, and a preset third loss weight. Limitations O-Q are all directed to a mathematical concept. The Subject Matter Eligibility analysis therefore remains unchanged. As per claims 5 and 14, the following limitations are added: R - obtaining a score matrix of a previous layer based on a query matrix and a key value matrix in a previous encoding layer; S - obtaining a score matrix of a current layer based on a query matrix and a key value matrix in a current encoding layer; T - merging the score matrix of the previous layer and the score matrix of the current layer to obtain the encoding feature, and U - inputting the encoding feature into a next encoding layer. Limitations R-U are all directed to a mathematical concept. The Subject Matter Eligibility analysis therefore remains unchanged. As per claim 6, the following limitations are added: V - obtaining a first total encoding loss based on the second encoding feature output by the second encoding layer and the encoding label; W - inputting the second encoding feature output by the second encoding layer into the decoding network of the model to obtain a first total decoding loss; X - obtaining a first total model loss based on the first total encoding loss, the first total decoding loss and a preset second loss weight; Y - updating the parameters in the encoding network and the decoding network based on the first total model loss, training the model until a preset condition is reached, and obtaining a pre-trained model; Z - using the parameters of the pre-trained model as initial parameters for the encoding network and the decoding network. Limitations V-Z are all directed to a mathematical concept. The Subject Matter Eligibility analysis therefore remains unchanged. As per claim 7, the following limitations are added: A1 - updating the parameters in the encoding network and the decoding network based on the total model loss and a regularization term during a parameter update phase. Limitation A1 is directed to a mathematical concept. The Subject Matter Eligibility analysis therefore remains unchanged. As per claim 8, the following limitations are added: B1 - obtaining the speech sample and segmenting the speech sample to obtain speech segments; C1 - annotating the speech segment that belongs to noise and obtaining a noise label. Limitations B1 & C1 are directed to a mathematical concept. The Subject Matter Eligibility analysis therefore remains unchanged. As per claim 9, the following limitations are added: D1 - inputting a to-be-recognized speech into a speech recognition model to obtain a speech recognition result of the to-be-recognized speech; E1 - wherein the to-be-recognized speech is a speech stream signal inputted from a telephone user end in real time, Limitations D1 & E1 are directed to a mathematical concept. The Subject Matter Eligibility analysis therefore remains unchanged. As per claim 21, the following limitations are added: F1 - the trained model is a speech recognition model; G1 - wherein the speech recognition model is deployed in an intelligent speech system; H1 - wherein the intelligent speech system is an intelligent customer service system or an intelligent sales system; I1 - the intelligent speech system is used to perform speech recognition on a to-be-recognized speech and generate a corresponding response speech; J1 - wherein the to-be- recognized speech is a speech stream signal inputted from a telephone user end in real time. Limitations F1-J1 are directed to details of the technological environment that the abstract idea is implemented in. The Subject Matter Eligibility analysis therefore remains unchanged. As per claim 22, the following limitations are added: K1 - performing intent judgment on the speech recognition result of the to-be- recognized speech to obtain an intention corresponding to the to-be-recognized speech; L1 - obtaining a response text corresponding to the to-be-recognized speech based on a judgment logic and the intention corresponding to the to-be-recognized speech; and M1 - performing speech synthesis on the response text corresponding to the to-be- recognized speech to obtain a response speech corresponding to the to-be-recognized speech; N1 - wherein the response speech corresponding to the to-be-recognized speech is used to achieve intelligent response to the to-be-recognized speech. Limitations K1-N1 are directed to a mental process. The Subject Matter Eligibility analysis therefore remains unchanged. As per claim 23, the following limitations are added: G1 - the speech recognition model is deployed in an intelligent speech system; H1 - wherein the intelligent speech system is an intelligent customer service system or an intelligent sales system; O1 - the intelligent speech system is used to perform speech recognition on the to-be-recognized speech and generate a corresponding response speech. Limitations G1, H1 and O1 are directed to details of the technological environment that the abstract idea is implemented in. The Subject Matter Eligibility analysis therefore remains unchanged. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 1-4, 6-7, 9-13, 15-16 and 18-20 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Hu et al. (U.S. Patent Application Publication 2024/0304185). As per claims 1, 9-10 and 18, Hu et al. discloses: A model training apparatus (Figure 6 and Paragraphs [0062-0067]), comprising: at least one processor (Figure 6, item 610 and Paragraphs [0062-0067]); and, a memory communicatively connected with the at least one processor, wherein, the memory stores one or more computer programs executed by the at least one processor (Figure 6, item 620 and Paragraphs [0062-0067]), the one or more computer programs are executed by the at least one processor to enable the at least one processor to: perform feature extraction from a speech sample to obtain a speech feature (Paragraph [0027]); input the speech feature into an encoding network in a model, wherein the encoding network comprises cascaded encoding layers, and the encoding layer comprises a first encoding layer and a second encoding layer (Figure 2, items 210 & 220 and paragraph [0035]); decode a first encoding feature to obtain an additional loss, wherein the first encoding feature is an encoding feature output by the first encoding layer (Equations 3 & 4 and Paragraphs [0047-0048]); obtain an encoding loss based on a second encoding feature output by the second encoding layer and an encoding label, and obtain a total encoding loss based on the additional loss, the encoding loss, and a preset first loss weight (Equations 3 & 4 and Paragraphs [0047-0048]); input the second encoding feature output by the second encoding layer into a decoding network for decoding processing to obtain a total decoding loss (Equations 3 & 4 and Paragraphs [0047-0048]); obtain a total model loss based on the total encoding loss, the total decoding loss, and a preset second loss weight (Equations 5-7 and Paragraphs [0047-0048]); update parameters in the encoding network and the decoding network based on the total model loss, and train the model according to the updated parameters until the total model loss converges, obtain a trained model (Paragraphs [0003] & [0048]). Claim 1 is directed to the method of using the apparatus of claim 10, so is rejected for similar reasons. Claim 9 is directed to the model trained by the method of using the apparatus of claim 10, so is rejected for similar reasons. Claim 18 is directed to a computer readable medium containing instructions to cause a processor to act as the apparatus of claim 10, so is rejected for similar reasons. As per claims 2, 11 and 19, Hu et al. discloses all of the limitations of claims 1, 10 and 18 above. Hu et al. further discloses: decoding the first encoding feature by using an additional decoding network to obtain an additional decoding feature; obtaining the additional loss based on the additional decoding feature and a preset additional decoding label (Equations 3-7 and Paragraphs [0047-0048] – there are two decoding networks). As per claims 3, 12 and 20, Hu et al. discloses all of the limitations of claims 1, 10 and 18 above. Hu et al. further discloses: the first feature at one-third of the encoding network, and/or the first encoding feature at two-thirds of the encoding network (Figure 2 and Paragraphs [0047-0048] – the language predictor can be construed as part of the encoding network, so the first feature is at one third of the way through). As per claims 4 and 13, Hu et al. discloses all of the limitations of claims 1 and 10 above. Hu et al. further discloses: obtaining a first loss based on a decoding feature output by a first decoding layer and a first decoding label, wherein the decoding network comprises the first decoding layer and a second decoding layer; obtaining a decoding loss based on the decoding feature output by the M-th decoding layer and a decoding label; obtaining the total decoding loss based on the first loss, the decoding loss, and a preset third loss weight (Equations 3-7 and Paragraphs [0047-0048]). As per claims 6 and 15, Hu et al. discloses all of the limitations of claims 1 and 10 above. Hu et al. further discloses: obtaining a first total encoding loss based on the second encoding feature output by the second encoding layer and the encoding label; inputting the second encoding feature output by the second encoding layer into the decoding network of the model to obtain a first total decoding loss; obtaining a first total model loss based on the first total encoding loss, the first total decoding loss and a preset second loss weight; updating the parameters in the encoding network and the decoding network based on the first total model loss, training the model until a preset condition is reached, and obtaining a pre-trained model; using the parameters of the pre-trained model as initial parameters for the encoding network and the decoding network (Equations 3-7 and Paragraphs [0047-0048]). As per claims 7 and 16, Hu et al. discloses all of the limitations of claims 1 and 10 above. Hu et al. further discloses: updating the parameters in the encoding network and the decoding network based on the total model loss and a regularization term during a parameter update phase (Equations 3-7 and Paragraphs [0047-0048]). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-7, 9-14 and 18-23 are rejected under 35 U.S.C. 103 as being unpatentable over Hu et al. (U.S. Patent Application Publication 2024/0304185) in view of Wang et al (European Patent Application Publication EP 4478241). As per claims 1, 10 and 18, Hu et al. discloses: A model training apparatus (Figure 6 and Paragraphs [0062-0067]), comprising: at least one processor (Figure 6, item 610 and Paragraphs [0062-0067]); and, a memory communicatively connected with the at least one processor, wherein, the memory stores one or more computer programs executed by the at least one processor (Figure 6, item 620 and Paragraphs [0062-0067]), the one or more computer programs are executed by the at least one processor to enable the at least one processor to: perform feature extraction from a speech sample to obtain a speech feature (Paragraph [0027]); input the speech feature into an encoding network in a model, wherein the encoding network comprises cascaded encoding layers, and the encoding layer comprises a first encoding layer and a second encoding layer (Figure 2, items 210 & 220 and paragraph [0035]); decode a first encoding feature to obtain an additional loss, wherein the first encoding feature is an encoding feature output by the first encoding layer (Equations 3 & 4 and Paragraphs [0047-0048]); obtain an encoding loss based on a second encoding feature output by the second encoding layer and an encoding label, and obtain a total encoding loss based on the additional loss, the encoding loss, and a preset first loss weight (Equations 3 & 4 and Paragraphs [0047-0048]); input the second encoding feature output by the second encoding layer into a decoding network for decoding processing to obtain a total decoding loss (Equations 3 & 4 and Paragraphs [0047-0048]); obtain a total model loss based on the total encoding loss, the total decoding loss, and a preset second loss weight (Equations 5-7 and Paragraphs [0047-0048]); update parameters in the encoding network and the decoding network based on the total model loss, and train the model according to the updated parameters until the total model loss converges, obtain a trained model (Paragraphs [0003] & [0048]). Hu et al. fails to disclose but Wang et al. the same field of endeavor teaches: the speech sample is a conversation speech between a customer and an agent in a customer service system (Paragraph [0052]) It would obvious for a person having ordinary skill in the art at the effective filing date of the invention to modify method, apparatus and computer readable storage medium of Hu et al. with the Customer Service System of Wang et al. because it is a case of combining prior art elements according to known methods to yield predictable results, as the results of such a combination would be predicable. Claim 1 is directed to the method of using the apparatus of claim 10, so is rejected for similar reasons. Claim 18 is directed to a computer readable medium containing instructions to cause a processor to act as the apparatus of claim 10, so is rejected for similar reasons. As per claims 2, 11 and 19, the combination of Hu et al. and Wang et al. discloses all of the limitations of claims 1, 10 and 18 above. Hu et al. in the combination further discloses: decoding the first encoding feature by using an additional decoding network to obtain an additional decoding feature; obtaining the additional loss based on the additional decoding feature and a preset additional decoding label (Equations 3-7 and Paragraphs [0047-0048] – there are two decoding networks). As per claims 3, 12 and 20, the combination of Hu et al. and Wang et al. discloses all of the limitations of claims 1, 10 and 18 above. Hu et al. in the combination further discloses: the first feature at one-third of the encoding network, and/or the first encoding feature at two-thirds of the encoding network (Figure 2 and Paragraphs [0047-0048] – the language predictor can be construed as part of the encoding network, so the first feature is at one third of the way through). As per claims 4 and 13, the combination of Hu et al. and Wang et al. discloses all of the limitations of claims 1 and 10 above. Hu et al. in the combination further discloses: obtaining a first loss based on a decoding feature output by a first decoding layer and a first decoding label, wherein the decoding network comprises the first decoding layer and a second decoding layer; obtaining a decoding loss based on the decoding feature output by the M-th decoding layer and a decoding label; obtaining the total decoding loss based on the first loss, the decoding loss, and a preset third loss weight (Equations 3-7 and Paragraphs [0047-0048]). As per claim 6, the combination of Hu et al. and Wang et al. discloses all of the limitations of claims 1 above. Hu et al. in the combination further discloses: obtaining a first total encoding loss based on the second encoding feature output by the second encoding layer and the encoding label; inputting the second encoding feature output by the second encoding layer into the decoding network of the model to obtain a first total decoding loss; obtaining a first total model loss based on the first total encoding loss, the first total decoding loss and a preset second loss weight; updating the parameters in the encoding network and the decoding network based on the first total model loss, training the model until a preset condition is reached, and obtaining a pre-trained model; using the parameters of the pre-trained model as initial parameters for the encoding network and the decoding network (Equations 3-7 and Paragraphs [0047-0048]). As per claim 7, the combination of Hu et al. and Wang et al. discloses all of the limitations of claims 1 above. Hu et al. in the combination further discloses: updating the parameters in the encoding network and the decoding network based on the total model loss and a regularization term during a parameter update phase (Equations 3-7 and Paragraphs [0047-0048]). As per claim 9, the combination of Hu et al. and Wang et al. discloses all of the limitations of claims 1 above. Wang et al. in the combination further discloses: inputting a to-be-recognized speech into a speech recognition model to obtain a speech recognition result of the to-be-recognized speech; wherein the to-be-recognized speech is a speech stream signal inputted from a telephone user end in real time (Paragraphs [0052-0053]) As per claim 21, the combination of Hu et al. and Wang et al. discloses all of the limitations of claims 1 above. Wang et al. in the combination further discloses: the trained model is a speech recognition model; wherein the speech recognition model is deployed in an intelligent speech system; wherein the intelligent speech system is an intelligent customer service system or an intelligent sales system; the intelligent speech system is used to perform speech recognition on a to-be-recognized speech and generate a corresponding response speech; wherein the to-be- recognized speech is a speech stream signal inputted from a telephone user end in real time (Paragraphs [0052-0053]). As per claim 22, the combination of Hu et al. and Wang et al. discloses all of the limitations of claims 9 above. Hu et al. in the combination further discloses: performing intent judgment on the speech recognition result of the to-be- recognized speech to obtain an intention corresponding to the to-be-recognized speech; obtaining a response text corresponding to the to-be-recognized speech based on a judgment logic and the intention corresponding to the to-be-recognized speech; and performing speech synthesis on the response text corresponding to the to-be- recognized speech to obtain a response speech corresponding to the to-be-recognized speech; wherein the response speech corresponding to the to-be-recognized speech is used to achieve intelligent response to the to-be-recognized speech (Paragraphs [0026], [0033] & [0071]). As per claim 23, the combination of Hu et al. and Wang et al. discloses all of the limitations of claims 9 above. Wang et al. in the combination further discloses: the speech recognition model is deployed in an intelligent speech system; wherein the intelligent speech system is an intelligent customer service system or an intelligent sales system; the intelligent speech system is used to perform speech recognition on the to-be-recognized speech and generate a corresponding response speech (Paragraphs [0052-0053]). Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Hu et al. (U.S. Patent Application Publication 2024/0304185) in view of Su et al (Chinese Patent Application Publication 114999455). As per claims 8 and 17, Hu et al. discloses all of the limitations of claims 1 and 10 above. Hu et al. fails to disclose but Su et al, in the same field of endeavor teaches: obtaining the speech sample and segmenting the speech sample to obtain speech segments; annotating the speech segment that belongs to noise and obtaining a noise label (“when the multi-task learning training, the step 1 to obtain the noise voice feature input built voice activity detection network,”). It would be obvious for a person having ordinary skill in the art at the effective filing date of the invention to modify the method and apparatus of Hu et al. with the noise labeling of Su et al. because it is a case of combining prior art elements according to known methods to yield predictable results. Response to Arguments Applicant’s arguments, see Remarks, filed 7/1/2026, with respect to the rejections of claims 1-7 and 9-20 under 35 U.S.C. 102 have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new grounds of rejection is made in view of Wang et al. Applicant's arguments filed 7/1/2026 regarding the rejection of claims 1-8 and 10-20 under 35 U.S.C. 101 have been fully considered but they are not persuasive. First the applicant argues that the claims are not directed to an abstract idea. The recited limitations of the independent claims are directed to mathematical steps of a mathematical concept and the Examiner respectfully holds that they are thus directed to an abstract idea, The applicant further argues that the claims are eligible under prong two of step 2A because they are integrated into practical application, but fails to point out what additional element transforms the abstract idea into a Examiner Notes The Examiner cites particular columns and line numbers in the references as applied to the claims above for the convenience of the Applicant. Although the specified citations are representative of the teachings in the art and are applied to the specific limitations within the individual claim, other passages and figures may apply as well. It is respectfully requested that, in preparing responses, the Applicant fully considers the references in its entirety as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior art or as disclosed by the Examiner. Communications via Internet e-mail are at the discretion of the applicant and require written authorization. Should the Applicant wish to communicate via e-mail, including the following paragraph in their response will allow the Examiner to do so: “Recognizing that Internet communications are not secure, I hereby authorize the USPTO to communicate with me concerning any subject matter of this application by electronic mail. I understand that a copy of these communications will be made of record in the application file.” Should e-mail communication be desired, the Examiner can be reached at Edwin.Leland@USPTO.gov Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to EDWIN S LELAND III whose telephone number is (571)270-5678. The examiner can normally be reached 8:00 - 5:00 M-F. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Hai Phan can be reached at 571-272-6338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /EDWIN S LELAND III/Primary Examiner, Art Unit 2654
Read full office action

Prosecution Timeline

Sep 26, 2024
Application Filed
Apr 08, 2026
Non-Final Rejection mailed — §101, §102, §103
Jul 01, 2026
Response Filed
Aug 25, 2026
Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748919
Near Real-Time Natural Language Sequence Generation
3y 6m to grant Granted Sep 29, 2026
Patent 12748933
SYSTEM AND METHOD FOR LANGUAGE TRANSLATION
2y 2m to grant Granted Sep 29, 2026
Patent 12744039
SPEECH RECOGNITION METHOD AND APPARATUS
2y 3m to grant Granted Sep 22, 2026
Patent 12743588
NATURAL LANGUAGE GENERATION USING KNOWLEDGE GRAPH INCORPORATING TEXTUAL SUMMARIES
1y 6m to grant Granted Sep 22, 2026
Patent 12738284
DETERMINATION OF THE SIGNIFICANCE OF SPATIAL AUDIO PARAMETERS AND ASSOCIATED ENCODING
2y 2m to grant Granted Sep 15, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
75%
Grant Probability
75%
With Interview (-0.5%)
2y 5m (~5m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 470 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month