Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claims 21-40 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 2-18 of U.S. Patent No. 12,008,329 and claims 2-6,8-17,19,20 of U.S. Patent No. 11,694,038. Although the claims at issue are not identical, they are not patentably distinct from each other because the additional limitations in the ‘329,‘038 patents (the further defined first and second loss entropy functions, and a combined third output) are not necessary to realize the functionality in the instant invention.
There are two mappings below.
The first table maps, claim numbers of the issued US Patents that contain the features
found in the claims of the current patent application.
The second table contains the claims for all applications involved; use the first table as a reference lookup to the second table.
Please see mapping below.
18/669060
12,008,329
11,694,038
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
2
2
11
3
4
5
6
12
13
14
15
16
17
18
2
10
11
7
8
9
2
2
11
3
4
5
6
12
19
20
13
14
15
16
2
10
11
17
8
9
18/669060
12,008,329
11,694,038
21. A method for generating dynamic conversational responses through aggregated outputs of machine learning models, the method comprising:
detecting, with a user interface, a user action during a conversational interaction; determining, using an ensemble model comprising:
a first machine learning model and a second machine learning model, based on the user action,
wherein the first machine learning model is trained using a first loss function and the second machine learning model is trained using a second loss function;
a first output from the first machine learning model and a second output from the second machine learning model,
identifying a subset of conversational responses for the conversational interaction based on a combination of the first output and the second output,
wherein the subset of conversational responses is selected from a plurality of conversational responses;
and updating the user interface, during the conversational interaction, to present at least some of the subset of conversational responses.
22. The method of claim 21, further comprising determining a third output based on a weighted average of the first output and the second output, wherein a first weight for the first output is greater than a second weight for the second output.
23. The method of claim 22, wherein the first weight is twice the second weight.
24. The method of claim 21, wherein the first output comprises a first plurality of probabilities that summed to one, wherein each of the first plurality of probabilities corresponds to a respective user intent.
25. The method of claim 21, wherein the second output comprises a second plurality of probabilities that summed do not sum to one, wherein each of the second plurality of probabilities corresponds to a respective user intent.
26. The method of claim 21, further comprising determining a first feature input for the first machine learning model, wherein the first feature input comprises a matrix, and wherein the first output corresponds to a prediction based on a column of the matrix and the second output corresponds to a row of the matrix.
27. The method of claim 21, wherein the first machine learning model comprises training a single classifier per class, wherein samples of the class are positive samples and all other samples are negative samples.
28. One or more non-transitory computer-readable media comprising of instructions that, when executed by one or more processors, cause operations comprising: detecting, with a user interface, a user action during a conversational interaction; determining, using an ensemble model comprising a first machine learning model and a second machine learning model, based on the user action, a first output from the first machine learning model and a second output from the second machine learning model, wherein the first machine learning model is trained using a first loss function and the second machine learning model is trained using a second loss function; identifying a subset of conversational responses for the conversational interaction based on a combination of the first output and the second output, wherein the subset of conversational responses is selected from a plurality of conversational responses; and updating the user interface, during the conversational interaction, to present at least some of the subset of conversational responses.
29. The one or more non-transitory computer-readable media of claim 28, wherein the instructions further cause the one or more processors to determine a third output based on a weighted average of the first output and the second output, wherein a first weight for the first output is greater than a second weight for the second output.
30. The one or more non-transitory computer-readable media of claim 29, wherein the first weight is twice the second weight.
31. The one or more non-transitory computer-readable media of claim 28, wherein the first output comprises a first plurality of probabilities that summed to one, wherein each of the first plurality of probabilities corresponds to a respective user intent.
32. The one or more non-transitory computer-readable media of claim 28, wherein the second output comprises a second plurality of probabilities that summed do not sum to one, wherein each of the second plurality of probabilities corresponds to a respective user intent.
33. The one or more non-transitory computer-readable media of claim 28, wherein the instructions further cause operations comprising determining a first feature input for the first machine learning model, wherein the first feature input comprises a matrix, and wherein the first output corresponds to a prediction based on a column of the matrix and the second output corresponds to a row of the matrix.
34. The one or more non-transitory computer-readable media of claim 28, wherein the first machine learning model comprises training a single classifier per class, wherein samples of the class are positive samples and all other samples are negative samples.
35. A method comprising: detecting a user action associated with a conversational interaction; determining, using an ensemble model comprising a first model and a second model, based on the user action, a first output from the first model and a second output from the second model, wherein the first model is trained using a first loss function and the second model is trained using a second loss function; identifying a subset of conversational responses for the conversational interaction based on a combination of the first output and the second output, wherein the subset of conversational responses is selected from a plurality of conversational responses; and generating at least some of the subset of conversational responses for use.
36. The method of claim 35, further comprising determining a third output based on a weighted average of the first output and the second output, wherein a first weight for the first output is greater than a second weight for the second output.
37. The method of claim 36, wherein the first weight is twice the second weight.
38. The method of claim 35, wherein the first model comprises a plurality of convolutional neural networks comprising a first convolutional neural network having a first column size and a second convolutional neural network having a second column size.
39. The method of claim 35, further comprising determining a first feature input for the first model, wherein the first feature input is generated using Bidirectional Encoder Representations from Transformers (“BERT”).
40. The method of claim 35, further comprising determining a first feature input for the first model, wherein the first feature input is generated based on textual data using natural language processing.
2. A method for generating dynamic conversational responses through aggregated outputs of machine learning models, the method comprising: receiving a first user action during a conversational interaction with a user interface; generating a first output from a first machine learning model based on the first user action, wherein the first machine learning model is trained using a multi-class cross entropy loss function; generating a second output from a second machine learning model based on the first user action, wherein the second machine learning model is trained using a binary cross entropy loss function; determining a third output based on a weighted average of the first output and the second output; selecting a subset of dynamic conversational responses from a plurality of dynamic conversational responses based on the third output; and generating, at the user interface, the subset of dynamic conversational responses during the conversational interaction.
3. The method of claim 2, wherein the first output comprises a first plurality of probabilities that summed to one, wherein each of the first plurality of probabilities corresponds to a respective user intent.
4. The method of claim 2, wherein the second output comprises a second plurality of probabilities that summed do not sum to one, wherein each of the second plurality of probabilities corresponds to a respective user intent.
5. The method of claim 2, further comprising determining a first feature input for the first machine learning model, wherein the first feature input comprises a matrix, and wherein the first output corresponds to a prediction based on a column of the matrix and the second out-put corresponds to a row of the matrix.
6. The method of claim 2, wherein the first machine learning model comprises training a single classifier per class, wherein samples of the class are positive samples and all other samples are negative samples.
7. The method of claim 2, wherein the first machine learning model comprises a plurality of convolutional neural networks comprising a first convolutional neural network having a first column size and a second convolutional neural network having a second column size.
8. The method of claim 2, further comprising determining a first feature input for the first machine learning model, wherein the first feature input is generated using Bidirectional Encoder Representations from Transformers (“BERT”).
9. The method of claim 2, further comprising determining a first feature input for the first machine learning model, wherein the first feature input is generated based on textual data using natural language processing.
10. The method of claim 2, wherein determining the third output based on the weighted average of the first output and the second output comprises determining a first weight for the first output and a second weight for the second output, wherein the first weight is greater than the second weight.
11. The method of claim 10, wherein the first weight is twice the second weight.
12. A non-transitory computer-readable medium comprising of instructions that, when executed by one or more processors, cause operations comprising: receiving a first user action during a conversational interaction with a user interface; generating a first output from a first machine learning model, wherein the first machine learning model is trained using a multi-class cross entropy loss function; generating a second output from a second machine learning model, wherein the second machine learning model is trained using a binary cross entropy loss function; determining a third output based on a weighted average of the first output and the second output; selecting a subset of dynamic conversational responses from a plurality of dynamic conversational responses based on the third output; and generating, at the user interface, the subset of dynamic conversational responses during the conversational interaction.
13. The non-transitory computer-readable medium of claim 12, wherein determining the third output based on the weighted average of the first output and the second output comprises determining a first weight for the first output and a second weight for the second out-put, wherein the first weight is greater than the second weight.
14. The non-transitory computer-readable medium of claim 13, wherein the first weight is twice the second weight.
15. The non-transitory computer-readable medium of claim 12, wherein the first output comprises a first plurality of probabilities that summed to one, wherein each of the first plurality of probabilities corresponds to a respective user intent.
16. The non-transitory computer-readable medium of claim 12, wherein the second output comprises a second plurality of probabilities that summed do not sum to one, wherein each of the second plurality of probabilities corresponds to a respective user intent.
17. The non-transitory computer-readable medium of claim 12, wherein the instructions further cause operations comprising determining a first feature input for the first machine learning model, wherein the first feature input comprises a matrix, and wherein the first output corresponds to a prediction based on a column of the matrix and the second output corresponds to a row of the matrix.
18. The non-transitory computer-readable medium of claim 12, wherein the first machine learning model comprises training a single classifier per class, wherein samples of the class are positive samples and all other samples are negative samples.
19. The non-transitory computer-readable medium of claim 12, wherein the first machine learning model comprises a plurality of convolutional neural networks comprising a first convolutional neural network having a first column size and a second convolutional neural network having a second column size.
20. The non-transitory computer-readable medium of claim 12, wherein the instructions further cause operations comprising determining a first feature input for the first machine learning model, and wherein the first feature input is generated based on textual data using natural language processing.
A system for generating dynamic conversational responses through aggregated outputs of machine learning models, the system comprising: storage circuitry configured to store:
control circuitry configured to: receive a first user action during a conversational interaction with a user interface;
a first machine learning model, wherein the first machine learning model is trained using a multi-class cross entropy loss function; and a second machine learning model, wherein the second machine learning model is trained using a binary cross entropy loss function;
generate a first output from the first machine learning model; generate a second output from the second machine learning model; determine a third output based on a weighted average of the first output and the second output; and select a subset of the dynamic conversational responses from a plurality of dynamic conversational responses based on the third output; and input/output circuitry configured to: generate, at the user interface, the subset of the dynamic conversational responses during the conversational interaction.
1.A system for generating dynamic conversational responses through aggregated outputs of machine learning models, the system comprising: storage circuitry configured to store:
a first machine learning model, wherein the first machine learning model is trained using a multi-class cross entropy loss function; and a second machine learning model, wherein the second machine learning model is trained using a binary cross entropy loss function; control circuitry configured to: receive a first user action during a conversational interaction with a user interface; determine, based on the first user action, a first feature input for the first machine learning model; determine, based on the first user action, a second feature input for the second machine learning model; input the first feature input into the first machine learning model to generate a first output from the first machine learning model; input the first feature input into the second machine learning model to generate a second output from the second machine learning model; determine a third output based on a weighted average of the first output and the second output; and select a subset of the dynamic conversational responses from a plurality of dynamic conversational responses based on the third output; and input/output circuitry configured to: generate, at the user interface, the subset of the dynamic conversational responses during the conversational interaction.
2.A method for generating dynamic conversational responses through aggregated outputs of machine learning models, the method comprising: receiving a first user action during a conversational interaction with a user interface; determining, based on the first user action, a first feature input for a first machine learning model, wherein the first machine learning model is trained using a multi-class cross entropy loss function; determining, based on the first user action, a second feature input for a second machine learning model, wherein the second machine learning model is trained using a binary cross entropy loss function; inputting the first feature input into the first machine learning model to generate a first output from the first machine learning model; inputting the first feature input into the second machine learning model to generate a second output from the second machine learning model; determining
a third output based on a weighted average of the first output and the second output; selecting a subset of dynamic conversational responses from a plurality of dynamic conversational responses based on the third output; and generating, at the user interface, the subset of dynamic conversational responses during the conversational interaction.
3. The method of claim 2, wherein the first output comprises a first plurality of probabilities that summed to one, wherein each of the first plurality of probabilities corresponds to a respective user intent.
4. The method of claim 2, wherein the second output comprises a second plurality of probabilities that summed do not sum to one, wherein each of the second plurality of probabilities corresponds to a respective user intent.
5. The method of claim 2, wherein the first feature input comprises a matrix, and wherein the first output corresponds to a prediction based on a column of the matrix and the second output corresponds to a row of the matrix.
6. The method of claim 2, wherein the first machine learning model comprises training a single classifier per class, wherein samples of the class are positive samples and all other samples are negative samples.
7. The method of claim 2, wherein the first machine learning model comprises a plurality of convolutional neural networks comprising a first convolutional neural network having a first column size and a second convolutional neural network having a second column size.
8. The method of claim 2, wherein the first feature input is generated using Bidirectional Encoder Representations from Transformers (“BERT”).
9. The method of claim 2, wherein the first feature input is generated based on textual data using natural language processing.
10. The method of claim 2, wherein determining the third output based on the weighted average of the first output and the second output comprises determining a first weight for the first output and a second weight for the second output, wherein the first weight is greater than the second weight.
11. The method of claim 10, wherein the first weight is twice the second weight.
12. A non-transitory computer-readable media for generating dynamic conversational responses through aggregated outputs of machine learning models, comprising of instructions that, when executed by one or more processors, cause operations comprising: receive a first user action during a conversational interaction with a user interface; determine, based on the first user action, a first feature input for a first machine learning model, wherein the first machine learning model is trained using a multi-class cross entropy loss function; determine, based on the first user action, a second feature input for a second machine learning model, wherein the second machine learning model is trained using a binary cross entropy loss function; input the first feature input into the first machine learning model to generate a first output from the first machine learning model; input the first feature input into the second machine learning model to generate a second output from the second machine learning model; determine a third output based on a weighted average of the first output and the second output; select a subset of the dynamic conversational responses from a plurality of dynamic conversational responses based on the third output; and generate, at the user interface, the dynamic conversational responses during the conversational interaction.
13. The non-transitory computer readable media of claim 12, wherein the first output comprises a first plurality of probabilities that summed to one, wherein each of the first plurality of probabilities corresponds to a respective user intent.
14. The non-transitory computer readable media of claim 12, wherein the second output comprises a second plurality of probabilities that summed do not sum to one, wherein each of the second plurality of probabilities corresponds to a respective user intent.
15. The non-transitory computer readable media of claim 12, wherein the first feature input comprises a matrix, and wherein the first output corresponds to a prediction based on a column of the matrix and the second output corresponds to a row of the matrix.
16. The non-transitory computer readable media of claim 12, wherein the first machine learning model comprises training a single classifier per class, wherein samples of the class are positive samples and all other samples are negative samples.
17. The non-transitory computer readable media of claim 12, wherein the first machine learning model comprises a plurality of convolutional neural networks comprising a first convolutional neural network having a first column size and a second convolutional neural network having a second column size.
18. The non-transitory computer readable media of claim 12, wherein the first feature input is generated based on textual data using natural language processing.
19. The non-transitory computer readable media of claim 12, wherein determining the third output based on the weighted average of the first output and the second output comprises determining a first weight for the first output and a second weight for the second output, wherein the first weight is greater than the second weight.
20. The non-transitory computer readable media of claim 19, wherein the first weight is twice the second weight.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 21-40 are rejected under 35 U.S.C. 103 as being unpatentable over Aguilar Alas et al (20210104245) in view of Wan et al (20200335083).
As per claim 21, Aguilar Alas et al (20210104245) teaches a method for generating dynamic conversational responses through aggregated outputs of machine learning models, the method comprising (as using machine learned models – para 0034, for conversations – para 0042, 0105):
detecting, with a user interface, a user action during a conversational interaction; determining, using an ensemble model comprising
a first machine learning model and a second machine learning model, based on the user action, a first output from the first machine learning model and a second output from the second machine learning model, wherein the first machine learning model is trained using a first loss function and the second machine learning model is trained using a second loss function
identifying a subset of conversational responses for the conversational interaction based on a combination of the first output and the second output, wherein the subset of conversational responses is selected from a plurality of conversational responses (as, updating the trained model – para 0035, based on the updated parameters from the loss function calculations showing in para 0034);
and updating the user interface, during the conversational interaction, to present at least some of the subset of conversational responses (and using the identified category to either highlight part of the conversation as belonging to that category, or displaying the category – para 0044).
Although Aguilar Alas et al (20210104245) teaches two different loss function calculations (on two different models representing different data types), Aguilar Alas et al (20210104245) does not explicitly teach two different types of loss functions; Wan et al (20200335083) teaches the use of tuple-loss and cross entropy loss calculations (see para 0014, when the tuple loss is for the full set of N-tuples, then it is a function of cross entropy loss). Therefore, it would have been obvious to one of ordinary skill in the art of loss function implementation in language model execution to modify the loss functions in Aguilar Alas et al (20210104245) with differing loss functions, as taught by Wan et al (20200335083) because it would take advantage of the positives of both models, which would result in a more accurate/robust model ( Wan et al (20200335083) para 0005, last sentence).
As per claims 22,23, the combination of Aguilar Alas et al (20210104245) in view of Wan et al (20200335083) teaches the method of claim 21, further comprising determining a third output based on a weighted average of the first output and the second output, wherein a first weight for the first output is greater than a second weight for the second output ((Aguilar Alas et al (20210104245), as using weights in the model – para 0034, and further, developing and using a weighted summation for the model – see para 0119). Furthermore, the weights disclosed in Aguilar Alas et al (20210104245) are in a range that covers the claim scope of claim 37, towards twice the weight (see Aguilar Alas et al (20210104245) , see above, as well as, para 0136, wherein a percentage is chosen for the inter-relationship between the weights).
As per claims 24,25, the combination of Aguilar Alas et al (20210104245) in view of Wan et al (20200335083) teaches the method of claim 21, wherein the first output comprises a first plurality of probabilities that summed to one/not summer to one, wherein each of the first plurality/second plurality of probabilities corresponds to a respective user intent (Aguilar Alas et al (20210104245), as see the attention models re-concentrating the probabilities over the words that capture the sentiment states – para 0119 – 0120; para 0121 teaches the cross-entropy loss value to be between 0-1; the net loss value between the acoustic view ML model and the lexical ML value is not defined to be restricted between 0-1).
As per claim 26, the combination of Aguilar Alas et al (20210104245) in view of Wan et al (20200335083) teaches the method of claim 21, further comprising determining a first feature input for the first machine learning model, wherein the first feature input comprises a matrix, and wherein the first output corresponds to a prediction based on a column of the matrix and the second output corresponds to a row of the matrix ( Aguilar Alas et al (20210104245), as see para 0099, disclosing the use of CNN’s with input/intermediate/output layers; examiner notes that by definition, cascading neural networks contain a variable number of columns – see para 154, defining a feature vector of X values, and a D-dimensional vector in the other direction, for the feature space).
As per claim 27, the combination of Aguilar Alas et al (20210104245) in view of Wan et al (20200335083) teaches the method of claim 21, wherein the first machine learning model comprises training a single classifier per class, wherein samples of the class are positive samples and all other samples are negative samples (Aguilar Alas et al (20210104245), as the sentiment categories are represented by the machine learning models – see mappings above, wherein the categories include positive and negative – see para 0038).
As per claim 35, the combination of Aguilar Alas et al (20210104245) in view of Wan et al (20200335083) teaches a method comprising:
detecting a user action associated with a conversational interaction (Aguilar Alas et al (20210104245), as performing the recognition on conversations – para 0042, 0105);
determining, using an ensemble model comprising a first model and a second model, based on the user action, a first output from the first model and a second output from the second model, wherein the first model is trained using a first loss function and the second model is trained using a second loss function
identifying a subset of conversational responses for the conversational interaction based on a combination of the first output and the second output (as, updating the trained model – para 0035, based on the updated parameters from the loss function calculations showing in para 0034),
wherein the subset of conversational responses is selected from a plurality of conversational responses; and generating at least some of the subset of conversational responses for use (the combined sentiment detection model – para 0035, is used to determine the sentiment category for the conversation – para 0039, and using the identified category to either highlight part of the conversation as belonging to that category, or displaying the category – para 0044).
Although Aguilar Alas et al (20210104245) teaches two different loss function calculations (on two different models representing different data types), Aguilar Alas et al (20210104245) does not explicitly teach two different types of loss functions; Wan et al (20200335083) teaches the use of tuple-loss and cross entropy loss calculations (see para 0014, when the tuple loss is for the full set of N-tuples, then it is a function of cross entropy loss). Therefore, it would have been obvious to one of ordinary skill in the art of loss function implementation in language model execution to modify the loss functions in Aguilar Alas et al (20210104245) with differing loss functions, as taught by Wan et al (20200335083) because it would take advantage of the positives of both models, which would result in a more accurate/robust model ( Wan et al (20200335083) para 0005, last sentence).
As per claims 36,37, the combination of Aguilar Alas et al (20210104245) in view of Wan et al (20200335083) teaches the method of claim 35, further comprising determining a third output based on a weighted average of the first output and the second output, wherein a first weight for the first output is greater than a second weight for the second output (Aguilar Alas et al (20210104245), as using weights in the model – para 0034, and further, developing and using a weighted summation for the model – see para 0119). Furthermore, the weights disclosed in Aguilar Alas et al (20210104245) are in a range that covers the claim scope of claim 37, towards twice the weight (see Aguilar Alas et al (20210104245) , see above, as well as, para 0136, wherein a percentage is chosen for the inter-relationship between the weights).
As per claim 38, the combination of Aguilar Alas et al (20210104245) in view of Wan et al (20200335083) teaches the method of claim 35, wherein the first model comprises a plurality of convolutional neural networks comprising a first convolutional neural network having a first column size and a second convolutional neural network having a second column size (Aguilar Alas et al (20210104245), see para 0099, disclosing the use of CNN’s with input/intermediate/output layers; examiner notes that by definition, cascading neural networks contain variable column sizes).
As per claim 39, the combination of Aguilar Alas et al (20210104245) in view of Wan et al (20200335083) teaches the method of claim 35, further comprising determining a first feature input for the first model, wherein the first feature input is generated using Bidirectional Encoder Representations from Transformers (“BERT”) (Aguilar Alas et al (20210104245), see para 0132, using bidirectional LSTM’s).
As per claim 40, the combination of Aguilar Alas et al (20210104245) in view of Wan et al (20200335083) teaches the method of claim 35, further comprising determining a first feature input for the first model, wherein the first feature input is generated based on textual data using natural language processing (Aguilar Alas et al (20210104245), as the features, that are operated on, are derived from audio conversations, naturally language processed into features represented by text – see para 0042, audio data of speech conversations, determining sentiment, by processed text – para 0044).
Claims 28-34 are non-transitory computer readable medium claims whose steps are found throughout method claims 21-27,35-40 and as such, claims 28-34 are similar in scope and content to method claims 21-27,35-40; therefore, claims 28-34 are rejected under similar rationale as presented against claims 21-27, 35-40 above. Furthermore, Aguilar Alas et al (20210104245) teaches computer readable media storing instructions when processed by a processor, perform the disclosed steps – see para 0168, to be processed by processors in para 0169.
Response to Arguments
Applicant’s arguments with respect to the claim(s) have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Examiner notes the use of the Wan et al (20200335083) reference teaching sub-tuple loss and cross entropy loss (with full N-tuple). Furthermore, examiner notes Refaat et al (20210133582) teaching 2 separate loss functions in para 0058-0061. Also see Merler (20230259716) teaches logit loss as well as entropy loss.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Please see related art listed on the PTO-892 form.
The following references were found, to be pertinent, to features found in applicants disclosure/claims:
Refaat et al (20210133582) teaching 2 separate loss functions in para 0058-0061.
Merler (20230259716) teaches logit loss as well as entropy loss.
Chen et al (20200334520) teaches cross entropy loss (para 0071, 0073) with ML models with language/query input
Muffat (20200279105) teaches convolutional neural networks in a BERT system operating on queries (para 0025, 0033, 0037)
Desjardins (20190236482) teaches machine learning models on speech/conversations (para 0028) using cross entropy loss models (para 0048)
Barad (20190362269) teaches weighted cross entropy losses (para 0040) in training ML/AI models (para 0037).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Michael Opsasnick, telephone number (571)272-7623, who is available Monday-Friday, 9am-5pm.
If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Mr. Richemond Dorvil, can be reached at (571)272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free).
/Michael N Opsasnick/Primary Examiner, Art Unit 2658 08/20/2026