Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under, including the fee set forth in 37 CFR1.17(e), was filed in this application after final rejection. Since this application is eligiblefor continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e)has been timely paid, the finality of the previous Office action has been withdrawnpursuant to 37 CFR 1.114. Applicant's submission filed on 09/02/2026 has been entered.
Status of the Claims
Claims 1-14 are pending.
Response to Applicant’s Arguments
Rejection under 35 USC 101 has been withdrawn for the following reason:
Under Prong (2) of Step 2A, the goal is to determine whether the claim is directed to the recited exception by evaluating whether the claim as a whole integrates the recited judicial exception into a practical application of the exception. See MPEP 2106.04II(A).
In particular, evaluating integration into a practical application requires identifying whether there are any additional elements recited in the claim beyond the judicial exception and evaluating those additional elements, individually and in combination, to determine whether they integrate the exception into a practical application, using one or more of the considerations laid out by the Supreme Court and the Federal Circuit (“CAFC”). See MPEP 2106.04(d).
According to the Supreme Court, a patent may issue for the means or method of producing a certain result, or effect, and not for the result or effect produced. Diamond v. Diehr, 450 U.S. 175, 182 n. 7 (1981). Therefore, the focus is on whether the claim “focus on a specific means or method that improves the relevant technology or are instead directed to a result or effect that itself is the abstract idea and merely invoke generic processes and machinery”. Enfish, L.L.C. v. Microsoft Corp., 822 F.3d 1327, 1336 (Fed. Cir. 2016).
For example, in Enfish, the CAFC found it relevant to ask whether claims were directed to an improvement to computer functionality versus being directed to an abstract idea. Enfish, 822 F.3d at 1335. To that extent, the CAFC found that the claims were specifically directed to a self-referential table for a computer database. Id. at 1337. In particular, the claim language required a four step algorithm specifically directed to a self-referential table for a computer database that improved upon prior art information search and retrieval systems by employing a flexible, self-referential table to store data. Id. at 1336-37.
Therefore, the focus of the claims was on a specific asserted improvement in computer capabilities (i.e., the self-referential table for a computer database), not on economic or other tasks for which a computer was used in its ordinary capacity. Id. at 1336. See also MPEP 2106.04(d)I (“an improvement in the functioning of a computer or an improvement to other technology or technical field, as discussed in MPEP 2106.04(d)(1) and 2106.05(a)”).
Claim 1 recites a method for text data processing, wherein the method comprises:
obtaining a to-be-processed text, wherein the to-be-processed text comprises a plurality of characters; and
processing the to-be-processed text by using a target model to obtain a prediction result, wherein the target model is a machine learning model for natural language processing in artificial intelligence field, the prediction result indicates to split the to-be-processed text into a plurality of target character sets, each target character set of the plurality of target character sets comprises at least one character, the prediction result further comprises a plurality of first labels, each first label of the plurality of first labels indicates semantics of a respective target character set of the plurality of target character sets, the plurality of first labels are used to determine an intention of the to-be- processed text, and the prediction result indicates a target splitting manner corresponding to the to-be-processed text and the target splitting manner is obtained based on degrees of matching between the plurality of target character sets and the plurality of first labels, wherein:
the target model comprises an encoder and a decoder; and the processing the to-be-processed text by using the target model to obtain the prediction result comprises:
performing feature extraction by using the encoder, to generate a vector representation corresponding to each target character set of the plurality of target character sets; and
generating the prediction result based on the vector representation by using the decoder.
Claim 6 recites a method for model training, wherein the method comprises:
processing a to-be-processed text by using a target model to obtain a prediction result, wherein the target model is a machine learning model for natural language processing in artificial intelligence field, the to-be-processed text comprises a plurality of characters, the prediction result indicates to split the to-be-processed text into a plurality of target character sets, each target character set of the plurality of target character sets comprises at least one character, the prediction result further comprises a plurality of first labels, each first label of the plurality of first labels indicates semantics of a respective target character set of the plurality of target character sets, the plurality of first labels are used to determine a predicted intention of the to-be-processed text, and the prediction result indicates a target splitting manner corresponding to the to-be-processed text and the target splitting manner is obtained based on degrees of matching between the plurality of target character sets and the plurality of first labels, wherein:
the target model comprises an encoder and a decoder; and
the processing the to-be-processed text by using the target model to obtain the prediction result comprises:
performing feature extraction by using the encoder, to generate a vector representation corresponding to each target character set of the plurality of target character sets; and
generating the prediction result based on the vector representation by using the decoder; and
training the target model according to a target loss function to obtain a trained target model, wherein the target loss function indicates a similarity between the prediction result and an expected result corresponding to the to-be-processed text, the expected result corresponding to the to-be- processed text indicates to split the to-be-processed text into a plurality of second character sets, each second character set comprises at least one character, the expected result corresponding to the to-be-processed text further comprises a plurality of expected labels, one expected label indicates semantics of one second character set, and the plurality of expected labels are used to determine a correct intention of the to-be-processed text.
Claim 10 recites a corresponding apparatus comprising processors and memories.
Like the specifically asserted self-referential table for improving database search and retrieval in Enfish, claims 1, 6 and 10 specifically asserted an encoder-decoder structure for a natural language processing machine learning model to perform natural language understanding of to-be-processed text to determine user intentions; i.e., a specifically asserted natural language processing machine learning model providing a natural language understanding method with a stronger generalization capability (specification at ¶6 and ¶173).
Therefore, claims 1, 6, and 10 recited a specifically asserted natural language processing machine learning model structure that integrated / implemented intention understanding process into a practical application.
In response to “Regarding the feature of "the target splitting manner is obtained based on degrees of matching between the plurality of target character sets and the plurality of first labels," the Examiner has acknowledged that Zhang does not disclose this feature and relies on Ramaswamy for it, mapping Ramaswamy's T=0/T=1 complete/incomplete-command decisions to the claimed first labels and splitting manners. Office Action, pages 24-25. Applicant respectfully disagrees” and “The training entries the Examiner maps to "labels" (e.g., "check new mail // T=1"; "check new mail show // T=0") merely flag command completeness. Ramaswamy (US 2001/0056344 A1), [0029]. This is not a label that "indicates semantics of a respective target character set" as recited by claim 1 (emphasis added). Ramaswamy's "T" does not indicate the meaning or semantics of a character set it merely indicates whether a string of words is a complete command”.
Semantics, or meaning of a target character set is the idea that is conveyed or intended to be conveyed to the mind by the target character set.1
Ramaswamy teaches automatically decomposing sentences including multiple commands to recognize multiple commands in the same sentence (¶25) that does not require the user to indicate an end of a command and thereby avoiding unwanted delays (¶4). In one example, for the sentence “check for new mail show me the first one”, decompose the sentence into multiple commands comprising “check new mail” // T = 1 and “show me the first one” // T = 1 (¶¶79-81).
In other words, Ramaswamy automatically determines that one or more words comprising recognized text were conveyed or intended as complete commands without requiring the user to define the words as such.
Specifically, Ramaswamy teaches a target model built from training data for generating conditional probability values (i.e., degree of matching) P(T|S) based on feature functions (¶28), the feature functions include one or more words from processed training data along with the correct decision of T = 1 or 0 (¶¶47-48, equation 1).
To determine if one or more words comprising recognized text 30 or S is a complete command (¶27), determine which feature functions are present by calculating P(T=1|S) such that S is a complete command if and only if P(T=1|S) > P(T=0|S) (¶¶51-52, equation 4; i.e., P(T|S) corresponds to “degree of matching” of S with T = 0 or T = 1).
Therefore, given a sentence of target character sets (¶27, one or more words that comprises recognized text 30 denoted as S), Ramaswamy determines the meaning or semantics of the sentence as comprising target character sets intended to convey complete commands (¶27, T = 1 if S is a complete command) with labels T=1 corresponding to respective target character sets at least indicate the meaning or semantics of respective target character sets as conveying complete commands.
This automatic recognition of multiple commands in the same sentence does not require the user to indicate that the user intended the one or more words in the recognized text to be complete commands.
In response to “Furthermore, neither Zhang nor Ramaswamy discloses selecting a target splitting manner "based on degrees of matching between the plurality of target character sets and the plurality of first labels." Ramaswamy evaluates one hypothesized command boundary at a time and computes a probability P(TIS) to decide completeness. See Ramaswamy, FIG. 3 description. Ramaswamy does not evaluate multiple whole-text splitting manners and select one based on degrees of matching between character sets and their semantic labels. As such, Ramaswamy has not been shown to remedy at least the above admitted deficiency of Zhang”.
Ramaswamy teaches using command boundary identification process to automatically decompose (i.e., split) sentences including multiple commands to recognize multiple commands (i.e., target splitting manner) in the same sentence (¶25).
Specifically, Ramaswamy’s boundary identifier uses a target model (Fig. 1 and ¶¶28-29, maximum entropy model constructed from training data 70 for boundary identifier 40) to automatically determine whether one or more words comprising recognized text is meant to be a complete command (¶27) rather than asking the user to indicate a complete command, which is cumbersome and results in unwanted delays (¶4).
Ramaswamy provided an established function of outputting target character sets (i.e., one or more words comprising recognized text 30 per ¶27) labelled as complete commands to a natural language understanding system (¶¶25-26; note that the natural language understanding system is like the intent recognition model of Zhang to perform natural language understanding of words representing complete command as formal command) in order to automatically decompose (i.e., split) sentences including multiple commands to recognize multiple commands (i.e., target splitting manner) in the same sentence (¶25), where target splitting manner is obtained based on degrees of matching between the plurality of target character sets / words and the plurality of first labels (¶¶51-52, equation 4; i.e., P(T|S) corresponds to “degree of matching” of one or more words / recognized text denoted as S with T = 0 or T = 1).
In one example, for the sentence “check for new mail show me the first one”, decompose (i.e., split) the input utterance into multiple commands or target splitting manner comprising “check new mail” // T = 1 and “show me the first one” // T = 1 (¶¶79-81).
By automatically decomposing (i.e., splitting) sentences to recognize multiple commands (i.e., target splitting manner) in the same sentence (¶25), Ramaswamy predicts target splitting manners indicating the to be processed text into a plurality of target character sets with first labels indicating semantics of respective target character sets as complete commands, where a natural language understanding system / intent recognition model can use the first labels of complete commands to determine an intention of the to-be-processed text as formal commands (¶26, e.g., translate complete command “do I have any new messages” into CheckNewMessage ()).
In response to “Moreover, Zhang does not disclose the other claimed features such as "processing the to- be-processed text by using a target model to obtain a prediction result, wherein... the prediction result indicates to split the to-be-processed text into a plurality of target character sets,..., the prediction result indicates a target splitting manner corresponding to the to-be-processed text" as recited in claim 1. In Zhang, word segmentation is performed as a preprocessing step before the text is input to the intent recognition model, not as part of the prediction result output by the model. Zhang explicitly states: "In 5301, a to-be-recognized text is acquired" and "In 5302, word segmentation results of the to-be-recognized text are inputted to an intent recognition model." See Zhang, [0046]-[0047]. Thus, in Zhang, the word segmentation is performed prior to inputting the text into the model, and the model receives already-segmented words as input. The model's output (the prediction result) does not "indicate to split" the text the splitting has already occurred before the model processes the text. In contrast, the present claims require that the prediction result itself indicates how to split the to-be-processed text”.
In the established function of Zhang, a trained recognition model comprises a first recognition layer to recognize a sentence level intent of a text and a second recognition layer to recognize a word level intent of the text (Zhang, ¶22):
PNG
media_image1.png
497
716
media_image1.png
Greyscale
where word segmentation results / target splitting manner for “open the navigation app and take the highway” comprises “open”, “navigation app”, “take”, and “highway” (Zhang, ¶51). Said segmentation results are inputted into the intent recognition model / target model to produce four different word level intents or second intent results “NAVI” “NAVI” “HIGHWAY” “HIGHWAY” (Zhang, Fig 4). According to Zhang, recognizing word level intent of the text improves recognition performance of the intent recognition model (Zhang, ¶22).
In the established function of Ramaswamy, a target model was built from training data to recognize multiple commands in the same sentence by automatically decomposing sentences including multiple commands (Ramaswamy, ¶25) as output to a natural language understanding system (Ramaswamy, ¶26) like the intent recognition model of Zhang; i.e., to provide word level intent recognition.
For example, the established function of target model in Ramaswamy would decompose “open the navigation app and take the highway” into “open navi app” // T = 1 and “take highway” // T= 1 as inputs to the encoder in Fig. 4 of Zhang.
Therefore, instead of requiring the intent recognition model to determine four different word level intents “NAVI” “NAVI” “HIGHWAY” and “HIGHWAY” corresponding to “open” “navigation app” “take” and “highway”, applying the feature functions of Ramaswamy to determine a degree / probability of matching “open navigation app” as a complete command and a degree / probability of matching “take highway” as another complete command, the intent recognition model of Zhang / natural language understanding of “open navigation app” and “take highway” would reduce the delay from separately processing “open” “navigation app” “take” and “highway”.
By providing a plurality of first labels indicating semantics of a respective target character set of a plurality of target character sets / words as respective complete commands (¶¶79-81 “check new mail” // T=1 and “show me the first message” // T = 1) to the intent recognition model of Zhang / natural language understanding system, unwanted delays can be avoided (Ramaswamy, ¶4).
Finally, in Zhang, the word segmentation is indeed performed prior to inputting the text into the model.
However, the predictable use of prior art elements (i.e., the model built from training data comprising one or more words along with the correct decision / label T in Ramaswamy, ¶28 and ¶47) according to their established function (Zhang, Abstract, training neural network model according to word segmentation results of a plurality of training texts) merely required one ordinarily skilled in the art to train the neural network model according to feature functions (Ramaswamy, ¶28) corresponding to word segmentation results along with the correct corresponding labels of T indicating which word segmentation results are complete and which word segmentation results are incomplete.
The result neural network model can automatically decompose / split / segment words of target character sets into multiple complete commands indicating a target splitting manner that is obtained based on degrees of matching P(T|S) between the target character sets S and corresponding first labels T.
Claim Rejections - 35 USC § 103
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 103 that form the basis for the rejections under this section made in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 3, 6, 8-10, and 12 are rejected under 35 USC 103(a) as being unpatentable over Zhang et al. (US 2023/0004798 A1) in view of Ramaswamy et al. (US 2014/0156268 A1).
Regarding Claims 1 and 10, Zhang discloses an apparatus for text data processing (Figs. 6-7, Intent Recognition Apparatus 600 being implemented by device 700), wherein the apparatus comprises:
one or more processors (¶76, computing unit 701 / CPU); and
one or more memories coupled to the one or more processors and storing programming instructions for execution by the one or more processors (¶74 and ¶76, computer program stored in storage unit 708 as computer software program for intent recognition methods) to:
obtain a to-be-processed text, wherein the to-be-processed text comprises a plurality of characters (¶46, step S301, acquire to be recognized text; e.g., ¶51, “Open the navigation app and take the highway”); and
process the to-be-processed text by using a target model to obtain a prediction result (¶47, S302, input word segmentation results of the to be recognized text into an intent recognition model to output intent results), wherein the target model is a machine learning model for natural language processing in artificial intelligence field (¶51 and Fig. 4, intent recognition model being a trained neural network comprising a first recognition layer outputs scores between word segmentation results in the to be recognized text and candidate intents and a second recognition layer processing word segmentation results to obtain second intent results), the prediction result indicates to split the to-be-processed text into a plurality of target character sets, each target character set of the plurality of target character sets comprises at least one character (¶47, input word segmentation results of the to be processed text into the intent recognition model; ¶51, input word segmentation results of “Open the navigation app and take the highway” comprising “open”, “navigation app”, “take”, and “highway” into the intent recognition model to obtain first intent results through the first recognition layer “NAVI” and “HIGHWAY”), the prediction result further comprises a plurality of first labels, each first label of the plurality of first labels indicates semantics of a respective target character set of the plurality of target character sets (¶51 and Fig. 4, for semantic vectors h1 “open”, h2 “navigation app”, h3 “take”, and h4 “highway”, intent recognition model outputs first intent results “NAVI” and “HIGHWAY”), the plurality of first labels are used to determine an intention of the to-be-processed text (¶51, first recognition layer of the intent recognition model outputs scores between the word segmentation results in the to be recognized text and the candidate intents; per ¶48, that is, intent recognition is performed on the to be recognized text by using the intent recognition model), and the prediction result indicates a target splitting manner corresponding to the to-be-processed text and degrees of matching between the plurality of target character sets and the plurality of first labels (¶51, to be processed text “Open the navigation app and take the highway” were segmented into first semantic vectors h1, h2, h3, and h4 and the first recognition layer outputted a score matrix indicating scores between the word segmentation results in the to be recognized text and the candidate intents), wherein:
the target model comprises an encoder and a decoder (¶51, intent recognition model comprising an encoder layer and a decoder layer); and
the processing the to-be-processed text by using the target model to obtain the prediction result comprises:
performing feature extraction by using the encoder, to generate a vector representation corresponding to each target character set of the plurality of target character sets (¶27, perform embedding processing on segmented words to obtain word vector of each segmented words; ¶51, pass word vector of each word segmentation result through an encoder layer to obtain encoded word vectors); and
generating the prediction result based on the vector representation by using the decoder (¶27, splice encoding result and attention calculation result of each segmented word and decode the splicing result to generate first semantic vector of each segmented word; ¶51, obtain first semantic vector h1, h2, h3, and h4 corresponding to “open”, “navigation app”, “take”, and “highway”).
Zhang does not disclose the target splitting manner is obtained based on degrees of matching between the plurality of target character sets and the plurality of first labels.
Ramaswamy discloses a conversational natural language system (Fig. 1) using a trained target model (¶¶28-29, maximum entropy model constructed from training data 70 for boundary identifier 40) to perform natural language processing of to be processed text (¶49, maximum entropy models for natural language processing) to obtain a prediction result (¶¶27-28, boundary identifier 40 uses the model to generate conditional probability P(T|S) by selecting T which maximizes P(T|S)), wherein the target model is a machine learning model for natural language processing in AI field (¶29, a maximum entropy model constructed from training data 70 including a large number of training utterances relevant to a domain corresponding to complete commands; ¶¶46-49, process training data to produce feature functions 41 for Maximum Entropy Models; per ¶47, feature functions include one or more words from processed training data along with the correction decision), the prediction result indicates to split the to be processed text into a plurality of target character sets (¶25, apply command boundary identification process to automatically decompose sentences including multiple commands to recognize multiple commands in the same sentence; e.g., recognize complete commands “check new mail” and “show me the next message”), the prediction result further comprises a plurality of first label indicating semantics of respective target character set of the plurality of target character sets (¶27, given S corresponding to one or more words comprising the recognized text 30, boundary identifier 40 decides if S is a complete command: set T to T=1 if S is a complete command or T=0 if S is not a complete command; i.e., determining the meaning / semantics of one or more words as complete command or incomplete command), the plurality of first labels are used to determine an intention of the to be processed text (¶¶25-26, output recognized text 30 that is a complete command to natural language understanding system 61 to generate a formal command), and the prediction result indicates a target splitting manner corresponding to the to be processed text and the target splitting manner is obtained based on a degree of matching between the plurality of target character sets and the plurality of first labels (¶¶51-57, determine which feature functions 41 are present in a given processed utterance by first calculating P(T=1|S) equation 3 and S is a complete command if and only if P(T=1|S) > P(T=0|S)).
Ramaswamy provided an established function of outputting target character sets (i.e., one or more words comprising recognized text 30 per ¶27) labelled as complete commands to the natural language understanding system (¶¶25-26; note that the natural language understanding system is like the intent recognition model of Zhang to perform natural language understanding of words representing complete command as formal command) in order to automatically decompose sentences including multiple commands to recognize multiple commands in the same sentence (¶25); i.e., predict splits indicating the to be processed text into a plurality of target character sets with first labels indicating semantics of respective target character sets as complete commands, where a natural language understanding system / intent recognition model can use the first labels of complete commands to determine an intention of the to-be-processed text.
It would’ve been obvious to one ordinarily skilled in the art before the effective filing date of the invention to obtain the target splitting manner based on degrees of matching between the plurality of target character sets and the plurality of first labels (e.g., implement Ramaswamy’s feature functions including one or more words with correct decision (i.e., labels indicating words comprising recognized text are complete commands) and determine which feature functions are present in a given processed utterance per Ramaswamy, ¶51) in order to split the to-be-processed text into a plurality of target character sets (Ramaswamy, ¶25, automatically decompose sentences including multiple commands to recognize multiple commands in the same sentence) indicating a target splitting manner corresponding to the to be processed text (Ramaswamy, ¶¶79-81, for “check for new mail show me the first one”, decompose the input utterance into multiple commands comprising “check new mail” and “show me the first one”).
Regarding Claims 3 and 12, Zhang as modified by Ramaswamy discloses
wherein there are N splitting manners corresponding to the to-be-processed text, N is an integer greater than or equal to 1, and the target splitting manner belongs to the N splitting manners (Ramaswamy, ¶¶37-44 and ¶¶54-57, processed training data showing complete commands (T = 1) or incomplete commands (T = 0) (i.e., different manners of splitting “check new mail show me”) are used to produce feature functions per ¶46).
Regarding Claim 6, Zhang discloses a method for model training (Fig. 1), wherein the method comprises:
processing a to-be-processed text by using a target model to obtain a prediction result (¶19, acquire training text and first annotation intents of the plurality of training text; ¶20, configuring a neural network model including a first recognition layer to output a first intent result of the training text according to a first semantic vector of each segmented word in the training text outputted by a feature extraction layer of the neural network), wherein the target model is a machine learning model for natural language processing in artificial intelligence field (¶51 and Fig. 4, intent recognition model being a trained neural network comprising a first recognition layer outputs scores between word segmentation results in the to be recognized text and candidate intents and a second recognition layer processing word segmentation results to obtain second intent results), the to-be-processed text comprises a plurality of characters, the prediction result indicates to split the to-be-processed text into a plurality of target character sets, each target character set of the plurality of target character sets comprises at least one character (¶24, one example of training text is “Open the navigation app and take the highway” and word segmentation results corresponding to the training text are “open”, “navigation app”, “take”, and “highway”; ¶31, output a first intent result of the training text according to a first semantic vector of each segmented word in the training text), the prediction result further comprises a plurality of first labels, each first label of the plurality of first labels indicates semantics of a respective target character set of the plurality of target character sets (¶26, preset a plurality of candidate intents and a semantic vector corresponding to each candidate intent; ¶31, output a first intent result of the training text and a score between each segmented word in the training text and the candidate intent), the plurality of first labels are used to determine a predicted intention of the to-be-processed text (¶31, a candidate intent whose score exceeds a preset threshold is selected as the first intent result of the training text), and the prediction result indicates a target splitting manner corresponding to the to-be-processed text (¶40, train neural network model according to word segmentation results of training texts, first annotation intents, and second annotation intents; e.g., ¶24, training text “open the navigation app and take the highway” with word segmentations “open”, “navigation app”, “take”, and “highway”, first annotation intents include “NAVI” and “HIGHWAY”, and second annotation intents include “NAVI”, “NAVI”, “HIGHWAY”, and “HIGHWAY”), wherein:
the target model comprises an encoder (¶27, Bi-LSTM encoder) and a decoder (¶27, a LSTM decoder); and
the processing the to-be-processed text by using the target model to obtain the prediction result comprises:
performing feature extraction by using the encoder, to generate a vector representation corresponding to each target character set of the plurality of target character sets (¶27, step S102 of an intent recognition model training method performs embedding processing on segmented words to obtain word vector of each segmented word, and input word vector of each segmented word of training text into Bi-LSTM encoder to obtain encoding results); and
generating the prediction result based on the vector representation by using the decoder (¶27, splice the encoding result and attention calculation result of each segmented word and input the splicing result to a LSTM decoder to obtain decoding results); and
training the target model according to a target loss function to obtain a trained target model, wherein the target loss function indicates a similarity between the prediction result and an expected result corresponding to the to-be-processed text (¶35, calculating a loss function value according to the first intent results of the plurality of training texts and the first annotation intents (“expected result”) of the plurality of training texts, and using the loss function value to obtain the intent recognition model), the expected result corresponding to the to-be-processed text indicates to split the to-be-processed text into a plurality of second character sets, each second character set comprises at least one character, the expected result corresponding to the to-be-processed text further comprises a plurality of expected labels, one expected label indicates semantics of one second character set, and the plurality of expected labels are used to determine a correct intention of the to-be-processed text (¶24, in the example training text “Open the navigation app and take the highway”, word segmentation results are “open”, “navigation app”, “take” and “highway”, the annotation intent of the training text includes “NAVI” corresponding to “open” and “navigation app”, “HIGHWAY” corresponding to “take” and “highway”; i.e., second character set comprises characters “open” and “navigation app” corresponding to the annotation intent / label “NAVI”, and characters “take” and “highway” corresponding to the annotation intent / label “HIGHWAY”).
Zhang does not disclose the target splitting manner is obtained based on degrees of matching between the plurality of target character sets and the plurality of first labels.
Ramaswamy discloses a conversational natural language system (Fig. 1) using a trained target model (¶¶28-29, maximum entropy model constructed from training data 70 for boundary identifier 40) to perform natural language processing of to be processed text (¶49, maximum entropy models for natural language processing) to obtain a prediction result (¶¶27-28, boundary identifier 40 uses the model to generate conditional probability P(T|S) by selecting T which maximizes P(T|S)), wherein the target model is a machine learning model for natural language processing in AI field (¶29, a maximum entropy model constructed from training data 70 including a large number of training utterances relevant to a domain corresponding to complete commands; ¶¶46-49, process training data to produce feature functions 41 for Maximum Entropy Models; per ¶47, feature functions include one or more words from processed training data along with the correction decision), the prediction result indicates to split the to be processed text into a plurality of target character sets (¶25, apply command boundary identification process to automatically decompose sentences including multiple commands to recognize multiple commands in the same sentence; e.g., recognize complete commands “check new mail” and “show me the next message”), the prediction result further comprises a plurality of first label indicating semantics of respective target character set of the plurality of target character sets (¶27, given S corresponding to one or more words comprising the recognized text 30, boundary identifier 40 decides if S is a complete command: set T to T=1 if S is a complete command or T=0 if S is not a complete command; i.e., determining the meaning / semantics of one or more words as complete command or incomplete command), the plurality of first labels are used to determine an intention of the to be processed text (¶¶25-26, output recognized text 30 that is a complete command to natural language understanding system 61 to generate a formal command), and the prediction result indicates a target splitting manner corresponding to the to be processed text and the target splitting manner is obtained based on a degree of matching between the plurality of target character sets and the plurality of first labels (¶¶51-57, determine which feature functions 41 are present in a given processed utterance by calculating P(T=1|S) according to equation 3 and declare S is a complete command if and only if P(T=1|S) > P(T=0|S)).
Ramaswamy provided an established function of outputting target character sets (i.e., one or more words comprising recognized text 30 per ¶27) labelled as complete commands to the natural language understanding system (¶¶25-26; note that the natural language understanding system is like the intent recognition model of Zhang to perform natural language understanding of words representing complete command as formal command) in order to automatically decompose sentences including multiple commands to recognize multiple commands in the same sentence (¶25); i.e., predict splits indicating the to be processed text into a plurality of target character sets with first labels indicating semantics of respective target character sets as complete commands, where a natural language understanding system / intent recognition model can use the first labels of complete commands to determine an intention of the to-be-processed text.
It would’ve been obvious to one ordinarily skilled in the art before the effective filing date of the invention to obtain the target splitting manner based on degrees of matching between the plurality of target character sets and the plurality of first labels (e.g., implement Ramaswamy’s feature functions including one or more words with correct decision (i.e., labels indicating words comprising recognized text are complete commands) and determine which feature functions are present in a given processed utterance per Ramaswamy, ¶51) in order to split the to-be-processed text into a plurality of target character sets (Ramaswamy, ¶25, automatically decompose sentences including multiple commands to recognize multiple commands in the same sentence) indicating a target splitting manner corresponding to the to be processed text (Ramaswamy, ¶¶79-81, for “check for new mail show me the first one”, decompose the input utterance into multiple commands comprising “check new mail” and “show me the first one”).
Regarding Claim 8, Zhang discloses wherein before the processing a to-be-processed text by using a target model, the method further comprises:
obtaining a target data subset, wherein the target data subset comprises a first subset and a second subset, the first subset comprises a first character string and a first expected label corresponding to the first character string, and the second subset comprises a second character string and a second expected label corresponding to the second character string (¶38, acquire training data including the plurality of training texts, the first annotation intents of the plurality of training texts and second annotation intents of the plurality of training texts); and
determining, based on the target data subset, the to-be-processed text and the expected result corresponding to the to-be-processed text, wherein the to-be-processed text comprises the first character string and the second character string, and the expected result comprises the first expected label corresponding to the first character string and the second expected label corresponding to the second character string (¶38, training texts corresponding to the first annotation intents (i.e., first expected label corresponding to first character string) and training texts corresponding to the second annotation intents (i.e., second expected label corresponding to second character string)).
Regarding Claim 9, Zhang discloses wherein a quality score corresponding to the to-be-processed text meets a preset condition, and the quality score indicates quality of the to-be-processed text (¶31, obtaining, for each training text according to a first semantic vector of each segmented word in the training text and the semantic vector of the candidate intent, a second semantic vector of each segmented word and a score between each segmented word and the candidate intent, wherein the score between each segmented word and the candidate intent may be an attention score between the two; determine candidate intent whose score exceeds a preset threshold for selection as the first intent result of the training text).
Claims 2, 7, and 11 are rejected under 35 USC 103(a) as being unpatentable over Zhang et al. (US 2023/0004798 A1) and Ramaswamy et al. (US 2014/0156268 A1) as applied to claims 1, 6, and 10, in view of Pitschel et al. (US 9922642 B2).
Regarding Claims 2, 7, and 11, Zhang does not disclose wherein the plurality of first labels comprise at least two levels of labels, the at least two levels of labels comprise a parent label and a child label, and a belonging relationship exists between the parent label and the child label.
Pitschel teaches a device for training a digital assistant (Abstract), the device uses a model process to-be-processed text to obtain a prediction result comprising a plurality of first labels (Col 10, Rows 6-15, natural language processing module 332 associates sequence of words or tokens with one or more actionable intent), wherein the plurality of first labels comprise at least two levels of labels, the at least two levels of labels comprise a parent label and a child label, and a belonging relationship exists between the parent label and the child label (Col 11, Rows 12-33, an actionable intent node along with its linked property nodes; in one example, actionable intent node (i.e., parent labels) “restaurant reservation” with property nodes (i.e., child labels) “restaurant”, “date/time”, “party size”).
It would’ve been obvious to one ordinarily skilled in the art before the effective filing date of the invention to obtain prediction result comprising a plurality of first labels, the first labels comprising at least two levels of labels comprising parent label and child label, in order to determine an intention of the to be processed text (Pitschel, Col 10, Rows 6-11).
Claims 4-5 and 13-14 are rejected under 35 USC 103(a) as being unpatentable over Zhang et al. (US 2023/0004798 A1) and Ramaswamy et al. (US 2014/0156268 A1) as applied to claims 3 and 12, in further view of Xu et al. (CN 108959257 B, see IP.com translation).
Regarding Claims 4-5 and 13-14, Zhang does not teach wherein the processing the to-be-processed text by using a target model to obtain a prediction result comprises match each target character set with a plurality of character strings in a first data set, to determine a target character string that matches the target character set.
Xu discloses a natural language parsing device for processing to-be-processed text by splitting the to-be-processed text into a plurality of target character sets and to obtain corresponding intent prediction result (Abstract, cutting words of a natural language text into word cutting segment, carry out concept labeling on each word segments to obtain concept labels, arranging and combining the concept labels into concept label sequences to carry out intention deduction) comprising:
match each target character set with a plurality of character strings in a first data set, to determine a target character string that matches the target character set (p. 9, ¶4, match each word segment with a pre-established knowledge word list; p. 9, ¶5, “the knowledge word list comprises a plurality of concept labels and phrases corresponding to the concept labels”; i.e., match each word segment with phrases in the pre-established knowledge word list);
obtain, from the first data set, at least one second label corresponding to the target character string, wherein one character string comprises at least one character (p. 9, ¶5, “one concept label in the knowledge word list is a team, and the phrases corresponding to the lower side of the team comprise the names of teams such as a Chinese team, a British team, a Germany team and the like”); and
match, based on each target character set and the second label by using the target model, each target character set with a plurality of labels in the first data set, to obtain a first label that matches each target character set (p. 9, ¶6, “And S230, if the matched concept label exists in the knowledge word list, using the matched concept label as the at least one concept label”);
wherein the at least one second label comprises at least two second labels (p. 9, ¶5, “The knowledge word list comprises a plurality of concept labels and phrases corresponding to the concept labels”), after the obtaining, from the first data set, at least one second label corresponding to the target character string, the method further comprises
generate target indication information based on the to-be-processed text, the target character set, and at least two second labels by using the target model, wherein the target indication information indicates that each second label matches or does not match the target character set (p. 9, ¶7, “if the concept label corresponding to “beijing” in the knowledge vocabulary is “city”, and “road” does not having a matching result in the knowledge vocabulary, then the word segmentation of “road” is not labeled, and if the word of “Beijing road” also exists in the knowledge vocabulary and corresponds to the concept label of “road”, then the concept label of “road” is also labeled on the combination of the adjacent word segmentation of “Beijing road”.”);
screen the at least two second labels based on the target indication information, to obtain at least one screened label (p. 9, ¶7, “Therefore, all word segmentation segments and the combination of adjacent word segmentation segments with matched concept labels in the knowledge word list can be labeled”); and
the matching, based on each target character set and the at least one second label by using the target model, each target character set with a plurality of labels in the first data set comprises match, based on the target character set and the at least one screened label by using the target model, the target character set with the plurality of labels in the first data set (p. 9, ¶8, “And S240, arranging and combining the at least one concept label to obtain a plurality of concept label sequences, wherein word cutting boundaries covered by the concept labels in each concept label sequence are not overlapped among different concept label sequences”).
It would’ve been obvious to one ordinarily skilled in the art before the effective filing date of the invention to match each target character set with a plurality of character strings in a first data set to determine a target character string that matches the target character set, obtain, from the first data set, at least one second label corresponding to the target character string, wherein one character string comprises at least one character, and match, based on each target character set and the second label by using the target model, each target character set with a plurality of labels in the first data set, to obtain a first label that matches each target character set in order to optimize natural language parsing (Xu, p. 9, ¶2, “In this embodiment, optimization is performed based on the above embodiment”).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to examiner Richard Z. Zhu whose telephone number is 571-270-1587 or examiner’s supervisor Hai Phan whose telephone number is 571-272-6338. Examiner Richard Zhu can normally be reached on M-Th, 0730:1700.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/RICHARD Z ZHU/Primary Examiner, Art Unit 2654 09/05/2026
1 www.merriam-webster.com/dictionary/meaning