Prosecution Insights
Last updated: October 01, 2026
Application No. 19/022,026

DIALOG SYSTEM AND METHOD WITH IMPROVED HUMAN-MACHINE DIALOG CONCEPTS

Non-Final OA §101§102§103
Filed
Jan 15, 2025
Priority
Sep 29, 2022 — EU PCT/EP2022/077210 +1 more
Examiner
LAM, PHILIP HUNG FAI
Art Unit
Tech Center
Assignee
Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e.V.
OA Round
1 (Non-Final)
84%
Grant Probability
Favorable
1-2
OA Rounds
10m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 84% — above average
84%
Career Allowance Rate
130 granted / 155 resolved
+23.9% vs TC avg
Strong +51% interview lift
Without
With
+50.9%
Interview Lift
resolved cases with interview
Typical timeline
2y 6m
Avg Prosecution
28 currently pending
Career history
177
Total Applications
across all art units

Statute-Specific Performance

§101
24.1%
-15.9% vs TC avg
§103
54.7%
+14.7% vs TC avg
§102
10.5%
-29.5% vs TC avg
§112
4.5%
-35.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 155 resolved cases

Office Action

§101 §102 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-9, 11-12, 16 and 20-28 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Claim 1 recites a system that, under the broadest reasonable interpretation, claims limitations that cover performance of the limitations in the human mind with the assistance of physical aids (e.g., pen and paper), but for the recitation of generic or well-known or conventional computer components. That is, other than reciting “input interface, a preprocessor, two or more extraction processors, an output interface”, nothing in these claim limitations precludes the steps from practically being performed in the mind. As a whole, claim 1 pertains to dialog understanding, which is a mental process that a human can do. Individually, each of the limitations also pertains to a mental process, and/or insignificant extra solution activity, for example: an input interface for acquiring an input representation of an input by receiving the input and deriving the input representation from the input or by receiving the input representation, the input representation being an audio signal representation or a speech representation or a text representation, wherein the input representation comprises a plurality of input representation elements, (e.g., listening to a user talk, and/or reading a user write, and with assistance of paper and pen, write down the words or break them down into phonemes.) a preprocessor for preprocessing the input representation to generate preprocessed information, such that the preprocessed information comprises a plurality of preprocessed information elements, and such that each of two or more of the plurality of preprocessed information elements depends on at least two of the plurality of information representation elements, (e.g., reading raw text and listening to voice request, look for keywords and context and group or combine them using pen and paper.) two or more information extraction processors, wherein each of the two or more information extraction processors is suitable to generate derived information from the preprocessed information according to an information extraction rule specific for the information extraction processor, and different from an information extraction rule of any other one of the two or more information extraction processors, (e.g., evaluate or determine different aspect of the information provided or received using set of rule or guideline, determine intent can be one task (what does the user want), and another could be domain identification (what is the category or topic).) and an output interface for generating an output, being an audio output and/or a textual output and/or visual output and/or being a signal for steering a machine, depending on the derived information from one or more of the two or more information extraction processors. (e.g., verbally saying aloud the command, writing them the command using pen and paper or drawing a picture to a machine that understands them.) The judicial exception is not integrated into a practical application. In particular, the claims only recites generic computing components. Such generic computing components are recited at a high-level of generality (i.e., as a generic processor performing a generic computer function of receiving, determining, or outputting information) such that they amount to no more than mere instructions to apply the exception using generic computer components. Accordingly, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea. Claim 1 does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional limitations of using generic computer components amount to no more than mere instructions to apply the exception using generic computer components. Mere instructions to apply an exception using generic computer components cannot provide an inventive concept. Claim 1 is not patent eligible. The examiner further notes that the use of claimed generic computer components (“input interface, a preprocessor, two or more extraction processors, an output interface”) to obtain, extract, and/or generate data invokes such generic computer components “merely as a tool to perform an existing process”. MPEP 2106.05(f). MPEP 2106.05(f) further explains: Use of a computer or other machinery in its ordinary capacity for economic or other tasks (e.g., to receive, store, or transmit data) or simply adding a general purpose computer or computer components after the fact to an abstract idea (e.g., a fundamental economic practice or mathematical equation) does not integrate a judicial exception into a practical application or provide significantly more. See Affinity Labs v. DirecTV, 838 F.3d 1253, 1262, 120 USPQ2d 1201, 1207 (Fed. Cir. 2016) (cellular telephone); TLI Communications LLC v. AV Auto, LLC, 823 F.3d 607, 613, 118 USPQ2d 1744, 1748 (Fed. Cir. 2016) (computer server and telephone unit). Similarly, "claiming the improved speed or efficiency inherent with applying the abstract idea on a computer" does not integrate a judicial exception into a practical application or provide an inventive concept. Intellectual Ventures I LLC v. Capital One Bank (USA), 792 F.3d 1363, 1367, 115 USPQ2d 1636, 1639 (Fed. Cir. 2015). Claim 1 recites generic computer components (“input interface, a preprocessor, two or more extraction processors, an output interface”), with respect to performing tasks. MPEP 2106.05(d) and (f) further provides examples of court decisions where the courts found generic computing components to be mere instructions to apply a judicial exception, and further explains “increased speed” (e.g., using a computer to increase the speed of an otherwise mental process) does not provide an inventive concept. For example: A commonplace business method or mathematical algorithm being applied on a general purpose computer, Alice Corp. Pty. Ltd. V. CLS Bank Int’l, 573 U.S. 208, 223, 110 USPQ2d 1976, 1983 (2014); Gottschalk v. Benson, 409 U.S. 63, 64, 175 USPQ 673, 674 (1972); Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015). A process for monitoring audit log data that is executed on a general-purpose computer where the increased speed in the process comes solely from the capabilities of the general-purpose computer, FairWarning IP, LLC v. Iatric Sys., 839 F.3d 1089, 1095, 120 USPQ2d 1293, 1296 (Fed. Cir. 2016) (emphasis added). Performing repetitive calculations. Bancorp Services v. Sun Life, 687 F.3d 1266, 1278, 103 USPQ2d 1425, 1433 (Fed. Cir. 2012) ("The computer required by some of Bancorp’s claims is employed only for its most basic function, the performance of repetitive calculations, and as such does not impose meaningful limits on the scope of those claims.") Claim 25 recites a method that corresponds to the system of claim 1 and is therefore rejected under the same/similar rationale or grounds as claim 1. Claim 25 is not patent eligible. Claim 27 recites a CRM claim that corresponds to the method of claim 1 and is therefore rejected under the same grounds as claim 1 above. While claim 27 further recites “non-transitory computer readable medium comprising a computer program”, these are merely generic computer components recited at a high-level of generality such that they amount to no more than mere instructions to apply the exception using a generic computer component. Therefore, none of these limitations (a) integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea or (b) amount to significantly more than the judicial exception, because in either case the additional limitations merely utilize generic computer components that amounts to no more than mere instructions to apply the exception using generic computer function. Claim 27 is not patent eligible. Regarding Claim 20, the analysis is similar to claim 1, an input interface for acquiring an input representation of an input by receiving the input and deriving the input representation from the input or by receiving the input representation, the input representation being an audio signal representation or a speech representation or a text representation, wherein the input representation comprises a plurality of input representation elements, (e.g., listening to a user talk, and/or reading a user write, and with assistance of paper and pen, write down the words or break them down into phonemes.) and two or more information extraction processors, wherein each of the two or more information extraction processors is suitable to generate derived information depending on the input representation according to an information extraction rule specific for the information extraction processor, and different from an information extraction rule of any other one of the two or more information extraction processors, (e.g., evaluate or determine different aspect of the information provided or received using set of rule or guideline, determine intent can be one task (what does the user want), and another could be domain identification (what is the category or topic).) and an output interface for generating an output, being an audio output and/or a textual output and/or visual output and/or being a signal for steering a machine, depending on the derived information from one or more of the two or more information extraction processors, (e.g., verbally saying aloud the command, writing them the command using pen and paper or drawing a picture to a machine that understands them.) wherein at least two information extraction processors of the two or more information extraction processors are dialog-state-dependent, (e.g., listening and read dialog or conversation and extracting information from them, like noting keywords and context of the words.) wherein the dialog system is configured to select one or more information extraction processors of the at least two information extraction processors, which are dialog-state-dependent, depending on a current state of the dialog, such that only those of the at least two information extraction processors, which are associated with the current state of the dialog, are selected, (e.g., selecting or choosing to read and extract information or to listen and then extract information from the dialog.) and wherein the one or more information extraction processors that have been selected are configured to generate the derived information depending on their information extraction rules. (e.g., read out aloud or write out a command according to rules.) Claim 26 recites a method that corresponds to the system of claim 1 and is therefore rejected under the same/similar rationale or grounds as claim 20. Claim 26 is not patent eligible. Claim 28 recites a CRM claim that corresponds to the method of claim 20 and is therefore rejected under the same grounds as claim 1 above. While claim 28 further recites “non-transitory computer readable medium comprising a computer program”, these are merely generic computer components recited at a high-level of generality such that they amount to no more than mere instructions to apply the exception using a generic computer component. Therefore, none of these limitations (a) integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea or (b) amount to significantly more than the judicial exception, because in either case the additional limitations merely utilize generic computer components that amounts to no more than mere instructions to apply the exception using generic computer function. Claim 28 is not patent eligible. Claims 2-9, 11-12, 16 and 21-24 depend from independent claims 1, and 20 respectively, do not remedy any of the deficiencies of claims 1 and 20, and therefore are rejected on the same grounds as claim 1, and 20 from above. Claim 2 further comprising: wherein the dialog system is configured to select at least one of the two or more information extraction processors, such that only those of the two or more information extraction processors that have been selected, are to generate, depending on their information extraction rules, the derived information. (e.g., selecting or choosing to read and extract information or to listen and then extract information from the dialog, read out aloud or write out a command according to rules.) Claim 3 further recite: wherein at least two of the two or more information extraction processors are to generate the derived information from the preprocessed information depending on their information extraction rules. (e.g., verbally saying aloud the command, writing them the command using pen and paper or drawing a picture according to information extraction rules.) Claim 4 further comprising: wherein said at least two of the two or more information extraction processors are configured to generate the derived information from the preprocessed information in parallel. (e.g., reading out loud in voice and handing out written text in a paper around the same time.) Claim 5 further recites: wherein the dialog system is configured to select the at least one of the two or more information extraction processors depending on a current state of a dialog, such that those of the two or more information extraction processors that have been selected, are to generate, depending on their information extraction rules, the derived information. (e.g., selecting type of output according to the state of dialog and information extraction rules, and provide output, either verbally or written in text format.) Claim 6 further recites: wherein at least two information extraction processors of the two or more information extraction processors are dialog-state-dependent, wherein the dialog system is configured to select one or more information extraction processors of the at least two information extraction processors, which are dialog-state-dependent, depending on the current state of the dialog, such that only those of the at least two information extraction processors, which are associated with the current state of the dialog, are selected, and wherein the one or more information extraction processors that have been selected are configured to generate the derived information depending on their information extraction rules. (e.g., selecting or choosing to read and extract information or to listen and then extract information from the dialog, read out aloud or write out a command according to rules.) Claim 7 further recites: wherein the dialog system comprises three or more information extraction processors as the two or more information extraction processors, wherein at least one information extraction processor of the three or more information extraction processors is dialog-state-independent, wherein the at least one information extraction processor, which is dialog- state-independent, is configured to always generate, depending on its information extraction rule, the derived information, independent from the current state. (e.g., listen or look for specific keywords in the conversation regardless of what the current topic of the discussion may be.) Claim 8 further recites: wherein each of at least two information extraction processors of the two or more information extraction processors is suitable to generate specific information being specific for said information extraction processor according to a modification rule, wherein said information extraction processor is suitable to generate the derived information from the specific information for said information extraction processor according to the information extraction rule specific for the information extraction processor, wherein said information extraction processor is suitable to generate the specific information for said information extraction processor according to the modification rule, such that the specific information for said information extraction processor is different from any specific information of any other information extraction processor of the at least two information extraction processors. (e.g., listen or read a conversation, apply specific rule to process the input data, extract specific piece of information using dedicated rule.) Claim 9 further recites: wherein each of at least one of the at least two information extraction processors is configured to generate the specific information for said information extraction processor using the derived information of another one of the at least two information extraction processors. (e.g., listen or read dialogue to extract information such as a user’s name, then use the name to lookup additional information) Claim 11 further recites: wherein each of the two or more information extraction processors is a classification unit, wherein each of the two or more classification units is suitable to generate the derived information from the preprocessed information such that the derived information indicates whether or not the input representation is associated with a class or indicates a probability that the input representation s associated with the class. (e.g., listen or read dialogue to extract information then make determining which categories or domain the information is in or make an educated guess or prediction the probability or chance that the information belongs to certain category.) Claim 12 further recites: wherein the preprocessed information comprises a numerical feature vector, wherein the plurality of preprocessed information elements comprises a plurality of numerical vector components of the feature vector. (e.g., listen or read dialogue to extract information then break down the text into words and use number to represent them.) Claim 16 further recites: wherein the pre-processor is configured to generate the preprocessed information such that each of the plurality of information elements depends on each of the plurality of information representation elements. (e.g., listen or read dialogue to extract and map information from the multiple information elements.) Claim 21 further recites: wherein the dialog system comprises three or more information extraction processors as the two or more information extraction processors, wherein at least one information extraction processor of the three or more information extraction processors is dialog-state-independent, wherein the at least one information extraction processor, which is dialog-state-independent, is configured to always generate, depending on its information extraction rule, the derived information, independent from the current state. (e.g., listen or read and extract information from the dialog, read out aloud or write out a command according to rules without regard to context of the current dialog state.) The analysis of claim 22 corresponds to claim 8 and therefore similar rationale of rejection is applied to the claim. Claim 23 further recites: wherein the input interface is configured to receive the input being a speech signal or an audio signal, wherein the input interface is configured to apply a speech recognition algorithm on the speech signal or on the audio signal to acquire a text representation of the speech signal or of the audio signal as the input representation. (e.g., listen to the dialogue and write them what is said, basically provide transcription.) The analysis of Claims 24 corresponds to claims 23 and therefore similar rationale of rejection is applied to the claim. In sum, claims 2-9, 11-12, 16 and 21-24 depend from claims 1, and 20 respectively, and further recite mental processes as explained above. None of the additional limitations recited in claims 2-9, 11-12, 16 and 21-24 amount to anything more than the same or a similar abstract idea as recited in claims 1 and 20. Nor do any limitations in claims 2-9, 11-12, 16 and 21-24: (a) integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea or (b) amount to significantly more than the judicial exception because the additional limitations of using generic computer components amounts to no more than mere instructions to apply the exception using generic computer components. Claims 2-9, 11-12, 16 and 21-24 are not patent eligible. Patent eligible claims Claims 10, 13, 14, 15, 17, 18, 19 are determined to be patent eligible because they contain sufficient technical details that overcomes the abstract idea and additionally may contain practical application or provide improvement to the technology field. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 1-3, 11-13, 16, 23, 25, and 27 are rejected under 35 U.S.C. 102 (a)(2) as being anticipated by Vu (US 20220229993). Regarding Claim 1, Vu discloses:1. A dialog system ([0006] chatbot system), comprising: an input interface ([0051] user interface via master bot, also see fig. 2, preprocessing sybsystem (210)) for acquiring an input representation of an input by receiving the input and deriving the input representation from the input or by receiving the input representation, the input representation being an audio signal representation or a speech representation or a text representation, wherein the input representation comprises a plurality of input representation elements, ([0085] Pre-processing subsystem 210 receives an utterance “A” 202 from a user and processes the utterance through a language detector 212 and a language parser 214. As indicated above, an utterance can be provided in various ways including audio or text. The utterance 202 can be a sentence fragment, a complete sentence, multiple sentences, and the like. Utterance 202 can include punctuation. For example, if the utterance 202 is provided as audio, the pre-processing subsystem 210 may convert the audio to text using a speech-to-text converter (not shown) that inserts punctuation marks into the resulting text, e.g., commas, semicolons, periods, etc.) a preprocessor (see fig. 2, preprocessing (210)) for preprocessing the input representation to generate preprocessed information, such that the preprocessed information comprises a plurality of preprocessed information elements, and such that each of two or more of the plurality of preprocessed information elements depends on at least two of the plurality of information representation elements, ([0087] Language parser 214 parses the utterance 202 to extract part of speech (POS) tags for individual linguistic units (e.g., words) in the utterance 202. POS tags include, for example, noun (NN), pronoun (PN), verb (VB), and the like. Language parser 214 may also tokenize the linguistic units of the utterance 202 (e.g., to convert each word into a separate token) and lemmatize words. A lemma is the main form of a set of words as represented in a dictionary (e.g., “run” is the lemma for run, runs, ran, running, etc.). Other types of pre-processing that the language parser 214 can perform include chunking of compound expressions, e.g., combining “credit” and “card” into a single expression “credit card.” Language parser 214 may also identify relationships between the words in the utterance 202. For example, in some embodiments, the language parser 214 generates a dependency tree that indicates which part of the utterance (e.g., a particular noun) is a direct object, which part of the utterance is a preposition, and so on. The results of the processing performed by the language parser 214 form extracted information 205 and are provided as input to MIS 220 together with the utterance 202 itself.) two or more information extraction processors (see fig. ref 242 320, 400 and 4000 in figs 2, 3 and 4A and 4B), wherein each of the two or more information extraction processors is suitable to generate derived information from the preprocessed information according to an information (intent 322, intent 475, entity/constraint 480) extraction rule specific (252, 254, 352 354 in figs 2 and 3) for the information extraction processor (see para 0092 for master bot and para 106 for skills bot), and different from an information extraction rule of any other one of the two or more information extraction processors,(see para 0092-0093, also see fig. 4A, as skill bot invocation, intent prediction, and entity detection all inherently use specific information extraction rule) and an output interface (see fig. 3, conversation manager (330) for generating an output (Dialog output to user (335), being an audio output and/or a textual output and/or visual output and/or being a signal for steering a machine ([0048] Each skill associated with a digital assistant helps a user of the digital assistant complete a task through a conversation with the user, where the conversation can include a combination of text or audio inputs provided by the user and responses provided by the skill bots. These responses may be in the form of text or audio messages to the user and/or using simple user interface elements (e.g., select lists) that are presented to the user for the user to make selections.) also see para 0108, depending on the derived information from one or more of the two or more information extraction processors. (ref 242, 320, see figs 2 and 3) Regarding Claim 2, Vu discloses all the element of claim 1, Vu further discloses: wherein the dialog system is configured to select at least one of the two or more information extraction processors, such that only those of the two or more information extraction processors that have been selected, are to generate, depending on their information extraction rules, the derived information. ([0106] Intent classifier 320 is configured to match a received utterance (e.g., utterance 306 or 308) to an intent associated with skill bot system 300. As explained above, a skill bot can be configured with one or more intents, each intent including at least one example utterance that is associated with the intent and used for training a classifier. In the embodiment of FIG. 2, the intent classifier 242 of the master bot system 200 is trained to determine confidence scores for individual skill bots and confidence scores for system intents. Similarly, intent classifier 320 can be trained to determine a confidence score for each intent associated with the skill bot system 300. Whereas the classification performed by intent classifier 242 is at the bot level, the classification performed by intent classifier 320 is at the intent level and therefore finer grained. The intent classifier 320 has access to intents information 354. The intents information 354 includes, for each intent associated with the skill bot system 300, a list of utterances that are representative of and illustrate the meaning of the intent and are typically associated with a task performable by that intent. The intents information 354 can further include parameters produced as a result of training on this list of utterances.) Also see ref (252, 254, 352 354 in figs 2 and 3). Regarding Claim 3, Vu discloses all the element of claim 1, Vu further discloses: wherein at least two of the two or more information extraction processors are to generate the derived information from the preprocessed information depending on their information extraction rules. ([0119] Alternatively, the skill bot invocation stage 415, the intent prediction stage 420, and the entity detection stage 422 may be conducted sequentially with one stage using the outputs of the other as inputs or one stage being invokes in a particular manner for a specific skill bot based on the outputs of the other. For instance, for a given text data 405, a skill bot invoker can invoke a skill bot through implicit invocation using the skill bot invocation stage 415 and the task prediction models 460. The task prediction models 460 can be trained, using machine-learning and/or rules-based training techniques, to determine a likelihood that an utterance is representative of a task that a particular skill bot 470 is configured to perform. Then for an identified or invoked skill bot and a given text data 405, the intent prediction stage 420 and intent prediction models 465 and/or the entity detection stage 422 and the entity extraction models 467 can be used to match a received utterance (e.g., utterance within given data asset 445) to an intent 475 associated with skill bot. As explained herein, a skill bot can be configured with one or more intents, each intent including at least one example utterance that is associated with the intent and used for training a classifier. In some embodiments, the skill bot invocation stage 415, the task prediction models 460, and the entity extraction models 467 for the master bot system are trained to determine confidence scores for individual skill bots and confidence scores for system intents. Similarly, the intent prediction stage 420 and intent prediction models 465 and/or the entity detection stage 422 and the entity extraction models 467 can be trained to determine a confidence score for each intent associated with the skill bot system. Whereas the classification performed by the skill bot invocation stage 415, the task prediction models 460, and the entity extraction models 467 are at the bot level, the classification performed by the intent prediction stage 420 and intent prediction models 465 and/or the entity detection stage 422 and the entity extraction models 467 are at the intent level and therefore finer grained.) Also see fig. 4A. Regarding Claim 11, Vu discloses all the element of claim 1, Vu further discloses: wherein each of the two or more information extraction processors is a classification unit, wherein each of the two or more classification units is suitable to generate the derived information from the preprocessed information such that the derived information indicates whether or not the input representation is associated with a class or indicates a probability that the input representation s associated with the class. ([0092] EIS 230 determines whether the utterance that it receives (e.g., utterance 206 or utterance 208) contains an invocation name of a skill bot. In certain embodiments, each skill bot in a chatbot system is assigned a unique invocation name that distinguishes the skill bot from other skill bots in the chatbot system. A list of invocation names can be maintained as part of skill bot information 254 in data store 250. An utterance is deemed to be an explicit invocation when the utterance contains a word match to an invocation name. If a bot is not explicitly invoked, then the utterance received by the EIS 230 is deemed a non-explicitly invoking utterance 234 and is input to an intent classifier (e.g., intent classifier 242) of the master bot to determine which bot to use for handling the utterance. In some instances, the intent classifier 242 will determine that the master bot should handle a non-explicitly invoking utterance. In other instances, the intent classifier 242 will determine a skill bot to route the utterance to for handling.) Regarding Claim 12, Vu discloses all the element of claim 1, Vu further discloses: wherein the preprocessed information comprises a numerical feature vector, wherein the plurality of preprocessed information elements comprises a plurality of numerical vector components of the feature vector. ([0115] In some examples, feature engineering 435 may include transforming data assets 445 into feature vectors and/or creating new features will be created using the data assets 445.) Regarding Claim 13, Vu discloses all the element of claim 12, Vu further discloses: wherein the input interface is configured to acquire a raw input text as the input representation, being a sequence of words, ([0125] The BERT model 4400 is a pre-trained algorithm that accepts one or more sequences of words from a user utterance(s) or system query(ies) as an input and generates one or more feature vectors (word embeddings) for each of the one or more words of the one or more sequences of words) wherein the preprocessor is configured to tokenize the raw input text using a tokenization method to acquire a plurality of tokens, ([0125] In some examples, the input sequence of words is tokenized to generate a plurality of word tokens.) wherein the preprocessor is configured to generate a multi-dimensional numerical vector for each of the plurality of tokens to acquire a plurality of multi-dimensional numerical vectors, wherein the preprocessor is configured to generate the numerical feature vector of the preprocessed information by combining the plurality of multi-dimensional numerical vectors for the plurality of tokens. ([0125] The BERT model 4400 is a pre-trained algorithm that accepts one or more sequences of words from a user utterance(s) or system query(ies) as an input and generates one or more feature vectors (word embeddings) for each of the one or more words of the one or more sequences of words. For example, as shown in FIG. 4B, for an input sequence of words “I would like to pay Merchant $10,” BERT model 4400 generates a separate word embedding for each individual word of the sequence (“I,” “would,” “like,” “to,” “pay,” “Merchant,” and “$10”). In some examples, BERT model 4400 includes at least one transformer layer for receiving the input sequence of words. In some examples, the at least one transformer layer includes a plurality of encoders. In some examples, each encoder includes a plurality of attention mechanisms and a plurality of feed-forward networks. In some examples, the input sequence of words is tokenized to generate a plurality of word tokens. In some examples, the plurality of attention mechanisms operates directly on the words of the input of sequence of words. In some examples, the plurality of attention mechanisms operates on the plurality of word tokens. In some examples, the plurality of attention mechanisms generates an attention score for each word of the input sequence of words or each token of the plurality of word tokens. In some examples, the input sequence of words (or the plurality of word tokens) and the attention scores are input into the plurality of feed-forward networks. In some examples, the plurality of feed-forward networks encodes the input sequence of words (or the plurality of word tokens) into a plurality of word embeddings.) Regarding Claim 16, Vu discloses all the element of claim 1, Vu further discloses: wherein the pre-processor is configured to generate the preprocessed information such that each of the plurality of information elements depends on each of the plurality of information representation elements. ([0085] Pre-processing subsystem 210 receives an utterance “A” 202 from a user and processes the utterance through a language detector 212 and a language parser 214. As indicated above, an utterance can be provided in various ways including audio or text. The utterance 202 can be a sentence fragment, a complete sentence, multiple sentences, and the like. Utterance 202 can include punctuation. For example, if the utterance 202 is provided as audio, the pre-processing subsystem 210 may convert the audio to text using a speech-to-text converter (not shown) that inserts punctuation marks into the resulting text, e.g., commas, semicolons, periods, etc.) also see para 0087. Regarding Claim 23, Vu discloses all the elements of claim 1. Vu further discloses: wherein the input interface is configured to receive the input being a speech signal or an audio signal, wherein the input interface is configured to apply a speech recognition algorithm on the speech signal or on the audio signal to acquire a text representation of the speech signal or of the audio signal as the input representation. ([0040] User inputs 110 are generally in a natural language form and are referred to as utterances. A user utterance 110 can be in text form, such as when a user types in a sentence, a question, a text fragment, or even a single word and provides it as input to digital assistant 106. In some examples, a user utterance 110 can be in audio input or speech form, such as when a user says or speaks something that is provided as input to digital assistant 106. The utterances are typically in a language spoken by the user. For example, the utterances may be in English, or some other language. When an utterance is in speech form, the speech input is converted to text form utterances in that particular language and the text utterances are then processed by digital assistant 106. Various speech-to-text processing techniques may be used to convert a speech or audio input to a text utterance, which is then processed by digital assistant 106. In some examples, the speech-to-text conversion may be done by digital assistant 106 itself.) Regarding Claim 25, it is a method claim that corresponds to the system of claim 1 and is therefore rejected under the same grounds as claim 1 above. Regarding Claim 27, Vu discloses: A non-transitory computer-readable medium comprising a computer program ([0013] Some embodiments of the present disclosure include a system including one or more data processors and a non-transitory computer readable storage medium containing instructions which, when executed on the one or more data processors, cause the one or more data processors to perform part or all of one or more methods and/or part or all of one or more processes disclosed herein.) As for the rest of the claim, they claim elements from the claim 1, therefore the rationale applied in rejection of claim 1 is also applicable to claim 27. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 4 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Vu, in view of Vig (US 20180113854). Regarding Claim 4, Vu discloses all the element of claim 3, However, its not clear if Vu discloses parallel data processing. Vig in the related art discloses: wherein said at least two of the two or more information extraction processors are configured to generate the derived information from the preprocessed information in parallel. ([0034] FIG. 3 presents a block diagram illustrating an exemplary architecture of a conversational structure system utilizing the conversational structure extraction method, according to embodiments of the present invention. A conversational structure system 300 may divide a conversation and extract conversational structure, according to embodiments, in parallel with multiple processors.) Vu and Vig are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Vu to combine the teaching of Vig, because the method described would enables the user to extract conversational structure on a larger scale and at a finer level of detail than previous systems, and can feed a comprehensive analytics and business intelligence platform. (Vig, [0034]). Regarding Claim 15, Vu and Vig discloses all the element of claim 4, Vu further discloses: wherein each information extraction processor of the two or more information extraction processors comprises a neural network, wherein the neural network comprises at least one of an attention layer, a pooling layer and a fully-connected layer, wherein the neural network is configured to receive the preprocessed information as input, and is configured to output the derived information; or wherein the neural network is configured to receive the specific information for said information extraction processor as input, and is configured to output the derived information. ([0128] In some examples, the first set of vector representations is input into the CNN/BiLSTM model 4700. Based on the first set of vector representations, the CNN of the CNN/BiLSTM model 4700 generates one or more character-level vector representations for each character of each word of the input sequence of words. The one or more character-level vector representations is then concatenated and/or interpolated with the first set of vector representations and input into the BiLSTM network to generate one or more sentence-level vector representations for the input sequence of words. In some examples, the one or more sentence-level vector representations represents named entity tag scores. In some examples, one or more vectors generated by the context tag vectorizer 4600 are concatenated and/or interpolated with one or more sentence-level vector representations generated by the CNN/BiLSTM model 4700 to generate a second set of vector representations. In some examples, the second set of vector representations represent named entity tag scores. In some examples, the named entity tag scores are decoded into named entities using CRF model 4800.) Claim 17 is rejected under 35 U.S.C. 103 as being unpatentable over Vu, in view of Johns (US 20190132334). Regarding Claim 17, Vu discloses all the element of claim 1, However, its not clear if Vu disclose using neural network as part of preprocessing. Johns in the related art discloses: wherein the pre-processor comprises a neural network which is configured to receive the plurality of input representation elements as input, and which is configured to output the plurality of preprocessed information elements as output, wherein the neural network comprises at least two of an attention layer, a pooling layer and a fully-connected layer. ([0068] Herein, as shown in FIG. 3A, operating as an input layer to the CNN, a pre-processor (shown at 420 in FIG. 4A) of the cyber-security system extracts a section of binary code from the received executable file (operation 305). Additionally, the pre-processor generates an input, namely a representation of the binary code (operation 310). The input is provided to CNN-based logic (shown at 430 of FIG. 4A), which includes the convolution logic, the pooling logic and the FCN logic as described herein. The convolution logic, pooling logic and FCN logic are trained using supervised learning and generates an output in response to that input (operation 315).) Vu and Johns are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Vu to combine the teaching of Johns, because the CNN is efficient and effective in detecting patterns in raw data (Johns, [0068]). Claims 18-19 are rejected under 35 U.S.C. 103 as being unpatentable over Vu, in view of Ramani (US 20220327488). Regarding Claim 18, Vu discloses all the element of claim 1, Although Vu discloses BERT see para 0128 and dense vector involving utterance, Vu does not explicitly discloses sentence embedding involving three or more numerical vector elements. Ramani in the related art discloses: wherein the input representation comprises a numerical multi-dimensional sentence representation vector, or wherein the preprocessor is configured to generate the numerical multi-dimensional sentence representation vector from the input representation, ([0100] discloses Sentence-BERT model which generates sentence vector or embeddings.) wherein the multi-dimensional sentence representation vector comprises three or more numerical vector elements, wherein each of the three or more numerical vector elements is associated with one of a plurality of dimensions. ([0092] Each of the above items is referred to below as a“feature”. In order to accurately extract the above features, the system represents each line as a 300-dimensional vector of numbers.) Vu and Rami are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Vu to combine the teaching of Rami, because there is a need for extracting information using deep learning and natural language processing techniques (Rami, [Background]). Regarding Claim 19, Vu discloses all the element of claim 1, Vu does not disclose sentence similarity comparison. Rami in the related art discloses: wherein, for each two pairs of the plurality of numerical multi-dimensional sentence representation vectors for a plurality of sentences of the input representation, two numerical multi-dimensional sentence representation vectors of a first one of the two pairs of the numerical multi-dimensional; sentence representation vectors that identify two first sentences with semantically related meaning comprise a smaller spatial distance in a multi- dimensional space, in which the plurality of numerical multi-dimensional sentence representation vectors is defined, than two numerical multi-dimensional sentence representation vectors of a second one of the two pairs of the numerical multi-dimensional sentence representation vectors that identify two second sentences with semantically non-related meaning; ([0100] As may be seen from the above examples, the algorithm uses contextual inference, as opposed to mere phrase extraction. To achieve this, recent developments in the field of Natural Language Processing, specifically—Deep Learning are relied upon. In an exemplary embodiment, the Sentence-BERT model, which is an improvement to the earlier BERT model, is employed by the algorithm. While the latter provides rich representations that capture context for words, the former is able to provide representations optimized for the task of sentence similarity. As a consequence, this model generates similar representations for similar sentences. In this context, a “representation” is expressed as an array of numbers, and in a mathematical sense, a “representation” is a vector.) or, the preprocessor is configured to generate the plurality of numerical multi- dimensional sentence representation vectors, such that for each two pairs of the plurality of numerical multi-dimensional sentence representation vectors, two numerical multi-dimensional sentence representation vectors of a first one of the two pairs of the numerical multi-dimensional sentence representation vectors that identify two first sentences with semantically related meaning comprise a smaller spatial distance in a multi-dimensional space, in which the plurality of numerical multi-dimensional sentence representation vectors is defined, than two numerical multi-dimensional sentence representation vectors of a second one of the two pairs of the numerical multi-dimensional sentence representation vectors that identify two second sentences with semantically non-related meaning. ([0100] As may be seen from the above examples, the algorithm uses contextual inference, as opposed to mere phrase extraction. To achieve this, recent developments in the field of Natural Language Processing, specifically—Deep Learning are relied upon. In an exemplary embodiment, the Sentence-BERT model, which is an improvement to the earlier BERT model, is employed by the algorithm. While the latter provides rich representations that capture context for words, the former is able to provide representations optimized for the task of sentence similarity. As a consequence, this model generates similar representations for similar sentences. In this context, a “representation” is expressed as an array of numbers, and in a mathematical sense, a “representation” is a vector.) Where the rationale for the combination would be similar to the one already provided. Claims 20, 24, 26 and 28 are rejected under 35 U.S.C. 103 as being unpatentable over Vu, in view of Suendermann (US 20110046951). Regarding Claim 20, Vu discloses: 20. A dialog system([0006] chatbot system), comprising: an input interface ([0051] user interface via master bot, also see fig. 2, preprocessing sybsystem (210)) for acquiring an input representation of an input by receiving the input and deriving the input representation from the input or by receiving the input representation, the input representation being an audio signal representation or a speech representation or a text representation, wherein the input representation comprises a plurality of input representation elements, ([0085] Pre-processing subsystem 210 receives an utterance “A” 202 from a user and processes the utterance through a language detector 212 and a language parser 214. As indicated above, an utterance can be provided in various ways including audio or text. The utterance 202 can be a sentence fragment, a complete sentence, multiple sentences, and the like. Utterance 202 can include punctuation. For example, if the utterance 202 is provided as audio, the pre-processing subsystem 210 may convert the audio to text using a speech-to-text converter (not shown) that inserts punctuation marks into the resulting text, e.g., commas, semicolons, periods, etc.) and two or more information extraction processors (see fig. ref 242 320, 400 and 4000 in figs 2, 3 and 4A and 4B), wherein each of the two or more information extraction processors is suitable to generate derived information (intent 322, intent 475, entity/constraint 480) extraction rule specific (252, 254, 352 354 in figs 2 and 3) depending on the input representation according to an information extraction rule specific (252, 254, 352 354 in figs 2 and 3) for the information extraction processor (see para 0092 for master bot and para 106 for skills bot), and different from an information extraction rule of any other one of the two or more information extraction processors, (see para 0092-0093, also see fig. 4A, as skill bot invocation, intent prediction, and entity detection all inherently use specific information extraction rule) and an output interface (see fig. 3, conversation manager (330) for generating an output (Dialog output to user (335), being an audio output and/or a textual output and/or visual output and/or being a signal for steering a machine ([0048] Each skill associated with a digital assistant helps a user of the digital assistant complete a task through a conversation with the user, where the conversation can include a combination of text or audio inputs provided by the user and responses provided by the skill bots. These responses may be in the form of text or audio messages to the user and/or using simple user interface elements (e.g., select lists) that are presented to the user for the user to make selections.) also see para 0108, depending on the derived information from one or more of the two or more information extraction processors, (ref 242, 320, see figs 2 and 3) and wherein the one or more information extraction processors that have been selected are configured to generate the derived information depending on their information extraction rules. ([0106] Intent classifier 320 is configured to match a received utterance (e.g., utterance 306 or 308) to an intent associated with skill bot system 300. As explained above, a skill bot can be configured with one or more intents, each intent including at least one example utterance that is associated with the intent and used for training a classifier. In the embodiment of FIG. 2, the intent classifier 242 of the master bot system 200 is trained to determine confidence scores for individual skill bots and confidence scores for system intents. Similarly, intent classifier 320 can be trained to determine a confidence score for each intent associated with the skill bot system 300. Whereas the classification performed by intent classifier 242 is at the bot level, the classification performed by intent classifier 320 is at the intent level and therefore finer grained. The intent classifier 320 has access to intents information 354. The intents information 354 includes, for each intent associated with the skill bot system 300, a list of utterances that are representative of and illustrate the meaning of the intent and are typically associated with a task performable by that intent. The intents information 354 can further include parameters produced as a result of training on this list of utterances.) Also see ref (252, 254, 352 354 in figs 2 and 3). Vu does not appear to disclose the following features. Suendermann in the related art discloses: wherein at least two information extraction processors of the two or more information extraction processors are dialog-state-dependent, ([0118] Again, in production, i.e., when a dialog system using state-dependent classifiers takes live calls, at every state, the rows in the table whose variables match the current state variables are selected and the semantic classifier belonging to the row with the highest performance will be used for classification in the current state.) wherein the dialog system is configured to select one or more information extraction processors of the at least two information extraction processors, which are dialog-state-dependent, depending on a current state of the dialog, such that only those of the at least two information extraction processors, which are associated with the current state of the dialog, are selected, ([0118] Again, in production, i.e., when a dialog system using state-dependent classifiers takes live calls, at every state, the rows in the table whose variables match the current state variables are selected and the semantic classifier belonging to the row with the highest performance will be used for classification in the current state.) Vu and Suendermann are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Vu to combine the teaching of Suendermann, because there is a need for providing state-dependent semantic classifier in a spoken dialog system (Suendermann, [Background]). Regarding Claim 24, Vu and Suendermann disclose all the elements of claim 20. Vu further discloses: wherein the input interface is configured to receive the input being a speech signal or an audio signal, wherein the input interface is configured to apply a speech recognition algorithm on the speech signal or on the audio signal to acquire a text representation of the speech signal or of the audio signal as the input representation. ([0040] User inputs 110 are generally in a natural language form and are referred to as utterances. A user utterance 110 can be in text form, such as when a user types in a sentence, a question, a text fragment, or even a single word and provides it as input to digital assistant 106. In some examples, a user utterance 110 can be in audio input or speech form, such as when a user says or speaks something that is provided as input to digital assistant 106. The utterances are typically in a language spoken by the user. For example, the utterances may be in English, or some other language. When an utterance is in speech form, the speech input is converted to text form utterances in that particular language and the text utterances are then processed by digital assistant 106. Various speech-to-text processing techniques may be used to convert a speech or audio input to a text utterance, which is then processed by digital assistant 106. In some examples, the speech-to-text conversion may be done by digital assistant 106 itself.) Regarding Claim 26, it is a method claim that corresponds to the system of claim 20 and is therefore rejected under the same grounds as claim 20 above. Regarding Claim 28, Vu discloses: A non-transitory computer-readable medium comprising a computer program ([0013] Some embodiments of the present disclosure include a system including one or more data processors and a non-transitory computer readable storage medium containing instructions which, when executed on the one or more data processors, cause the one or more data processors to perform part or all of one or more methods and/or part or all of one or more processes disclosed herein.) As for the rest of the claim, they claim elements from the claim 20, therefore the rationale applied in rejection of claim 1 is also applicable to claim 28. Claim 21 is rejected under 35 U.S.C. 103 as being unpatentable over Vu, in view of Suendermann (US 20110046951), and further in view of Cheng (US 20200395007). Regarding Claim 21, Vu and Suendermann disclose all the elements of claim 20. Vu and Suendermann do not appear to teach dialog-state-independent information extraction. Cheng in the related art discloses: wherein the dialog system comprises three or more information extraction processors as the two or more information extraction processors, wherein at least one information extraction processor of the three or more information extraction processors is dialog-state-independent, wherein the at least one information extraction processor, which is dialog-state-independent, is configured to always generate, depending on its information extraction rule, the derived information, independent from the current state. ([0206] FIG. 18 shows a diagram 1800 of an action agent selection method utilizing a concept ontology classification and a text-based intent classification, in accordance with example embodiments of the disclosure. The end user's utterance 1802 is processed to determine a functional command that is processing state independent to handle common conversation flow control commands (e.g., repeat, hold/wait/pause, etc.). For example, if the user utterance is “Hold on”, the masterbot can assign a special utility action agent to respond with something like “Sure, take your time”, or alternatively, simply remain silent for the end user to proceed with the next turn.) Vu/Suendermann/Cheng are considered analogous art. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Vu and Suendermann to combine the teaching of Cheng, because there is an unsolved need to provide an enhanced virtual assistant platform and natural language interface for user interactions with computing systems that provides for a flexible and fluid conversation experience (Cheng, [Background]). Potentially Allowable Subject Matter Claims 5-9, and 22 would be potentially allowable if amended to overcome the pertinent rejections under section 35 U.S.C. 101. (reason for them being potentially allowable will be provided when the claims are in condition for allowance) Notwithstanding, said aforementioned teachings of prior art cited is respectfully reconsidered and found to fail to teach or fairly suggest either individually or in a reasonable combination the presented limitations in claims 5-9, and 22, as specifically recited. Allowable Subject Matter Claims 10 and 14 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Notwithstanding, said aforementioned teachings of prior art cited is respectfully reconsidered and found to fail to teach or fairly suggest either individually or in a reasonable combination the presented limitations in claims 10 and 14, as specifically recited. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Su (US 11335346) – discloses natural language understanding in dialog system. “As illustrated in FIG. 5, statistical models 520 may be implemented at least partially in parallel to the FSTs 510. FIG. 12 illustrates how statistical models (e.g., named entity recognition models, intent classification models, and domain classification models) may be used as part of NLU processing.” See Abstract, column 24 and fig. 5 for additional details. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Philip H Lam whose telephone number is (571)272-1721. The examiner can normally be reached 9 AM-3 PM Pacific time. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bhavesh Mehta can be reached on 571-272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /PHILIP H LAM/ Examiner, Art Unit 2656
Read full office action

Prosecution Timeline

Jan 15, 2025
Application Filed
Aug 19, 2026
Non-Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12688847
ERROR-CORRECTION AND EXTRACTION IN REQUEST DIALOGS
4y 1m to grant Granted Jul 21, 2026
Patent 12682164
CHAT SUPPORT PLATFORM HAVING AUTOMATIC KEYWORD CORRECTION
3y 3m to grant Granted Jul 14, 2026
Patent 12670519
CONTENT RECOMMENDATION USING RETRIEVAL AUGMENTED ARTIFICIAL INTELLIGENCE
3y 2m to grant Granted Jun 30, 2026
Patent 12657395
METHODS AND SYSTEMS FOR AVOIDING OFFENSIVE LANGUAGE BASED ON PERSONAS
2y 9m to grant Granted Jun 16, 2026
Patent 12639529
ENHANCING LARGE LANGUAGE MODELS USING IN-CONTEXT LEARNING AND ONLINE KNOWLEDGE
2y 6m to grant Granted May 26, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
84%
Grant Probability
99%
With Interview (+50.9%)
2y 6m (~10m remaining)
Median Time to Grant
Low
PTA Risk
Based on 155 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month