Prosecution Insights
Last updated: October 02, 2026
Application No. 17/521,262

RECOGNITION OF USER INTENTS AND ASSOCIATED ENTITIES USING A NEURAL NETWORK IN AN INTERACTION ENVIRONMENT

Non-Final OA §103
Filed
Nov 08, 2021
Examiner
BLANKENAGEL, BRYAN S
Art Unit
2658
Tech Center
2600 — Communications
Assignee
NVIDIA Corporation
OA Round
7 (Non-Final)
67%
Grant Probability
Favorable
7-8
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 67% — above average
67%
Career Allowance Rate
262 granted / 390 resolved
+5.2% vs TC avg
Strong +33% interview lift
Without
With
+33.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 8m
Avg Prosecution
23 currently pending
Career history
418
Total Applications
across all art units

Statute-Specific Performance

§101
25.1%
-14.9% vs TC avg
§103
50.7%
+10.7% vs TC avg
§102
11.8%
-28.2% vs TC avg
§112
7.4%
-32.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 390 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114 was filed in this application after appeal to the Patent Trial and Appeal Board, but prior to a decision on the appeal. Since this application is eligible for continued examination under 37 CFR 1.114 and the fee set forth in 37 CFR 1.17(e) has been timely paid, the appeal has been withdrawn pursuant to 37 CFR 1.114 and prosecution in this application has been reopened pursuant to 37 CFR 1.114. Applicant’s submission filed on 06/16/2026 has been entered. Response to Arguments Applicant’s arguments with respect to claim(s) 21-40 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 21-26, 28-29, 31-34, 36-38, and 40 is/are rejected under 35 U.S.C. 103 as being unpatentable over Assa et al. (US 2023/0067305 A1), hereinafter referred to as Assa, in view of Sharifi et al. (US 2021/0406260 A1), hereinafter referred to as Sharifi. Regarding claim 21, Assa teaches: A method comprising: instantiating a virtual interaction environment including one or more objects (Fig. 7A, para [0129], [0132], where a GUI of an AR system is displayed, including objects in the schema); causing presentation of the virtual interaction environment on a display (Fig. 7A, para [0129], [0132], where a GUI of an AR system is displayed); during the presentation, receiving, from a user, an initial individual voice input corresponding to one or more desired changes to the one or more objects (Fig. 7A, para [0129], [0132], where a speech input from a user is processed for changing an object); processing, using one or more language models, first data corresponding to the initial individual voice input and second data corresponding to the virtual interaction environment, the second data providing context for the initial individual voice input (Fig. 7A, para [0108], [0125], where the conversation-based AR system uses a trained speech understanding model to process the raw speech input and the schema, which provides the context for the input); based at least on the processing, determining that an intent of the initial individual voice input is for one or more changes to the one or more objects within the virtual interaction environment (Fig. 7A, para [0129], [0132], where an intent is determined for changing at least one object); based at least on the determined intent for the one or more changes to the one or more objects, iteratively generating a follow on question to obtain, from the initial individual voice input, additional information as to the one or more changes to the one or more objects (Fig. 7A, para [0129], [0132], where based on the intent, each corresponding slot is filled based on information found in the initial individual voice input, the slots to be filled being interpreted as the questions); updating the virtual interaction environment to reflect the one or more changes to the one or more objects (Fig. 7A, para [0129], [0132], where the change is reflected in the updated GUI shown to the user); and causing presentation of the updated virtual interaction environment, including the one or more changes to the one or more objects, on the display (Fig. 7A, para [0129], [0132], where the change is reflected in the updated GUI shown to the user). Assa does not explicitly teach that follow on questions are generated. based at least on the determined intent for the one or more changes to the one or more objects, iteratively generating a follow on question to obtain, from the initial individual voice input, additional information as to the one or more changes to the one or more objects; Sharifi teaches: generating a follow on question (para [0029], where a combined search query is generated based on previously submitted queries that share a line of inquiry); It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Assa by using the question generation of Sharifi (Sharifi para [0029]) for the information retrieval of Assa (Assa Fig. 7, para [0129]), in order to reduce likelihood of the user being presented with identical results during different turns, and to relieve the user of having to manually formulate their own ever-growing search query that captures all the parameters used during different turns (Sharifi para [0002]). Regarding claim 22, Assa in view of Sharifi teaches: The method of claim 21, wherein the initial individual voice input is received by a conversational artificial intelligence (AI) associated with the virtual interaction environment (Assa para [0017], [0059-60], where the conversation-based AR system uses a speech understanding model, such as a neural network). Regarding claim 23, Assa in view of Sharifi teaches: The method of claim 21, wherein the one or more language models generate a natural language formulation of the intent for the processing (Assa Fig. 7A, para [0125], where the intent is determined, and Sharifi para [0017], where the intent is in natural language). Regarding claim 24, Assa in view of Sharifi teaches: The method of claim 23, further comprising: determining the natural language formulation requires additional information from the initial individual voice input (Assa Fig. 7A, para [0125], where both intents and slots are determined from the input); and identifying, from the initial individual voice input, the additional information required by natural language formulation (Assa Fig. 7A, para [0125], where the slot values are the additional information corresponding to the intent). Regarding claim 25, Assa in view of Sharifi teaches: The method of claim 21, further comprising: selecting, from a predetermined list, one or more actions supported by the virtual interaction environment determined to correspond to the one or more changes to the one or more objects (Assa para [0058], [0063], where changes to property values of the configurable entity are considered the actions, the configurable entities including lists of objects and corresponding properties); and providing the one or more actions to the virtual interaction environment to cause the updating (Assa para [0063], where the updates with the changes are sent to the client device). Regarding claim 26, Assa in view of Sharifi teaches: The method of claim 21, wherein the intent is determined at least in part by a trained neural network (Assa para [0017], [0059-60], where the conversation-based AR system uses a speech understanding model, such as a neural network). Regarding claim 28, Assa teaches: At least one processor, comprising: processing circuitry (para [0135], where a processor is used) to: instantiate a virtual interaction environment including one or more objects (Fig. 7A, para [0129], [0132], where a GUI of an AR system is displayed, including objects in the schema); present the virtual interaction environment on a display (Fig. 7A, para [0129], [0132], where a GUI of an AR system is displayed); during the presenting, receive, from a user, an initial individual voice input corresponding to one or more desired changes to the one or more objects (Fig. 7A, para [0129], [0132], where a speech input from a user is processed for changing an object); process, using one or more language models, first data corresponding to the initial individual voice input and second data corresponding to the virtual interaction environment, the second data providing context for the initial individual voice input (Fig. 7A, para [0108], [0125], where the conversation-based AR system uses a trained speech understanding model to process the raw speech input and the schema, which provides the context for the input); based at least on the processing, determine that an intent of the initial individual voice input is for one or more changes to the one or more objects within the virtual interaction environment (Fig. 7A, para [0129], [0132], where an intent is determined for changing at least one object); based at least on the determined intent for the one or more changes, iteratively generate a follow on question to obtain, from the initial individual voice input, additional information as to the one or more changes to the one or more objects (Fig. 7A, para [0129], [0132], where based on the intent, each corresponding slot is filled based on information found in the initial individual voice input, the slots to be filled being interpreted as the questions); update the virtual interaction environment to reflect the one or more changes to the one or more objects (Fig. 7A, para [0129], [0132], where the change is reflected in the updated GUI shown to the user); and present the updated virtual interaction environment, including the one or more changes to the one or more objects, on the display (Fig. 7A, para [0129], [0132], where the change is reflected in the updated GUI shown to the user). Assa does not explicitly teach that follow on questions are generated. Sharifi teaches: generate a follow on question (para [0029], where a combined search query is generated based on previously submitted queries that share a line of inquiry); It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Assa by using the question generation of Sharifi (Sharifi para [0029]) for the information retrieval of Assa (Assa Fig. 7, para [0129]), in order to reduce likelihood of the user being presented with identical results during different turns, and to relieve the user of having to manually formulate their own ever-growing search query that captures all the parameters used during different turns (Sharifi para [0002]). Regarding claim 29, Assa in view of Sharifi teaches: The at least one processor of claim 28, wherein the processing circuitry is further to execute a trained entailment neural network, wherein the processing circuitry is further to determine the intent of the initial individual voice input using the trained entailment neural network (Assa para [0038-39], where a speech understanding model uses a neural network to generate intents from the raw speech input). Regarding claim 31, Assa in view of Sharifi teaches: The at least one processor of claim 28, wherein the one or more language models generate a natural language formulation of the intent for the processing (Assa Fig. 7A, para [0125], where the intent is determined, and Sharifi para [0017], where the intent is in natural language). Regarding claim 32, Assa in view of Sharifi teaches: The at least one processor of claim 31, further comprising: determine the natural language formulation requires additional information from the initial individual voice input (Assa Fig. 7A, para [0125], where both intents and slots are determined from the input); and identify, from the initial individual voice input, the additional information required by natural language formulation (Assa Fig. 7A, para [0125], where the slot values are the additional information corresponding to the intent). Regarding claim 33, Assa in view of Sharifi teaches: The at least one processor of claim 28, wherein the processing circuitry is further to select, from a predetermined list, one or more actions supported by the virtual interaction environment determined to correspond to the one or more changes to the one or more objects, wherein the processing circuitry is further to provide the one or more actions to the virtual interaction environment to cause the updating (Assa para [0058], [0063], where changes to property values of the configurable entity are considered the actions, the configurable entities including lists of objects and corresponding properties, and where the updates with the changes are sent to the client device). Regarding claim 34, Assa in view of Sharifi teaches: The at least one processor of claim 28, wherein the intent is determined at least in part by a trained neural network (Assa para [0017], [0059-60], where the conversation-based AR system uses a speech understanding model, such as a neural network). Regarding claim 36, Assa teaches: A system, comprising one or more processors (para [0135], where a processor is used) to cause performance of operations comprising: displaying an instance of a virtual interaction environment including one or more content elements (Fig. 7A, para [0129], [0132], where a GUI of an AR system is displayed, including objects in the schema); receiving, from a user, an initial individual utterance associated with one or more specified modifications to be made to the content elements as displayed in the instance (Fig. 7A, para [0129], [0132], where a speech input from a user is processed for changing an object); generating a representation of an intent of the initial individual utterance by using natural language understanding to process at least part of the initial individual utterance and information associated with the virtual interaction environment (Fig. 7A, para [0108], [0125], where the conversation-based AR system uses a trained speech understanding model to process the raw speech input and the schema, which provides the context for the input, and para [0129], [0132], where an intent is determined for changing at least one object); determining the generated representation of the intent is related to one or more modifications, supported by the virtual interaction environment, to the content elements (Fig. 7A, para [0129], [0132], where an intent is determined for changing at least one object); based at least on the determined representation of the intent related to the one or more modifications, generating an operation to obtain, from the initial individual utterance, additional information as to the one or more modifications (Fig. 7A, para [0129], [0132], where based on the intent, each corresponding slot is filled based on information found in the initial individual voice input, the slots to be filled being interpreted as the questions); causing the one or more modifications to the content elements to be made within the virtual interaction environment to update the instance (Fig. 7A, para [0129], [0132], where the change is reflected in the updated GUI shown to the user); and displaying the updated instance of the virtual interaction environment including the one or more modifications to the content elements (Fig. 7A, para [0129], [0132], where the change is reflected in the updated GUI shown to the user). Assa does not explicitly teach that follow on questions are generated. Sharifi teaches: generating an operation to obtain additional information (para [0029], where a combined search query is generated based on previously submitted queries that share a line of inquiry); It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Assa by using the question generation of Sharifi (Sharifi para [0029]) for the information retrieval of Assa (Assa Fig. 7, para [0129]), in order to reduce likelihood of the user being presented with identical results during different turns, and to relieve the user of having to manually formulate their own ever-growing search query that captures all the parameters used during different turns (Sharifi para [0002]). Regarding claim 37, Assa in view of Sharifi teaches: The system of claim 36, wherein the determining is based, at least in part, on a pre-determined list of entities associated with capabilities of the virtual interaction environment (Assa para [0108], where lists of properties or slot values that the slots can receive are used). Regarding claim 38, Assa in view of Sharifi teaches: The system of claim 36, further comprising: extracting, from the initial individual utterance, one or more portions associated with the intent to provide additional information to answer the intent (Assa Fig. 7A, para [0125], where both intents and slots are determined from the input); and identifying, using the one or more extracted portions, the one or more modifications to the content elements (Assa Fig. 7A, para [0125], where the slot values are the additional information corresponding to the intent). Regarding claim 40, Assa in view of Sharifi teaches: The system of claim 36, further comprising: selecting the intent from a list of intents, individual intents of the list of intents corresponding to a respective intent label (Assa para [0108], where lists of intents are used and selected from, each intent corresponding to a label such as put_mascara). Claim(s) 27 is/are rejected under 35 U.S.C. 103 as being unpatentable over Assa, in view of Sharifi, and further in view of Duong et al. (US 2023/0376696 A1), hereinafter referred to as Duong. Regarding claim 27, Assa in view of Sharifi teaches: The method of claim 21, further comprising: determining a plurality of pre-determined intent labels based at least in part on the initial individual voice input (Assa Fig. 6B element 632, para [0058], where a list of intents is stored); Assa in view of Sharifi does not teach: determining a probability, for individual labels of the plurality of pre-determined intent labels, of corresponding to the initial individual voice input; and selecting one or more of the individual labels, having the probability exceeding a threshold, to be used as the intent. Duong teaches: determining a probability, for individual labels of the plurality of pre-determined intent labels, of corresponding to the initial individual voice input (para [0101], where a confidence score is used for identifying an intent); and selecting one or more of the individual labels, having the probability exceeding a threshold, to be used as the intent (para [0101], where a threshold confidence score must be met for selection). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Assa in view of Sharifi by using the confidence scores of Duong (Duong para [0101]) in the intent determination of Assa in view of Sharifi (Assa para [0058]), in order to determine a particular skill bot to be invoked to handle the utterance (Duong para [0101]). Claim(s) 30 and 35 is/are rejected under 35 U.S.C. 103 as being unpatentable over Assa, in view of Sharifi, and further in view of Avijeet (US 2022/0130378 A1). Regarding claim 30, Assa in view of Sharifi teaches: The at least one processor of claim 28, Assa in view of Sharifi does not teach: wherein the processing circuitry is further to execute a trained extractive question and answer neural network model, wherein the processing circuitry is further to determine, at least in part, the one or more changes to the one or more objects using the trained extractive question and answer neural network model. Avijeet teaches: wherein the processing circuitry is further to execute a trained extractive question and answer neural network model, wherein the processing circuitry is further to determine, at least in part, the one or more changes to the one or more objects using the trained extractive question and answer neural network model (para [0072], where a neural network performs question answering by fetching data from external sources). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Assa in view of Sharifi by using the question answering of Avijeet (Avijeet para [0072]) in the neural networks of Assa in view of Sharifi (Assa para [0038-39]), in order to determine a proper response to user queries (Avijeet para [0072]). Regarding claim 35, Assa in view of Sharifi teaches: The at least one processor of claim 28, wherein the processing circuitry is further to: Assa in view of Sharifi does not teach: receive a second individual voice input from the user; determine a second intent of the second individual voice input cannot be identified; and provide a response to the user that includes a request for additional intent information. Avijeet teaches: receive a second voice input from the user (para [0072], where a user queries about the weather); determine a second intent of the second voice input cannot be identified (para [0072], where the system determines a query to resolve ambiguities); and provide a response to the user that includes a request for additional intent information (para [0072], where the system outputs the query to the user). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Assa in view of Sharifi by using the question answering of Avijeet (Avijeet para [0072]) in the neural networks of Assa in view of Sharifi (Assa para [0038-39]), in order to determine a proper response to user queries (Avijeet para [0072]). Claim(s) 39 is/are rejected under 35 U.S.C. 103 as being unpatentable over Assa, in view of Sharifi, and further in view of Zhao (US 2020/0251091 A1). Regarding claim 39, Assa in view of Sharifi teaches: The system of claim 36, further comprising: Assa in view of Sharifi does not teach: determining the intent based, at least in part, on one or more machine learning systems using a zero-shot approach. Zhao teaches: determining the intent based, at least in part, on one or more machine learning systems using a zero-shot approach (para [0037-38], [0070], where the Zero-shot intent recognition model is trained using machine learning tools to recognize intents). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Assa in view of Sharifi by using the dynamic intent list of Zhao (Zhao para [0038], [0076]) in the intent determination of Assa in view of Sharifi (Assa para [0125]), to achieve rich semantic parsing results including matching results with all the intents, without the need for retraining (Zhao para [0038]). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. US 2019/0392828 A1 para [0026] teaches generating a new query in response to existing record data. Any inquiry concerning this communication or earlier communications from the examiner should be directed to BRYAN S BLANKENAGEL whose telephone number is (571)270-0685. The examiner can normally be reached 8:00am-5:30pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached at 571-272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /BRYAN S BLANKENAGEL/Primary Examiner, Art Unit 2658
Read full office action

Prosecution Timeline

Show 23 earlier events
Dec 16, 2025
Final Rejection mailed — §103
Jan 23, 2026
Notice of Allowance
Mar 19, 2026
Response after Non-Final Action
Apr 06, 2026
Response after Non-Final Action
Apr 16, 2026
Response after Non-Final Action
Jun 16, 2026
Request for Continued Examination
Jun 18, 2026
Response after Non-Final Action
Jul 14, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12724962
PRE-TRAINING LANGUAGE MODELS USING NATURAL LANGUAGE EXPRESSIONS EXTRACTED FROM STRUCTURED DATABASES
4y 3m to grant Granted Sep 01, 2026
Patent 12718826
AUDIO SIGNAL ENCODING AND DECODING METHOD AND APPARATUS
2y 7m to grant Granted Aug 25, 2026
Patent 12711947
METHOD FOR TRAINING A NEURAL NETWORK AND A DATA PROCESSING DEVICE
2y 8m to grant Granted Aug 18, 2026
Patent 12711969
METHODS AND APPARATUS FOR SUPPLEMENTING PARTIALLY READABLE AND/OR INACCURATE CODES IN MEDIA
2y 2m to grant Granted Aug 18, 2026
Patent 12711983
METHOD OF DETECTING SPEECH AND SPEECH DETECTOR FOR LOW SIGNAL-TO-NOISE RATIOS
2y 1m to grant Granted Aug 18, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

7-8
Expected OA Rounds
67%
Grant Probability
99%
With Interview (+33.3%)
2y 8m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 390 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month