Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Acknowledgement
Acknowledgement is made of applicant’s amendment made on 05/15/2026. Applicant’s submission filed has been entered and made of record.
Status of the Claims
Claims 1-20 are pending.
Response to Applicant’s Arguments
In view of amendment to claims 9-14, rejection under 35 USC 101 has been withdrawn.
In response to “That is, the solution focuses on, after triggering creation of a session window through speech input, how to adjust the duration of the session window via subsequent speech input, i.e., it involves only a single interaction mode (speech) and does not involve multiple interaction modes, nor does it involve obtaining complete semantic information based on multiple interaction modes”, “That is, the solution focuses on, during the same interaction phase, obtaining semantic information acquired using different interaction modes over multiple rounds of interaction, and using that semantic information to obtain complete semantic information, i.e., the solution is aimed at accurately determining the user's true control intent by leveraging different interaction modes”, and “Third, the Examiner believes that Weinberg discloses the feature of determining semantic completeness. Applicant does not agree. The reasons are as follows. In the solution of the present disclosure, the completeness of the target semantic information can be detected, and when the target semantic information is incomplete, complete semantic information can be generated based on the target semantic information and the cached to-be-combined semantic information. Therefore, it can be unequivocally inferred that the completeness detection in this solution essentially detects whether semantic components are missing. That is, the completeness detection is at the semantic content level, ensuring that the semantic content is complete. In contrast, the solution of Weinberg detects whether the user is looking at the screen or recognizes a specific wake-up word to determine whether the user is speaking to the machine, or uses the user's tone, etc., to determine whether to extend the window. If it is determined that the user is speaking to the machine or that the window should be extended, it can be determined that the dialogue is not finished, i.e., the dialogue is incomplete. That is, this solution detects whether the user's dialogue has ended, not the completeness of the semantic content. Thus, Weinberg only delays temporally to wait for the user to finish speaking and cannot ensure semantic completeness. As such, the two solutions involve completeness detection for different objects and different dimensions”.
Weinberg discloses a digital assistant identifying user’s intent expressed in a natural language input received from the user, actively eliciting and obtaining information needed to fully infer the user’s intent (“first interaction mode” being speech interaction between digital assistant and the user), determining task flow for fulfilling the inferred intent, and execute the task flow to fulfill the inferred intent (¶208).
In particularly, the digital assistant obtains contextual information associated with the user input from the user device, along with or shortly after the receipt of the user input (¶209) to clarify, supplement, and further define information contained in text representation of user’s natural language input (¶217).
In one example, after receiving “Make me a dinner reservation at a sushi place at 7”, the digital assistant’s NLP module identified “restaurant reservation” as user’s actionable intent, generate a structured query (“target semantic information”) corresponding to the actionable intent in an ontology in the restaurant reservation domain, and determine that user’s utterance contains insufficient information to complete the structured query because the user utterance did not specify parameters “party size” and “date” (¶230; i.e., “determining the completeness of the target semantic information” and “in response to the target semantic information being incomplete”).
Therefore, a partial structured query is generated for the actionable intent (¶231) and the digital assistant initiates additional dialogue with the user to obtain additional information and disambiguate potentially ambiguous utterances (¶234) to generate dialogue response that at least partially fulfill user’s intent (¶238; see further ¶283, provide a first response based on first speech input).
In particular, during a target interaction phase (¶244, digital assistant in an available state being invoked into a listening state to begin processing speech input), when the digital assistant starts processing commands (i.e., determining “make me a dinner reservation at a sushi place at 7” contains insufficient information and that the corresponding “restaurant reservation” structured query / target semantic information is incomplete), the user interrupts the processing of the speech input (¶244; see further ¶283, while the first response is being provided, a third speech input is received and allowing the user to interrupt the digital assistant) by looking in a direction of the digital assistant (“second interaction mode” being different than the “first interaction mode” via speech) to re-engage the digital assistant with relevant additional speech (¶245; see further ¶284, adjust session window variable speech threshold based on user gaze to focus on a relevant window of time to capture additional speech from the user) directed at the digital assistant (¶292).
However, Weinberg does not disclose the to-be-combined semantic information cached in a preset semantic state record library that was obtained when the target user interacts with the digital assistant / target device according to the second interaction mode during the target interaction phase”, the second interaction mode is different from the first interaction mode since Weinberg does not make it clear that the gaze can be used as to-be-combined semantic information and it is cached in the preset semantic state record library.
Accordingly, rejection of claim 1 under Weinberg and Pitschel has been withdrawn. Upon further search and consideration, please see details of a new combination of references set forth below.
Claim Rejections - 35 USC § 103
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 103 that form the basis for the rejections under this section made in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-6, 8-13, and 15-19 are rejected under 35 USC 103(a) as being unpatentable over Weinberg et al. (US 2022/0293124 A1) in view of Jarosz et al. (US 11455982 B2) and Pitschel et al. (US 9922642 B2).
Regarding Claim 1, Weinberg discloses a human-machine interaction method (¶243, systems and processes for continuous dialog with a digital assistant), including:
in response to receiving target interaction information collected when a target user interacts with a target device according to a first interaction mode during a target interaction phase (¶244, upon being invoked by “Hey Siri”, the digital assistant enters a listening state 804), performing semantic recognition on the target interaction information to obtain target semantic information (¶244, digital assistant samples speech input including questions and commands and begin processing the questions and commands in processing state 806; per ¶208, convert speech input into text, identifying a user’s intent expressed in the natural language input, actively obtaining information needed to fully infer the user’s intent, and determining a task flow for fulfilling the inferred intent);
determining completeness of the target semantic information (¶208, processing includes actively eliciting and obtaining information needed to fully infer the user’s intent);
in response to the target semantic information being incomplete (¶230, determine “Make me a dinner reservation at a sushi place at 7” corresponds to restaurant reservation domain to generate a structured query and that user utterance contains insufficient information to complete the structured query’s necessary parameters “Party Size” and “date”), determining whether there is to-be-combined semantic information cached in a preset semantic state record library (¶217, natural language processing module 728 uses contextual information including prior interactions / dialogues between the digital assistant and the user to clarify, supplement, and further define information);
in response to presence of the to-be-combined semantic information in the semantic state record library, generating complete semantic information based on the target semantic information and the to-be-combined semantic information (¶230, natural language processing module 732 populates some parameters of the structured query with received contextual information); and
based on the complete semantic information, determining a target controlled object and a control mode for controlling the target controlled object, and generating a control instruction corresponding to the control mode (¶245, if the digital assistant completes processing of the speech input to obtain one or more results, then the digital assistant enters a response state 808 where digital assistant provides one or more results; e.g., ¶235, task flow processing module 736 performs steps (1)-(4) to make a restaurant reservation for restaurant reservation structured query at ABC Café, on 3/12/12, at 7pm, for a party of 5).
Weinberg does not disclose wherein the to-be-combined semantic information is semantic information obtained when the target user interacts with the target device according to at least one second interaction mode different from the first interaction mode during the target interaction phase and updating the to-be-combined semantic information in the semantic state record library based on the target semantic information.
Jarosz teaches a natural language processing device (Fig. 2 and Col 8, Rows 52-60) performing semantic recognition on target interaction information from a target user interacting with the device according to a first interaction mode during a target interaction phase to obtain target semantic information (Col 7, Rows 7-15, natural language processor interpreting driver vocal requests / first interaction mode during a target interaction phase such as “lower the rear passenger side window, please” or “lower the window”), determining a completeness of the target semantic information (Col 7, Rows 10-15, in the example, recognize a complete command to lower the rear passenger side window; Col 11, Rows 16-20, determine whether there is ambiguity regarding a current utterance) being incomplete (Col 7, Rows 30-38, for ambiguous command “lower the window” or simply “lower”, the NLP processor recognizes the word “lower” but unable to determine upon which of the plurality of windows, seat, or volume of an infotainment center to carry out the command; Col 11, Rows 31-35, determine that ambiguities remain in the utterance), and in response to the target semantic information being incomplete, determining whether there is to-be-combined semantic information cached in a preset semantic state record library (Col 7, Rows 44-52, establishing a context for disambiguation by storing data related to a speaker and to the environment within which utterances are made for recall to establish the context of the utterance to disambiguate the utterance), the to-be-combined semantic information is semantic information obtained when the target user interacts with the target device according to at least one second interaction mode different from the first interaction mode during the target interaction phase (Col 10, Rows 54-65 and Col 12, Rows 35-39, tracking user’s gaze as “recent antecedent” context factor corresponding to most recent interaction a user has had with the system).
In response to presence of the to-be-combined semantic information in the semantic state record library, generating complete semantic information based on the target semantic information and the to-be-combined semantic information (Col 13, Rows 44-51, determine that the current gaze target context factor is applicable to address remaining utterance ambiguity; e.g., Col 14, Rows 25-30, employ the driver’s gaze target context factor and eye tracking input mode to resolve the ambiguity of which window to lower for “lower the window”);
and based on the complete semantic information, determining a target controlled object and a control mode for controlling the target controlled object, and generating a control instruction corresponding to the control mode (Col 14, Rows 28-30, the system resolves the question of which window to lower; in view of Col 8, Rows 20-21, executes commands in response to the driver’s command).
It would’ve been obvious to one ordinarily skilled in the art before the effective filing date of the invention to determine there is to-be-combined semantic information cached in a preset semantic state record library, the to-be-combined semantic information is semantic information obtained when the target user interacts with the target device according to at least one second interaction mode different from the first interaction mode during the target interaction phase, in order to apply applicable context factor to address ambiguity in user’s utterance / target interaction information (Jarosz, Col 13, Rows 44-51; compare Weinstein, ¶209 and ¶217, obtain contextual information associated with user input to clarify, supplement, and further define user request).
Weinberg does not disclose updating the to-be-combined semantic information in the semantic state record library based on the target semantic information.
Pitschel teaches a human machine interaction method (Col 9, Rows 25-31, I/O processing module 328 interacts with user to obtain user input and to provide response to user input) performing semantic recognition on target interaction information to obtain target semantic information (Col 10, Rows 6-14, natural language processing module 332 takes words / tokens of speech to text processed user input and associate the token sequence with one or more actionable intent), determining completeness of the target semantic information (Col 13, Rows 31-56, generate a structured query to represent the identified actionable intent and determine that the user utterance contains insufficient information to complete the structured query associated with the domain), using to-be-combined semantic information cached in a preset semantic state record library to generate complete semantic information for the incomplete target semantic information (Col 13, Rows 54-59, use context information to populate parameters of the structured query; per Col 10, Rows 33-38, context information includes prior interaction / dialogue between the digital assistant and the user), and updating the to be combined semantic information in the semantic state record library based on the target semantic information (Col 17, Rows 26-35, digital assistant maintains a user log 370 based on user requests and interactions to store information such as user requests received, context information, responses provided to the user, clarification inputs, the parameters used by digital assistant to generate and provide the response).
It would’ve been obvious to one ordinarily skilled in the art before the effective filing date of the invention to update the to be combined semantic information in the semantic state record library based on the target semantic information to provide a searchable semantic state record library (Pitschel, Col 17, Rows 35-38).
Regarding Claim 2, Weinberg discloses after the determining the completeness of the target semantic information, in response to determining that the target semantic information is complete, generating a control instruction corresponding to the target semantic information based on the target semantic information (¶235, once task flow processing module 736 has completed the structured query for an actionable intent, proceed to perform the ultimate task associated with the actionable intent; see e.g., task flow steps (1)-(4)); and
as modified by Pitschel, updating the to-be-combined semantic information in the semantic state record library based on the target semantic information (Pitschel, Col 17, Rows 29-38, user log stores context information surrounding the user requests, responses provided to the user, the parameters used by the digital assistant to generate and provide the response).
Regarding Claim 3, Weinberg discloses wherein, the method further includes:
after the determining whether there is to-be-combined semantic information cached in a predetermined semantic state record library, in response to determining that there is not the to-be-combined semantic information in the semantic state record library, generating the to-be-combined semantic information based on the target semantic information (¶230, populating some parameters of the structured query with contextual information means not all parameters can be populated with contextual information (i.e., there is not contextual information to complete the structure query); e.g., ¶227, for “invite my friends to my birthday party”, access user data 748 to determine who the “friends” are and when and where the “birthday party” would be held), and as modified by Pitschel, storing the to-be-combined semantic information into the semantic state record library (Pitschel, Col 17, Rows 29-38, user log stores context information surrounding the user requests, responses provided to the user, the parameters used by the digital assistant to generate and provide the response).
Regarding Claim 4, Weinberg discloses wherein the receiving target interaction information collected when the target user interacts with the target device according to the first interaction mode during the target interaction phase includes:
receiving first interaction information collected from the target user in response to an interaction start signal for triggering a first-round interaction with the target device during the target interaction phase (¶244, user utters “hey Siri” to invoke the digital assistant; once invoked, the digital assistant enters listening state 804 to sample audio including speech input from the user); and
in response to determining that a type of the first interaction information is speech interaction information, determining the first interaction information as the target interaction information (¶244, the digital assistant samples speech input from the user while in the listening state 804 and begin processing the speech input in processing state 806; per ¶208, converting speech input into text, identifying a user’s intent expressed in a natural language input received from the user, determining the task flow for fulling the intent).
Regarding Claim 5, Weinberg discloses wherein in response to receiving target interaction information collected when a target user interacts with a target device according to a first interaction mode during a target interaction phase, the performing semantic recognition on the target interaction information to obtain target semantic information includes:
receiving second interaction information collected when interacting with the target user according to any of predetermined at least one interaction mode, in response to an interaction start signal for triggering a non-first-round interaction with the target device during the target interaction phase (¶284, at block 1406 (after receiving first speech and third speech at blocks 1402 and 1404), initiate a session window associated with a user gaze directed to a displayed digital assistant object and receive a second speech input; the system focuses on a relevant window of time to capture relevant additional speech from the user and improves user experience by capturing additional speech from the user);
generating the target interaction information based on the second interaction information (¶285, determine that the second speech input includes speech directed to the digital assistant according to detected user gaze being directed to a display of the digital assistant electronic device and detect a command within the second speech input); and
performing semantic recognition on the target interaction information according to a semantic recognition mode corresponding to the target interaction information to obtain the target semantic information (¶285, identify a command within the second speech input and determine that the second speech input includes speech directed to the digital assistant; i.e., ¶208, converting speech input into text, identifying a user’s intent expressed in a natural language input received from the user, determining the task flow for fulling the intent).
Regarding Claim 6, Weinberg discloses wherein the performing semantic recognition on the target interaction information according to the semantic recognition mode corresponding to the target interaction information to obtain the target semantic information includes:
in response to determining that the target interaction information is line-of-sight interaction information obtained by utilizing a line-of-sight interaction mode among the at least one interaction mode, performing recognition on the line-of-sight interaction information according to the line-of-sight recognition mode to obtain target controlled object information, and determining the target controlled object information as target semantic information (¶285, determine that the second speech input includes speech directed to the digital assistant according to detected user gaze being directed to a display of the digital assistant electronic device and detect a command within the second speech input; i.e., perform semantic processing per ¶208); and
in response to determining that the target interaction information is interaction information obtained by utilizing a non-line-of-sight interaction mode among the at least one interaction mode, performing semantical recognition on the target interaction information according to a semantic recognition mode corresponding to the target interaction information to obtain the target controlled object information and/or target instruction information (¶283, at blocks 1402-1404, perform processing of first speech input and third speech input to identify predefined words therein to determine speech inputs include speech directed to the digital assistant in a non-line of sight mode prior to line of sight mode / gaze detection at block 1406), and
determining the target controlled object information and/or the target instruction information as target semantic information (¶283, for non-gaze / non-line of sight mode, determine speech is directed to digital assistant and determine corresponding task flows per ¶235; for gaze / line of sight mode, ¶286, determine second speech input includes speech directed to the digital assistant and determine corresponding task flows per ¶235).
Regarding Claim 8, Weinberg as modified by Pitschel discloses wherein the updating the to-be-combined semantic information in the semantic state record library based on the target semantic information includes:
extracting target controlled object information and/or target instruction information from the target semantic information (Pitschel, Col 17, Rows 29-35, user log stores the responses provided to the user (e.g., Col 15, Rows 15-23, restaurant reservation at ABC Café, 3/12/2012, at 7pm, for party of 5); compare Weinberg, ¶235, restaurant reservation at ABC Café, 3/12/2012, at 7pm, for party of 5); and
updating to-be-combined controlled object information and/or to-be-combined instruction information included in the to-be-combined semantic information by using the target controlled object information and/or the target instruction information (Pitschel, Col 17, Rows 29-35, user log stores the parameters and the procedures used by the digital assistant to generate and provide the response).
Regarding Claim 9, Weinberg discloses a non-transitory computer-readable storage medium (¶¶198-99, digital assistant system 700 includes memory 702 / non-transitory computer readable medium), in which a computer program is stored, the computer program is configured for being executed by a processor (¶197, software instructions for execution by one or more processors; ¶198, digital assistant 700 includes processors 704) to implement the method according to claim 1 (¶197, software instructions for execution by one or more processors).
Regarding Claim 10, Weinberg discloses wherein the method further includes:
after the determining the completeness of the target semantic information, in response to determining that the target semantic information is complete, generating a control instruction corresponding to the target semantic information based on the target semantic information (¶235, once task flow processing module 736 has completed the structured query for an actionable intent, proceed to perform the ultimate task associated with the actionable intent; see e.g., task flow steps (1)-(4)); and
as modified by Pitschel, updating the to-be-combined semantic information in the semantic state record library based on the target semantic information (Pitschel, Col 17, Rows 29-38, user log stores context information surrounding the user requests, responses provided to the user, the parameters used by the digital assistant to generate and provide the response).
Regarding Claim 11, Weinberg discloses wherein, the method further includes: after the determining whether there is to-be-combined semantic information cached in a predetermined semantic state record library, in response to determining that there is not the to-be-combined semantic information in the semantic state record library, generating the to-be-combined semantic information based on the target semantic information (¶230, populating some parameters of the structured query with contextual information means not all parameters can be populated with contextual information (i.e., there is not contextual information to complete the structure query); e.g., ¶227, for “invite my friends to my birthday party”, access user data 748 to determine who the “friends” are and when and where the “birthday party” would be held), and as modified by Pitschel, storing the to-be-combined semantic information into the semantic state record library (Pitschel, Col 17, Rows 29-38, user log stores context information surrounding the user requests, responses provided to the user, the parameters used by the digital assistant to generate and provide the response).
Regarding Claim 12, Weinberg discloses wherein the receiving target interaction information collected when the target user interacts with the target device according to the first interaction mode includes:
receiving first interaction information collected from the target user in response to an interaction start signal for triggering a first-round interaction with the target device during the target interaction phase (¶244, user utters “hey Siri” to invoke the digital assistant; once invoked, the digital assistant enters listening state 804 to sample audio including speech input from the user); and
in response to determining that a type of the first interaction information is speech interaction information, determining the first interaction information as the target interaction information (¶244, the digital assistant samples speech input from the user while in the listening state 804 and begin processing the speech input in processing state 806; per ¶208, converting speech input into text, identifying a user’s intent expressed in a natural language input received from the user, determining the task flow for fulling the intent).
Regarding Claim 13, Weinberg discloses wherein in response to receiving target interaction information collected when a target user interacts with a target device according to a first interaction mode during a target interaction mode, the performing semantic recognition on the target interaction information to obtain target semantic information includes:
receiving second interaction information collected when interacting with the target user according to any of predetermined at least one interaction mode, in response to an interaction start signal for triggering a non-first-round interaction with the target device during the target interaction phase (¶284, at block 1406 (after receiving first speech and third speech at blocks 1402 and 1404), initiate a session window associated with a user gaze directed to a displayed digital assistant object and receive a second speech input; the system focuses on a relevant window of time to capture relevant additional speech from the user and improves user experience by capturing additional speech from the user);
generating the target interaction information based on the second interaction information (¶285, determine that the second speech input includes speech directed to the digital assistant according to detected user gaze being directed to a display of the digital assistant electronic device and detect a command within the second speech input); and
performing semantic recognition on the target interaction information according to a semantic recognition mode corresponding to the target interaction information to obtain the target semantic information (¶285, identify a command within the second speech input and determine that the second speech input includes speech directed to the digital assistant; i.e., ¶208, converting speech input into text, identifying a user’s intent expressed in a natural language input received from the user, determining the task flow for fulling the intent).
Regarding Claim 15, Weinberg discloses an electronic device (¶198, digital assistant system 700), including:
a processor (¶198, digital assistant system 700 includes one or more processors 704); and
a memory configured for storing processor-executable instructions (¶197 and ¶199, memory 702 includes non-transitory computer readable medium for software instructions); wherein the processor is configured for reading the executable instructions from the memory and executing the instructions to implement the method according to claim 1 (¶197, software instructions for execution by one or more processors).
Regarding Claim 16, Weinberg discloses wherein the method further includes:
after the determining the completeness of the target semantic information, in response to determining that the target semantic information is complete, generating a control instruction corresponding to the target semantic information based on the target semantic information (¶235, once task flow processing module 736 has completed the structured query for an actionable intent, proceed to perform the ultimate task associated with the actionable intent; see e.g., task flow steps (1)-(4)); and
as modified by Pitschel, updating the to-be-combined semantic information in the semantic state record library based on the target semantic information (Pitschel, Col 17, Rows 29-38, user log stores context information surrounding the user requests, responses provided to the user, the parameters used by the digital assistant to generate and provide the response).
Regarding Claim 17, Weinberg discloses wherein, the method further includes:
after the determining whether there is to-be-combined semantic information cached in a predetermined semantic state record library, in response to determining that there is not the to-be-combined semantic information in the semantic state record library, generating the to-be-combined semantic information based on the target semantic information (¶230, populating some parameters of the structured query with contextual information means not all parameters can be populated with contextual information (i.e., there is not contextual information to complete the structure query); e.g., ¶227, for “invite my friends to my birthday party”, access user data 748 to determine who the “friends” are and when and where the “birthday party” would be held), and
as modified by Pitschel, storing the to-be-combined semantic information into the semantic state record library (Pitschel, Col 17, Rows 29-38, user log stores context information surrounding the user requests, responses provided to the user, the parameters used by the digital assistant to generate and provide the response).
Regarding Claim 18, Weinberg discloses wherein the receiving target interaction information collected when the target user interacts with the target device according to the first interaction mode during the target interaction phase includes:
receiving first interaction information collected from the target user in response to an interaction start signal for triggering a first-round interaction with the target device during the target interaction phase (¶244, user utters “hey Siri” to invoke the digital assistant; once invoked, the digital assistant enters listening state 804 to sample audio including speech input from the user); and
in response to determining that a type of the first interaction information is speech interaction information, determining the first interaction information as the target interaction information (¶244, the digital assistant samples speech input from the user while in the listening state 804 and begin processing the speech input in processing state 806; per ¶208, converting speech input into text, identifying a user’s intent expressed in a natural language input received from the user, determining the task flow for fulling the intent).
Regarding Claim 19, Weinberg discloses wherein in response to receiving target interaction information collected when a target user interacts with a target device according to a first interaction mode during a target interaction phase, the performing semantic recognition on the target interaction information to obtain target semantic information includes:
receiving second interaction information collected when interacting with the target user according to any of predetermined at least one interaction mode, in response to an interaction start signal for triggering a non-first-round interaction with the target device during the target interaction phase (¶284, at block 1406 (after receiving first speech and third speech at blocks 1402 and 1404), initiate a session window associated with a user gaze directed to a displayed digital assistant object and receive a second speech input; the system focuses on a relevant window of time to capture relevant additional speech from the user and improves user experience by capturing additional speech from the user);
generating the target interaction information based on the second interaction information (¶285, determine that the second speech input includes speech directed to the digital assistant according to detected user gaze being directed to a display of the digital assistant electronic device and detect a command within the second speech input); and
performing semantic recognition on the target interaction information according to a semantic recognition mode corresponding to the target interaction information to obtain the target semantic information (¶285, identify a command within the second speech input and determine that the second speech input includes speech directed to the digital assistant; i.e., ¶208, converting speech input into text, identifying a user’s intent expressed in a natural language input received from the user, determining the task flow for fulling the intent).
Claims 7, 14, and 20 are rejected under 35 USC 103(a) as being unpatentable over Weinberg et al. (US 2022/0293124 A1) in view of Jarosz et al. (US 11455982 B2) and Pitschel et al. (US 9922642 B2) as applied to claim 1, in further view of Weinsten et al. (US 9990925 B2).
Regarding Claims 7, 14, and 20 Weinberg discloses wherein the method further includes: in response to triggering a sleep signal for causing the target device to enter an interactive sleep state, controlling the target device to enter the interactive sleep state and exit a target interaction phase (¶288, toggling the device between speech recognition states based on speech thresholds, the system conserves processing resources by transitioning to low power states when appropriate).
Weinberg and Pitschel do not disclose deleting the to-be-combined semantic information from the semantic state record library.
Weinstein discloses processing speech inputs for semantic information that are likely sensitive data (Col 3, Rows 63-66, speech processing system 106 analyzes speech recognition request 107 to generate text transcription; Col 4, Rows 35-38 and Col 5, Rows 55-59, determine at least a portion of the data in the request 107 is likely sensitive data) and deleting the semantic information from a semantic state record library (Col 5, Rows 64-66, portions of request 107 that are sensitive data are not to be logged and may be deleted).
It would’ve been obvious to one ordinarily skilled in the art before the effective filing date of the invention to delete the to-be-combined semantic information from the semantic state record library if it is determined that the to be combined semantic information is sensitive in order to address the challenge that certain to be combined semantic information may not be loggable and need to be removed from the system quickly (Weinstein, Col 2, Rows 28-32).
Conclusion
Applicant's amendment necessitated the new grounds of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to examiner Richard Z. Zhu whose telephone number is 571-270-1587 or examiner’s supervisor Hai Phan whose telephone number is 571-272-6338. Examiner Richard Zhu can normally be reached on M-Th, 0730:1700.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/RICHARD Z ZHU/Primary Examiner, Art Unit 2654 07/22/2026