Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
Claims 6 and 11 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claims 6 and 11 recite the limitation "first sensing information" in the body of the respective claims . There is insufficient antecedent basis for this limitation in the claims.
Claims 6 and 11 depend directly from claim 1; however, claim 1 does not introduce “first sensing information” and instead recites “second sensing information”. Although, claim 3 introduces “ first sensing information,” claims 6 and 11 do not depend from claim 3. Therefore, it is unclear what information is intended by “first sensing information” rendering the scope of claims 6 and 11 indeterminate.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention
was made.
Claim(s) 1-9, 11-12, and 14-15 are rejected under 35 U.S.C. 103 as being unpatentable over Amimoto (Made of reference in ids: WO 2022102432 A1) in view of Osuala ( US-20240111794-A1) and in further view of Chiang (US-20190146590-A1) .
Regarding claim 1, Amimoto teaches “an action processing device for evaluating an action related to communication of a group of action subjects performing the communication with each other” (Amimoto, Para 31, 44–48, 103–105, and 250–253—describing an information-processing apparatus that senses, analyzes, scores, and displays information concerning a dialogue between multiple speakers). Amimoto further teaches “a storage unit configured to store a knowledge database” (Amimoto, Para 54, 59–60, 65–68, 72, 75–77, 99, and 151—describing storage containing previously prepared reference information, scoring example sentences, images, and scoring conditions used to analyze a dialogue). Amimoto further teaches “an input unit configured to receive second sensing information indicating an action related to communication in a second group” (Amimoto, Para 54–56, 69–75, and 153–170—describing cameras, microphones, and other sensors that acquire speech, images, facial expressions, movements, gaze, and physiological information concerning the speakers). Amimoto further teaches “a group state index calculation unit configured to calculate a second group state index related to a state of the communication in the second group based on the second sensing information” (Amimoto, Para 66–79, 92–98, and 173–181—describing analysis and scoring units that calculate evaluation results from sensed speech, images, actions, and other speaker information). Amimoto’s evaluation scores correspond generally to the claimed group-state index, although Amimoto does not expressly use the term “group state index.” Amimoto further teaches “a vector information generation unit configured to generate visible information indicating a state related to the communication of the second group based on the second group state index” (Amimoto, Para:78–79, 113–138, and 173–194—describing the generation and display of scoring information, evaluation-axis scores, bar graphs, checklists, and dialogue-progress information based on the calculated dialogue evaluation). Amimoto further teaches “an action evaluation unit configured to evaluate the action of the second group” (Amimoto, Para:72–79, 92–99, and 250–253—describing evaluation of the speakers’ utterances and behavior by comparing sensed dialogue information with stored reference information and predetermined scoring conditions).
However, Amimoto fails to teach “a knowledge database in which a first group state index related to a state of communication in a first group is associated with first vector information generated based on the first group state index and indicating a feature of the action.” Although Amimoto stores reference examples and scoring conditions, it does not expressly disclose generating first vector information from a first group state index or storing an association between that vector information and the first group state index. Amimoto also fails to teach “to generate second vector information by vectorizing the visible information.” Amimoto generates visible scoring information such as graphs and checklists, but it does not expressly disclose vectorizing that visible information to generate the claimed second vector information. Amimoto further fails to teach “an action evaluation unit configured to evaluate the action of the second group based on a first group state index associated with a piece of first vector information whose comparison result with the second vector information satisfies a predetermined condition among pieces of the first vector information included in the knowledge database.” Amimoto compares recognized dialogue information with stored examples and applies predetermined scoring conditions, but it does not expressly disclose selecting a stored first vector based on its comparison with the second vector and then using the first group state index associated with that selected vector to evaluate the second group.
Osuala teaches “a knowledge database” including “pieces of first vector information” by disclosing a data source represented by stored data-object embeddings in a vector space. Each embedding is a numerical vector representing an associated data object, including text, images, video, sound, or time-series information. (Osuala, claims 1–2 and 16.)
Osuala further teaches “to generate second vector information by vectorizing the visible information” by inputting a received data object into a trained embedding-generation model and receiving an embedding representing the data object in the vector space. Because the data object may comprise text, image, video, or time-series data, the embedding model can be applied to Amimoto’s visible textual and graphical communication-state information. (Osuala, claims 1–2, 10, and 16.)
Osuala further teaches “a piece of first vector information whose comparison result with the second vector information satisfies a predetermined condition among pieces of the first vector information included in the knowledge database” by searching the stored data-object embeddings using the newly generated query embedding, calculating similarity scores between the query embedding and corresponding stored embeddings, and selecting matching embeddings using nearest-neighbor searching or a threshold-similarity condition. (Osuala, claims 1, 5, 7–8, and 16–18; Para 41)
Therefore, Osuala is being relied upon for: Storing multiple reference vectors in a data source; Vectorizing Amimoto’s visible information; and Comparing the resulting second vector with the stored first vectors and selecting a vector satisfying a predetermined similarity condition. It would have been obvious to one of ordinary skill in the art to modify Amimoto’s communication-evaluation system using Osuala’s embedding generation and vector-similarity techniques to efficiently identify stored reference communication information that is semantically similar to the sensed communication information, thereby improving the accuracy and efficiency of Amimoto’s dialogue evaluation.
Chiang teaches “a first group state index … associated with first vector information generated based on the first group state index and indicating a feature of the action” by disclosing storage containing standard action labels and corresponding raw-data sets, generating standard action feature vectors using the labeled raw data, and selecting a representative action feature vector corresponding to each standard action label. Chiang further discloses that the associated classification-vector components represent respective action labels and score levels. In the combination, Chiang’s standard action label and score level correspond to Amimoto’s communication-state index. (Chiang, Fig. 2A; claims 1 and 4.)
Chiang further teaches “an action evaluation unit configured to evaluate the action of the second group based on a first group state index associated with a piece of first vector information” by generating a to-be-evaluated action feature vector from new sensor data, identifying the corresponding action label and score level, selecting the representative action vector corresponding to that action label, and performing an inner-product operation between the new action vector and the representative action vector to generate an evaluation score. (Chiang, Fig. 4; claims 4–5.)
Therefore, Chiang is being relied upon for: Associating each stored representative action vector with a corresponding action label/state index and score level; and Using that associated label/state index and representative vector to evaluate a newly detected action. Clean separation Osuala: stored vectors, vectorizing visible information, vector comparison, and predetermined similarity condition. Chiang: association between a representative vector and its corresponding state/action label, and evaluation based on that associated label/vector. Amimoto supplies the underlying group-communication context, sensing information, calculated communication-state scores, visible scoring information, and general communication evaluation. It would have been obvious to further modify the combined system using Chiang’s technique of associating representative action feature vectors with corresponding action labels and score levels so that a communication action matched to a stored vector could be consistently classified and quantitatively evaluated according to its associated communication-state index.
Regarding Claim 2, Amimoto in view of Osuala and Chiang teaches the action-processing device of Claim 1, as mapped above. Amimoto teaches “wherein the vector information generation unit generates an explanatory text or video information indicating a state related to communication of a group as the visible information” (Amimoto, Para:103–105, 133–138, and 185–194—describing the generation and display of video showing the dialogue participants together with visible scoring information, dialogue-state information, bar graphs, checklists, and progress information indicating the current state of the group’s communication).
Regarding Claim 3, Amimoto in view of Osuala and Chiang teaches the action-processing device of Claim 2, as mapped above. Amimoto teaches “wherein the group state index calculation unit calculates a first group state index related to the state of the communication in the first group based on first sensing information related to the communication of the first group” (Amimoto, Para 69–79, 92–99, and 153–181—describing the acquisition of speech, images, facial expressions, movements, gaze, and other sensing information and the calculation of dialogue-evaluation scores from that sensing information). The calculated dialogue-evaluation score corresponds to the claimed first group state index. Amimoto further teaches “the vector information generation unit generates a first explanatory text or first video information indicating a state related to the communication of the first group based on the first sensing information or the first group state index” (Amimoto, Para 103–105, 133–138, and 185–194—describing the generation and display of dialogue video together with scoring information, dialogue-state information, bar graphs, checklists, and progress information based on sensed dialogue information and calculated scores). However, Amimoto fails to expressly teach “generates the first vector information by vectorizing the first explanatory text or the first video information, and the storage unit stores the generated first vector information.” Osuala teaches “generates the first vector information by vectorizing the first explanatory text or the first video information” (Osuala, claims 1–2, 10, and 16—describing applying an embedding model to a data object to generate a vector-space embedding, where the data object may include text or video information). Osuala further teaches “and the storage unit stores the generated first vector information” (Osuala, claims 1–2 and 16—describing a data source containing stored embeddings representing respective data objects, thereby teaching storage of the generated vector information).
Regarding Claim 4, Amimoto in view of Osuala and Chiang teaches the action-processing device of Claim 1, as mapped above. Amimoto teaches “wherein the action evaluation unit identifies the action of the second group based on first sensing information” (Amimoto, Para: 69–77, 92–99, and 153–181—describing analyzing sensed speech, images, facial expressions, body movements, gaze, and other speaker information to identify and evaluate the speakers’ communication actions). Amimoto further teaches “the action processing device further includes an output unit configured to output the action of the second group” (Amimoto, Para 78–79, 133–138, 185–194, and 293—describing a display and speaker that output the results of the dialogue and action analysis as images, dialogue information, scores, graphs, checklists, or sound). However, Amimoto fails to expressly teach identifying the action “based on … the first group state index related to the communication of the first group corresponding to the first vector information satisfying the condition.” Osuala teaches identifying “the first vector information satisfying the condition” (Osuala, claims 1, 5, 7–8, and 16–18—describing comparing an input embedding with stored embeddings using similarity scores or nearest-neighbor searching and selecting a stored embedding that satisfies a similarity threshold or other search condition). Chiang teaches “the action evaluation unit identifies the action of the second group based on first sensing information and the first group state index related to the communication of the first group corresponding to the first vector information” (Chiang, Fig. 4 and claims 4–5—describing generating an evaluated action-feature vector from newly received sensor data and identifying the action label and score using the classification information and representative action-feature vector associated with that label). Chiang’s stored action label or score level corresponds to the claimed first group state index. Thus, in the combination, Osuala supplies the selection of stored first vector information satisfying the comparison condition, while Chiang supplies the use of the action label or state index associated with that vector to identify the sensed action. Amimoto supplies the communication-sensing context and the output of the identified or evaluated action.
Regarding Claim 5, Amimoto in view of Osuala and Chiang teaches the action-processing device of Claim 4, as mapped above. Amimoto teaches “wherein the output unit outputs an evaluation result obtained by the action evaluation unit” (Amimoto, Para 78–79, 133–138, 173–194, and 313–317—describing the action-analysis and scoring units generating dialogue-evaluation results and the display unit outputting those results as real-time scores, intermediate information, bar graphs, checklists, and final dialogue scores).
Regarding Claim 6, Amimoto in view of Osuala and Chiang teaches the action-processing device of Claim 1, as mapped above. Amimoto teaches “wherein first sensing information related to the communication of the first group and the second sensing information are nonverbal information indicating a nonverbal action” (Amimoto, Para 74–77, 153–170, 281–285, and 324–326—describing cameras and other sensors that obtain and analyze captured images showing the speakers’ facial expressions, body movements, line of sight or gaze, gestures, presentation behavior, and physiological information). These sensed characteristics constitute nonverbal information indicating nonverbal communication actions.
Regarding Claim 7, Amimoto in view of Osuala and Chiang teaches the action-processing device of Claim 1, as mapped above. Amimoto teaches “wherein the first group state index and the second group state index include at least one of a synchrony index and an information flow index of the action subject in the action” (Amimoto, Para 111–118, particularly Para,115—describing a “synchronization” evaluation axis that evaluates whether one dialogue participant attempts to match the speaking pace of the other participant; Para 118 further explains that values for the evaluation axes are calculated and displayed as scoring information).
Regarding Claim 8, Amimoto in view of Osuala and Chiang teaches the action-processing device of Claim 1, as mapped above. Amimoto teaches “wherein the knowledge database further includes additional information on a nonverbal response technique in the state of the communication of the first group” (Amimoto, Para 66–68, 74–76, and 94–99—describing a point-addition target image-information database that stores reference image information used to recognize and score facial expressions, body movements, gaze or line of sight, gestures, and presentation behavior during interpersonal communication).
Regarding Claim 9, Amimoto in view of Osuala and Chiang teaches the action-processing device of Claim 4, as mapped above. Amimoto teaches “an intervention action generation unit configured to generate an intervention action for an action subject in the second group in accordance with the action identified by the action evaluation unit” (Amimoto, Para 195–210 and 217–227—describing feedback processing that identifies a problem in the participant’s dialogue behavior and, based on that identified behavior, generates a goal notification, sound, screen display, or tactile notification intended to cause the participant to change the dialogue strategy or paraphrase an utterance).
Regarding Claim 11, Amimoto in view of Osuala and Chiang teaches the action-processing device of Claim 1, as mapped above. Amimoto teaches “wherein the action evaluation unit evaluates the action of the second group by further using first sensing information” (Amimoto, Para 72–77 and 92–99—describing evaluation of a current dialogue action using previously stored reference language and image information concerning speech, facial expressions, movements, gaze, and presentation behavior). However, Amimoto fails to expressly teach that the first sensing information is “associated with the first vector information within a threshold value.” Osuala teaches “first sensing information associated with the first vector information within a threshold value” (Osuala, claims 1–2, 5, 7–8, and 16–18—describing stored embeddings representing respective data objects, comparing an input embedding with the stored embeddings, identifying an embedding that satisfies a similarity threshold or nearest-neighbor condition, and retrieving or using the data object associated with the qualifying embedding). The stored data object corresponds to the claimed first sensing information, and its stored embedding corresponds to the first vector information.
Regarding Claim 12, Amimoto in view of Osuala and Chiang teaches the action-processing device of Claim 1, as mapped above. Osuala teaches “wherein the condition is at least one of a condition that a difference between the first vector information and the second vector information is within a predetermined threshold value” (Osuala, claims 1, 5, 7–8, and 16–18—describing calculating similarity scores or vector distances between an input embedding and stored embeddings and selecting a stored embedding that satisfies a predetermined similarity threshold or nearest-neighbor condition).
Regarding claim 14, Amimoto teaches “An action processing method executed by an action processing device for evaluating an action related to communication of a group of action subjects performing the communication with each other, the method comprising” (Amimoto, Para 31, 44–48, 103–105, and 250–253—describing an information-processing apparatus that senses, analyzes, scores, and displays information concerning a dialogue between multiple speakers). Amimoto further teaches “a storage unit configured to store a knowledge database” (Amimoto, Para 54, 59–60, 65–68, 72, 75–77, 99, and 151—describing storage containing previously prepared reference information, scoring example sentences, images, and scoring conditions used to analyze a dialogue). Amimoto further teaches “an input unit configured to receive second sensing information indicating an action related to communication in a second group” (Amimoto, Para 54–56, 69–75, and 153–170—describing cameras, microphones, and other sensors that acquire speech, images, facial expressions, movements, gaze, and physiological information concerning the speakers). Amimoto further teaches “a group state index calculation unit configured to calculate a second group state index related to a state of the communication in the second group based on the second sensing information” (Amimoto, Para 66–79, 92–98, and 173–181—describing analysis and scoring units that calculate evaluation results from sensed speech, images, actions, and other speaker information). Amimoto’s evaluation scores correspond generally to the claimed group-state index, although Amimoto does not expressly use the term “group state index.” Amimoto further teaches “a vector information generation unit configured to generate visible information indicating a state related to the communication of the second group based on the second group state index” (Amimoto, Para:78–79, 113–138, and 173–194—describing the generation and display of scoring information, evaluation-axis scores, bar graphs, checklists, and dialogue-progress information based on the calculated dialogue evaluation). Amimoto further teaches “an action evaluation unit configured to evaluate the action of the second group” (Amimoto, Para:72–79, 92–99, and 250–253—describing evaluation of the speakers’ utterances and behavior by comparing sensed dialogue information with stored reference information and predetermined scoring conditions).
However, Amimoto fails to teach “a knowledge database in which a first group state index related to a state of communication in a first group is associated with first vector information generated based on the first group state index and indicating a feature of the action.” Although Amimoto stores reference examples and scoring conditions, it does not expressly disclose generating first vector information from a first group state index or storing an association between that vector information and the first group state index. Amimoto also fails to teach “to generate second vector information by vectorizing the visible information.” Amimoto generates visible scoring information such as graphs and checklists, but it does not expressly disclose vectorizing that visible information to generate the claimed second vector information. Amimoto further fails to teach “an action evaluation unit configured to evaluate the action of the second group based on a first group state index associated with a piece of first vector information whose comparison result with the second vector information satisfies a predetermined condition among pieces of the first vector information included in the knowledge database.” Amimoto compares recognized dialogue information with stored examples and applies predetermined scoring conditions, but it does not expressly disclose selecting a stored first vector based on its comparison with the second vector and then using the first group state index associated with that selected vector to evaluate the second group.
Osuala teaches “a knowledge database” including “pieces of first vector information” by disclosing a data source represented by stored data-object embeddings in a vector space. Each embedding is a numerical vector representing an associated data object, including text, images, video, sound, or time-series information. (Osuala, claims 1–2 and 16.)
Osuala further teaches “to generate second vector information by vectorizing the visible information” by inputting a received data object into a trained embedding-generation model and receiving an embedding representing the data object in the vector space. Because the data object may comprise text, image, video, or time-series data, the embedding model can be applied to Amimoto’s visible textual and graphical communication-state information. (Osuala, claims 1–2, 10, and 16.)
Osuala further teaches “a piece of first vector information whose comparison result with the second vector information satisfies a predetermined condition among pieces of the first vector information included in the knowledge database” by searching the stored data-object embeddings using the newly generated query embedding, calculating similarity scores between the query embedding and corresponding stored embeddings, and selecting matching embeddings using nearest-neighbor searching or a threshold-similarity condition. (Osuala, claims 1, 5, 7–8, and 16–18; Para 41)
Therefore, Osuala is being relied upon for: Storing multiple reference vectors in a data source; Vectorizing Amimoto’s visible information; and Comparing the resulting second vector with the stored first vectors and selecting a vector satisfying a predetermined similarity condition. It would have been obvious to one of ordinary skill in the art to modify Amimoto’s communication-evaluation system using Osuala’s embedding generation and vector-similarity techniques to efficiently identify stored reference communication information that is semantically similar to the sensed communication information, thereby improving the accuracy and efficiency of Amimoto’s dialogue evaluation.
Chiang teaches “a first group state index … associated with first vector information generated based on the first group state index and indicating a feature of the action” by disclosing storage containing standard action labels and corresponding raw-data sets, generating standard action feature vectors using the labeled raw data, and selecting a representative action feature vector corresponding to each standard action label. Chiang further discloses that the associated classification-vector components represent respective action labels and score levels. In the combination, Chiang’s standard action label and score level correspond to Amimoto’s communication-state index. (Chiang, Fig. 2A; claims 1 and 4.)
Chiang further teaches “an action evaluation unit configured to evaluate the action of the second group based on a first group state index associated with a piece of first vector information” by generating a to-be-evaluated action feature vector from new sensor data, identifying the corresponding action label and score level, selecting the representative action vector corresponding to that action label, and performing an inner-product operation between the new action vector and the representative action vector to generate an evaluation score. (Chiang, Fig. 4; claims 4–5.)
Therefore, Chiang is being relied upon for: Associating each stored representative action vector with a corresponding action label/state index and score level; and Using that associated label/state index and representative vector to evaluate a newly detected action. Clean separation Osuala: stored vectors, vectorizing visible information, vector comparison, and predetermined similarity condition. Chiang: association between a representative vector and its corresponding state/action label, and evaluation based on that associated label/vector. Amimoto supplies the underlying group-communication context, sensing information, calculated communication-state scores, visible scoring information, and general communication evaluation. It would have been obvious to further modify the combined system using Chiang’s technique of associating representative action feature vectors with corresponding action labels and score levels so that a communication action matched to a stored vector could be consistently classified and quantitatively evaluated according to its associated communication-state index.
Regarding claim 15, Amimoto teaches “A storage medium storing an action processing program causing an action processing device that is a computer for evaluating an action related to communication of a group of action subjects performing the communication with each other to function as” (Amimoto, Para 31, 44–48, 103–105, and 250–253—describing an information-processing apparatus that senses, analyzes, scores, and displays information concerning a dialogue between multiple speakers). Amimoto further teaches “a storage unit configured to store a knowledge database” (Amimoto, Para 54, 59–60, 65–68, 72, 75–77, 99, and 151—describing storage containing previously prepared reference information, scoring example sentences, images, and scoring conditions used to analyze a dialogue). Amimoto further teaches “an input unit configured to receive second sensing information indicating an action related to communication in a second group” (Amimoto, Para 54–56, 69–75, and 153–170—describing cameras, microphones, and other sensors that acquire speech, images, facial expressions, movements, gaze, and physiological information concerning the speakers). Amimoto further teaches “a group state index calculation unit configured to calculate a second group state index related to a state of the communication in the second group based on the second sensing information” (Amimoto, Para 66–79, 92–98, and 173–181—describing analysis and scoring units that calculate evaluation results from sensed speech, images, actions, and other speaker information). Amimoto’s evaluation scores correspond generally to the claimed group-state index, although Amimoto does not expressly use the term “group state index.” Amimoto further teaches “a vector information generation unit configured to generate visible information indicating a state related to the communication of the second group based on the second group state index” (Amimoto, Para:78–79, 113–138, and 173–194—describing the generation and display of scoring information, evaluation-axis scores, bar graphs, checklists, and dialogue-progress information based on the calculated dialogue evaluation). Amimoto further teaches “an action evaluation unit configured to evaluate the action of the second group” (Amimoto, Para:72–79, 92–99, and 250–253—describing evaluation of the speakers’ utterances and behavior by comparing sensed dialogue information with stored reference information and predetermined scoring conditions).
However, Amimoto fails to teach “a knowledge database in which a first group state index related to a state of communication in a first group is associated with first vector information generated based on the first group state index and indicating a feature of the action.” Although Amimoto stores reference examples and scoring conditions, it does not expressly disclose generating first vector information from a first group state index or storing an association between that vector information and the first group state index. Amimoto also fails to teach “to generate second vector information by vectorizing the visible information.” Amimoto generates visible scoring information such as graphs and checklists, but it does not expressly disclose vectorizing that visible information to generate the claimed second vector information. Amimoto further fails to teach “an action evaluation unit configured to evaluate the action of the second group based on a first group state index associated with a piece of first vector information whose comparison result with the second vector information satisfies a predetermined condition among pieces of the first vector information included in the knowledge database.” Amimoto compares recognized dialogue information with stored examples and applies predetermined scoring conditions, but it does not expressly disclose selecting a stored first vector based on its comparison with the second vector and then using the first group state index associated with that selected vector to evaluate the second group.
Osuala teaches “a knowledge database” including “pieces of first vector information” by disclosing a data source represented by stored data-object embeddings in a vector space. Each embedding is a numerical vector representing an associated data object, including text, images, video, sound, or time-series information. (Osuala, claims 1–2 and 16.)
Osuala further teaches “to generate second vector information by vectorizing the visible information” by inputting a received data object into a trained embedding-generation model and receiving an embedding representing the data object in the vector space. Because the data object may comprise text, image, video, or time-series data, the embedding model can be applied to Amimoto’s visible textual and graphical communication-state information. (Osuala, claims 1–2, 10, and 16.)
Osuala further teaches “a piece of first vector information whose comparison result with the second vector information satisfies a predetermined condition among pieces of the first vector information included in the knowledge database” by searching the stored data-object embeddings using the newly generated query embedding, calculating similarity scores between the query embedding and corresponding stored embeddings, and selecting matching embeddings using nearest-neighbor searching or a threshold-similarity condition. (Osuala, claims 1, 5, 7–8, and 16–18; Para 41)
Therefore, Osuala is being relied upon for: Storing multiple reference vectors in a data source; Vectorizing Amimoto’s visible information; and Comparing the resulting second vector with the stored first vectors and selecting a vector satisfying a predetermined similarity condition. It would have been obvious to one of ordinary skill in the art to modify Amimoto’s communication-evaluation system using Osuala’s embedding generation and vector-similarity techniques to efficiently identify stored reference communication information that is semantically similar to the sensed communication information, thereby improving the accuracy and efficiency of Amimoto’s dialogue evaluation.
Chiang teaches “a first group state index … associated with first vector information generated based on the first group state index and indicating a feature of the action” by disclosing storage containing standard action labels and corresponding raw-data sets, generating standard action feature vectors using the labeled raw data, and selecting a representative action feature vector corresponding to each standard action label. Chiang further discloses that the associated classification-vector components represent respective action labels and score levels. In the combination, Chiang’s standard action label and score level correspond to Amimoto’s communication-state index. (Chiang, Fig. 2A; claims 1 and 4.)
Chiang further teaches “an action evaluation unit configured to evaluate the action of the second group based on a first group state index associated with a piece of first vector information” by generating a to-be-evaluated action feature vector from new sensor data, identifying the corresponding action label and score level, selecting the representative action vector corresponding to that action label, and performing an inner-product operation between the new action vector and the representative action vector to generate an evaluation score. (Chiang, Fig. 4; claims 4–5.)
Therefore, Chiang is being relied upon for: Associating each stored representative action vector with a corresponding action label/state index and score level; and Using that associated label/state index and representative vector to evaluate a newly detected action. Clean separation Osuala: stored vectors, vectorizing visible information, vector comparison, and predetermined similarity condition. Chiang: association between a representative vector and its corresponding state/action label, and evaluation based on that associated label/vector. Amimoto supplies the underlying group-communication context, sensing information, calculated communication-state scores, visible scoring information, and general communication evaluation. It would have been obvious to further modify the combined system using Chiang’s technique of associating representative action feature vectors with corresponding action labels and score levels so that a communication action matched to a stored vector could be consistently classified and quantitatively evaluated according to its associated communication-state index.
Claim(s) 10 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Amimoto (Made of reference in ids: WO 2022102432 A1 ) in view of Osuala ( US-20240111794-A1) and in further view of Chiang (US-20190146590-A1) and Veltrop (US-20170120446-A1) .
Regarding Claim 10, Amimoto in view of Osuala and Chiang teaches the device of Claim 9. However, the three-reference combination does not appear to teach the newly added limitation. Amimoto teaches generating feedback or an intervention based on an evaluated communication action (Amimoto, Para 195–210 and 217–227—describing goal notifications, sounds, and screen displays that prompt a speaker to change a dialogue strategy).
However, Amimoto’s action subjects are real speakers communicating through displays, even when a speaker performs a simulated patient or customer role (Amimoto, Para 103–105 and 195–201). Accordingly, Amimoto fails to teach: “wherein the action subject in the second group includes a virtual human, and the action processing device further includes a control command unit configured to create a control command for controlling the virtual human in accordance with the intervention action.” Osuala’s embedding-based similarity searching does not teach a virtual human or generating commands to control one. Chiang’s action-classification system likewise does not teach controlling a virtual human according to a generated intervention.
Veltrop teaches “wherein the action subject in the second group includes a virtual human” (Veltrop, Para [0036]–[0042], [0053], and claim 1—describing a humanoid robot that senses and communicates with humans through speech, gestures, movements, lights, and other social behaviors). The humanoid robot corresponds to a virtual human, particularly because the present application expressly treats a robot as one type of virtual human.
Veltrop further teaches “a control command unit configured to create a control command for controlling the virtual human in accordance with the intervention action” (Veltrop, Para [0042], [0049], [0053], and [0061], and claim 1—describing a Mind module that evaluates sensed interaction conditions, selects an appropriate activity, and commands execution of that activity by the humanoid robot). For example, detecting a human smile causes selection of an empathic robot response such as smiling or speaking kind words (Para [0061]). The Mind module corresponds to the claimed control-command unit, and the selected empathic behavior corresponds to an intervention action.
It would have been obvious to further modify Amimoto in view of Osuala and Chiang using Veltrop’s humanoid-robot control system so that the intervention action generated from the evaluated communication could be converted into an executable robot activity and corresponding actuator commands, thereby allowing the system to automatically provide responsive communication feedback through speech, gestures, facial expressions, or movement.
Regarding claim 13, The action processing device according to claim 10, wherein an actuator is controlled based on the created control command (Veltrop, Para [0037], Para [0042], and claims 1 and 12—describing actuator services that operate joint and base motors, lights, speech, gestures, and other robot behaviors, with the Mind module commanding execution of the selected activity by the actuators).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Wang et al (20220269713): relevant for generating vector representations of text and visual information, comparing vectors using pairwise inner products, and selecting the most similar stored information.
Rao (20200134038): relevant for generating word vectors, storing information in a database and mapping vectors to stored information using a deep similarity network.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to LATRELL ANTHONY CREARY whose telephone number is (703)756-1219. The examiner can normally be reached Mon - Fri 7:30am - 4:30pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Xiao WU can be reached on (571) 272-7761. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/LATRELL ANTHONY CREARY/Examiner, Art Unit 2613
/XIAO M WU/Supervisory Patent Examiner, Art Unit 2613