DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of Claims
This office action is in response to the amendment filed on 7/22/2026.
Claims 1-8 have been amended.
Claims 9-17 have been added.
Claims 1-17 are pending and have been examined.
Priority
Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed in parent Application No. JP2024-183669 filed on 10/18/2024.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-17 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: Claims 1-17 are directed to a system. Thus, on their face they fall within the four statutory categories of patentable subject matter.
Step 2A prong 1:
The following limitations, when considered individually and as an ordered combination, are merely descriptive of abstract concepts:
Claim 1:
analyze a video;
generate images and descriptive text based on a content of the video
create a manual based on the images and descriptive text generated
provide an entity that responds to a query regarding the manual
estimate an emotion of a user by inputting user data based on at least one of a facial expression or a voice of the user to a pre-learned model that maps user data to an emotion map; and
adjust, based on the estimated emotion, at least one of speed, a priority, a length, or a structure of the analyzing, the generating, the creating, or the entity responding.
The following dependent claim limitations, when considered individually and as an ordered combination, are merely further descriptive of abstract concepts:
Claim 2:
detect a specific object in a frame of the video
Claim 3:
generate descriptive text including interactive elements based on the content of the video
Claim 4:
determine a structure and format of the manual based on the images and descriptive text
Claim 5:
when a user asks a question about the content of the manual, the entity provides an answer.
Claim 6:
perform version control by always provide the latest manual
Claim 7:
adjust a video analysis method used to analyze the video based on the estimated emotion.
Claim 8:
detect specific actions or gestures during video analysis and to detail the analysis results based on the detection.
Claim 9:
wherein the specific object comprises at least one of a tool or a screen element shown in the video.
Claim 10:
wherein the specific action or gesture comprises at least one of a hand movement, a body movement, or a facial expression of a person shown in the video.
Claim 11:
adjust the video analysis method by adjusting a speed at which the video is analyzed based on the estimated emotion.
Claim 12:
adjust the video analysis method by adjusting a priority of analysis results based on the estimated emotion.
Claim 13:
adjust the video analysis method by adjusting a display method of analysis results based on the estimated emotion.
Claim 14:
adjust a length of the images and descriptive text generated based on the estimated emotion.
Claim 15:
adjust a method of structuring the manual based on the estimated emotion.
Claim 16:
adjust a response method of the entity based on the estimated emotion.
Claim 17:
estimate the emotion of the user by inputting the user data to the pre-learned model, the pre-learned model outputting, from the emotion map in which a plurality of emotions are arranged, an emotion value for each of the plurality of emotions, and determining the emotion of the user based on the emotion values.
The claims provide a manner of analyzing a video, generating images and text based on the video, create a manual based on the images and text, provide an entity that responds to queries about the manual, estimates an emotion, and adjusts the analyzing, generating, creating, or entity responding based on the estimated emotion. But for the inclusion of generic computing devices (i.e. processor), the limitations of the claims could be performed mentally or with pen and paper. A human analog would be able to analyze a video, generate image and text based on the video, create a manual based on the images and text, provide an entity that responds to queries about the manual, estimate emotion from facial expression or voice, and adjust the analyzing, generating, creating, or entity responding based on the estimated emotion. As a result, the claims fall within the mental process grouping of abstract ideas.
Step 2A prong 2: This judicial exception is not integrated into a practical application. The claims recite the following additional elements: processor (claim 1-8, 11-17); chatbot (claim 1, 5, 16); generative AI (claim 1, 2-4, 6); pre-learned neural network (claim 1, 17)
The processor is recited at a high level of generality and merely “apply it” (the abstract idea) using a generic computing device ([0018]). The processor is used to process data (analyze, generate, create, provide, estimate, adjust, detect, determine, estimate). The limitations are merely result based claiming and fail to recite details of how a solution to a problem is accomplished. Nothing in the claims improves upon technology or a technical field. Therefore, they do not go beyond the “apply it” level of implementation (See MPEP 2106.05(f)).
The chatbot is recited at a high level of generality and amounts to mere computer implementation. The limitation regarding the creation of the chatbot is merely result based claiming without any details of how said chatbot is created. Nothing in the claims improves upon chatbot technology or a technical field. (See MPEP 2106.05(f)).
The generative AI and pre-learned neural network are merely high-level results based claiming without providing any meaningful details as to how the generative AI or the pre-trained neural network perform the claimed features. The claims amount to little more than “do it” with generative AI. Thus, nothing in the claims improves upon generative AI technology, neural network technology, or a technical field. Thus, the limitations do not go beyond the “apply it” level of implementation (see MPEP 2106.05(f)).
Accordingly, when considered both individually and as an ordered combination, the additional elements do not impose any meaningful limits on practicing the abstract idea.
Step 2B: The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception. Similarly, as above with regard to practical application, the additional elements when considered both individually and as an ordered combination, do not provide an inventive concept as they merely provide generic computing components used as a tool to implement the abstract idea.
As a result, the claims are not patent eligible.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1, 2, 4, 5, 7, 8, 9, 11, 12, 13, 15, and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over McKenzie-Kelly et al (US 2023/0350947) hereafter “Kelly” in view of KR102073928 hereafter “chatbot” in view of Angeles (US 2026/0037863)
As per claim 1:
Kelly teaches:
A system comprising:
A processor configured to ([0005]):
analyze a video by inputting the video to a generative AI; ([0022] The method may receive 101 a video file with an associated transcript of the audio associated with the video file. The associated transcript may be generated by converting from speech in the video file to text using known conversion methods. The associated transcript may include timestamps corresponding to the audio of the video file. [0023] The method may identify 102 task action terms in the transcript. Natural language processing (NLP) is used to locate terminology that is commonly used for explanations and instructions of task actions. Identifying the task action terms may be configured and improved through learning. [0024] The method may locate 103 a visual section of the video file corresponding to the task action term. In one embodiment, locating a visual section of the video may be achieved by determining a timestamp of the identified task action term in the video file and locating a visual section of the video file corresponding to the timestamp. In another embodiment, locating a visual section may involve visually analyzing the video file to identify visual elements matching the task action terms. [0030] In the described method, the sound file 222 undergoes a speech-to-text conversion 201 to obtain a transcript 223 of the narration. This may be carried out using existing speech-to-text converters. Timestamps are added to the transcript 223 corresponding to the video file 221, for example, for each word in the transcript 223. [0056] The visual section component may include a visual analysis component 422 for visually analyzing the video file to identify visual elements matching the task action terms. This may be used when a timestamp is not available or where the timestamp provides a range including multiple visual elements.)
generate images and descriptive text based on the content of the video by means of generative AI; ([0022] The associated transcript may be generated by converting from speech in the video file to text using known conversion methods. The associated transcript may include timestamps corresponding to the audio of the video file. [0023] The method may identify 102 task action terms in the transcript. Natural language processing (NLP) is used to locate terminology that is commonly used for explanations and instructions of task actions. Identifying the task action terms may be configured and improved through learning. [0025] The method may capture 104 at least a portion of the visual section of the video file. The captured visual section may be a screenshot, a video frame, or a video clip or exert of the video file. Capturing the visual section may store a visual section from the video file in a repository and may provide a link to the visual section in the repository for use when generating the task instruction document. [0030] In the described method, the sound file 222 undergoes a speech-to-text conversion 201 to obtain a transcript 223 of the narration. This may be carried out using existing speech-to-text converters. Timestamps are added to the transcript 223 corresponding to the video file 221, for example, for each word in the transcript 223. [0031] NLP 202 may be is used to distinguish the actions, or steps, within the transcript and to identify elements of the interface that are used for those actions. The NLP may identify task action terms or references to actions. For example, this may identify terms such as instructions of “click on”, “navigate to”, “open” and sequencing words like “now”, “next”, “then”, “before”, and/or image references such as “here”, “as you can see”, “shown”, etc. [0032] When a task action term is identified in the transcript, a screenshot or frame 224 from the video file 221 at the corresponding timestamp is captured 203 and added 206 to the narrative text. Screenshots 224 are therefore added 206 where certain words are detected that indicate a reference to or an instruction of an action.)
create a manual based on the images and descriptive text by means of the generative AI; and ([0018] Embodiments of a method, system, and computer program product are provided for automatically creating task content. The described method receives a video file with an associated transcript and converts this into a task instruction document including task action terms augmented with interface element information and captured visual sections from the video file. [0027] The method may generate 106 a task instruction document including the task action terms augmented with interface element information and with at least a portion of the captured visual section. The method extracts and enriches the transcript content with visual information from the video file to build an accurate documentation task topic.)
Chatbot teaches:
provide a chatbot that responds to a query regarding the manual, “Manual search service providing server 100 using a chatbot according to an embodiment of the present invention may provide a chatbot service related to the manual search service to the user terminal (10). The manual may be one of various kinds of manuals, such as a manual for a plurality of products and a manual for specific instructions.” “To this end, the manual search service providing server 100 using the chatbot according to an embodiment of the present invention receives a manual to provide a chatbot service, extracts a key word from the manual, and stores the key word in advance, and through the chatbot in the user terminal 10. By using the received query items and pre-stored keywords, the user terminal 10 may provide a response to the question, that is, an answer.” “For example, the communication interface 110 may receive a query related to the manual search service from the user terminal 10 through the chatbot.” “In addition, the communication interface 110 may transmit one or more sentences (the sentences retrieved from the manual) to the user terminal 10 through the chatbot as a response to the received inquiry.” “In addition, the processor 130 may provide a chatbot service to the user terminal 10 requesting a chatbot service by executing a chatbot program stored in the memory 120. The processor 130 extracts the keywords (ie, query keywords) from the query received from the user terminal 10 through the chatbot, compares the extracted query keywords with manual keywords stored in the memory 120, After searching for the answer related to the question, the user terminal 10 may be processed to provide the answer.”
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include a support unit configured to provide a chatbot that responds to the manual created by the creation unit as taught by Chatbot with the instruction document creator of Kelly in order to allow a user to freely ask the user’s desired question and decrease the time and effort to obtain the desired answer (“Therefore, the user may feel a limitation in freely asking the user's desired question, and as a result, the user may not get the desired answer or it may require a lot of time and effort to obtain the desired answer.”)
Kelly in view of Chatbot does not expressly teach estimate an emotion of a user by inputting user data based on at least one of a facial expression or a voice of the user to a pre-learned neural network that maps the user data to an emotion map: and adjust, based on the estimated emotion, at least one of a speed, a priority, a length, or a structure of at least one of the analyzing, the generating, the creating, or the chatbot responding.
Angeles teaches:
estimate an emotion of a user by inputting user data based on at least one of a facial expression or a voice of the user to a pre-learned neural network that maps the user data to an emotion map; and ([0555] Emotional Recognition and Response Algorithm: Utilizes deep neural networks to analyze real-time data from facial recognition, voice intonation, and physiological sensors integrated into VR/AR interfaces. The algorithm processes these data points to ascertain emotional states and adjust digital content accordingly, such as modifying difficulty levels, narrative elements, or interactive features. This algorithm goes beyond simple emotion detection by actively altering digital and virtual environments in response to detected emotional cues. It's crafted to foster emotional well-being and engagement, using sophisticated model training that incorporates psychological and behavioral science principles. Interfaces with devices capable of detecting user emotions (e.g., facial expression recognition in VR/AR environments). It adapts chatbot responses based on emotional cues, improving user interaction by providing empathetic and contextually appropriate responses to emotional states related to health monitoring. See also [0200]-[0202], [0204], [0207], [0342], [0344], [0349], [0527], [0536])
adjust, based on the estimated emotion, at least one of a speed, a priority, a length, or a structure of at least one of the analyzing, the generating, the creating, or the chatbot responding ([0555] Emotional Recognition and Response Algorithm: Utilizes deep neural networks to analyze real-time data from facial recognition, voice intonation, and physiological sensors integrated into VR/AR interfaces. The algorithm processes these data points to ascertain emotional states and adjust digital content accordingly, such as modifying difficulty levels, narrative elements, or interactive features. This algorithm goes beyond simple emotion detection by actively altering digital and virtual environments in response to detected emotional cues. It's crafted to foster emotional well-being and engagement, using sophisticated model training that incorporates psychological and behavioral science principles. Interfaces with devices capable of detecting user emotions (e.g., facial expression recognition in VR/AR environments). It adapts chatbot responses based on emotional cues, improving user interaction by providing empathetic and contextually appropriate responses to emotional states related to health monitoring. See also [0200]-[0202], [0204], [0207], [0342], [0344], [0349], [0527], [0536]))
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include estimate an emotion of a user by inputting user data based on at least one of a facial expression or a voice of the user to a pre-learned neural network that maps the user data to an emotion map: and adjust, based on the estimated emotion, at least one of a speed, a priority, a length, or a structure of at least one of the analyzing, the generating, the creating, or the chatbot responding as taught by Angeles with the instruction document creator of Kelly in view of Chatbot in order to foster emotional well-being and engagement, using sophisticated model training that incorporates psychological and behavioral science principles ([0555]).
Kelly in view of Chatbot in view of Angeles teaches the limitations of claim 1. As per claim 2:
Kelly further teaches:
wherein the processor is further configured to detect a specific object in a frame of the video by means of the generative AI. ([0022] The method may receive 101 a video file with an associated transcript of the audio associated with the video file. The associated transcript may be generated by converting from speech in the video file to text using known conversion methods. The associated transcript may include timestamps corresponding to the audio of the video file. [0023] The method may identify 102 task action terms in the transcript. Natural language processing (NLP) is used to locate terminology that is commonly used for explanations and instructions of task actions. Identifying the task action terms may be configured and improved through learning. [0024] The method may locate 103 a visual section of the video file corresponding to the task action term. In one embodiment, locating a visual section of the video may be achieved by determining a timestamp of the identified task action term in the video file and locating a visual section of the video file corresponding to the timestamp. In another embodiment, locating a visual section may involve visually analyzing the video file to identify visual elements matching the task action terms. [0030] In the described method, the sound file 222 undergoes a speech-to-text conversion 201 to obtain a transcript 223 of the narration. This may be carried out using existing speech-to-text converters. Timestamps are added to the transcript 223 corresponding to the video file 221, for example, for each word in the transcript 223.)
Kelly in view of Chatbot in view of Angeles teaches the limitations of claim 1. As per claim 4:
Kelly teaches:
wherein the processor is further configured to determine a structure and format of the manual based on the images and descriptive text generated by means of the generative AI. ([0022] The method may receive 101 a video file with an associated transcript of the audio associated with the video file. The associated transcript may be generated by converting from speech in the video file to text using known conversion methods. The associated transcript may include timestamps corresponding to the audio of the video file. [0023] The method may identify 102 task action terms in the transcript. Natural language processing (NLP) is used to locate terminology that is commonly used for explanations and instructions of task actions. Identifying the task action terms may be configured and improved through learning. [0024] The method may locate 103 a visual section of the video file corresponding to the task action term. In one embodiment, locating a visual section of the video may be achieved by determining a timestamp of the identified task action term in the video file and locating a visual section of the video file corresponding to the timestamp. In another embodiment, locating a visual section may involve visually analyzing the video file to identify visual elements matching the task action terms. [0027] The method may generate 106 a task instruction document including the task action terms augmented with interface element information and with at least a portion of the captured visual section. The method extracts and enriches the transcript content with visual information from the video file to build an accurate documentation task topic. [0030] In the described method, the sound file 222 undergoes a speech-to-text conversion 201 to obtain a transcript 223 of the narration. This may be carried out using existing speech-to-text converters. Timestamps are added to the transcript 223 corresponding to the video file 221, for example, for each word in the transcript 223.)
Kelly in view of Chatbot in view of Angeles teaches the limitations of claim 1. As per claim 5:
Kelly in view of Angeles does not expressly teach wherein the support unit is configured such that when a user asks a question about the content of the manual, the chatbot provides an answer.
Chatbot teaches:
wherein the processor is further configured such that when a user asks a question about the content of the manual, the chatbot provides an answer. “Manual search service providing server 100 using a chatbot according to an embodiment of the present invention may provide a chatbot service related to the manual search service to the user terminal (10). The manual may be one of various kinds of manuals, such as a manual for a plurality of products and a manual for specific instructions.” “To this end, the manual search service providing server 100 using the chatbot according to an embodiment of the present invention receives a manual to provide a chatbot service, extracts a key word from the manual, and stores the key word in advance, and through the chatbot in the user terminal 10. By using the received query items and pre-stored keywords, the user terminal 10 may provide a response to the question, that is, an answer.” “For example, the communication interface 110 may receive a query related to the manual search service from the user terminal 10 through the chatbot.” “In addition, the communication interface 110 may transmit one or more sentences (the sentences retrieved from the manual) to the user terminal 10 through the chatbot as a response to the received inquiry.” “In addition, the processor 130 may provide a chatbot service to the user terminal 10 requesting a chatbot service by executing a chatbot program stored in the memory 120. The processor 130 extracts the keywords (ie, query keywords) from the query received from the user terminal 10 through the chatbot, compares the extracted query keywords with manual keywords stored in the memory 120, After searching for the answer related to the question, the user terminal 10 may be processed to provide the answer.”
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include wherein the support unit is configured such that when a user asks a question about the content of the manual, the chatbot provides an answer as taught by Chatbot with the instruction document creator of Kelly in order to allow a user to freely ask the user’s desired question and decrease the time and effort to obtain the desired answer (“Therefore, the user may feel a limitation in freely asking the user's desired question, and as a result, the user may not get the desired answer or it may require a lot of time and effort to obtain the desired answer.”)
Kelly in view of Chatbot in view of Angeles teaches the limitations of claim 1. As per claim 7:
Kelly in view of Chatbot does not expressly teach wherein the processor is further configured to adjust a video analysis method used to analyze the video based on the estimated emotion.
Angeles teaches:
wherein the processor is further configured to adjust a video analysis method used to analyze the video based on the estimated emotion. ([0555] Emotional Recognition and Response Algorithm: Utilizes deep neural networks to analyze real-time data from facial recognition, voice intonation, and physiological sensors integrated into VR/AR interfaces. The algorithm processes these data points to ascertain emotional states and adjust digital content accordingly, such as modifying difficulty levels, narrative elements, or interactive features. This algorithm goes beyond simple emotion detection by actively altering digital and virtual environments in response to detected emotional cues. It's crafted to foster emotional well-being and engagement, using sophisticated model training that incorporates psychological and behavioral science principles. Interfaces with devices capable of detecting user emotions (e.g., facial expression recognition in VR/AR environments). It adapts chatbot responses based on emotional cues, improving user interaction by providing empathetic and contextually appropriate responses to emotional states related to health monitoring. See also [0200]-[0202], [0204], [0207], [0342], [0344], [0349], [0527], [0536]))
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include wherein the processor is further configured to adjust a video analysis method used to analyze the video based on the estimated emotion. as taught by Angeles with the instruction document creator of Kelly in view of Chatbot in order to foster emotional well-being and engagement, using sophisticated model training that incorporates psychological and behavioral science principles ([0555]).
Kelly in view of Chatbot in view of Angeles teaches the limitations of claim 1. As per claim 8:
Kelly teaches:
wherein the processor is further configured to detect specific actions or gestures during video analysis and to detail the analysis results based on the detection. ([0056] The visual section component may include a visual analysis component 422 for visually analyzing the video file to identify visual elements matching the task action terms. This may be used when a timestamp is not available or where the timestamp provides a range including multiple visual elements. [0057] The capturing component 414 may include a storing component 423 for storing a visual section from the video file in a repository and a link component 424 for providing a link in the transcript to the visual section in the repository. See also [0044])
Kelly in view of Chatbot in view of Angeles teaches the limitations of claim 2. As per claim 9:
Kelly teaches:
wherein the specific object comprises at least one of a tool or a screen element shown in the video. ([0024] In another embodiment, locating a visual section may involve visually analyzing the video file to identify visual elements matching the task action terms. [0026] The method may identify 105 information relating to one or more interface elements that are being interacted with in the visual section by using image recognition. This may include identifying or verifying interface elements in the visual section corresponding to the task action term by visual analysis. [0044] The analysis may also include movement detection 305 of where the mouse pointer is or where typing is happening indicating an interface element to which an action is being applied. Once the rectangles on the screen are identified, movement detection 305 may be used to determine the focus on the screen. Motion detection may be used to work out what is moving on the screen, in particular a mouse cursor, typing, and text input. Any changes in the detected rectangles may also suggest windows opening and closing. To detect the differences in movement a machine learning may be used. For example, to detect a difference between a mouse cursor or typing compared to a loading bar.
Kelly in view of Chatbot in view of Angeles teaches the limitations of claim 7. As per claim 11:
Kelly in view of Chatbot does not expressly teach wherein the processor is configured to adjust the video analysis method by adjusting a speed at which the video is analyzed based on the estimated emotion.
Angeles teaches:
wherein the processor is configured to adjust the video analysis method by adjusting a speed at which the video is analyzed based on the estimated emotion. ([0555] Emotional Recognition and Response Algorithm: Utilizes deep neural networks to analyze real-time data from facial recognition, voice intonation, and physiological sensors integrated into VR/AR interfaces. The algorithm processes these data points to ascertain emotional states and adjust digital content accordingly, such as modifying difficulty levels, narrative elements, or interactive features. This algorithm goes beyond simple emotion detection by actively altering digital and virtual environments in response to detected emotional cues. It's crafted to foster emotional well-being and engagement, using sophisticated model training that incorporates psychological and behavioral science principles. Interfaces with devices capable of detecting user emotions (e.g., facial expression recognition in VR/AR environments). It adapts chatbot responses based on emotional cues, improving user interaction by providing empathetic and contextually appropriate responses to emotional states related to health monitoring. [0202] It can adjust the difficulty and pace of the curriculum based on the owner's progress and feedback. [0344] The chatbot tracks the owner's progress towards their learning goals, providing regular updates and motivational feedback to keep the owner engaged and focused. It can adjust the difficulty and pace of the curriculum based on the owner's progress and feedback.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include wherein the processor is configured to adjust the video analysis method by adjusting a speed at which the video is analyzed based on the estimated emotion as taught by Angeles with the instruction document creator of Kelly in view of Chatbot in order to foster emotional well-being and engagement, using sophisticated model training that incorporates psychological and behavioral science principles ([0555]).
Kelly in view of Chatbot in view of Angeles teaches the limitations of claim 7. As per claim 12:
Kelly in view of Chatbot does not expressly teach wherein the processor is configured to adjust the video analysis method by adjusting a priority of analysis results based on the estimated emotion.
Angeles teaches:
wherein the processor is configured to adjust the video analysis method by adjusting a priority of analysis results based on the estimated emotion. ([0555] Emotional Recognition and Response Algorithm: Utilizes deep neural networks to analyze real-time data from facial recognition, voice intonation, and physiological sensors integrated into VR/AR interfaces. The algorithm processes these data points to ascertain emotional states and adjust digital content accordingly, such as modifying difficulty levels, narrative elements, or interactive features. This algorithm goes beyond simple emotion detection by actively altering digital and virtual environments in response to detected emotional cues. It's crafted to foster emotional well-being and engagement, using sophisticated model training that incorporates psychological and behavioral science principles. Interfaces with devices capable of detecting user emotions (e.g., facial expression recognition in VR/AR environments). It adapts chatbot responses based on emotional cues, improving user interaction by providing empathetic and contextually appropriate responses to emotional states related to health monitoring. [0200] The chatbot assistant can also provide skill gap analysis, by continuously monitoring the owner's interactions and progress, to identify skill gaps and adjusts the curriculum in real-time to address these gaps, ensuring a focused and efficient learning path. [0201] Adaptive learning environment: The chatbot assistant can provide content customization, by curating and recommending learning materials, such as articles, videos, and online courses, that specifically match the owner's current level of understanding and interest areas. It can also generate custom content or exercises using its generative AI capabilities to address specific learning needs. The chatbot assistant can also provide interactive learning, where through conversational interfaces, the chatbot engages the owner in interactive learning sessions, quizzes, and problem-solving exercises, providing immediate feedback and explanations to foster understanding. See also [0200]-[0202], [0204], [0207], [0342], [0344], [0349], [0527], [0536]))
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include wherein the processor is configured to adjust the video analysis method by adjusting a priority of analysis results based on the estimated emotion as taught by Angeles with the instruction document creator of Kelly in view of Chatbot in order to foster emotional well-being and engagement, using sophisticated model training that incorporates psychological and behavioral science principles ([0555]).
Kelly in view of Chatbot in view of Angeles teaches the limitations of claim 7. As per claim 13:
Kelly in view of Chatbot does not expressly teach wherein the processor is configured to adjust the video analysis method by adjusting a display method of analysis results based on the estimated emotion.
Angeles teaches:
wherein the processor is configured to adjust the video analysis method by adjusting a display method of analysis results based on the estimated emotion ([0555] Emotional Recognition and Response Algorithm: Utilizes deep neural networks to analyze real-time data from facial recognition, voice intonation, and physiological sensors integrated into VR/AR interfaces. The algorithm processes these data points to ascertain emotional states and adjust digital content accordingly, such as modifying difficulty levels, narrative elements, or interactive features. This algorithm goes beyond simple emotion detection by actively altering digital and virtual environments in response to detected emotional cues. It's crafted to foster emotional well-being and engagement, using sophisticated model training that incorporates psychological and behavioral science principles. Interfaces with devices capable of detecting user emotions (e.g., facial expression recognition in VR/AR environments). It adapts chatbot responses based on emotional cues, improving user interaction by providing empathetic and contextually appropriate responses to emotional states related to health monitoring. [0349] 14-1. Emotional Intelligence Training. Sentiment Analysis: Utilize advanced NLP techniques to analyze user input for emotional cues and sentiment. This allows the chatbot to gauge the user's mood and emotional state. Emotionally Aware Responses: Train the chatbot to respond appropriately to the user's emotional state. For example, if a user seems frustrated, the chatbot can adopt a more soothing tone or offer assistance. Personalized Data Utilization: Incorporate personalized data to better understand individual user preferences and emotional triggers. This personalized approach enables the chatbot to tailor its interactions more effectively. [0571] Personalized Content Generation Algorithm. Hardware Interaction: Interacts with digital signage, personal computing devices, and mobile devices by dynamically generating and displaying content. Hardware Performance Improvement: By tailoring content to individual preferences, the algorithm increases the utilization efficiency of these displays and devices, leading to better user engagement and satisfaction. Measurement: This is measured by user interaction metrics such as increased time spent on digital platforms (percentage increase) and higher conversion rates (percentage increase in sales or desired actions). See also [0200]-[0202], [0204], [0207], [0342], [0344], [0349], [0527], [0536]))
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include wherein the processor is configured to adjust the video analysis method by adjusting a display method of analysis results based on the estimated emotion as taught by Angeles with the instruction document creator of Kelly in view of Chatbot in order to foster emotional well-being and engagement, using sophisticated model training that incorporates psychological and behavioral science principles ([0555]).
Kelly in view of Chatbot in view of Angeles teaches the limitations of claim 1. As per claim 15:
Kelly in view of Chatbot does not expressly teach wherein the processor is further configured to adjust a method of structuring the manual based on the estimated emotion.
Angeles teaches:
wherein the processor is further configured to adjust a method of structuring the manual based on the estimated emotion. ([0555] Emotional Recognition and Response Algorithm: Utilizes deep neural networks to analyze real-time data from facial recognition, voice intonation, and physiological sensors integrated into VR/AR interfaces. The algorithm processes these data points to ascertain emotional states and adjust digital content accordingly, such as modifying difficulty levels, narrative elements, or interactive features. This algorithm goes beyond simple emotion detection by actively altering digital and virtual environments in response to detected emotional cues. It's crafted to foster emotional well-being and engagement, using sophisticated model training that incorporates psychological and behavioral science principles. Interfaces with devices capable of detecting user emotions (e.g., facial expression recognition in VR/AR environments). It adapts chatbot responses based on emotional cues, improving user interaction by providing empathetic and contextually appropriate responses to emotional states related to health monitoring. [0200] Personalized curriculum development: The chatbot assistant can provide data-driven insights, by analyzing the owner's data, including past learning experiences, interests, and performance, to identify strengths and areas for improvement. Based on this analysis, it develops a personalized learning curriculum tailored to the owner's specific goals and learning style. The chatbot assistant can also provide skill gap analysis, by continuously monitoring the owner's interactions and progress, to identify skill gaps and adjusts the curriculum in real-time to address these gaps, ensuring a focused and efficient learning path [0527] 1. Adaptive Learning for Enhanced Memory: Software employs advanced machine learning algorithms that adapt content delivery based on the user's interaction history, enhancing memory retention and recall. It adjusts the difficulty and format of memory games and tasks on smart devices and wearables, based on real-time assessments of user performance. Unlike conventional static learning applications, this approach uses a dynamic adjustment mechanism that predicts and reacts to individual memory capacities, significantly improving personalized learning experiences. See also [0200]-[0202], [0204], [0207], [0342], [0344], [0349], [0527], [0536]))
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include wherein the processor is further configured to adjust a method of structuring the manual based on the estimated emotion as taught by Angeles with the instruction document creator of Kelly in view of Chatbot in order to foster emotional well-being and engagement, using sophisticated model training that incorporates psychological and behavioral science principles ([0555]).
Kelly in view of Chatbot in view of Angeles teaches the limitations of claim 1. As per claim 16:
Kelly in view of Chatbot does not expressly teach wherein the processor is further configured to adjust a response method of the chatbot based on the estimated emotion.
Angeles teaches:
wherein the processor is further configured to adjust a response method of the chatbot based on the estimated emotion. ([0555] Emotional Recognition and Response Algorithm: Utilizes deep neural networks to analyze real-time data from facial recognition, voice intonation, and physiological sensors integrated into VR/AR interfaces. The algorithm processes these data points to ascertain emotional states and adjust digital content accordingly, such as modifying difficulty levels, narrative elements, or interactive features. This algorithm goes beyond simple emotion detection by actively altering digital and virtual environments in response to detected emotional cues. It's crafted to foster emotional well-being and engagement, using sophisticated model training that incorporates psychological and behavioral science principles. Interfaces with devices capable of detecting user emotions (e.g., facial expression recognition in VR/AR environments). It adapts chatbot responses based on emotional cues, improving user interaction by providing empathetic and contextually appropriate responses to emotional states related to health monitoring. [0349] 14-1. Emotional Intelligence Training. Sentiment Analysis: Utilize advanced NLP techniques to analyze user input for emotional cues and sentiment. This allows the chatbot to gauge the user's mood and emotional state. Emotionally Aware Responses: Train the chatbot to respond appropriately to the user's emotional state. For example, if a user seems frustrated, the chatbot can adopt a more soothing tone or offer assistance. Personalized Data Utilization: Incorporate personalized data to better understand individual user preferences and emotional triggers. This personalized approach enables the chatbot to tailor its interactions more effectively. [0529] 3. Emotionally Intelligent Interfaces: Software integrates biometric sensors and affective computing models to gauge emotional states from physiological signals. The output modifies interaction strategies of digital assistants across devices (e.g., smartphones, VR/AR systems) to respond appropriately to the user's emotional cues, such as lowering voice tone or changing content. See also [0200]-[0202], [0204], [0207], [0342], [0344], [0349], [0527], [0536]))
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include wherein the processor is further configured to adjust a response method of the chatbot based on the estimated emotion as taught by Angeles with the instruction document creator of Kelly in view of Chatbot in order to foster emotional well-being and engagement, using sophisticated model training that incorporates psychological and behavioral science principles ([0555]).
Claim(s) 3 is/are rejected under 35 U.S.C. 103 as being unpatentable over McKenzie-Kelly et al (US 2023/0350947) hereafter “Kelly” in view of KR102073928 hereafter “chatbot” in view of Angeles (US 2026/0037863) in view of Sipe, III et al (US 2025/0182356) hereafter “Sipe”
Kelly in view of Chatbot Angeles teaches the limitations of claim 1. As per claim 3:
Kelly in view of chatbot in view of Angeles does not expressly teach wherein the processor is further is configured to generate descriptive text including an interactive element based on the content of the video by means of the generative AI.
Sipe teaches:
wherein the processor is further is configured to generate descriptive text including an interactive element based on the content of the video by means of the generative AI. ([0022] The generative AI presentation engine is responsible for providing generative AI image data or text data. In particular, the generative AI presentation engine provides composite image data that includes an image that is created by combining and layering multiple individual images and text elements or visual components. Composite image data can be created by combining photographs, graphics, illustrations, text or other visual elements to form a single, cohesive image. The composite image data can include non-generative AI data elements and generative AI data elements. The non-generative AI data elements can include pre-existing images or text that are not generated via the generative AI model, while generative AI elements include images or text that are originally created using the generative AI model. The composite image data can include generative AI image elements and generative AI item listing interface elements. The generative AI images elements can include different types of images (e.g., AI images and non-AI images) that are part of the composite image and the generative AI item listing interface elements can includes item listing features (e.g., text, price, title, descriptions, filters, buttons, color pallets, and actions) that are part of the generative composite image.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include wherein the processor is further is configured to generate descriptive text including an interactive element based on the content of the video by means of the generative AI as taught by Sipe with the instruction document creator of Kelly in view of Chatbot n view of Angeles in order to improve the presentation of generative AI content.
Claim(s) 6 is/are rejected under 35 U.S.C. 103 as being unpatentable over McKenzie-Kelly et al (US 2023/0350947) hereafter “Kelly” in view of KR102073928 hereafter “chatbot” in view of Angeles (US 2026/0037863) in view of Danilchenko (US 2026/0073244)
Kelly in view of Chatbot in view of Angeles teaches the limitations of claim 1. As per claim 6:
Kelly in view of Chatbot does not expressly teach wherein the processor is further configured to perform version control by means of the generative AI and always provide the latest manual.
Danilchenko teaches:
wherein the processor is further configured to perform version control by means of the generative AI and always provide the latest manual. ([0067] For example, the federated KB computing platform 112 may determine that the consumer 102 is requesting document A 364, and determine the snapshot reference 352 indicating a storage location of a certified version of the document A 364 (e.g., the third version that has been certified). In other examples, such as if the consumer 102 is requesting document D 372, the latest version of the document D 372 may be the same as the certified version (e.g., both of them may be version five). The federated KB computing platform 112 may determine the certified document snapshot reference associated with the requested document (e.g., the certified document snapshot reference A 352 associated with the requested document 364), and retrieve the snapshot reference from the document database 312. In some examples, as mentioned above, the federated KB computing platform 112 may use the search system 306 and/or the generative AI search system 308 to determine the certified document snapshot reference that is associated with the search request from the consumer device 104.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include wherein the processor is further configured to perform version control by means of the generative AI and always provide the latest manual as taught by Danilchenko with the instruction document creator of Kelly in view of Chatbot in view of Angeles in order to provide a certified document associated with a search request ([0067]).
Claim(s) 10 is/are rejected under 35 U.S.C. 103 as being unpatentable over McKenzie-Kelly et al (US 2023/0350947) hereafter “Kelly” in view of KR102073928 hereafter “chatbot” in view of in view of Angeles (US 2026/0037863) in view of Griffen et al (US 2025/0132941)
Kelly in view of Chatbot in view of Angeles teaches the limitations of claim 8. As per claim 10:
Kelly in view of Chatbot in view of Angeles does not expressly teach wherein the specific action or gesture comprises at least one of a hand movement, a body movement, or a facial expression of a person shown in the video.
Griffen teaches:
wherein the specific action or gesture comprises at least one of a hand movement, a body movement, or a facial expression of a person shown in the video. ([0049] Facial expression 576 may be captured, e.g., live-streamed and/or recorded, in a video 552b or video content by collaboration application 558 during collaboration session 560. Video 552b may be provided by collaboration application 558 to AI system 550. RMM 550b may process video 552b to identify the presence of facial expression 576 in video 552b, and to obtain at least one insight 554 into facial expression 576. RMM 550b may provide insight 554 to LLM 550a in a format suitable for processing by LLM 550a. LLM 550a may process insight 554 to generate an insight summary that relates to insight 554. AI system 550 may then generate an output 556 that includes a summary or other information, e.g., an interpretation or conclusion, relating to insight 554.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include wherein the specific action or gesture comprises at least one of a hand movement, a body movement, or a facial expression of a person shown in the video as taught by Griffen with the instruction document creator of Kelly in view of Chatbot in view of Angeles in order to effectively extract data from audio and visual cues, including, but not limited to including, a tone of voice, an intonation of a voice, a physical action such as a gesture, and a facial expression ([0020]).
Claim(s) 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over McKenzie-Kelly et al (US 2023/0350947) hereafter “Kelly” in view of KR102073928 hereafter “chatbot” in view of in view of Angeles (US 2026/0037863) in view of Nakahashi (US 2017/0199918)
Kelly in view of Chatbot in view of Angeles teaches the limitations of claim 1. As per claim 14:
Kelly in view of Chatbot in view of Angeles does not expressly teach wherein the processor is further configured to adjust a length of the images and descriptive text generated based on the estimated emotion.
Nakahashi teaches:
wherein the processor is further configured to adjust a length of the images and descriptive text generated based on the estimated emotion. ([0007] For example, if the text of presented information is fully displayed, it takes a long time to recognize the whole information. Thus, displaying only a list of titles or outlines is more convenient to a user in a hurry. If there is adequate time, displaying not only text, but also images allows a user to more specifically recognize the contents, achieving high convenience.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include wherein the processor is further configured to adjust a length of the images and descriptive text generated based on the estimated emotion as taught by Nakahashi with the instruction document creator of Kelly in view of Chatbot in view of Angeles in order to determine the granularity of information at the presentation of the information in accordance with a user status ([0008]).
Claim(s) 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over McKenzie-Kelly et al (US 2023/0350947) hereafter “Kelly” in view of KR102073928 hereafter “chatbot” in view of in view of Angeles (US 2026/0037863) in view of Tsuji et al (US 2023/0206691)
Kelly in view of Chatbot in view of Angeles teaches the limitations of claim 1. As per claim 17:
Kelly in view of Chatbot do not expressly teach wherein the processor is configured to estimate the emotion of the user by inputting the user data to the pre-learned neural network, the pre-learned neural network.
Angeles further teaches:
wherein the processor is configured to estimate the emotion of the user by inputting the user data to the pre-learned neural network, the pre-learned neural network ([0555] Emotional Recognition and Response Algorithm: Utilizes deep neural networks to analyze real-time data from facial recognition, voice intonation, and physiological sensors integrated into VR/AR interfaces. The algorithm processes these data points to ascertain emotional states and adjust digital content accordingly, such as modifying difficulty levels, narrative elements, or interactive features. This algorithm goes beyond simple emotion detection by actively altering digital and virtual environments in response to detected emotional cues. It's crafted to foster emotional well-being and engagement, using sophisticated model training that incorporates psychological and behavioral science principles. Interfaces with devices capable of detecting user emotions (e.g., facial expression recognition in VR/AR environments). It adapts chatbot responses based on emotional cues, improving user interaction by providing empathetic and contextually appropriate responses to emotional states related to health monitoring.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include wherein the processor is configured to estimate the emotion of the user by inputting the user data to the pre-learned neural network, the pre-learned neural network as taught by Angeles with the instruction document creator of Kelly in view of Chatbot in order to foster emotional well-being and engagement, using sophisticated model training that incorporates psychological and behavioral science principles ([0555]).
Tsuji teaches:
outputting, from the emotion map in which a plurality of emotions are arranged, an emotion value for each of the plurality of emotions, and determining the emotion of the user based on the emotion values. ([0012] Each of the first estimator and the second estimator may calculate a score indicating a likelihood of the user holding an emotion of a plurality of emotions. The determiner may determine the emotion of the user based on a score for each of the plurality of emotions calculated by the first estimator and a score for each of the plurality of emotions calculated by the second estimator. The scores are normalized by the same criteria and thus are in the same range for the emotion estimation using the face image and the emotion estimation using the temperature. For example, the maximum score is commonly a predetermined value of, for example, 100, in the emotion estimation using the face image and the emotion estimation using the temperature. The multiple emotions include calmness, anger, sadness, and joy. [0013] For example, the determiner may determine, as the emotion of the user, an emotion with a greatest score calculated by the first estimator matching an emotion with a greatest score calculated by the second estimator.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include outputting, from the emotion map in which a plurality of emotions are arranged, an emotion value for each of the plurality of emotions, and determining the emotion of the user based on the emotion values as taught by Tsuji with the with the instruction document creator of Kelly in view of Chatbot in view of Angeles in order to determine emotions with high accuracy using a simple structure with high usability ([0008]).
Response to Arguments
Applicant’s arguments regarding previous objections are found persuasive. As a result, such objections have been withdrawn.
The examiner has considered but does not find persuasive applicant’s arguments regarding rejections under 25 USC 101. With regard to using a neural network to estimate emotion the examiner respectfully disagrees. First, the neural network is an additional element. A human would be able to analyze facial expression or voice data to estimate a human emotion. The claims recite no meaningful details regarding how the neural network performs such tasks. The claims provide little more than “do it” with a neural network. Thus, the neural network is recited at a high level of generality and does not go beyond the “apply it” level of implementation as it merely provides result based claiming.
With regard to step 2A prong 2 the examiner respectfully disagrees. As presently claimed, adjusting content based on user emotion is merely part of the abstract idea and any alleged improvement is to the abstract idea itself and not technology or a technical field. Similarly, the same issues exist with regard to step 2B. The claims recite no meaningful details regarding how the neural network performs such tasks. The claims provide little more than “do it” with a neural network and thus they do not provide an inventive concept in accordance with step 2B. As a result, such rejections have been maintained.
The examiner has considered and finds persuasive applicant’s arguments regarding rejections under 35 USC 112. As a result, such rejections have been withdrawn.
Applicant’s arguments regarding rejections under 35 USC 103 are moot in light of new grounds of rejection which have been necessitated by amendment.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CHRISTOPHER STROUD whose telephone number is (571)272-7930. The examiner can normally be reached Mon. - Fri. 9AM-5PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Waseem Ashraff can be reached at (571) 270-3948. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
CHRISTOPHER STROUD
Primary Examiner
Art Unit 3621
/CHRISTOPHER STROUD/ Primary Examiner, Art Unit 3621