Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
CLAIM INTERPRETATION
The following is a quotation of 35 U.S.C. 112(f): (FP 7.30.03)
(f) ELEMENT IN CLAIM FOR A COMBINATION.—An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as "configured to" or "so that"; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. (FP 7.30.05)
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) (Claims 1-15) is/are:
Personalized feature amount inference unit configured to infer a second prompt token sequence in claim 1, 3, 4, 6, 14, 15;
User information acquisition unit configured to acquire attribute information of the first user in claim 2, 8;
Image generation unit configured to be capable of generating an image in claim 2, 11, 12;
Interface control unit that causes a generated image generated by the image generation unit in claim 9, 10;
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. (FP 7.30.06)
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-4, 6-7, 9-10, 13-15 are rejected under 35 U.S.C. 103 as being unpatentable over Dimitriadis et al (US20230297777) in view of Fink et al (US20240161743) further in view of Wang et al (US20230245651).
Regarding Claim 1. Dimitriadis teaches An information processing apparatus comprising
a personalized feature amount inference unit configured to infer a second prompt token sequence related to a second feature amount on the basis of data related to background information of a person having an attribute close to an attribute of a first user (Dimitriadis, abstract, the invention describes a personalized natural language processing system tokenizes a plurality of sets of raw text data to generate a plurality of sets of tokenized text data for the plurality of users, respectively. The tokenized text data includes a sequence of tokens corresponding to the raw text data, the tokens at least identifying distinct words or portions of words in the raw text. The system appends predetermined user-specific tokens to the sets of tokenized text data from the users, respectively. Each predetermined user-specific token corresponds to one of the users. The system processes the sets of tokenized text data using the NLP model in accordance with the appended predetermined user-specific tokens to predict a personalized classification for the sets of tokenized text data from each of the users, and outputs the personalized classifications of the tokenized text data for each of the users.
[0021] Turning now to the operation of personalized NLP system 10 during training time as illustrated in FIG. 1, a training data set 54 is stored in non-volatile memory. The training data set 54 is generated by gathering data from multiple users using multiple client computing devices 42. The data may be gathered from public sources such as social media posts, online reviews, and blogs, or private sources such as documents, emails, and instant messages upon receiving appropriate opt-in permission from each user to participate in the training. The training data set 54 includes multiple tuples of training data, each tuple including user identifier data 54A identifying a particular user, raw text data 54B from that particular user, and ground truth classification data 54C. The user identifier data 54A can be configured so as to not include any personally identifying information. The raw text data 54B may be utterances (i.e., recognized speech from an audio source), social media posts, electronic messages such as emails, chat messages, or text messages, or other type of text for which classification is desired, for example. The ground truth classification 54C may be, for example, a sentiment classification, or may be, for example, a prediction of a next word in a sequence, or a prediction of an output sequence, as described in more detail below. Three tuples are depicted for each of three different users: USER A, USER B, and USER N. Although three tuples are depicted, it will be understood that in practice, tens, hundreds, or thousands of such tuples may be included for each user, and training data for tens, hundreds, or thousands of users may be used to train one model.
[0029] The personalized NLP model 38 may be trained on a centralized personalized NLP computing device 12. Alternatively, a federated learning approach may be used, in which the personalized NLP model 38 is initially trained on a client computing device 42, which then shares the gradients or model updates with the centralized personalized NLP computing device 12, which then aggregates the gradients from different users and sends back an updated model back to the client computing device 42 for further training.
[0030] Once the personalized NLP model 38 is trained, the trained personalized model may be made available for inference computations, as shown in FIG. 5. In the illustrated example, two users, User A and User B, are shown in interacting with the personalized NLP model from their respective client computing devices 42. The client computing devices 42 each execute an application client 26A configured to send inference-time user input 48 in the form of raw text data 28 to the computing device 12, and subsequently receive personalized classifications 40 from the computing device 12 as output. The application client 26A may include a graphical user interface 46, and configured to display graphical output 50 based on the personalized classifications 40 outputted from the personalized NLP model 38 and received at the client program 26A.),
Dimitriadis fails to explicitly teach, however, Fink teaches second prompt (Fink, abstract, the invention teaches method of generating expanded responses that guide continuance of a human-to computer dialog that is facilitated by a client device and that is between at least one user and an automated assistant. The expanded responses are generated by the automated assistant in response to user interface input provided by the user via the client device, and are caused to be rendered to the user via the client device, as a response, by the automated assistant, to the user interface input of the user. An expanded response is generated based on at least one entity of interest determined based on the user interface input, and is generated to incorporate content related to one or more additional entities that are related to the entity of interest, but that are not explicitly referenced by the user interface input.
[0090] FIGS. 4, 5, and 6 each depict example dialogs between a user and an automated assistant, in which composite responses can be provided.
[0091-0092] Fig 4…In response, the user provides second natural language input 490B of "How do you configure them", that is guided by the content of the responsive natural language output 492A. In response to the first natural language input 490A, automated assistant provides another responsive natural language output 492B that is a composite response generated according to implementations disclosed herein. For example, the first sentence of the natural language output 492B can be based on a textual snippet from a given content agent, "It saves at least 5 minutes" can be based on an additional textual snippet from another content agent, and "is called Hypothetical App." can be based on a further textual snippet from a further content agent.).
Dimitriadis and Fink are analogous art because they both teach personalized natural language processing system generating response based on past user responses. Fink further teaches multiple user responses with the language system Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention, to modify the personalized natural language processing system (taught in Dimitriadis), to further consider the multiple interactive responses between user and system (taught in Fink), so as to provide a continuate human to machine communication system such as automated assistant (Fink, [0002]).
The combination of Dimitriadis and Fink fails to explicitly teach, however, Wang teaches among data included in a background information database accumulating background information in which attribute information related to
an attribute of a person,
a feature amount input by the person, and
a behavior history related to an action of the person with respect to a product generated on the basis of the feature amount are associated with each other, and a first feature amount input by the first user (Wang, abstract, the invention describes method for enabling contextually relevant conversational interaction. Environment data is received by an AI System which detects a plurality of physical objects in a physical environment and forms a contextual understanding of the plurality of physical objects and the physical environment and identifies a user relevant to the contextual understanding. A most relevant contextual information to the user is predicted by the AI system and transformed into a textual form. A set of intents and objectives is predicted by the AI system for user-centered interaction. The AI system and the user interact iteratively through the user centered interaction to determine an understanding of a most relevant intent and a most relevant objective which is validated by the AI system with the user until the user agrees. The validated most relevant intent and the most relevant objective is utilized to facilitate the user-centered and contextually relevant conversational interaction.
[0034] The present invention aims to deliver an improved user experience by considering individual preferences, conversational history, and communication styles. The AI system adapts its responses according to a user's past interaction or preferred tone, resulting in more effective and satisfying conversations.
[0035] Additionally, the invention incorporates contextual information, such as the user's current location, time, or activity, to provide more relevant and timely responses.
[0049] Referring to FIG. 1, the OKB 104 (Object Knowledge Base) is a structured system, but not limited to a database, a set of databases, a repository, a set of repositories, and the like, that stores different types of data and various types of information that the AI system can access and use to provide contextually relevant and personalized responses.
[0056] To add human-AI interaction data to the OKB, the AI system initially processes the data to extract significant information such as user intents, behaviors, and preferences. This processing often involves using NLP techniques to parse the conversation and identify the essential elements. After the data has been processed, it is added to the OKB in a structured format that is easily accessible and analyzed by the AI system. This format can include relevant contextual information such as the time and location of the interaction, the user’s input, and the AI system’s response.).
Dimitriadis, Fink and Wang are analogous art because they all teach personalized natural language processing system generating response based on past user responses. Wang further teaches the AI system will predict response based on contextual data such as user’s past behavior, user’s preferences, user’s input and user’s surrounding information. Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention, to modify the personalized natural language processing system (taught in Dimitriadis and Fink), to further consider the contextual data about the user (taught in Wang), so as to provide a user-centered, contextually relevant and personalized interaction system (Wang, [0032]).
Regarding Claim 2. The combination of Dimitriadis, Fink and Wang further teaches The information processing apparatus according to claim 1, further comprising:
a user information acquisition unit configured to acquire attribute information of the first user (Wang, [0046] The AI system 101 can also integrate with various sensors 113, smart devices 114, virtual reality (VR) / augmented reality (AR) headsets 115, and Internet of Things (IoT) 116. This allows the AI system 101 to identify physical objects 112 in the environment and gather data such as object location, relative positions, and temperature readings. To obtain contextual information about physical objects, the system can use a range of sensors, including temperature sensor 117, light sensor 118, noise sensor 119, motion sensor 120, and presence sensor 121. This information is then preprocessed and categorized before being displayed to users.
[0047] Moreover, smart devices 114 come equipped with various features such as microphones 122, speakers 123, touchscreens 124, cameras 125, Wi-Fi®2 126, Bluetooth®3 127, and near field communications (NFC®4) 128. These features enable the AI system to gather contextual information about the user's environment and interactions, allowing it to provide more personalized and relevant responses. For instance, the microphone can capture the user's voice input, while the camera can capture visual data such as facial expressions or object recognition.); and
an image generation unit configured to be capable of generating an image using the second feature amount as an input prompt (Wang, [0144] Generative AI is a subfield of artificial intelligence that focuses on the creation of new content or data, such as text, images, or audio, based on input data and context. This is achieved through advanced ML techniques, such as deep learning and neural networks. Generative AI models, such as Generative Adversarial Networks (GANs) or Variational Autoencoders (VAEs), can generate realistic and high-quality outputs by learning complex patterns and structures from large datasets during the training process.
[0145] Integrating the Generative AI module 206 within the NLG module can significantly enhance the capabilities of the AI system in generating contextually relevant, natural-sounding text responses during conversational interactions. The Generative AI can leverage its ability to learn complex patterns and structures from large language datasets to produce human-like responses that are not only coherent but also tailored to the specific context of the interaction.).
The reasoning for combination of Dimitriadis, Fink and Wang is the same as described in Claim 1.
Regarding Claim 3. The combination of Dimitriadis, Fink and Wang further teaches The information processing apparatus according to claim 1, wherein the personalized feature amount inference unit infers the personalized prompt token sequence in which personalized feature amount is tokenized by selecting a token sequence on the basis of a likelihood sequence of the second prompt token sequence output from an attribute information token sequence in which attribute information of the first user is tokenized and a first prompt token sequence in which the first feature amount is tokenized (Wang, [0437) A response can be generated using an encoder-decoder model and a language model AI. The encoder-decoder model is a type of deep neural network that consists of two main components: an encoder and a decoder. The encoder takes in the input sequence, such as a question or a statement, and converts it into a hidden state that captures the semantic meaning of the input. The decoder then uses this hidden state to generate a response sequence, such as an answer or a reply, by predicting one token at a time. The encoder-decoder model is commonly used for sequence-to-sequence tasks, such as machine translation or dialogue generation. The encoder-decoder model consists of two main components: an encoder network that processes the input sequence and generates a fixed-length context vector, and a decoder network that generates an output sequence based on the context vector. In the context of dialogue generation, the encoder-decoder model can be used to generate a preliminary response based on the input sequence.).
The reasoning for combination of Dimitriadis, Fink and Wang is the same as described in Claim 1.
Regarding Claim 4. The combination of Dimitriadis, Fink and Wang further teaches The information processing apparatus according to claim 1, wherein
the personalized feature amount inference unit uses a personalized feature amount model when inferring the second prompt token sequence from the first feature amount (Fink, abstract, the invention teaches method of generating expanded responses that guide continuance of a human-to computer dialog that is facilitated by a client device and that is between at least one user and an automated assistant. The expanded responses are generated by the automated assistant in response to user interface input provided by the user via the client device, and are caused to be rendered to the user via the client device, as a response, by the automated assistant, to the user interface input of the user. An expanded response is generated based on at least one entity of interest determined based on the user interface input, and is generated to incorporate content related to one or more additional entities that are related to the entity of interest, but that are not explicitly referenced by the user interface input.
[0090] FIGS. 4, 5, and 6 each depict example dialogs between a user and an automated assistant, in which composite responses can be provided.
[0091-0092] Fig 4…In response, the user provides second natural language input 490B of "How do you configure them", that is guided by the content of the responsive natural language output 492A. In response to the first natural language input 490A, automated assistant provides another responsive natural language output 492B that is a composite response generated according to implementations disclosed herein. For example, the first sentence of the natural language output 492B can be based on a textual snippet from a given content agent, "It saves at least 5 minutes" can be based on an additional textual snippet from another content agent, and "is called Hypothetical App." can be based on a further textual snippet from a further content agent.), and
the personalized feature amount model is capable of being trained on the basis of background information in which attribute information of the first user, the first feature amount, and a behavior history of the first user with respect to a product generated on the basis of the first feature amount or the second feature amount are associated with each other (Wang, [0475] FIG. 22 illustrates how the Al system can generate personalized recommendations or suggestions based on user preferences and behavior 2200.
[0476] The process of providing personalized recommendations or suggestions based on user preferences and behavior starts with the Al system receiving input from the user, such as a search query or selection 2201. This input serves as a starting point for the Al system to generate relevant options for the user.
[0477] Once the Al system has received user input, it retrieves user data such as preferences and behavior history 2202. This data is then filtered and analyzed by the Al system to identify relevant options that can be recommended to the user 2203. Based on this analysis, the Al system generates personalized recommendations or suggestions that are tailored to the user's preferences and behavior 2204. These recommendations can take the form of product suggestions, content recommendations, or any other relevant options based on the user's input and data.
[0478] The Al system then presents these personalized recommendations or suggestions to the user 2205, who can provide feedback on the options presented 2206. This feedback is valuable in improving the recommendations or suggestions for future interactions with the user. The Al system uses feedback from the user to update the recommendations or suggestions 2207. This iterative process of feedback and data analysis allows the Al system to continually learn and adapt to the user's preferences and behavior. The Al system uses the updated information to update the OKB to improve future recommendations or suggestions 2208).
The reasoning for combination of Dimitriadis, Fink and Wang is the same as described in Claim 1.
Regarding Claim 6. The combination of Dimitriadis, Fink and Wang further teaches The information processing apparatus according to claim 2, wherein
the user information acquisition unit acquires an identifier from a terminal operated by the first user, and inquires of an external server holding information related to the identifier to acquire the attribute information of the first user (Dimitriadis, [0021] Turning now to the operation of personalized NLP system 10 during training time as illustrated in FIG. 1, a training data set 54 is stored in non-volatile memory. The training data set 54 is generated by gathering data from multiple users using multiple client computing devices 42. The data may be gathered from public sources such as social media posts, online reviews, and blogs, or private sources such as documents, emails, and instant messages upon receiving appropriate opt-in permission from each user to participate in the training. The training data set 54 includes multiple tuples of training data, each tuple including user identifier data 54A identifying a particular user, raw text data 54B from that particular user, and ground truth classification data 54C.
Wang, [0237] Access control systems typically involve three main components: (1) Identification: Users should identify themselves to the AI system, typically through a usename or unique identifier. (2) Authentication: Users should prove their identity to the AI system, usually by providing a password, biometric data (e.g., fingerprint, facial recognition), or a security token. (3) Authorization: Once authenticated, the AI system checks the user's access permissions against the requested resource and determines if the user is allowed to access it based on the access control model in place.
[0499] Furthermore, the Al system can provide a personalized Ul for interacting with the conversational Al agent. In an embodiment, the Ul includes examples of natural language queries and responses. The Al system handles multiple user inputs and generates personalized responses based on user preferences and history by analyzing the user's input, identifying the user's intent, accessing the user's preferences and history from the database, generating a response, and sending the response to the user. The Al system can handle a variety of user inputs, such as voice commands, text input, and gesture recognition, and can use natural language processing to understand and interpret the user's intent. Based on the user's preferences and history, the Al system can generate personalized responses that are tailored to the user's specific needs and interests. The Al system can also adapt and learn from user feedback, improving the accuracy and relevance of its responses over time.).
The reasoning for combination of Dimitriadis, Fink and Wang is the same as described in Claim 1.
Regarding Claim 7. The combination of Dimitriadis, Fink and Wang further teaches The information processing apparatus according to claim 2, wherein the user information acquisition unit acquires information related to the first user from an account of a social network service (SNS) (Dimitriadis, [0021] Turning now to the operation of personalized NLP system 10 during training time as illustrated in FIG. 1, a training data set 54 is stored in non-volatile memory. The training data set 54 is generated by gathering data from multiple users using multiple client computing devices 42. The data may be gathered from public sources such as social media posts, online reviews, and blogs, or private sources such as documents, emails, and instant messages upon receiving appropriate opt-in permission from each user to participate in the training. The training data set 54 includes multiple tuples of training data, each tuple including user identifier data 54A identifying a particular user, raw text data 54B from that particular user, and ground truth classification data 54C.).
Regarding Claim 9. The combination of Dimitriadis, Fink and Wang further teaches The information processing apparatus according to claim 2, further comprising an interface control unit that causes a generated image generated by the image generation unit to be displayed on a terminal operated by the first user together with an evaluation input section for inputting evaluation on the generated image by the first user (Wang, [0144] Generative AI is a subfield of artificial intelligence that focuses on the creation of new content or data, such as text, images, or audio, based on input data and context. This is achieved through advanced ML techniques, such as deep learning and neural networks. Generative AI models, such as Generative Adversarial Networks (GANs) or Variational Autoencoders (VAEs), can generate realistic and high-quality outputs by learning complex patterns and structures from large datasets during the training process.
[0158] Referring to FIG. 2, the Multimodal Generation module 212 serves as an important part of the AI system 200, generating contextually relevant multimedia content based on input data and context. The Multimodal Generation module 212 is responsible for generating a variety of multimedia content, including images, drawings, audio, and video. This module operates in collaboration with other components of the AI system, such as Natural Language Understanding (NLU), context-aware modules, and other data processing components, to produce contextually relevant multimedia outputs.
[0478] The Al system then presents these personalized recommendations or suggestions to the user 2205, who can provide feedback on the options presented 2206. This feedback is valuable in improving the recommendations or suggestions for future interactions with the user. The Al system uses feedback from the user to update the recommendations or suggestions 2207. This iterative process of feedback and data analysis allows the Al system to continually learn and adapt to the user's preferences and behavior. The Al system uses the updated information to update the OKB to improve future recommendations or suggestions 2208).
The reasoning for combination of Dimitriadis, Fink and Wang is the same as described in Claim 1.
Regarding Claim 10. The combination of Dimitriadis, Fink and Wang further teaches The information processing apparatus according to claim 2, further comprising an interface control unit configured to display the second feature amount on a terminal operated by the first user (Fink, [0090] FIGS. 4, 5, and 6 each depict example dialogs between a user and an automated assistant, in which composite responses can be provided.
[0091-0092] Fig 4…In response, the user provides second natural language input 490B of "How do you configure them", that is guided by the content of the responsive natural language output 492A. In response to the first natural language input 490A, automated assistant provides another responsive natural language output 492B that is a composite response generated according to implementations disclosed herein. For example, the first sentence of the natural language output 492B can be based on a textual snippet from a given content agent, "It saves at least 5 minutes" can be based on an additional textual snippet from another content agent, and "is called Hypothetical App." can be based on a further textual snippet from a further content agent.).
The reasoning for combination of Dimitriadis, Fink and Wang is the same as described in Claim 1.
Regarding Claim 13. The combination of Dimitriadis, Fink and Wang further teaches The information processing apparatus according to claim 4, wherein the personalized feature amount model is capable of being trained by dividing the person included in the background information database into a plurality of user groups and aggregating feature amounts co-occurring in each of the user groups (Wang, [0354] Some examples of default assumptions or general patterns that the AI system can rely on to provide contextually relevant interaction if there are no user preferences or interaction history available in the OKB: (1) Demographic assumptions: The AI system can make assumptions based on general demographic information, such as age, gender, or cultural background, to predict user preferences and needs. (2) Environmental conditions: The AI system can use current environmental conditions to make informed decisions. For instance, if the temperature in a healthcare setting is unusually high, the AI system could assume that users would prefer a cooler environment and adjust the thermostat accordingly. (3) Common needs in healthcare settings: The AI system can assume that users in healthcare settings share common needs, such as privacy, comfort, and access to information about their health. Based on these assumptions, the AI system can prioritize actions that address these common needs. (4) Time-based assumptions: The AI system can make assumptions based on the time of day or day of the week. For example, it might assume that users would be more likely to require assistance or information during regular business hours. (5) Role-based assumptions: In a healthcare setting, the AI system can assume that different user roles have different needs and preferences. For example, it might assume that a doctor needs access to patient records and diagnostic tools, while a patient may require information about their treatment plan and recovery. (6) General patterns from similar users: The AI system can use data from similar users to make assumptions about an individual's preferences and needs. For example, if a majority of patients in a particular age group have shown a preference for a certain type of interaction or information, the AI system can assume that a new user from the same age group would have similar preferences.).
The reasoning for combination of Dimitriadis, Fink and Wang is the same as described in Claim 1.
Regarding Claim 14. The combination of Dimitriadis, Fink and Wang further teaches An information processing apparatus comprising a personalized feature amount inference unit configured to infer a second feature amount (Dimitriadis, abstract, the invention describes a personalized natural language processing system tokenizes a plurality of sets of raw text data to generate a plurality of sets of tokenized text data for the plurality of users, respectively. The tokenized text data includes a sequence of tokens corresponding to the raw text data, the tokens at least identifying distinct words or portions of words in the raw text. The system appends predetermined user-specific tokens to the sets of tokenized text data from the users, respectively. Each predetermined user-specific token corresponds to one of the users. The system processes the sets of tokenized text data using the NLP model in accordance with the appended predetermined user-specific tokens to predict a personalized classification for the sets of tokenized text data from each of the users, and outputs the personalized classifications of the tokenized text data for each of the users.
[0021] Turning now to the operation of personalized NLP system 10 during training time as illustrated in FIG. 1, a training data set 54 is stored in non-volatile memory. The training data set 54 is generated by gathering data from multiple users using multiple client computing devices 42. The data may be gathered from public sources such as social media posts, online reviews, and blogs, or private sources such as documents, emails, and instant messages upon receiving appropriate opt-in permission from each user to participate in the training. The training data set 54 includes multiple tuples of training data, each tuple including user identifier data 54A identifying a particular user, raw text data 54B from that particular user, and ground truth classification data 54C. The user identifier data 54A can be configured so as to not include any personally identifying information. The raw text data 54B may be utterances (i.e., recognized speech from an audio source), social media posts, electronic messages such as emails, chat messages, or text messages, or other type of text for which classification is desired, for example. The ground truth classification 54C may be, for example, a sentiment classification, or may be, for example, a prediction of a next word in a sequence, or a prediction of an output sequence, as described in more detail below. Three tuples are depicted for each of three different users: USER A, USER B, and USER N. Although three tuples are depicted, it will be understood that in practice, tens, hundreds, or thousands of such tuples may be included for each user, and training data for tens, hundreds, or thousands of users may be used to train one model.
[0029] The personalized NLP model 38 may be trained on a centralized personalized NLP computing device 12. Alternatively, a federated learning approach may be used, in which the personalized NLP model 38 is initially trained on a client computing device 42, which then shares the gradients or model updates with the centralized personalized NLP computing device 12, which then aggregates the gradients from different users and sends back an updated model back to the client computing device 42 for further training.
[0030] Once the personalized NLP model 38 is trained, the trained personalized model may be made available for inference computations, as shown in FIG. 5. In the illustrated example, two users, User A and User B, are shown in interacting with the personalized NLP model from their respective client computing devices 42. The client computing devices 42 each execute an application client 26A configured to send inference-time user input 48 in the form of raw text data 28 to the computing device 12, and subsequently receive personalized classifications 40 from the computing device 12 as output. The application client 26A may include a graphical user interface 46, and configured to display graphical output 50 based on the personalized classifications 40 outputted from the personalized NLP model 38 and received at the client program 26A.
Fink, abstract, the invention teaches method of generating expanded responses that guide continuance of a human-to computer dialog that is facilitated by a client device and that is between at least one user and an automated assistant. The expanded responses are generated by the automated assistant in response to user interface input provided by the user via the client device, and are caused to be rendered to the user via the client device, as a response, by the automated assistant, to the user interface input of the user. An expanded response is generated based on at least one entity of interest determined based on the user interface input, and is generated to incorporate content related to one or more additional entities that are related to the entity of interest, but that are not explicitly referenced by the user interface input.
[0090] FIGS. 4, 5, and 6 each depict example dialogs between a user and an automated assistant, in which composite responses can be provided.
[0091-0092] Fig 4…In response, the user provides second natural language input 490B of "How do you configure them", that is guided by the content of the responsive natural language output 492A. In response to the first natural language input 490A, automated assistant provides another responsive natural language output 492B that is a composite response generated according to implementations disclosed herein. For example, the first sentence of the natural language output 492B can be based on a textual snippet from a given content agent, "It saves at least 5 minutes" can be based on an additional textual snippet from another content agent, and "is called Hypothetical App." can be based on a further textual snippet from a further content agent.) by adding attribute information of a first user to a first feature amount input by the first user (Wang, abstract, the invention describes method for enabling contextually relevant conversational interaction. Environment data is received by an AI System which detects a plurality of physical objects in a physical environment and forms a contextual understanding of the plurality of physical objects and the physical environment and identifies a user relevant to the contextual understanding. A most relevant contextual information to the user is predicted by the AI system and transformed into a textual form. A set of intents and objectives is predicted by the AI system for user-centered interaction. The AI system and the user interact iteratively through the user centered interaction to determine an understanding of a most relevant intent and a most relevant objective which is validated by the AI system with the user until the user agrees. The validated most relevant intent and the most relevant objective is utilized to facilitate the user-centered and contextually relevant conversational interaction.
[0034] The present invention aims to deliver an improved user experience by considering individual preferences, conversational history, and communication styles. The AI system adapts its responses according to a user's past interaction or preferred tone, resulting in more effective and satisfying conversations.
[0035] Additionally, the invention incorporates contextual information, such as the user's current location, time, or activity, to provide more relevant and timely responses.
[0049] Referring to FIG. 1, the OKB 104 (Object Knowledge Base) is a structured system, but not limited to a database, a set of databases, a repository, a set of repositories, and the like, that stores different types of data and various types of information that the AI system can access and use to provide contextually relevant and personalized responses.
[0056] To add human-AI interaction data to the OKB, the AI system initially processes the data to extract significant information such as user intents, behaviors, and preferences. This processing often involves using NLP techniques to parse the conversation and identify the essential elements. After the data has been processed, it is added to the OKB in a structured format that is easily accessible and analyzed by the AI system. This format can include relevant contextual information such as the time and location of the interaction, the user’s input, and the AI system’s response.).
The reasoning for combination of Dimitriadis, Fink and Wang is the same as described in Claim 1.
Claim 15 is similar in scope as Claim 1, and thus is rejected under same rationale.
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Dimitriadis et al (US20230297777) in view of Fink et al (US20240161743), Wang et al (US20230245651) further in view of Kulkarni et al (US20240311652).
Regarding Claim 5. The combination of Dimitriadis, Fink and Wang fails to explicitly teach, however, Kulkarni teaches The information processing apparatus according to claim 4, wherein the personalized feature amount model is capable of being trained by inferring the second prompt token sequence on the basis of an attribute information token sequence obtained by tokenizing the attribute information of the first user and a first prompt token sequence obtained by tokenizing the first feature amount, calculating a cross entropy loss related to the second prompt token sequence, and updating the personalized feature amount model using an error backpropagation method (Kulkarni, abstract, the invention describes methods for prompt generation for generative models including utilizing a specialized markup language. A markup language transform can be utilized to augment user input data to generate a prompt that includes structure and/or wording that facilitates the generation of a generative output that reflects a user's intent. The systems and methods can leverage the specialized markup language and/or an integrated development environment interface to inform a user of the prompt parts and provide editing options.
[0102] The training computing system 150 can include a model trainer 160 that trains the machine-learned models 120 and/or 140 stored at the user computing device 102 and/or the server computing system 130 using various training or learning techniques, such as, for example, backwards propagation of errors. For example, a loss function can be backpropagated through the model(s) to update one or more parameters of the model (s) (e.g., based on a gradient of the loss function). Various loss functions can be used such as mean squared error, likelihood loss, cross entropy loss, hinge loss, and/or various other loss functions. Gradient descent techniques can be used to iteratively update the parameters over a number of training iterations.).
Dimitriadis, Fink, Wang and Kulkarni are analogous art because they all teach generating response based on past user responses. Kulkarni further teaches the AI training system using various training or learning techniques including backwards propagation of errors and cross entropy loss. Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention, to modify the personalized natural language processing system (taught in Dimitriadis, Fink and Wang), to further use the various training or learning techniques (taught in Kulkarni), so as to provide user with an easy and intuitive user interface to write prompts for large language models (Kulkarni, [0002]).
Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Dimitriadis et al (US20230297777) in view of Fink et al (US20240161743), Wang et al (US20230245651) further in view of Yagi (JP2022023767).
Regarding Claim 8. The combination of Dimitriadis, Fink and Wang fails to explicitly teach, however, Yagi teaches The information processing apparatus according to claim 2, wherein the user information acquisition unit acquires the attribute information of the first user on a basis of data stored in a terminal operated by the first user or data stored in an application included in the terminal (Yagi, abstract, the invention describes a program capable of providing coffee suitable for a user, an information processing method and a method for generating a learning model. In an information processing system, a server 10 has a function of a web server, and discloses an order site 12S for receiving an order for coffee via a network N via the network N. The server 10 performs various information processing of processing for receiving an order for coffee through the order site 12S, processing of identifying and proposing a type of coffee suitable for a user on the basis of information about the user of a user terminal 20, or the like.
[0030] When the control unit 21 of the user terminal 20 launches the order application 22AP in accordance with instructions from the user via the input unit 24, it displays a menu screen as shown in Figure 5A on the display unit 25 (S11). The menu screen has an "Order from Recommended Menu" button for ordering from the recommended menu generated by server 10, and displays a purchase history menu for ordering from purchase history, and a regular menu for ordering from the menu provided by the store. If user purchase history information and regular menu information are stored in the order application 22AP or in the storage unit 22, the control unit 21 reads the purchase history information and regular menu information and generates the purchase history menu and the regular menu. The control unit 21 then generates a menu screen as shown in Figure 5A by displaying the "Order from Recommendation Menu" button, the generated purchase history menu, and the regular menu. The purchase history menu displays product information for each item the user has previously purchased, including coffee type, purchase date and time, coffee description, and thumbnail image. It also includes an "Add to Cart" button to add the item to the purchase list and a "Enter Rating" button to enter a rating for the item. The purchase history menu in Figure 5A displays the product from one purchase history and a "See More" button for viewing other products. However, it may also be configured to display multiple products from multiple purchase histories. The regular menu displays product information for each item available for purchase in-store, including coffee type, description, and thumbnail image, and includes an "Add to Cart" button for that item. The control unit 21 may be configured to access the server 10 (order site 12S) to request a menu screen when the order application 22AP is launched, receive the menu screen generated by the server 10, and display it on the display unit 25. Furthermore, the control unit 21 may be configured to access the server 10 (order site 12S) and obtain from it either the purchase history menu or the regular menu, or both, to be displayed on the menu screen when the order application 22AP is launched.).
Dimitriadis, Fink, Wang and Yagi are analogous art because they all teach generating response based on past user responses. Yagi further teaches the AI system will predict response based on past user purchase history stored in local device. Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention, to modify the personalized natural language processing system (taught in Dimitriadis, Fink and Wang), to further consider the contextual data about the user stored in the local device (taught in Yagi), so as to provide user with customized recommendation based on personal preferences (Yagi, [0001-0007]).
Claim 11 is rejected under 35 U.S.C. 103 as being unpatentable over Dimitriadis et al (US20230297777) in view of Fink et al (US20240161743), Wang et al (US20230245651) further in view of Zhang et al (US20230230198).
Regarding Claim 11. The combination of Dimitriadis, Fink and Wang fails to explicitly teach, however, Zhang teaches The information processing apparatus according to claim 4, further comprising
an image generation unit capable of generating an image with the second feature amount as an input prompt, wherein
in the personalized feature amount model, a feature amount input by the first user after a generated image generated by the image generation unit is displayed on a terminal operated by the first user is set as a behavior history of the first user (Zhang, abstract, the invention describes methods that implement a neural network framework for interactive multi-round image generation from natural language inputs. Specifically, the disclosed systems provide an intelligent framework (i.e., a text-based interactive image generation model) that facilitates a multi-round image generation and editing workflow that comports with arbitrary input text and synchronous interaction. In particular embodiments, the disclosed systems utilize natural language feedback for conditioning a generative neural network that performs text-to-image generation and text-guided image modification. For example, the disclosed systems utilize a trained model to inject textual features from natural language feedback into a unified joint embedding space for generating text-informed style vectors. In tum, the disclosed systems can generate an image with semantically meaningful features that map to the natural language feedback. Moreover, the disclosed systems can persist these semantically meaningful features throughout a refinement process and across generated images.
[0026] In addition to improved accuracy, the interactive image generation system can also improve system flexibility. For example, unlike some conventional image systems, the interactive image generation system can perform multi-round image generation while persisting edits throughout the image generation rounds. To do so, the interactive image generation system selectively determines elements of a previous style vector to update based on the additional textual feedback. To illustrate, the interactive image generation system determines a similarity between a semantic feature change for each style element of the previous style vector and a desired semantic change based on the additional textual feedback. From a modified style vector with the updated style elements, a generative neural network can flexibly generate modified images that reflect iterative feedback plus prior feedback.).
Dimitriadis, Fink, Wang and Zhang are analogous art because they all teach generating response based on past user responses. Zhang further teaches the AI system generating image based on continuous user inputs. Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention, to modify the personalized natural language processing system (taught in Dimitriadis, Fink and Wang), to further generate image based on iterative user input and user history data (taught in Zhang), so as to provide user with efficient image generation method using limited computing resources (Zhang, [0001-0005]).
Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Dimitriadis et al (US20230297777) in view of Fink et al (US20240161743), Wang et al (US20230245651) further in view of Piramuthu et al (US20240095987).
Regarding Claim 12. The combination of Dimitriadis, Fink and Wang fails to explicitly teach, however, Piramuthu teaches The information processing apparatus according to claim 2, wherein the image generation model used by the image generation unit is capable of performing beam search during training and inference of the second feature amount (Piramuthu, abstract, the invention describes method for generating content associated with a user input/system. Natural language data associated with a user input may be generated. For each portion of the natural language data, ambiguous references to entities in the portion may be replaced with the corresponding entity. Entities included in the portion may be extracted, and image data representing the entity may be determined. Background image data associated with the entities and the portion may be determined, and attributes which modify the entities in the natural language sentence may be extracted. Spatial relationships between two or more of the entities may further be extracted. Image data representing the natural language data may be generated based on the background image data, the entities, the attributes, and the spatial relationships. Video data may be generated based on the image data, where the video data includes animations of the entities moving.
[0203] In some implementations, the ASR component 850 may process the audio data 811 using the ASR model 1050. The ASR model 1050 may be, for example, a recurrent neural network such as an RNN-T. An example RNN-T architecture is illustrated in FIG. 10. The ASR model 1050 may predict a probability (ylx) of labels y=(y1, ... , yu) given acoustic features x=(x1, ... , xt). During inference, the ASR model 1050 can generate an N-best list using, for example, a beam search decoding algorithm. The ASR model 1050 may include an encoder 1012, a prediction network 1020, a joint network 1030, and a softmax 1040. The encoder 1012 may be similar or analogous to an acoustic model (e.g., similar to the acoustic model 1053 described below), and may process a sequence of acoustic input features to generate encoded hidden representations. The prediction network 1020 may be similar or analogous to a language model (e.g., similar to the language model 1054 described below), and may process the previous output label predictions, and map them to corresponding hidden representations. The joint network 1030 may be, for example, a feed forward neural network (NN) that may process hidden representations from both the encoder 1012 and prediction network 1020, and predict output label probabilities. The softmax 1040 may be a function implemented (e.g., as a layer of the joint network 1030) to normalize the predicted output probabilities.).
Dimitriadis, Fink, Wang and Piramuthu are analogous art because they all teach inferring response based on user inputs. Piramuthu further teaches the AI system generating response using beam search decoding algorithm. Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention, to modify the personalized natural language processing system (taught in Dimitriadis, Fink and Wang), to further use the beam search decoding algorithm in generating inference response (taught in Piramuthu), so as to provide user with improved human-computer interactions (Piramuthu, [0002]).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. LaRhette et al (US20240281472), abstract, the invention teaches an interactive search method which utilizes a browser-based interface and generative artificial intelligence to enhance user search experiences. The method involves receiving an initial search query from a user, generating a proposed search result via hardware processors, and displaying the result to the user. To refine search accuracy, the method recommends clarifying questions, soliciting additional search parameters. Upon receiving an updated search query, the system interactively refines the initial query and displays an updated search result. This process allows for dynamic query adjustment and improved search result relevance in real-time.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to XIN SHENG whose telephone number is (571)272-5734. The examiner can normally be reached M-F 9:30AM-3:30PM 6:00PM-8:30PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jason Chan can be reached at 5712723022. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Xin Sheng/Primary Examiner, Art Unit 2619