Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
Claims 1-20 are pending.
Response to Arguments
Applicant’s arguments with respect to claim(s) 1-20 have been considered but are moot because the arguments are directed towards the newly amended claim limitations that change the scope of the claims as a whole and are open to new grounds of rejection/interpretation.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 7 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 7 claims generating a modified media content, but depends on claim 1. Amended claim 1 already claims a generated “modified media content.” Hence, it is unclear and indefinite as to if this is a different modified media content in addition to the previous modified media content or the same modified media content.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1-3, 5-9, 11-12, 14-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Xu et al. (US 20250014245) in view of Voss et al. (US 20210368039).
Re claim 1, Xu teaches a method comprising:
receiving an input associated with a first user having a user profile in a messaging application ([0040] Specifically, FIG. 1 depicts a chat window 100 displayed on a display device of a first user. The chat window 100 of FIG. 1 corresponds with a point in time after three messages are exchanged between the first user and a second user, i.e., a first message 110 from the second user to the first user sending a text message (e.g., “Hello”); a second message 120 from the first user to the second user responding to the first message 110 (e.g., “What's up”); and a third message 130 from the second user to the first user posing a question (e.g., “How did you like the movie last night?”). The chat window 100 includes a prompt field 140, which reflects, for example, characters typed by the first user during the communication. From a perspective of the first user, the prompt field 140 will populate in response to the first user typing or entering information via an input device. For example, the input device is at least one of a keyboard, a touchscreen, a mouse, a virtual input system (e.g., a virtual keyboard in virtual reality applications), an audio input device, a speech-to-text input device, a transcription device, a combination of the same, or the like), ([0041] In some embodiments, a determination is made that the first message 110 and the second message 120 are introductory in nature, and meme generation mode is not enabled; whereas, the third message 130 is a question, and a determination is made that the third message 130 is a suitable candidate for meme generation), ([0043] In the embodiment of FIG. 1, each of the three candidate memes is generated using a different process. A first meme candidate 160 is generated based on a template (e.g., one of a known group of templates, i.e., a “thumbs up” version of a “Hide the Pain Harold” template), which is chosen based on a determined context of the communication and with a caption (e.g., “IT'S GREAT!!!”) populated by any of the captioning processes disclosed herein. A second meme candidate 170 is generated using a text-to-image process. In some embodiments, the text-to-image process generates images and captions based on predicted answers and a determined context of the communication. A third meme candidate 180 is generated using the user's profile, a text-to-image process, and a determined context of the communication. In this example, the generated image has attributes extracted from the image from the user's profile. The extracted attributes include a facial appearance resembling that of the user, a hairstyle of the user, accessories worn by the user, clothing worn by the user, and the like. The caption (e.g., “Boring”) is generated by any of the captioning processes disclosed herein). Xu teaches receiving input associated with a first user having a user profile in a messaging application.
Analyzing, by an AI assistant, the input to determine an intent to generate media content ([0035] Meme generation leverages Al and ML to achieve a natural language understanding of message context. Large models are pre-trained and leveraged into the automatic meme generation. The automated meme generation includes determining a context, an intent, and/or a mood of an electronic communication. Suggestions are generated including responsive images and captions based on the determined context, intent, and/or mood. The suggestions convey emotion in the communication. Users are prompted for feedback regarding the generated suggestions, which inform the trained models and promote customization per user, group, or communication), ([0054] As shown in FIG. 1, candidate memes are generated based on a context of the communication. If a user continues typing, as shown in FIG. 4, a typed partial answer is used together with the communication context to complete the answer and generate one or more candidate memes. In this example, the user types “What if I told you” into a prompt field 440. In response to this partial sentence, substantially in real-time, candidate memes based on the partial text are generated and the generated meme selection field 450 is replaced with a generated meme selection field 450, in which candidates are instead based on the entered text of “What if I told you.” In the example of FIG. 4, a first candidate meme 460 is based on a template, a second candidate meme 470 is based on a text-to-image process and a description from the user profile, and a third candidate meme 480 is based on a text-to-image process and an image from the user's profile. The caption for each candidate meme includes the entered text of “What if I told you” and an automatically generated conclusion to the sentence, e.g., “It was boring” in the first candidate meme 460, “It's none of your business” in the second candidate meme 470, and “I don't like it” in the third candidate meme 480. Selection of the expand button 190 allows the user to browse through additional generated memes. As with the example of FIGS. 1 and 2, the user is prompted to select one of the candidate memes for inclusion in the communication. In some embodiments, selection of one of the candidate memes 460, 470, 480, which includes the entered text of “What if I told you” replaces the prompt field 440 with the selected meme. As with the example of FIGS. 1 and 2, selection of the configuration button 195 allows the user to change settings for the meme generation process. Although not shown in FIG. 4, it is understood that a database similar to the database 105 is populated for the example of FIG. 4. That is, autocompleted responses are associated with, e.g., at least one of the User 1 ID, the User 2 ID, the Session ID, the Context 1 ID, the Context 2 ID, the Meme 1 ID, the Meme 2 ID, the Meme 3 ID, or the like). Xu teaches analyzing the input that is assisted by AI/ML, that determines if there is an intent to generate the meme media content for electronic communication.
Generating, by the AI assistant, the media content based on the determined intent and content of the input ([0036] The determined content, context, intent, and/or mood of the communication, details from a user profile, user preferences, and the like are analyzed. Inferences are based on the analyzed information. The inferences include a decision of whether and when to generate and display a plurality of candidate memes for user selection. As a result of the analysis, memes are suggested at a point in time best suited for the communication. The analyzed information is utilized as an input for image generation. The analyzed information is input into a text-to-image diffusion model. The model is fine-tuned for generating meme images; that is, for example, weights of a pre-trained model are trained on data. The model is personalized with user-specific information, including a user's profile images), and (([0041] In some embodiments, a determination is made that the first message 110 and the second message 120 are introductory in nature, and meme generation mode is not enabled; whereas, the third message 130 is a question, and a determination is made that the third message 130 is a suitable candidate for meme generation), ([0043] In the embodiment of FIG. 1, each of the three candidate memes is generated using a different process. A first meme candidate 160 is generated based on a template (e.g., one of a known group of templates, i.e., a “thumbs up” version of a “Hide the Pain Harold” template), which is chosen based on a determined context of the communication and with a caption (e.g., “IT'S GREAT!!!”) populated by any of the captioning processes disclosed herein. A second meme candidate 170 is generated using a text-to-image process. In some embodiments, the text-to-image process generates images and captions based on predicted answers and a determined context of the communication. A third meme candidate 180 is generated using the user's profile, a text-to-image process, and a determined context of the communication. In this example, the generated image has attributes extracted from the image from the user's profile. The extracted attributes include a facial appearance resembling that of the user, a hairstyle of the user, accessories worn by the user, clothing worn by the user, and the like. The caption (e.g., “Boring”) is generated by any of the captioning processes disclosed herein). Xu teaches generating, by the AI assistant, media meme content based on the determination and content of the input, such as a meme response to a question, including aspects of a user’s profile image, information, etc.
and displaying the generated media content within the messaging application (see Fig. 1-2, 4, wherein the media content (meme) is generated within the messaging application and displayed).
Xu does not explicitly teach modifying, by the AI assistant, the generated media content to generate modified media content in response to determining a context of a second input, from the first user or from a second user in a group communication with the first user, within the messaging application, and displaying, based on the determined context of the second input, the generated modified media content within the messaging application.
However, Voss teaches modifying, by the AI assistant, the generated media content to generate modified media content in response to determining a context of a second input, from the first user or from a second user in a group communication with the first user, within the messaging application (see abstract: A system for customizing text messages in modifiable videos of a multimedia messaging application (MMA) is provided. In one example embodiment, the system includes a processor and a memory storing processor-executable codes, wherein the processor is configured to analyze recent messages of a user in the MMA to determine a context of the recent messages; determine, based on the context, a customized text message; select, based on the context, a list of relevant modifiable videos from a database configured to store modifiable videos, the modifiable videos being associated with preset text messages; replace the preset text messages in the relevant modifiable videos with the customized message; and render the list of relevant modifiable videos for viewing and selecting by the user, the rendering including displaying the customized text message in the relevant modifiable videos) ([0018] In some embodiments, the MMA can analyze recent messages of the user in the communication chat to determine context of the recent messages and emotional state of the user. The context and the emotional state can be determined with one of machine learning techniques (for example, a specially trained artificial neural network). The MMA may also generate a customized text message based on the context and the emotional state of the user), ([0019] Upon determining that the user has entered an option for typing in a text message, the MMA may provide a list of reels that are relevant to the context and the emotional state. The MMA may also prefill a field for entering the text message with one or more keywords corresponding to the context and the emotional state. In some embodiments, the MMA may prefill the field for entering the text message with a customized text message), ([0020] When the user selects the reels from the lists of relevant reels, the MMA may provide an option for modifying the customized message and the style of the customized message in the selected reel. The MMA may modify the customized message based on the context of the communication chat and the emotional state. The MMA may also modify the style of the customized message based on the context of the communication chat and the emotional state of the user. For example, if the emotional state of the user is “upset,” letters of the customized message may be depicted as frozen ice. In another example, if the user is “angry,” the letters of the customized message may be depicted as burning flames), ([0037] Each time when the user modifies the customized text message determined based on the context of the recent messages, the reel search module 310 may store, into the statistical log 150, information concerning the context, the customized text message, and the modified customized text message. The information from the statistical log 150 can be used to update an algorithm for determining the context and the customized text message. In some embodiments, the customized text message can be determined using a machine learning model based on historical data in the statistical log 150. The machine learning model can be based on an artificial neural network, decision trees, support vector machine, regression analysis, and so forth). Voss teaches modifying, by the AI assistant, the generated media content (modifiable videos that are customizable) in response to determining a context of a second input from the first user or from a second user in a group communication with the first user, within the messaging application (context and emotional state and a user’s selection to edit text or selecting a reel in the multimedia messaging application)
and displaying, based on the determined context of the second input, the generated modified media content within the multi messaging application (see Fig. 5A and 5B, in reference to [0053-0054], wherein modified media such selection and text input is displayed based on second input within the messaging application).
Xu and Voss teaches claim 1. It would have been obvious to one of ordinary skill in the art modify Xu’s display of generated media content by an AI assistant to explicitly include modifying, by the AI assistant, the generated media content to generate modified media content in response to determining a context of a second input from a user, as taught by Voss, as the references are in the analogous art of messaging application for user communication. An advantage of the modification is that it achieves the result of evaluating a context of a second input to further modify generated media content for displaying within a messaging application, to allow for more customization of messages sent between users.
Re claim 2, Xu and Voss teaches claim 1. Furthermore, Voss teaches the analyzing of the input comprises evaluating a conversation history between the first user and the AI assistant; and the input comprising textual input or a voice input ([0022] FIG. 1 shows an example environment 100, wherein a system and a method for customizing text messages in modifiable videos in MMAs can be practiced. The environment 100 may include a PCD 105, a user 102, a PCD 110, a user 104, a network 120, and messaging server system (MSS) 130. The PCD 105 and PCD 110 can refer to a mobile device such as a mobile phone, smartphone, or tablet computer. In further embodiments, however, the PCD 105 and PCD 110 can refer to a personal computer, laptop computer, netbook, set top box, television device, multimedia device, personal digital assistant, game console, entertainment system, infotainment system, vehicle computer, or any other computing device), ([0030] FIG. 2 is a block diagram showing an example embodiment of a PCD 105 (or PCD 110) for implementing methods for customizing text messages in modifiable videos in MMAs. In the example shown in FIG. 2, the PCD 105 includes both hardware components and software components. Particularly, the PCD 105 includes a camera 205 or any other image-capturing device or scanner to acquire digital images. The PCD 105 can further include a processor module 210 and a memory storage 215 for storing software components and processor-readable (machine-readable) instructions or codes, which, when performed by the processor module 210, cause the PCD 105 to perform at least some steps of methods for customizing text messages in modifiable videos as described herein. The PCD 105 may include graphical display system 230 and a communication module 240. In other embodiments, the PCD 105 may include additional or different components. Moreover, the PCD 105 can include fewer components that perform functions similar or equivalent to those depicted in FIG. 2), ([0031] The PCD 105 can further include an MMA 160. The MMA 160 may be implemented as software components and processor-readable (machine-readable) instructions or codes stored in the memory storage 215, which when performed by the processor module 210, cause the PCD 105 to perform at least some steps of methods for providing communication chats, generation of personalized videos, and customizing text messages in modifiable videos in MMAs as described herein. A user interface of the MMA 160 can be provided via the graphical display system 230. The communication chats can be enabled by MMA 160 via the communication module 240 and the network 120. The communication module 240 may include a GSM module, a WiFi module, a Bluetooth™ module and so forth), ([0057] In block 610, the method 600 may include determining, based on the context, a customized text message. The customized text message can be determined by a machine learning model based on historical data from a statistical log. The statistical log may store historical data concerning information on the context and the customized text messages. The method 600 may include selecting, based on the context, a customized soundtrack from a library. The method 600 may include selecting, based on the emotional state, a style of text animation from a list of styles of text animations. The method 600 may include selecting, based on the emotional state, a hairstyle from a list of hairstyles. The method may also include modifying the selfie by applying the hairstyle to an image of the hair. In block 615, the method 600 may include selecting, based on the context, a list of relevant modifiable videos from a database. The database is configured to store modifiable videos, and the modifiable videos being associated with preset text messages, preset styles of text animation, and preset soundtracks. In some embodiments, the list of relevant modifiable videos can be determined by searching, in the database, for modifiable videos having preset text messages semantically close to the customized text message), ([0037] Each time when the user modifies the customized text message determined based on the context of the recent messages, the reel search module 310 may store, into the statistical log 150, information concerning the context, the customized text message, and the modified customized text message. The information from the statistical log 150 can be used to update an algorithm for determining the context and the customized text message. In some embodiments, the customized text message can be determined using a machine learning model based on historical data in the statistical log 150. The machine learning model can be based on an artificial neural network, decision trees, support vector machine, regression analysis, and so forth) and ([0061] Upon receiving an indication that the user has selected the modifiable video from the list of modifiable videos, the method 600 may include providing options to the user to modify the style of text animation and the customized soundtrack in the selected modifiable video. The method 600 may include receiving an indication that the user has modified the selected style of text animation in the selected modifiable video and, in response to the indication, storing information concerning the context and the modified selected style of text animation into the statistical log. The method 600 may adjust, based on historical data in the statistical log, an algorithm for selecting the style of text animation from the list of styles of text animations). For motivation, see claim 1.
Re claim 3, Xu and Voss teaches claim 1. Furthermore, Xu teaches wherein the intent to generate the media content is determined without an explicit command from the user to generate the media content ([0054] As shown in FIG. 1, candidate memes are generated based on a context of the communication. If a user continues typing, as shown in FIG. 4, a typed partial answer is used together with the communication context to complete the answer and generate one or more candidate memes. In this example, the user types “What if I told you” into a prompt field 440. In response to this partial sentence, substantially in real-time, candidate memes based on the partial text are generated and the generated meme selection field 450 is replaced with a generated meme selection field 450, in which candidates are instead based on the entered text of “What if I told you.” In the example of FIG. 4, a first candidate meme 460 is based on a template, a second candidate meme 470 is based on a text-to-image process and a description from the user profile, and a third candidate meme 480 is based on a text-to-image process and an image from the user's profile. The caption for each candidate meme includes the entered text of “What if I told you” and an automatically generated conclusion to the sentence, e.g., “It was boring” in the first candidate meme 460, “It's none of your business” in the second candidate meme 470, and “I don't like it” in the third candidate meme 480. Selection of the expand button 190 allows the user to browse through additional generated memes. As with the example of FIGS. 1 and 2, the user is prompted to select one of the candidate memes for inclusion in the communication. In some embodiments, selection of one of the candidate memes 460, 470, 480, which includes the entered text of “What if I told you” replaces the prompt field 440 with the selected meme. As with the example of FIGS. 1 and 2, selection of the configuration button 195 allows the user to change settings for the meme generation process. Although not shown in FIG. 4, it is understood that a database similar to the database 105 is populated for the example of FIG. 4. That is, autocompleted responses are associated with, e.g., at least one of the User 1 ID, the User 2 ID, the Session ID, the Context 1 ID, the Context 2 ID, the Meme 1 ID, the Meme 2 ID, the Meme 3 ID, or the like).
Re claim 5, Xu and Voss teach claim 1. Furthermore, Xu teaches receiving a subsequent input associated with animating the generated media content ([0071] In some embodiments, a text prompt in the communication is analyzed to determine whether one or more of the participants of the communication is a likely subject for the generated meme. For example, as shown in FIG. 1, the text prompt from the third message 130 from the second user to the first user posing the question “How did you like the movie last night?” is analyzed in this context to determine that “you” likely refers to the first user, and one of the generated candidate memes (e.g., the third meme candidate 180) includes an image based on user profile of the first user) and ([0054] As shown in FIG. 1, candidate memes are generated based on a context of the communication. If a user continues typing, as shown in FIG. 4, a typed partial answer is used together with the communication context to complete the answer and generate one or more candidate memes. In this example, the user types “What if I told you” into a prompt field 440. In response to this partial sentence, substantially in real-time, candidate memes based on the partial text are generated and the generated meme selection field 450 is replaced with a generated meme selection field 450, in which candidates are instead based on the entered text of “What if I told you.” In the example of FIG. 4, a first candidate meme 460 is based on a template, a second candidate meme 470 is based on a text-to-image process and a description from the user profile, and a third candidate meme 480 is based on a text-to-image process and an image from the user's profile. The caption for each candidate meme includes the entered text of “What if I told you” and an automatically generated conclusion to the sentence, e.g., “It was boring” in the first candidate meme 460, “It's none of your business” in the second candidate meme 470, and “I don't like it” in the third candidate meme 480. Selection of the expand button 190 allows the user to browse through additional generated memes. As with the example of FIGS. 1 and 2, the user is prompted to select one of the candidate memes for inclusion in the communication. In some embodiments, selection of one of the candidate memes 460, 470, 480, which includes the entered text of “What if I told you” replaces the prompt field 440 with the selected meme. As with the example of FIGS. 1 and 2, selection of the configuration button 195 allows the user to change settings for the meme generation process. Although not shown in FIG. 4, it is understood that a database similar to the database 105 is populated for the example of FIG. 4. That is, autocompleted responses are associated with, e.g., at least one of the User 1 ID, the User 2 ID, the Session ID, the Context 1 ID, the Context 2 ID, the Meme 1 ID, the Meme 2 ID, the Meme 3 ID, or the like).
Xu teaches receiving a subsequent textual input associated with animating the generated media content (“What if I told you…” is subsequent to, “How did you like the movie last night?”).
generating an animated version of the generated media content based on the subsequent input; and displaying the animated version within the messaging application (see candidate memes in Fig. 4 as modified generated media content including animation content) and ([0072] In some embodiments, a further modification of the generated image is added. For example, a filter is added to the resulting image based on the user preference, or a different shape template is used to modify the generated image. In some embodiments, the above mentioned modification is embedded in the fine-tuned diffusion model, so that the output from the fine-tuned diffusion model will contain the modification directly. In some embodiments, the text-to-image diffusion model generates animated images or looped videos based on the prompt).
Re claim 6, Xu and Voss teach claim 5. Furthermore, Xu teaches modifying of the generated media content based on changing one or more visual attributes of the generated media content ([0072] In some embodiments, a further modification of the generated image is added. For example, a filter is added to the resulting image based on the user preference, or a different shape template is used to modify the generated image. In some embodiments, the above mentioned modification is embedded in the fine-tuned diffusion model, so that the output from the fine-tuned diffusion model will contain the modification directly. In some embodiments, the text-to-image diffusion model generates animated images or looped videos based on the prompt) and (([0054] As shown in FIG. 1, candidate memes are generated based on a context of the communication. If a user continues typing, as shown in FIG. 4, a typed partial answer is used together with the communication context to complete the answer and generate one or more candidate memes. In this example, the user types “What if I told you” into a prompt field 440. In response to this partial sentence, substantially in real-time, candidate memes based on the partial text are generated and the generated meme selection field 450 is replaced with a generated meme selection field 450, in which candidates are instead based on the entered text of “What if I told you.” In the example of FIG. 4, a first candidate meme 460 is based on a template, a second candidate meme 470 is based on a text-to-image process and a description from the user profile, and a third candidate meme 480 is based on a text-to-image process and an image from the user's profile. The caption for each candidate meme includes the entered text of “What if I told you” and an automatically generated conclusion to the sentence, e.g., “It was boring” in the first candidate meme 460, “It's none of your business” in the second candidate meme 470, and “I don't like it” in the third candidate meme 480. Selection of the expand button 190 allows the user to browse through additional generated memes. As with the example of FIGS. 1 and 2, the user is prompted to select one of the candidate memes for inclusion in the communication. In some embodiments, selection of one of the candidate memes 460, 470, 480, which includes the entered text of “What if I told you” replaces the prompt field 440 with the selected meme. As with the example of FIGS. 1 and 2, selection of the configuration button 195 allows the user to change settings for the meme generation process. Although not shown in FIG. 4, it is understood that a database similar to the database 105 is populated for the example of FIG. 4. That is, autocompleted responses are associated with, e.g., at least one of the User 1 ID, the User 2 ID, the Session ID, the Context 1 ID, the Context 2 ID, the Meme 1 ID, the Meme 2 ID, the Meme 3 ID, or the like).
Re claim 7, Xu and Voss teach claim 1. Furthermore, Xu teaches receiving a subsequent input associated with modifying the generated media content ([0071] In some embodiments, a text prompt in the communication is analyzed to determine whether one or more of the participants of the communication is a likely subject for the generated meme. For example, as shown in FIG. 1, the text prompt from the third message 130 from the second user to the first user posing the question “How did you like the movie last night?” is analyzed in this context to determine that “you” likely refers to the first user, and one of the generated candidate memes (e.g., the third meme candidate 180) includes an image based on user profile of the first user) and ([0054] As shown in FIG. 1, candidate memes are generated based on a context of the communication. If a user continues typing, as shown in FIG. 4, a typed partial answer is used together with the communication context to complete the answer and generate one or more candidate memes. In this example, the user types “What if I told you” into a prompt field 440. In response to this partial sentence, substantially in real-time, candidate memes based on the partial text are generated and the generated meme selection field 450 is replaced with a generated meme selection field 450, in which candidates are instead based on the entered text of “What if I told you.” In the example of FIG. 4, a first candidate meme 460 is based on a template, a second candidate meme 470 is based on a text-to-image process and a description from the user profile, and a third candidate meme 480 is based on a text-to-image process and an image from the user's profile. The caption for each candidate meme includes the entered text of “What if I told you” and an automatically generated conclusion to the sentence, e.g., “It was boring” in the first candidate meme 460, “It's none of your business” in the second candidate meme 470, and “I don't like it” in the third candidate meme 480. Selection of the expand button 190 allows the user to browse through additional generated memes. As with the example of FIGS. 1 and 2, the user is prompted to select one of the candidate memes for inclusion in the communication. In some embodiments, selection of one of the candidate memes 460, 470, 480, which includes the entered text of “What if I told you” replaces the prompt field 440 with the selected meme. As with the example of FIGS. 1 and 2, selection of the configuration button 195 allows the user to change settings for the meme generation process. Although not shown in FIG. 4, it is understood that a database similar to the database 105 is populated for the example of FIG. 4. That is, autocompleted responses are associated with, e.g., at least one of the User 1 ID, the User 2 ID, the Session ID, the Context 1 ID, the Context 2 ID, the Meme 1 ID, the Meme 2 ID, the Meme 3 ID, or the like).
Xu teaches receiving a subsequent textual input associated with modifying the generated media content (“What if I told you…” is subsequent to, “How did you like the movie last night?”).
generating a modified media content based on the subsequent input; and displaying the modified media content within the messaging application (see candidate memes in Fig. 4 generated as modified media content) and ([0072] In some embodiments, a further modification of the generated image is added. For example, a filter is added to the resulting image based on the user preference, or a different shape template is used to modify the generated image. In some embodiments, the above mentioned modification is embedded in the fine-tuned diffusion model, so that the output from the fine-tuned diffusion model will contain the modification directly. In some embodiments, the text-to-image diffusion model generates animated images or looped videos based on the prompt).
Re claim 8, Xu and Voss teach claim 1. Furthermore, Xu teach wherein the generating of the media content is performed using one or more multiple media content generation models selected based on the determined intent ([0043] In the embodiment of FIG. 1, each of the three candidate memes is generated using a different process. A first meme candidate 160 is generated based on a template (e.g., one of a known group of templates, i.e., a “thumbs up” version of a “Hide the Pain Harold” template), which is chosen based on a determined context of the communication and with a caption (e.g., “IT'S GREAT!!!”) populated by any of the captioning processes disclosed herein. A second meme candidate 170 is generated using a text-to-image process. In some embodiments, the text-to-image process generates images and captions based on predicted answers and a determined context of the communication. A third meme candidate 180 is generated using the user's profile, a text-to-image process, and a determined context of the communication. In this example, the generated image has attributes extracted from the image from the user's profile. The extracted attributes include a facial appearance resembling that of the user, a hairstyle of the user, accessories worn by the user, clothing worn by the user, and the like. The caption (e.g., “Boring”) is generated by any of the captioning processes disclosed herein).
Re claim 9, Xu and Voss teach claim 1. Furthermore, Xu teaches prior to the displaying the generated media content, providing one or more suggestion messages that indicates a suggestion to generate media content based on the input in the messaging application; and receiving user approval to generate and display the media content ([0054] As shown in FIG. 1, candidate memes are generated based on a context of the communication. If a user continues typing, as shown in FIG. 4, a typed partial answer is used together with the communication context to complete the answer and generate one or more candidate memes. In this example, the user types “What if I told you” into a prompt field 440. In response to this partial sentence, substantially in real-time, candidate memes based on the partial text are generated and the generated meme selection field 450 is replaced with a generated meme selection field 450, in which candidates are instead based on the entered text of “What if I told you.” In the example of FIG. 4, a first candidate meme 460 is based on a template, a second candidate meme 470 is based on a text-to-image process and a description from the user profile, and a third candidate meme 480 is based on a text-to-image process and an image from the user's profile. The caption for each candidate meme includes the entered text of “What if I told you” and an automatically generated conclusion to the sentence, e.g., “It was boring” in the first candidate meme 460, “It's none of your business” in the second candidate meme 470, and “I don't like it” in the third candidate meme 480. Selection of the expand button 190 allows the user to browse through additional generated memes. As with the example of FIGS. 1 and 2, the user is prompted to select one of the candidate memes for inclusion in the communication. In some embodiments, selection of one of the candidate memes 460, 470, 480, which includes the entered text of “What if I told you” replaces the prompt field 440 with the selected meme. As with the example of FIGS. 1 and 2, selection of the configuration button 195 allows the user to change settings for the meme generation process. Although not shown in FIG. 4, it is understood that a database similar to the database 105 is populated for the example of FIG. 4. That is, autocompleted responses are associated with, e.g., at least one of the User 1 ID, the User 2 ID, the Session ID, the Context 1 ID, the Context 2 ID, the Meme 1 ID, the Meme 2 ID, the Meme 3 ID, or the like).
Re claim 11, Xu and Voss teach claim 1. Furthermore, Xu teaches wherein the input is received in a group chat session, and wherein the generated media content is displayed to multiple users in the group chat session ([0132] Throughout the specification the terms “communication,” “electronic communication,” “online communication” and the like are provided. It is understood that the communication is at least one of an electronic communication, an online communication, an online chat, instant messaging, an iMessage, an electronic conversation, a group chat, a text message, an internet forum, an online message board, an email, a blog, an online article, a comments or posting section of a website (including a social media website), an offline communication, a voice message, a voicemail, combinations of the same, or the like. For instance, in some embodiments, the communication is a companion chat window and/or communication application for livestreaming video, such as YouTube, Zoom, Twitch, and the like. The communication may be text-based or audio-based. In some embodiments, for an audio-based communication, a speech-to-text module converts the audio to text for processing. The communication includes at least one of raw data from the communication, text, an entirety of a communication session, a session identifier, a context identifier, a user profile, a user preference, a user-provided image, an image of a face of a user engaged in the communication, an audio file, a recording, a transcription, a portion of the same, a combination of the same, or the like) and ([0040] Specifically, FIG. 1 depicts a chat window 100 displayed on a display device of a first user. The chat window 100 of FIG. 1 corresponds with a point in time after three messages are exchanged between the first user and a second user, i.e., a first message 110 from the second user to the first user sending a text message (e.g., “Hello”); a second message 120 from the first user to the second user responding to the first message 110 (e.g., “What's up”); and a third message 130 from the second user to the first user posing a question (e.g., “How did you like the movie last night?”). The chat window 100 includes a prompt field 140, which reflects, for example, characters typed by the first user during the communication. From a perspective of the first user, the prompt field 140 will populate in response to the first user typing or entering information via an input device. For example, the input device is at least one of a keyboard, a touchscreen, a mouse, a virtual input system (e.g., a virtual keyboard in virtual reality applications), an audio input device, a speech-to-text input device, a transcription device, a combination of the same, or the like), ([0041] In some embodiments, a determination is made that the first message 110 and the second message 120 are introductory in nature, and meme generation mode is not enabled; whereas, the third message 130 is a question, and a determination is made that the third message 130 is a suitable candidate for meme generation), ([0043] In the embodiment of FIG. 1, each of the three candidate memes is generated using a different process. A first meme candidate 160 is generated based on a template (e.g., one of a known group of templates, i.e., a “thumbs up” version of a “Hide the Pain Harold” template), which is chosen based on a determined context of the communication and with a caption (e.g., “IT'S GREAT!!!”) populated by any of the captioning processes disclosed herein. A second meme candidate 170 is generated using a text-to-image process. In some embodiments, the text-to-image process generates images and captions based on predicted answers and a determined context of the communication. A third meme candidate 180 is generated using the user's profile, a text-to-image process, and a determined context of the communication. In this example, the generated image has attributes extracted from the image from the user's profile. The extracted attributes include a facial appearance resembling that of the user, a hairstyle of the user, accessories worn by the user, clothing worn by the user, and the like. The caption (e.g., “Boring”) is generated by any of the captioning processes disclosed herein).
Claim 12 claims limitations in scope to claim 1 and is rejected for at least the reasons above.
Claim 14 claims limitations in scope to claim 2 and is rejected for at least the reasons above.
Claim 15 claims limitations in scope to claim 5 and is rejected for at least the reasons above.
Claim 16 claims limitations in scope to claim 8 and is rejected for at least the reasons above.
Claim 17 claims limitations in scope to claim 9 and is rejected for at least the reasons above.
Claim 18 claims limitations in scope to claim 1 and is rejected for at least the reasons above.
Claim 19 claims limitations in scope to claim 8 and is rejected for at least the reasons above.
Claim 20 claims limitations in scope to claim 9 and is rejected for at least the reasons above.
Claim(s) 4, 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Xu et al. (US 20250014245) in view of Voss et al. (US 20210368039) and Kim et al. (US 20250007870).
Re claim 4, Xu and Voss teaches claim 1. Xu and Voss do not explicitly teach
teach wherein analyzing of the input comprises determining one or more keywords associated with media content generation.
However, Kim teaches wherein analyzing of the input comprises determining one or more keywords associated with media content generation ([0067] More simply, an LLM is configured to determine what word, phrase, number, whitespace, nonalphanumeric character, or punctuation is most statistically likely to be next in a sequence, given the context of the sequence itself. The sequence may be initialized by the input prompt provided to the LLM. In this manner, output of an LLM is a continuation of the sequence of words, characters, numbers, whitespace, and formatting provided as the prompt input to the LLM), and ([0269] In the example of FIG. 8, the user input provided to the command prompt interface 802 includes reference to a project name (“Project Hercules”). The editor assistant service or other similar service may process the user input before or while generating the prompt for processing by the generative output engine. For example, the editor assistant service may parse the input text to screen for proper names, key words, and other grammatical content that indicates the user input is referencing a particular user, object, or grouping of objects. Triggering words or phrases may include “my.” “project.” “team,” “standup,” “weekly meeting.” “recent,” “last week” or the use of capitalized or proper nouns. In this example, the editor assistant service parses the user input and recognizes the term “Hercules,” which is a proper noun and used in conjunction with a triggering word “project.” In response, the editor assistant service may substitute the text of the user input with a link or identifier to a platform that is predicted to contain content related to the named project. In this example, the editor assistant service identifies an external issue tracking platform (example other collaboration platform), which includes issues and other content associated with the named project. When constructing the prompt, the system may either provide a link (if the content is able to be accessed by the generative output engine) or may retrieve content from the source and add the retrieved content to the prompt for processing. As shown in the example of FIG. 8, both the generated table of preview window 820 and the list of selectable object links of window 830 relate to content (e.g., issues) that are hosted on or provided by the issue tracking platform or system).
Xu, Voss, and Kim teaches claim 4. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Xu and Voss’s content generation system of analyzing prompts to explicitly include analyzing of keywords associated with media content generation, as explicitly taught by Kim, as the references are in the analogous art of AI aided content generation systems. An advantage of the modification is that it achieves the result of using keywords to help appropriate media content generation based on learned keywords and phrases.
Claim 13 claims limitations in scope to claim 4 and is rejected for at least the reasons above.
Claim(s) 10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Xu et al. (US 20250014245) in view of Voss et al. (US 20210368039) and Kim (US 20130201211, hereinafter “K. Kim”).
Re claim 10, Xu and Voss teach claim 9. Furthermore, Xu and Voss teaches receiving a user approval with at least one vote (as taught in claim 9), but does not explicitly teach receiving two or more votes to approve (i.e. an additional vote).
However, K. Kim teaches receiving two or more votes to approve (i.e. an additional vote) ([0169] In the course of taking the video, if the user intends to insert side information at a specific position, referring to FIG. 7(b), the user may be able to apply a touch input of a prescribed pattern to a desired position 720 in the preview image. Hence, referring to FIG. 7(c), a popup window 730 for enabling the user to confirm whether to input the side information may be displayed. If the user selects `input now`, a user interface (not shown in the drawing) for receiving an input of side information from the user may be displayed while the video taking continues. Once the video taking is completed, the controller 180 may be able to insert QR code including the inputted side information into the selected position. Of course, as mentioned in the foregoing description, `QR code insertion` may mean that the QR code is inserted to configure one portion of image information of the video. Alternatively, after the QR code has been created as a separate file, `QR code insertion` may mean that the QR code is inserted into the corresponding position in the corresponding view in a manner of being overlaid on the corresponding video that is being played). K. Kim teaches receiving an additional vote to approve, such as receiving a popup window to further confirm a user’s selection.
Xu, Voss, and K. Kim teach claim 10. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Xu and Voss’s media content generation system including one vote of a user selection to explicitly include an additional vote to approve (two or more votes), as taught by K. Kim, as the references are in the analogous art of selection-based user interface systems. An advantage of the modification is that it achieves the result of allowing for a confirmation selection by a user, in order to avoid incorrect selections and allow for more accurate selection by a user by including a popup window for an additional user confirmation to further perform a process.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Peter Hoang whose telephone number is (571)270-1346. The examiner can normally be reached Monday-Friday 8:00 am - 5:00 pm PST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Hajnik F. Daniel can be reached at (571) 272-7642. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PETER HOANG/ Primary Examiner, Art Unit 2616