Prosecution Insights
Last updated: October 02, 2026
Application No. 18/503,741

DEVICE, METHOD, AND PROGRAM FOR ENHANCING OUTPUT CONTENT THROUGH ITERATIVE GENERATION

Final Rejection §103
Filed
Nov 07, 2023
Priority
Dec 04, 2019 — RE 10-2019-0160008 +2 more
Examiner
AGAHI, DARIOUSH
Art Unit
2656
Tech Center
2600 — Communications
Assignee
Samsung Electronics Co., Ltd.
OA Round
4 (Final)
84%
Grant Probability
Favorable
5-6
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 84% — above average
84%
Career Allowance Rate
154 granted / 184 resolved
+21.7% vs TC avg
Strong +30% interview lift
Without
With
+30.1%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
21 currently pending
Career history
209
Total Applications
across all art units

Statute-Specific Performance

§101
23.7%
-16.3% vs TC avg
§103
55.0%
+15.0% vs TC avg
§102
10.7%
-29.3% vs TC avg
§112
6.7%
-33.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 184 resolved cases

Office Action

§103
DETAILED ACTION This office action is in response to Applicant’s submission filed on 7/20/2026. Claims 37, 49, and 60 were amended. Claims 37-75 are pending in the application of which Claims 37, 49, and 61 are independent and have been examined. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement(s)(IDS) submitted on 8/7/2026 has been considered by the examiner. Response to Arguments Applicant’s arguments filed in the Amendment filed 7/20/2026 (herein “Amendment”) with respect to the double patenting rejection has been fully considered and concur with the Applicant’s request with respect to the double patenting rejection to be held in abeyance until all substantive issues in the application are addressed. Applicants’ amendments with respect to the 35 USC §103 rejection raised in the previous office action have been fully considered but they are not persuasive. Applicant on pages 11 and 12 present arguments with references to MPEP§2141, 2142, 2143.01. In response to the said arguments and clarity of records the Examiner would like to make the following statements: One cannot state arguments against the references individually, in other words, one cannot show nonobviousness by attacking references individually where the rejections are based on combinations of references. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981); In re Merck & Co., 800 F.2d 1091, 231 USPQ 375 (Fed. Cir. 1986). The test for obviousness is not whether the features of a secondary reference may be bodily incorporated into the structure of the primary reference; nor is it that the claimed invention must be expressly suggested in any one or all of the references. Rather, the test is what the combined teachings of the references would have suggested to those of ordinary skill in the art. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981). MPEP 2141.01(a) Analogous and Non-analogous Art [R-01.2024] where it recites:” In order for a reference to be proper for use in an obviousness rejection under 35 U.S.C. 103, the reference must be analogous art to the claimed invention. In re Bigio, 381 F.3d 1320, 1325, 72 USPQ2d 1209, 1212 (Fed. Cir. 2004). A reference is analogous art to the claimed invention if: (1) the reference is from the same field of endeavor as the claimed invention (even if it addresses a different problem); or (2) the reference is reasonably pertinent to the problem faced by the inventor (even if it is not in the same field of endeavor as the claimed invention). Note that "same field of endeavor" and "reasonably pertinent" are two separate tests for establishing analogous art; it is not necessary for a reference to fulfill both tests in order to qualify as analogous art. See Bigio, 381 F.3d at 1325, 72 USPQ2d at 1212. .... When more than one prior art reference is used as the basis of an obviousness rejection, it is not required that the references be analogous art to each other. See Sanofi-Aventis Deutschland GMbH v. Mylan Pharms. Inc., 66 F.4th 1373, 1380, 2023 USPQ2d 552 (Fed. Cir. 2023) and Corephotonics, Ltd. v. Apple Inc., 84 F.4th 990, 1007, 2023 USPQ2d 1202 (Fed. Cir. 2023). Therefore, Examiner does not find the Applicant’s argument persuasive. Particularly, each and every limitation as presented in the previous Office action was mapped to teach a certain feature of the instant application. That does not mean that the entire body of the given prior art MUST be in line with the applicant’s disclosed invention. Also, the mapping/rejection is based on a nonobviousness (35 U.S.C. §103) rejection, which means the invention would have been obvious to a person skilled in the art at the time of invention. Therefore, one cannot attack references individually where the rejections are based on combinations of references. Applicant set forth on page 13 of the Amendment: PNG media_image1.png 328 378 media_image1.png Greyscale Applicant continues in page 14: PNG media_image2.png 304 582 media_image2.png Greyscale Examiner disagrees, and asserts the Cohen teaches the said limitation: Cohen, Par. 0095:” …The device directs the conversation by asking “What would you like to replace?”, to which the user answers “The boring sky”. The device again directs the conversation by narrowing the parameters of the replacement task, and asks “What would you like to replace the boring sky with?” “, and Par. 0096:” Continuing with the example directed user conversation 204 in FIG. 2, the user responds “A cloudy sky”, indicating to the device that the boring sky should be replaced by a cloudy sky. In response, the device generates harmonized image 206, which includes a cloudy sky.”, and Par. 0147:” … harmonized image 528 can be generated and displayed on a user interface of one of computing devices …”) As taught and recited above, Cohen switches the boring sky with a cloudy sky. Applicant continues on page 15 with a similar argument switching from prior content to the modified content, where Examiner still does not find the argument as persuasive since Cohen already teaches that in multiple figures, such as figure 2: (modifying boring sky with a cloudy sky), Figure 3: (removing woman in the image), Figure 5: (removing the fire hydrant). Pages 16 thru 19 are the depiction of Cohen’s PGPUB pages. Applicant on pages 20 and 21 state: PNG media_image3.png 192 586 media_image3.png Greyscale The Examiner does not find the argument persuasive. To remind the Applicant, this is a 103 rejection, and when the content is modified and a modified content is shown, it already teaches switching from a non-modified to a modified content. Applicants still argue the same point on pages 21& 22 which is moot per above discussions. Applicant on page 23 under the heading of the deficiencies of Cohen repeats the same points already discussed. Applicant on page 24 argues the simultaneous presentation of the before and after image modification. The Examiner likes to remind the Applicant that a 103 rejection one cannot attack references individually where the rejections are based on combinations of references. Applicant on page 26 argues: PNG media_image4.png 140 590 media_image4.png Greyscale Once again, Examiner like to stress the above limitation is covered by the combination of references. The rest of the argument with respect to Ohara and Gupta is already mentioned that each and every limitation as presented in the previous Office action was mapped to teach a certain feature of the instant application. That does not mean that the entire body of the given prior art MUST be in line with the applicant’s disclosed invention. Also, the mapping/rejection is based on a nonobviousness (35 U.S.C. §103) rejection, which means the invention would have been obvious to a person skilled in the art at the time of invention. Applicant further argues the rejection of claims 73-75 on pages 29-33. The Examiner did not find the argument persuasive since the main argument was centered around the alleged deficiencies of the primary references with respect to the independent claims. Applicant on page 30 and 31 states: PNG media_image5.png 248 580 media_image5.png Greyscale Applicant furthers on page 32: PNG media_image6.png 136 582 media_image6.png Greyscale The Examiner disagrees since the claim language is properly mapped by Gaash, where he teaches Gaash, Par. 0036:” Block 214 comprises selecting an image attribute modification from the set of image attribute modifications determined in block 210. Block 216 comprises applying the selected image attribute modification to the jth copy of the selected seed image xi. As indicated by the looping arrow, blocks 214 and 216 may be carried out for each image attribute modification. Once all image attribute modifications in the set have been applied to the copy of the appropriate seed image, the modified image is defined. Generating the modified image therefore comprises, at block 218, applying each image attribute modification in the set to the copy of the appropriate seed image.”) Claim language is mapped once the first attribute is carried out. Therefore, while all of the Applicant’s arguments filed in the Amendment have been fully considered, they are not persuasive. Please see below for more detail including updated citations and obviousness rationale. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 37-72 are rejected under 35 U.S.C. 103 as being unpatentable over Cohen (US20190196698A1), and in further view of Gupta et al. (US20200250453A1)(herein "Gupta"). Chen and Gupta were applied in the previous Office Action. Regarding claims 37, 49, and 61, Cohen teaches [One or more non-transitory computer-readable storage media storing one or more computer programs including computer-executable instructions that, when executed by one or more processors of an electronic device individually or collectively, cause the electronic device to perform operations, the operations comprising: - claim 37], [A method performed by an electronic device for modifying content, the method comprising: - claim 49], and [An electronic device for modifying content, the electronic device comprising: memory, comprising one or more storage media, storing instructions; and one or more processors communicatively coupled to the memory, wherein the instructions, when executed by the one or more processors individually or collectively, cause the electronic device to: - claim 61] (Cohen, Par. 0042:” Image enhancement system 110 also includes processors 124. Hence, image enhancement system 110 may be implemented at least partially by executing instructions stored on storage 126 on processors 124. For instance, processors 124 may execute portions of image enhancement application 120.”, and Par. 0184:” Computer-readable storage media 806 is illustrated as including memory/storage 812. Storage 126 in FIG. 1 is an example of memory/storage included in memory/storage 812.”, and Par. 0189:” “Computer-readable storage media” refers to media, devices, or combinations thereof that enable persistent or non-transitory storage of information in contrast to mere signal transmission, ...”) present a base content; (Cohen, Figure 2, image 202; Fig. 5, image 510; Par. 0094:” … In one example, directed user conversation 204 is initiated in response to the image to be edited 202 being flagged for editing, such as by a user selecting the image to be edited 202 for editing, the image to be edited 202 being loaded into an image editing application (e.g., image enhancement application 120), and the like.”) presenting an indication, on the base content being presented, corresponding to the target area; (Cohen, Fig. 5, Par. 0142:” Indicator 512 can be any suitable indicator [indication], such as a lasso, circle, shading, pattern, mask, overlay, arrow, proximate text, and the like. In the example in FIG. 5, indicator 512 includes a dashed outline [indication] encompassing the fire hydrant [target area] together with the dog's head.”, and Par. 0143:” … In the example in FIG. 5, the user selects indicator 512, (e.g., by pointing with a mouse and clicking a mouse button), which is denoted by a hand representation 514. The user also moves the indicator 512 to adjust content that it indicates (e.g., by holding a mouse button down and moving or adjusting indicator 512).”) while the base content is being presented, receiving a natural language input, including at least one attribute information, for generating output content; and (Cohen, Par. 0143:” … Any suitable selection with a tool in user interface 500 can be used to provide multi-modal user input, such as user clicking on a center of an object (e.g., clicking on the center of the fire hydrant in intermediate image 510) while speaking (e.g., “No, you've selected the dog's head, too. This is the fire hydrant” in the directed user conversation of representation 504) represents a multi-modal user input. “, and Par. 0170:” … Confirming that the candidate object image matches the object can include receiving a multi-modal user input to correct the candidate object. One of the modes can be speech, and another mode can be input from a keyboard, mouse, stylus, gesture, or touchscreen, and the like.”, and Par. 0095:” …The device directs the conversation by asking “What would you like to replace?”, to which the user answers “The boring sky”. The device again directs the conversation by narrowing the parameters of the replacement task, and asks “What would you like to replace the boring sky with?” “, and Par. 0096:” Continuing with the example directed user conversation 204 in FIG. 2, the user responds “A cloudy sky”, indicating to the device that the boring sky should be replaced by a cloudy sky. In response, the device generates harmonized image 206, which includes a cloudy sky.”) Note: boring and cloudy are the attribute of the sky. switching, from presenting the base content with the indication corresponding to the target area being presented thereon, to [[presenting, i) instead of, and without presenting the base content with the indication corresponding to the target area being presented thereon and ii) in a same region, of an entire display area, which the base content with the indication corresponding to the target area being presented thereon occupies when being presented before the switching,]] modified base content in which the base content is modified to include the output content, in the target area, that is generated based on a detected object in the target area and as having at least one same attribute as the at least one attribute information included in the natural language input, (Cohen, Par. 0143:” … Any suitable selection with a tool in user interface 500 can be used to provide multi-modal user input, such as user clicking on a center of an object (e.g., clicking on the center of the fire hydrant in intermediate image 510) while speaking (e.g., “No, you've selected the dog's head, too. This is the fire hydrant” in the directed user conversation of representation 504) represents a multi-modal user input. “, and Par. 0170:” … Confirming that the candidate object image matches the object can include receiving a multi-modal user input to correct the candidate object. One of the modes can be speech, and another mode can be input from a keyboard, mouse, stylus, gesture, or touchscreen, and the like.”, and Par. 0095:” …The device directs the conversation by asking “What would you like to replace?”, to which the user answers “The boring sky”. The device again directs the conversation by narrowing the parameters of the replacement task, and asks “What would you like to replace the boring sky with?” “, and Par. 0096:” Continuing with the example directed user conversation 204 in FIG. 2, the user responds “A cloudy sky”, indicating to the device that the boring sky should be replaced by a cloudy sky. In response, the device generates harmonized image 206, which includes a cloudy sky.”, and Par. 0147:” … harmonized image 528 can be generated and displayed on a user interface of one of computing devices …”) Note: boring and cloudy are the attribute of the sky. Furthermore, once the harmonized image is displayed it reads on switching from base content to the target area where the enhancement/modification is conducted. Also, Cohen Fig. 5 depict fire hydrant with the arrow (indication) and subsequently switched to display Panel 526 which shows the modified image. wherein the output content is generated by using at least one artificial intelligence (Al) model, and (Cohen, Par. 0131:” … Additionally or alternatively, harmonizing module 154 may harmonize a composite image with a neural network trained specifically for the type of object removed from or replaced in an image to be edited used to form the composite image, a background of the image to be edited, or combinations thereof. For instance, harmonizing module 154 can use a neural network trained to harmonize persons in a beach scene when removing or replacing a person in an image with a beach scene. “, and Par. 0005:” … In one example, a vision module specific to the object is used, such as using a sky vision module including a neural network trained to identify skies when satisfying the replace request “Replace the boring sky with a cloudy sky”. … In another example, intermediate results are exposed to the user, and multi-modal input is received during a directed user conversation.”) Cohen, does not teach, however, Gupta teaches while the base content is being presented without a target area of the base content having been selected, receiving a user input, related to one or more coordinates, to the base content for selecting the target area of the base content based on the user input related to the one or more coordinates; Gupta, Par. 0045:” … an image editing program can include a content-aware selection system. The content-aware selection system can enable a user to select an area of an image using a label or a tag that identifies object in the image, ... For example, for an image that includes a dog and a cat, the content-aware selection system can enable a user to input the label “dog,” upon which the content-aware selection system will generate a selection area around the pixels that represent the dog. As a further example, the system can enable the user to input the label “animals,” which will generate a selection area including the pixels for both the dog and the cat.”, and Par. 0046:” … Instead of having to draw a selection boundary around an object, or painting over the area that contains the object, users can click or tap on the object, and the content-aware selection system can automatically draw a selection area around the object. The content-aware selection system may be particularly useful when an image editing program supports voice input. With voice input, the user can speak a phrase such as “select the dog,” and the content-aware selection system will generate a selection area around the dog, without the user needing to provide any physical input. “) Note: per as-filed Spec. Par. 0060:” … The user input may be related to one or more coordinates, but is not limited thereto. For example, the user input may be an audio input, voice input, text input, or a combination thereof. An input related to a coordinate may be a touch input, click input, gesture input, etc.” wherein the user input related to the one or more coordinates is different than the natural language input. (Gupta, Par. 0046:” … Instead of having to draw a selection boundary around an object, or painting over the area that contains the object, users can click or tap on the object, and the content-aware selection system can automatically draw a selection area around the object. “)Note: per as-filed Spec. Par. 0060:” … The user input may be related to one or more coordinates, but is not limited thereto. For example, the user input may be an audio input, voice input, text input, or a combination thereof. An input related to a coordinate may be a touch input, click input, gesture input, etc.” Gupta is considered to be analogous to the claimed invention because it is in the same field of endeavor. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Cohen further in view of Gupta to while the base content is being presented without a target area of the base content having been selected, receiving a user input, related to one or more coordinates, to the base content for selecting the target area of the base content based on the user input related to the one or more coordinates, wherein the user input related to the one or more coordinates is different than the natural language input. Motivation to do so would improve the image editing process, in terms of speed and accuracy (Gupta, Par. 0046). Cohen, as modified above, does not teach, however, Ohara teaches [[switching, from presenting the base content with the indication corresponding to the target area being presented thereon, to ]]presenting, i) instead of, and without presenting the base content with the indication corresponding to the target area being presented thereon and ii) in a same region, of an entire display area, which the base content with the indication corresponding to the target area being presented thereon occupies when being presented before the switching,[[ modified base content in which the base content is modified to include the output content, in the target area, that is generated based on a detected object in the target area and as having at least one same attribute as the at least one attribute information included in the natural language input,]] (Ohara, Col. 20, ll. 58-65:” In the controller 40, by supplying the image data before correction and the image data after correction, each for one field, being made as a pair, to the display control section 55, through displaying both the radiation image based on the image data before correction and the radiation image based on the image data after correction simultaneously on the image surface of the image display device 56 as shown in FIG. 14A, …”). PNG media_image7.png 476 546 media_image7.png Greyscale Ohara is considered to be analogous to the claimed invention because it is in the same field of endeavor. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Cohen, as modified above, further in view of Ohara to present, instead of the base content with the indication corresponding to the target area being presented thereon and in a same display area in which the base content with the indication corresponding to the target area being presented thereon was presented. The motivation to so would provide instant, accurate, and detailed visual comparison of the changes and provide verification of the desired changes. Regarding claims 38, 50, and 62, Cohen, as modified above, teaches the media, the method, and the electronic device of claims 37, 49, and 61 respectively. Cohen, as modified above, further teaches detecting the object, wherein the object is detected by using the at least one Al model. (Cohen, Par. 0021:” … For instance, a sky vision module including a neural network trained to identify [detect] skies [object] is used to ascertain pixels of a sky in an image when an object to be replaced in the image is identified as a sky, such as for the replace request “Replace the boring sky with a cloudy sky”.) Regarding claims 39, 51, and 63, Cohen, as modified above, teaches the media, the method, and the electronic device of claims 37, 49, and 61 respectively. Cohen, as modified above, further teaches wherein the base content and the modified base content are images. (Cohen, Fig. 2, Par. 0021:” … For instance, a sky vision module including a neural network trained to identify skies [images] is used to ascertain pixels of a sky in an image when an object to be replaced in the image is identified as a sky, such as for the replace request “Replace the boring sky with a cloudy sky”.) Note: both boring sky and cloudy sky are images. Regarding claims 40, 52, and 64, Cohen, as modified above, teaches the media, the method, and the electronic device of claims 37, 49, and 61 respectively. Cohen, as modified above, further teaches wherein the target area is less than an entire area of the base content. (Cohen, Fig. 2, Par. 0021:” … a sky in an image when an object to be replaced in the image is identified as a sky, such as for the replace request “Replace the boring sky with a cloudy sky”.) Note: Sky is less than the entire area of the base content. Regarding claims 41, 53, and 65, Cohen, as modified above, teaches the media, the method, and the electronic device of claims 37, 49, and 61 respectively. Cohen, as modified above, further teaches wherein at least one of a size or a shape of the target area is user adjustable. (Cohen, Par. 0129:” … For instance, compositing module 152 can extract fill material or replacement material from an image obtained by image search module 150, filter the material (e.g., adjust color, brightness, contrast, apply a filter, and the like), re-size the material (e.g., interpolate between pixels of the material, decimate pixels of the material, or both, to stretch or squash the material), rotate the material, crop the material, composite the material with itself or other fill or replacement material, and the like.”, and Par. 0133:” … In one example, a user may adjust a border of a background segmentation generated by vision module 146, such as by moving a water line separating a beach and ocean that defines a background scene of an image …”). Regarding claims 42, 54, and 66, Cohen, as modified above, teaches the media, the method, and the electronic device of claims 37, 49, and 61 respectively. Cohen, as modified above, further teaches wherein the natural language input includes at least one of a voice input or a text input, and wherein in case the natural language input includes the voice input, the voice input is converted into text by using an automatic speech recognition (ASR) model. (Cohn, Par. 0027:" … In one example, computing devices 104 include speech recognition, identification, and synthesis functionalities, microphones, and speakers that allow computing devices 104 to communicate with user 102 in a conversation, e.g., a directed user conversation.", and Par. 0048:” … A user conversation can include any suitable type of communication, such as verbal communication (e.g., with microphones and speakers of conversation module 144), written communication (e.g., a user may type into a keyboard or provide a document to conversation module 144), or combinations of verbal communication and written communication.”, and Par. 0056:” … In one example, an editing query includes a transcript of a directed user conversation (e.g., text in ASCII format).”, and Par. 0143:” … Any suitable selection with a tool in user interface 500 can be used to provide multi-modal user input, such as user clicking on a center of an object (e.g., clicking on the center of the fire hydrant in intermediate image 510) while speaking (e.g., “No, you've selected the dog's head, too.”) Regarding claims 43, 55, and 67, Cohen, as modified above, teaches the media, the method, and the electronic device of claims 37, 49, and 61 respectively. Cohen, as modified above, further teaches at least one of: wherein the output content is generated based on the base content, or wherein the output content is generated to match the base content. (Cohen, Par. 0005:” … Based on the directed user conversation indicating a remove request or replace request, an object is removed and fill material is added [output content] in its place, or an object is replaced with replacement material, to produce a plurality of composite images that are harmonized to make the editing appear natural.”) Note: harmonized output maps to the output content are generated based on the base content or output content is generated to match the base content. Regarding claims 44, 56, and 68, Cohen, as modified above, teaches the media, the method, and the electronic device of claims 37, 49, and 61 respectively. Cohen, as modified above, further teaches at least one of: wherein the output content is generated based on compositing content into the target area of the base content, or wherein the base content is modified based on compositing the output content into the base content. (Cohen, Par. 0022:” … The composite images are harmonized to make them look natural (e.g., so that the editing is not easily detected). In one example, harmonizing includes adjusting lighting of a composite image to match times of day between image materials.”) Note: when lighting of a composite image is adjusted, implies base content is also modified. Regarding claims 45, 57, and 69, Cohen, as modified above, teaches the media, the method, and the electronic device of claims 37, 49, and 61 respectively. Cohen, as modified above, further teaches wherein the output content is generated based on an object corresponding to content information included in the natural language input, or wherein the output content is generated to include the detected object in the target area with at least one attribute thereof changed so as to have the at least one same attribute as the at least one attribute information included in the natural language input. (Cohen, Fig. 2, Par. 0021:” … For instance, a sky vision module including a neural network trained to identify skies is used to ascertain pixels of a sky in an image when an object to be replaced in the image is identified as a sky, such as for the replace request “Replace the boring sky with a cloudy sky”.) Note: content information for the output is the “cloudy sky” and the output has cloudy sky. Regarding claims 46, 58, and 70, Cohen, as modified above, teaches the media, the method, and the electronic device of claims 37, 49, and 61 respectively. Cohen, as modified above, further teaches wherein the output content is generated as having at least one same attribute as the detected object in the target area. (Cohen, Par. 0146:” As a result of a user selecting an image in images panel 518, an image in display panel 526 is exposed. … In one example, display panel 526 is displayed in user interface 500 responsive to a user selection of one of the harmonized images 520 displayed in images panel 518.”) Note: As depicted in Fig. 5, all of the output content has the same dog attribute, as the dog in the detected object in the target area of 506. As noted only the background is being replace but the dog stayed the same between the target area and the output. Regarding claims 47, 59, and 71, Cohen, as modified above, teaches the media, the method, and the electronic device of claims 37, 49, and 61 respectively. Cohen, as modified above, further teaches wherein the at least one attribute information included in the natural language input is obtained from the natural language input by using the at least one AI model. (Cohn, Par. 0027:" … In one example, computing devices 104 include speech recognition, identification, and synthesis functionalities, microphones, and speakers that allow computing devices 104 to communicate with user 102 in a conversation, e.g., a directed user conversation.", and Par. 0095:” …The device directs the conversation by asking “What would you like to replace?”, to which the user answers “The boring sky”. The device again directs the conversation by narrowing the parameters of the replacement task, and asks “What would you like to replace the boring sky with?” “, and Par. 0096:” Continuing with the example directed user conversation 204 in FIG. 2, the user responds “A cloudy sky”, indicating to the device that the boring sky should be replaced by a cloudy sky. In response, the device generates harmonized image 206, which includes a cloudy sky.”) Note: boring and cloudy are the attribute of the sky. Also, user is conversing thru a speech recognition module, which a speech recognizer is considered an AI model. Regarding claims 48, 60, and 72, Cohen, as modified above, teaches the media, the method, and the electronic device of claims 37, 49, and 61 respectively. Cohen, as modified above, further teaches presenting a user interface for selecting among a plurality of contents that each correspond to a different version of the output content. (Cohen, Par. 0005:” … Based on the directed user conversation indicating a remove request or replace request, an object is removed and fill material is added in its place, or an object is replaced with replacement material, to produce a plurality of composite images that are harmonized to make the editing appear natural. Multiple harmonized images are exposed in a user interface. Thus, a user is presented a plurality of options (e.g., harmonized images with different versions of fill material or replacement material) that satisfy the editing query based on a directed user conversation. In one example, the plurality of images are presented to the user automatically and without user intervention once a directed user conversation is completed and an image to be edited is obtained. In another example, intermediate results are exposed to the user, and multi-modal input is received during a directed user conversation.”) Claims 73-75 are rejected under 35 U.S.C. 103 as being unpatentable over Cohen, Gupta, and Ohara, and in further view of Gaash (US 20210334612 A1). Gaash was applied in the previous Office Action. Regarding claims 73-75, Cohen, as modified above, teaches the media, the method, and the electronic device of claims 37, 49, and 61 respectively. Cohen, as modified above, further teaches included in the natural language input (Cohen, Par. 0017:” … user input during a directed user conversation in addition to speech input during the directed user conversation.”, and Par. 0018:” … directing a user conversation to obtain an editing query, and providing a plurality of images that have been enhanced by fulfilling a remove request or a replace request with different content, … based on the editing query. . Received user responses are processed to determine parameters of an editing query, such as whether the user conversation indicates a remove request or replace request, objects to be removed, objects to be replaced, objects to replace objects, modifiers of objects, …”, and Par. 0030:” … user conversation 108 includes an editing query for the image to be edited 106, such as “Replace the rainy background with a sunny day” …”) Cohen, as modified above, does not teach, however, Gaash teaches wherein the at least one attribute information [[included in the natural language input]] is a plurality of attribute information included in the natural language input, and (Gaash, Par. 0016:” … how to apply a particular image attribute modification to a seed image, thereby generating a modified image, and a rule relating to the placement of the modified image in a collage (or print) area.”, and Par. 0034:” … determining a set of image attribute modifications. … The set of image attribute modifications may therefore be the same for each individual seed image, ... Accordingly, any single seed may have the same set of image attribute modifications applied to it differently, the application of which is determined by the unique identifier selected in block 204. The set of images attribute modifications may comprise a single image attribute modification, or a plurality of image attribute modifications. … an output may be used to indicate a set of image attribute modifications to be applied.”) wherein the output content is generated based on the detected object in the target area and as having at least one same attribute as only less than all of the plurality of attribute information [[included in the natural language input]]. (Gaash, Par. 0036:” Block 214 comprises selecting an image attribute modification from the set of image attribute modifications determined in block 210. Block 216 comprises applying the selected image attribute modification to the jth copy of the selected seed image xi. As indicated by the looping arrow, blocks 214 and 216 may be carried out for each image attribute modification. Once all image attribute modifications in the set have been applied to the copy of the appropriate seed image, the modified image is defined. Generating the modified image therefore comprises, at block 218, applying each image attribute modification in the set to the copy of the appropriate seed image.”) Note: As an image attribute is applied to the image, once one attribute is implemented, the claim language is fulfilled. Gaash is considered to be analogous to the claimed invention because it is in the same field of endeavor. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Cohen, as modified above, further in view of Gaash to wherein the at least one attribute information is a plurality of attribute information included in the natural language input, and wherein the output content is generated based on the detected object in the target area and as having at least one same attribute as only less than all of the plurality of attribute information. Motivation to do so would provide multiple images that have a degree of consistency or commonality (Gaash, Par. 0008). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. Zhang et al. (Text-to-Image Synthesis via Visual-Memory Creative Adversarial Network”, Advances in Multimedia Information Processing–PCM 2018: 19th Pacific-Rim …, 2018) teaches in ABS:” … propose a method named visual-memory Creative Adversarial Network (vmCAN) to generate images depending on their corresponding narrative sentences.”, and Section 3.1:” Region Proposal Network ranks and refines region boxes called anchors to generate high-quality region proposals which most likely contain an object.”, and Section 2.2:” … method to edit a given image with specific textual description. … proposed vmCAN attempts to synthesize images conditioned on the textual description and multiple relevant sub-images, …” Examiner's Note: Examiner has cited particular columns and line numbers and/or paragraph numbers in the references applied to the claims above for the convenience of the applicant. Although the specified citations are representative of the teachings of the art and are applied to specific limitations within the individual claim, other passages and figures may apply as well. It is respectfully requested from the applicant in preparing responses, to fully consider the references in entirety as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior art or disclosed by the Examiner. In the case of amending the Claimed invention, Applicant is respectfully requested to indicate the portion(s) of the specification which dictate(s) the structure relied on for proper interpretation and also to verify and ascertain the metes and bounds of the claimed invention. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to DARIOUSH AGAHI, P.E. whose telephone number is (408)918-7689. The examiner can normally be reached Monday - Thursday and alternate Fridays, 7:30-4:30 PT. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bhavesh Mehta can be reached at 571-272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. DARIOUSH AGAHI, P.E. Primary Examiner /DARIOUSH AGAHI/Primary Examiner, Art Unit 2656
Read full office action

Prosecution Timeline

Show 4 earlier events
Nov 10, 2025
Applicant Interview (Telephonic)
Dec 08, 2025
Response Filed
Jan 15, 2026
Final Rejection mailed — §103
Feb 26, 2026
Request for Continued Examination
Feb 27, 2026
Response after Non-Final Action
May 07, 2026
Non-Final Rejection mailed — §103
Jul 20, 2026
Response Filed
Sep 14, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749478
ENHANCED SPEECH-TO-TEXT PERFORMANCE WITH MIXED LANGUAGES
1y 12m to grant Granted Sep 29, 2026
Patent 12718030
LARGE LANGUAGE MODELS PROVIDING EVIDENCE MAPPINGS FOR GENERATED OUTPUT
3y 1m to grant Granted Aug 25, 2026
Patent 12718017
SYSTEM AND METHOD FOR PROVIDING LARGE LANGUAGE MODEL FOR SANCTIONS ARTIFICIAL INTELLIGENCE ASSISTED AUTOMATION
2y 11m to grant Granted Aug 25, 2026
Patent 12718819
INTERRUPTION DETECTION AND HANDLING BY DIGITAL ASSISTANTS
2y 3m to grant Granted Aug 25, 2026
Patent 12710915
VOICE MODIFICATION FOR WEARABLE DEVICE
3y 9m to grant Granted Aug 18, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
84%
Grant Probability
99%
With Interview (+30.1%)
2y 7m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 184 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month