DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed in parent Application No. CN202211098607.9, filed on 09/06/2022.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 01/27/2025 has been considered by the examiner.
Preliminary Amendment
The Preliminary Amendment submitted on 12/20/2024 has been entered and made of record.
Status of Claims
Currently pending Claim(s):
Amended claim(s):
Canceled claim(s):
New Claim(s):
1-11 and 12-21
5, 7, and 13-14
12
15-21
Claim Objections
Claims 1, 3, 13, 14, 16, and 21 are objected to because of the following informalities:
Claim 1 lines 7-8 the phrase “obtain a second image after inpainted.” is grammatically unclear, the Examiner suggests amending the claim to recite “obtain a second image after inpainting”.
Claim 3 line 2 the phrase “obtain a second image after inpainted.” is grammatically unclear, the Examiner suggests amending the claim to recite “obtain a second image after inpainting”.
Claim 13 lines 8-9 the phrase “obtain a second image after inpainted.” is grammatically unclear, the Examiner suggests amending the claim to recite “obtain a second image after inpainting”.
Claim 14 lines 10-11 the phrase “obtain a second image after inpainted.” is grammatically unclear, the Examiner suggests amending the claim to recite “obtain a second image after inpainting”.
Claim 16 line 2 the phrase “obtain a second image after inpainted.” is grammatically unclear, the Examiner suggests amending the claim to recite “obtain a second image after inpainting”.
Claim 21 lines 2-3 the phrase “obtain a second image after inpainted.” is grammatically unclear, the Examiner suggests amending the claim to recite “obtain a second image after inpainting”.
Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-2, 13, and 14-15 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The limitations, under their broadest reasonable interpretation, cover mental process (concept performed in a human mind, including as observation, evaluation, judgment, opinion, organizing human activity). The independent claims 1, 13, and 14 recite a method, a non-transitory computer-readable storage medium, and a device. This judicial exception is not integrated into a practical application because the steps do not add meaningful limitations to be considered specifically applied to a particular technological problem to be solved .The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the steps of the claimed invention can be done mentally and no additional features in the claims would preclude them from being performed as such except for the generic computer elements at high level of generality (i.e., processor, memory).
According to the USPTO guidelines, a claim is directed to non-statutory subject matter if:
STEP 1: the claim does not fall within one of the four statutory categories of invention (process, machine, manufacture or composition of matter), or
STEP 2: the claim recites a judicial exception, e.g. an abstract idea, without reciting additional elements that amount to significantly more than the judicial exception, as determined using the following analysis:
STEP 2A (PRONG 1): Does the claim recite an abstract idea, law of nature, or natural phenomenon?
STEP 2A (PRONG 2): Does the claim recite additional elements that integrate the judicial exception into a practical application?
STEP 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception?
Using the two-step inquiry, it is clear that the independent claims 1, 13, and 14 are directed to an abstract idea as shown below:
STEP 1: Do the claims fall within one of the statutory categories? YES. Independent claims 1, 13, and 14 are directed to a process, manufacture, and machine
STEP 2A (PRONG 1): Is the claim directed to a law of nature, a natural phenomenon or an abstract idea? YES, the claims are directed toward a mental process (i.e. abstract idea).
With regard to STEP 2A (PRONG 1), the guidelines provide three groupings of subject matter that are considered abstract ideas:
Mathematical concepts – mathematical relationships, mathematical formulas or equations, mathematical calculations;
Certain methods of organizing human activity – fundamental economic principles or practices (including hedging, insurance, mitigating risk); commercial or legal interactions (including agreements in the form of contracts; legal obligations; advertising, marketing or sales activities or behaviors; business relations); managing personal behavior or relationships or interactions between people (including social activities, teaching, and following rules or instructions); and
Mental processes – concepts that are practicably performed in the human mind (including an observation, evaluation, judgment, opinion).
Independent claims 1, 13, and 14 comprise a mental process that can be practicably performed in the human mind (or generic computers or components configured to perform the method) and, therefore, an abstract idea.
Regarding independent claim(s) 1: the limitations recite:
An image inpainting method, comprising:
acquiring a first image, wherein the first image is obtained by processing a target object in an original image (data gathering);
determining a first area to be inpainted in the first image, wherein the first area is at least a partial area of the target object (mental process including observation and evaluation, and can be done mentally in the human mind);
acquiring a target semantic graph corresponding to the first image (data gathering); and
inpainting the first area based on the target semantic graph to obtain a second image after inpainted (extra solution activity).
Regarding independent claim(s) 13: the limitations recite:
A non-transitory computer-readable storage medium storing instructions that cause a processor to (generic computer component):
acquire a first image, wherein the first image is obtained by processing a target object in an original image (data gathering);
determine a first area to be inpainted in the first image, wherein the first area is at least a partial area of the target object (mental process including observation and evaluation, and can be done mentally in the human mind);
acquire a target semantic graph corresponding to the first image (data gathering); and
inpaint the first area based on the target semantic graph to obtain a second image after inpainted (extra solution activity).
Regarding independent claim(s) 14: the limitations recite:
An electronic device, comprising a processor and a non-transitory memory with instructions thereon (generic computer component),
wherein the instructions upon execution by the processor, cause the processor to (generic computer component):
acquire a first image, wherein the first image is obtained by processing a target object in an original image (data gathering);
determine a first area to be inpainted in the first image, wherein the first area is at least a partial area of the target object (mental process including observation and evaluation, and can be done mentally in the human mind);
acquire a target semantic graph corresponding to the first image (data gathering); and
inpaint the first area based on the target semantic graph to obtain a second image after inpainted (extra solution activity).
These limitations, as drafted, is a simple process that, under their broadest reasonable interpretation, covers performance of the limitations in the mind or by a human. The Examiner notes that under MPEP 2106.04(a)(2)(III), the courts consider a mental process (thinking) that “can be performed in the human mind, or by a human using a pen and paper" to be an abstract idea. CyberSource Corp. v. Retail Decisions, Inc., 654 F.3d 1366, 1372, 99 USPQ2d 1690, 1695 (Fed. Cir. 2011). As the Federal Circuit explained, "methods which can be performed mentally, or which are the equivalent of human mental work, are unpatentable abstract ideas the ‘basic tools of scientific and technological work’ that are open to all.’" 654 F.3d at 1371, 99 USPQ2d at 1694 (citing Gottschalk v. Benson, 409 U.S. 63, 175 USPQ 673 (1972)). See also Mayo Collaborative Servs. v. Prometheus Labs. Inc., 566 U.S. 66, 71, 101 USPQ2d 1961, 1965 ("‘[M]ental processes[] and abstract intellectual concepts are not patentable, as they are the basic tools of scientific and technological work’" (quoting Benson, 409 U.S. at 67, 175 USPQ at 675)); Parker v. Flook, 437 U.S. 584, 589, 198 USPQ 193, 197 (1978) (same).
As such, a person could mentally determine a first area to be inpainted in the image based on the target object. The mere nominal recitation that the various steps are being executed by the generic computer component(s), for example, a non-transitory computer-readable storage medium, a processor, device, etc does not take the limitations out of the mental process grouping. Thus, the claims recite a mental process.
STEP 2A (PRONG 2): Does the claim recite additional elements that integrate the judicial exception into a practical application? NO, the claims do not recite additional elements that integrate the judicial exception into a practical application.
With regard to STEP 2A (prong 2), whether the claim recites additional elements that integrate the judicial exception into a practical application, the guidelines provide the following exemplary considerations that are indicative that an additional element (or combination of elements) may have integrated the judicial exception into a practical application:
an additional element reflects an improvement in the functioning of a computer, or an improvement to other technology or technical field;
an additional element that applies or uses a judicial exception to affect a particular treatment or prophylaxis for a disease or medical condition;
an additional element implements a judicial exception with, or uses a judicial exception in conjunction with, a particular machine or manufacture that is integral to the claim;
an additional element effects a transformation or reduction of a particular article to a different state or thing; and
an additional element applies or uses the judicial exception in some other meaningful way beyond generally linking the use of the judicial exception to a particular technological environment, such that the claim as a whole is more than a drafting effort designed to monopolize the exception.
While the guidelines further state that the exemplary considerations are not an exhaustive list and that there may be other examples of integrating the exception into a practical application, the guidelines also list examples in which a judicial exception has not been integrated into a practical application:
an additional element merely recites the words “apply it” (or an equivalent) with the judicial exception, or merely includes instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea;
an additional element adds insignificant extra-solution activity to the judicial exception; and
an additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use.
Independent claims 1, 13, and 14 do not recite any of the exemplary considerations that are indicative of an abstract idea having been integrated into a practical application. Independent claims 1, 13, and 14 discloses an a non-transitory computer-readable storage medium, a processor, device, and inpainting the first area based on the target semantic graph to obtain a second image after inpainted, which are generic computer components and/or insignificant pre/post-solution extra activity that do not add a meaningful limitation to the abstract idea because they amount to simply implementing the abstract idea in a method.
These limitations are recited at a high level of generality (i.e. as a general action or change being taken based on the results of the acquiring step) and amounts to mere post solution actions, which is a form of insignificant extra-solution activity. Further, the claims are claimed generically and are operating in their ordinary capacity such that they do not use the judicial exception in a manner that imposes a meaningful limit on the judicial exception. Accordingly, even in combination, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea.
STEP 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception? No, the claims do not recite additional elements that amount to significantly more than the judicial exception.
With regard to STEP 2B, whether the claims recite additional elements that provide significantly more than the recited judicial exception, the guidelines specify that the pre-guideline procedure is still in effect. Specifically, that examiners should continue to consider whether an additional element or combination of elements:
adds a specific limitation or combination of limitations that are not well-understood, routine, conventional activity in the field, which is indicative that an inventive concept may be present; or
simply appends well-understood, routine, conventional activities previously known to the industry, specified at a high level of generality, to the judicial exception, which is indicative that an inventive concept may not be present.
Independent claim(s) 1, 13, and 14 do not recite any additional elements that are not well-understood, routine or conventional. The use of a generic computer elements are routine, well-understood and conventional process that is performed by computers.
Thus, since independent claims 1, 13, and 14 are: (a) directed toward an abstract idea, (b) do not recite additional elements that integrate the judicial exception into a practical application, and (c) do not recite additional elements that amount to significantly more than the judicial exception, it is clear that independent claims 1, 13, and 14 are not eligible subject matter under 35 U.S.C 101.
Regarding claims 2 and 15: the additional elements do not integrate the mental process into practical application or add significantly more to the mental process.
In detail claim 2 depends on claim 1, and claim 15 depends on claim 14 and add:
removing the target object (claims 2 and 15).
Regarding claim 3: the additional limitations do integrate the mental process into practical application or add significantly more to the mental process. The limitation: “acquiring a first feature graph corresponding to the first image; regenerating features corresponding to the first area by features corresponding to a second area in the first feature graph based on the target semantic graph, so as to obtain a second feature graph, wherein the second area is an area except the first area in the first image; and acquiring the second image based on the second feature graph." integrates the mental process into a practical application.
Regarding dependent claims 4-11: claims 4-11 are similarly eligible under 35 USC 101 due to their dependency on claim 3.
Regarding claim 16: the additional limitations do integrate the mental process into practical application or add significantly more to the mental process. The limitation: “acquiring a first feature graph corresponding to the first image; regenerating features corresponding to the first area by features corresponding to a second area in the first feature graph based on the target semantic graph, so as to obtain a second feature graph, wherein the second area is an area except the first area in the first image; and acquiring the second image based on the second feature graph." integrates the mental process into a practical application.
Regarding dependent claims 17-20: claims 17-20 are similarly eligible under 35 USC 101 due to their dependency on claim 16.
Regarding claim 21: the additional limitations do integrate the mental process into practical application or add significantly more to the mental process. The limitation: “acquiring a first feature graph corresponding to the first image; regenerating features corresponding to the first area by features corresponding to a second area in the first feature graph based on the target semantic graph, so as to obtain a second feature graph, wherein the second area is an area except the first area in the first image; and acquiring the second image based on the second feature graph." integrates the mental process into a practical application.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-8 and 13-21 are rejected under 35 U.S.C. 103 as being unpatentable over Li et al. (US 2021/0158491 A1) (hereinafter, “Li”) in view of Dhamo et al. (Dhamo, Helisa, et al. "Semantic image manipulation using scene graphs." 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2020.) (hereinafter; “Dhamo”).
Regarding claim 1, Li discloses an image inpainting method (Abstract “An inpainting method includes retrieving image information at an electronic device, where the image information identifies an area within an image.”), comprising:
acquiring a first image, wherein the first image (input image in Paragraph [0058] equates to first image) is obtained by processing a target object in an original image (Paragraph [0058] “The input image 208 represents an image in which one or more portions of the image (in this case a circular area) are being removed. Each portion of the input image 208 being removed is often referred to as a “hole.””; Note that portion of the image being removed equates to processing a target object);
determining a first area to be inpainted in the first image, wherein the first area is at least a partial area of the target object (unwanted person or other object in Paragraph [0128] equates to target object) (Paragraph [0128] “an image 1402 is being presented on an electronic device, and a user has used an electronic pen 1404 or his or her finger to define an area 1406 of the image 1402 to be removed. The area 1406 may, for instance, include an unwanted person or other object in the image 1402.”);
acquiring a target semantic [graph] corresponding to the first image (Paragraph [0059] “The filled semantic mask 212 represents an initial estimation of the semantic classes associated with the input image 208, including labels for one or more semantic classes estimated for each hole in which content is being removed in the input image 208.”); and
inpainting the first area based on the target semantic [graph] to obtain a second image (final output image in Paragraph [0130] equates to second image) after inpainted (Paragraph [0130] “ the electronic device has used the distribution of the semantic class or classes 1408 in the defined area 1406 to generate a final output image 1414. As can be seen here, the area 1406 has been filled with replacement content of both the “ground” and “lawn” semantic classes in the specific arrangement as defined by the user.”).
However, Li fails to teach a [semantic] graph.
Dhamo teaches a [semantic] graph (Page 5214 Figure 2 Overview “Given an image, we predict its scene graph and reconstruct the input from a masked representation. a) The graph nodes oi (blue) are enriched with bounding boxes xi (green) and visual features φi (violet) from cropped objects. We randomly mask boxes xi, object visual features φi and the source image; the model then reconstructs the same graph and image utilizing the remaining information.”; Page 5216 right column paragraph 3 “A node is removed entirely from the graph together with all the edges that connect this object with others. The source image region corresponding to the object is occluded.”).
Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Li’s reference to include a [semantic] graph taught by Dhamo’s reference. The motivation for doing so would have been to make modifications with respect to visual entities in the image and the way they interact with each other, both spatially and semantically as suggested by Dhamo (see Dhamo, Page 5213 left column paragraph 3).
Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Dhamo with Li to obtain the invention specified in claim 1.
Regarding claim 2, which claim 1 is incorporated, Li discloses wherein the processing comprises an operation of removing the target object (Paragraph [0058] “The input image 208 represents an image in which one or more portions of the image (in this case a circular area) are being removed. Each portion of the input image 208 being removed is often referred to as a “hole.””).
Regarding claim 3, which claim 1 is incorporated, Li discloses wherein inpainting the first area based on the target semantic [graph] to obtain the second image after inpainted (Paragraph [0130] “the electronic device has used the distribution of the semantic class or classes 1408 in the defined area 1406 to generate a final output image 1414. As can be seen here, the area 1406 has been filled with replacement content of both the “ground” and “lawn” semantic classes in the specific arrangement as defined by the user.”), comprises:
acquiring a first feature graph (initial image feature map in Paragraph [0138] equates to first feature map) corresponding to the first image (Paragraph [0138] “An initial image feature map is generated at step 1608. This may include, for example, the processor 120 of the electronic device 101 providing at least one collection 216 a-216 b of semantic code vectors 218 and the filled semantic mask 212 to the initial decoder 220 in order to generate the initial image feature map 222.”);
regenerating features corresponding to the first area by features corresponding to a second area in the first feature graph based on the target semantic [graph], so as to obtain a second feature graph (resulting feature map in Paragraph [0117] equates to second feature map), wherein the second area is an area except the first area in the first image (Paragraph [0039] “the semantic code vectors identify various semantic classes of content contained in different portions of the input image. The semantic code vectors are generated using the input image, a semantic mask with at least one hole associated with unwanted content being removed, and a filled semantic mask with the at least one hole filled with one or more semantic classes.”; Paragraph [0117] “The initial set of image features 1202 is processed using a ResBlock operation 1208…The resulting feature map and the semantic codes 1204 are provided to a location- and class-wise adaptive instance normalization operation 1212, which applies an affine transformation to the semantic codes 1204 and normalizes the feature map based of the transformed semantic codes 1204.”); and
acquiring the second image based on the second feature graph (Paragraph [0121] “The feature map from the aggregation operation 1220 is provided to an upsampling operation 1222, which upsamples the feature map from the aggregation operation 1220 in order to generate a set of processed image features 1224…Eventually, the last iteration of the processing operation 1102 produces a set of processed image features 1224 at the desired resolution.”).
However, Li fails to teach a [semantic] graph.
Dhamo teaches a [semantic] graph (Page 5214 Figure 2 Overview “Given an image, we predict its scene graph and reconstruct the input from a masked representation. a) The graph nodes oi (blue) are enriched with bounding boxes xi (green) and visual features φi (violet) from cropped objects. We randomly mask boxes xi, object visual features φi and the source image; the model then reconstructs the same graph and image utilizing the remaining information.”; Page 5216 right column paragraph 3 “A node is removed entirely from the graph together with all the edges that connect this object with others. The source image region corresponding to the object is occluded.”).
Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Li’s reference to include a [semantic] graph taught by Dhamo’s reference. The motivation for doing so would have been to make modifications with respect to visual entities in the image and the way they interact with each other, both spatially and semantically as suggested by Dhamo (see Dhamo, Page 5213 left column paragraph 3).
Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Dhamo with Li to obtain the invention specified in claim 3.
Regarding claim 4, which claim 3 is incorporated, Li discloses wherein acquiring the first feature graph corresponding to the first image (Paragraph [0138] “An initial image feature map is generated at step 1608. This may include, for example, the processor 120 of the electronic device 101 providing at least one collection 216 a-216 b of semantic code vectors 218 and the filled semantic mask 212 to the initial decoder 220 in order to generate the initial image feature map 222.”), comprises:
performing mask processing on the first image using the first area (Paragraph [0099] “A multiplier operation 908 multiplies or scales the contents of the raw initial image feature map 906 by a first masked version 212 a of the filled semantic mask 212."), and
performing down-sampling processing on an image subjected to the mask processing (Paragraph [0077] “the input image 208 is provided to a downsampling and convolution operation 502. The downsampling and convolution operation 502 downsamples the input image 208 by reducing the resolution of the input image 208.”), and performing semantic correction on a result of the down-sampling processing based on the target semantic [graph], so as to obtain the first feature graph (Paragraph [0077] ”A number of convolutional layers may be used here, where the first convolutional layer receives and processes the downsampled input image 208 and each remaining convolutional layer receives and processes the outputs from the prior convolutional layer. The output of each convolutional layer has a lower resolution than its input. The output of the downsampling and convolution operation 502 is a downsampled feature map 504, which represents the high-level features of the downsampled version of the input image 208.”; Paragraph [0078] “A contextual information extraction operation 506 uses the downsampled feature map 504 and different versions of the semantic mask 210 and the filled semantic mask 212 to generate the collections 216 a-216 b of semantic code vectors 218.”).
However, Li fails to teach a [semantic] graph.
Dhamo teaches a [semantic] graph (Page 5214 Figure 2 Overview “Given an image, we predict its scene graph and reconstruct the input from a masked representation. a) The graph nodes oi (blue) are enriched with bounding boxes xi (green) and visual features φi (violet) from cropped objects. We randomly mask boxes xi, object visual features φi and the source image; the model then reconstructs the same graph and image utilizing the remaining information.”; Page 5216 right column paragraph 3 “A node is removed entirely from the graph together with all the edges that connect this object with others. The source image region corresponding to the object is occluded.”).
Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Li’s reference to include a [semantic] graph taught by Dhamo’s reference. The motivation for doing so would have been to make modifications with respect to visual entities in the image and the way they interact with each other, both spatially and semantically as suggested by Dhamo (see Dhamo, Page 5213 left column paragraph 3).
Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Dhamo with Li to obtain the invention specified in claim 4.
Regarding claim 5, which clam 3 is incorporated, Li discloses wherein regenerating the features corresponding to the first area by the features corresponding to the second area in the first feature graph based on the target semantic [graph] (Paragraph [0039] “the semantic code vectors identify various semantic classes of content contained in different portions of the input image. The semantic code vectors are generated using the input image, a semantic mask with at least one hole associated with unwanted content being removed, and a filled semantic mask with the at least one hole filled with one or more semantic classes.”; Paragraph [0117] “The initial set of image features 1202 is processed using a ResBlock operation 1208…The resulting feature map and the semantic codes 1204 are provided to a location- and class-wise adaptive instance normalization operation 1212, which applies an affine transformation to the semantic codes 1204 and normalizes the feature map based of the transformed semantic codes 1204.”), comprises:
determining a first cell corresponding to the first area, and determining at least one second cell with same semantics as the first cell based on the target semantic [graph], wherein the second cell corresponds to the second area (Paragraph [0087] “A chunk 704 associated with the selected invalid semantic code vector 602 in the segmented semantic mask 210″ is examined in order to identify the semantic class to be used with the selected invalid semantic code vectors 602. In this case, the chunk 704 associated with the selected invalid semantic code vector 602 indicates that the selected invalid semantic code vector 602 should have the semantic class associated with the lower half of the input image 208.”; Paragraph [0088] “the masked adaptive pooling operation 510 searches for valid semantic code vectors 218 that are (i) relatively near the selected invalid semantic code vector 602 and (ii) associated with the same semantic class as the selected invalid semantic code vector 602… These valid semantic code vectors 218 are associated with four chunks 708 in the segmented semantic mask 210″. These four valid semantic code vectors 218 are near the selected invalid semantic code vector 602, and all of these valid semantic code vectors 218 have the same semantic class as the selected invalid semantic code vector 602. Other semantic code vectors near the selected invalid semantic code vector 602 are either invalid or belong to a different semantic class.”; See FIGS. 7A, 7B, and 7C the features are associated with 4 chunks (i.e. the first and second cell); and
regenerating features of the first cell according to features corresponding to the second cell in the first feature graph (Paragraph [0090] “the selected valid semantic code vectors 218 are pooled in order to produce a new valid semantic code vector 218′ in place of the selected invalid semantic code vector 602. This results in the creation of an updated collection 702′ of semantic code vectors.”).
However, Li fails to teach a [semantic] graph.
Dhamo teaches a [semantic] graph (Page 5214 Figure 2 Overview “Given an image, we predict its scene graph and reconstruct the input from a masked representation. a) The graph nodes oi (blue) are enriched with bounding boxes xi (green) and visual features φi (violet) from cropped objects. We randomly mask boxes xi, object visual features φi and the source image; the model then reconstructs the same graph and image utilizing the remaining information.”; Page 5216 right column paragraph 3 “A node is removed entirely from the graph together with all the edges that connect this object with others. The source image region corresponding to the object is occluded.”).
Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Li’s reference to include a [semantic] graph taught by Dhamo’s reference. The motivation for doing so would have been to make modifications with respect to visual entities in the image and the way they interact with each other, both spatially and semantically as suggested by Dhamo (see Dhamo, Page 5213 left column paragraph 3).
Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Dhamo with Li to obtain the invention specified in claim 5.
Regarding claim 6, which claim 5 is incorporated, Li discloses wherein regenerating the features corresponding to the first cell according to the features corresponding to the second cell in the first feature graph (Paragraph [0090] “the selected valid semantic code vectors 218 are pooled in order to produce a new valid semantic code vector 218′ in place of the selected invalid semantic code vector 602. This results in the creation of an updated collection 702′ of semantic code vectors.”; Paragraph [0117] “The initial set of image features 1202 is processed using a ResBlock operation 1208…The resulting feature map and the semantic codes 1204 are provided to a location- and class-wise adaptive instance normalization operation 1212, which applies an affine transformation to the semantic codes 1204 and normalizes the feature map based of the transformed semantic codes 1204.”), comprises:
acquiring a first feature corresponding to the first cell in the first feature graph and respective second features corresponding to each second cell in the first feature graph (Paragraph [0088] “the masked adaptive pooling operation 510 identifies four valid semantic code vectors 218 satisfying these criteria. These valid semantic code vectors 218 are associated with four chunks 708 in the segmented semantic mask 210″. These four valid semantic code vectors 218 are near the selected invalid semantic code vector 602, and all of these valid semantic code vectors 218 have the same semantic class as the selected invalid semantic code vector 602.; See FIGS. 7A, 7B, and 7C there are multiple identified semantic code features (i.e. the first and second features)); and
regenerating features of the first cell according to the first feature and second features (Paragraph [0090] “the selected valid semantic code vectors 218 are pooled in order to produce a new valid semantic code vector 218′ in place of the selected invalid semantic code vector 602. This results in the creation of an updated collection 702′ of semantic code vectors.”).
Regarding claim 7, which claim 3 is incorporated, Li discloses wherein acquiring the second image based on the second feature graph, comprises: generating the second image based on the target semantic [graph] and the second feature graph (Paragraph [0039] “An image decoder uses the semantic code vectors and the initial image feature map to precisely apply the semantic code vectors and reconstruct an output image, where the unwanted content in the input image has been removed and replaced with other content in the output image.”; Paragraph [0121] “The feature map from the aggregation operation 1220 is provided to an upsampling operation 1222, which upsamples the feature map from the aggregation operation 1220 in order to generate a set of processed image features 1224…Eventually, the last iteration of the processing operation 1102 produces a set of processed image features 1224 at the desired resolution.”).
However, Li fails to teach a [semantic] graph.
Dhamo teaches a [semantic] graph (Page 5214 Figure 2 Overview “Given an image, we predict its scene graph and reconstruct the input from a masked representation. a) The graph nodes oi (blue) are enriched with bounding boxes xi (green) and visual features φi (violet) from cropped objects. We randomly mask boxes xi, object visual features φi and the source image; the model then reconstructs the same graph and image utilizing the remaining information.”; Page 5216 right column paragraph 3 “A node is removed entirely from the graph together with all the edges that connect this object with others. The source image region corresponding to the object is occluded.”).
Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Li’s reference to include a [semantic] graph taught by Dhamo’s reference. The motivation for doing so would have been to make modifications with respect to visual entities in the image and the way they interact with each other, both spatially and semantically as suggested by Dhamo (see Dhamo, Page 5213 left column paragraph 3).
Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Dhamo with Li to obtain the invention specified in claim 7.
Regarding claim 8, which claim 7 is incorporated, Li discloses wherein generating the second image based on the target semantic [graph] and the second feature graph Paragraph [0039] “An image decoder uses the semantic code vectors and the initial image feature map to precisely apply the semantic code vectors and reconstruct an output image, where the unwanted content in the input image has been removed and replaced with other content in the output image.”; Paragraph [0121] “The feature map from the aggregation operation 1220 is provided to an upsampling operation 1222, which upsamples the feature map from the aggregation operation 1220 in order to generate a set of processed image features 1224…Eventually, the last iteration of the processing operation 1102 produces a set of processed image features 1224 at the desired resolution.”, comprises:
performing up-sampling processing on the second feature graph, and performing semantic correction on a result of the up-sampling processing based on the target semantic [graph] so as to obtain the second image (Paragraph [0120] “In FIG. 12, the adaptive instance normalization operations 1212 and 1218 are used to fuse such information back into intermediate feature maps within the network. This is also why the same semantic codes are provided in the sets of semantic codes 1204 and 1206 to the adaptive instance normalization operations 1212 and 1218. In this way, the input condition (the semantic codes) are provided as inputs to the initial decoder 220, and the semantic codes are maintained via fusing using the adaptive instance normalization operations 1212 and 1218 in the image decoder 224.”).
However, Li fails to teach a [semantic] graph.
Dhamo teaches a [semantic] graph (Page 5214 Figure 2 Overview “Given an image, we predict its scene graph and reconstruct the input from a masked representation. a) The graph nodes oi (blue) are enriched with bounding boxes xi (green) and visual features φi (violet) from cropped objects. We randomly mask boxes xi, object visual features φi and the source image; the model then reconstructs the same graph and image utilizing the remaining information.”; Page 5216 right column paragraph 3 “A node is removed entirely from the graph together with all the edges that connect this object with others. The source image region corresponding to the object is occluded.”).
Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Li’s reference to include a [semantic] graph taught by Dhamo’s reference. The motivation for doing so would have been to make modifications with respect to visual entities in the image and the way they interact with each other, both spatially and semantically as suggested by Dhamo (see Dhamo, Page 5213 left column paragraph 3).
Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Dhamo with Li to obtain the invention specified in claim 8.
Regarding claim 13, Li discloses a non-transitory computer-readable storage medium storing instructions that cause a processor to (Paragraph [0010] “various functions described below can be implemented or supported by one or more computer programs, each of which is formed from computer readable program code and embodied in a computer readable medium”):
acquire a first image, wherein the first image (input image in Paragraph [0058] equates to first image) is obtained by processing a target object in an original image (Paragraph [0058] “The input image 208 represents an image in which one or more portions of the image (in this case a circular area) are being removed. Each portion of the input image 208 being removed is often referred to as a “hole.””; Note that portion of the image being removed equates to processing a target object);
determine a first area to be inpainted in the first image, wherein the first area is at least a partial area of the target object (unwanted person or other object in Paragraph [0128] equates to target object) (Paragraph [0128] “an image 1402 is being presented on an electronic device, and a user has used an electronic pen 1404 or his or her finger to define an area 1406 of the image 1402 to be removed. The area 1406 may, for instance, include an unwanted person or other object in the image 1402.”);
acquire a target semantic [graph] corresponding to the first image (Paragraph [0059] “The filled semantic mask 212 represents an initial estimation of the semantic classes associated with the input image 208, including labels for one or more semantic classes estimated for each hole in which content is being removed in the input image 208.”); and
inpaint the first area based on the target semantic graph to obtain a second image (final output image in Paragraph [0130] equates to second image) after inpainted (Paragraph [0130] “ the electronic device has used the distribution of the semantic class or classes 1408 in the defined area 1406 to generate a final output image 1414. As can be seen here, the area 1406 has been filled with replacement content of both the “ground” and “lawn” semantic classes in the specific arrangement as defined by the user.”).
However, Li fails to teach a [semantic] graph.
Dhamo teaches a [semantic] graph (Page 5214 Figure 2 Overview “Given an image, we predict its scene graph and reconstruct the input from a masked representation. a) The graph nodes oi (blue) are enriched with bounding boxes xi (green) and visual features φi (violet) from cropped objects. We randomly mask boxes xi, object visual features φi and the source image; the model then reconstructs the same graph and image utilizing the remaining information.”; Page 5216 right column paragraph 3 “A node is removed entirely from the graph together with all the edges that connect this object with others. The source image region corresponding to the object is occluded.”).
Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Li’s reference to include a [semantic] graph taught by Dhamo’s reference. The motivation for doing so would have been to make modifications with respect to visual entities in the image and the way they interact with each other, both spatially and semantically as suggested by Dhamo (see Dhamo, Page 5213 left column paragraph 3).
Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Dhamo with Li to obtain the invention specified in claim 13.
Regarding claim 14, Li discloses an electronic device, comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor, cause the processor to (Paragraph [0043] “an electronic device 101 is included in the network configuration 100. The electronic device 101 can include at least one of a bus 110, a processor 120, a memory 130, an input/output (I/O) interface 150, a display 160, a communication interface 170, and a sensor 180.”):
acquire a first image, wherein the first image (input image in Paragraph [0058] equates to first image) is obtained by processing a target object in an original image (Paragraph [0058] “The input image 208 represents an image in which one or more portions of the image (in this case a circular area) are being removed. Each portion of the input image 208 being removed is often referred to as a “hole.””; Note that portion of the image being removed equates to processing a target object);
determine a first area to be inpainted in the first image, wherein the first area is at least a partial area of the target object (unwanted person or other object in Paragraph [0128] equates to target object) (Paragraph [0128] “an image 1402 is being presented on an electronic device, and a user has used an electronic pen 1404 or his or her finger to define an area 1406 of the image 1402 to be removed. The area 1406 may, for instance, include an unwanted person or other object in the image 1402.”);
acquire a target semantic [graph] corresponding to the first image (Paragraph [0059] “The filled semantic mask 212 represents an initial estimation of the semantic classes associated with the input image 208, including labels for one or more semantic classes estimated for each hole in which content is being removed in the input image 208.”); and
inpaint the first area based on the target semantic [graph] to obtain a second image (final output image in Paragraph [0130] equates to second image) after inpainted (Paragraph [0130] “ the electronic device has used the distribution of the semantic class or classes 1408 in the defined area 1406 to generate a final output image 1414. As can be seen here, the area 1406 has been filled with replacement content of both the “ground” and “lawn” semantic classes in the specific arrangement as defined by the user.”).
However, Li fails to teach a [semantic] graph.
Dhamo teaches a [semantic] graph (Page 5214 Figure 2 Overview “Given an image, we predict its scene graph and reconstruct the input from a masked representation. a) The graph nodes oi (blue) are enriched with bounding boxes xi (green) and visual features φi (violet) from cropped objects. We randomly mask boxes xi, object visual features φi and the source image; the model then reconstructs the same graph and image utilizing the remaining information.”; Page 5216 right column paragraph 3 “A node is removed entirely from the graph together with all the edges that connect this object with others. The source image region corresponding to the object is occluded.”).
Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Li’s reference to include a [semantic] graph taught by Dhamo’s reference. The motivation for doing so would have been to make modifications with respect to visual entities in the image and the way they interact with each other, both spatially and semantically as suggested by Dhamo (see Dhamo, Page 5213 left column paragraph 3).
Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Dhamo with Li to obtain the invention specified in claim 14.
Regarding claim 15 (drawn to a device), claim 15 is rejected the same as claim 2 and the arguments similar to that presented above for claim 2 are equally applicable to the claim 15, and all the other limitations similar to claim 2 are not repeated herein, but incorporated by reference.
Regarding claim 16 (drawn to a device), claim 16 is rejected the same as claim 3 and the arguments similar to that presented above for claim 3 are equally applicable to the claim 16, and all the other limitations similar to claim 3 are not repeated herein, but incorporated by reference.
Regarding claim 17 (drawn to a device), claim 17 is rejected the same as claim 4 and the arguments similar to that presented above for claim 4 are equally applicable to the claim 17, and all the other limitations similar to claim 4 are not repeated herein, but incorporated by reference.
Regarding claim 18 (drawn to a device), claim 18 is rejected the same as claim 5 and the arguments similar to that presented above for claim 5 are equally applicable to the claim 18, and all the other limitations similar to claim 5 are not repeated herein, but incorporated by reference.
Regarding claim 19 (drawn to a device), claim 19 is rejected the same as claim 6 and the arguments similar to that presented above for claim 6 are equally applicable to the claim 19, and all the other limitations similar to claim 6 are not repeated herein, but incorporated by reference.
Regarding claim 20 (drawn to a device), claim 20 is rejected the same as claim 7 and the arguments similar to that presented above for claim 7 are equally applicable to the claim 20, and all the other limitations similar to claim 7 are not repeated herein, but incorporated by reference.
Regarding claim 21 (drawn to a non-transitory computer-readable storage medium), claim 21 is rejected the same as claim 3 and the arguments similar to that presented above for claim 3 are equally applicable to the claim 21, and all the other limitations similar to claim 3 are not repeated herein, but incorporated by reference.
Claims 9-11 are rejected under 35 U.S.C. 103 as being unpatentable over Li et al. (US 11,526,967 B2) (hereinafter, “Li”) in view of Dhamo et al. (Dhamo, Helisa, et al. "Semantic image manipulation using scene graphs." 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2020.) (hereinafter; “Dhamo”) as applied to claim 6 above; and further in view of Weng et al. (Weng, Shuchen, et al. "Conditional image repainting via semantic bridge and piecewise value function." European Conference on Computer Vision. Cham: Springer International Publishing, 2020.) (hereinafter, “Weng”).
Regarding claim 9, which claim 6 is incorporated, Li discloses wherein regenerating the features of the first cell according to the first feature and second features (Paragraph [0090] “the selected valid semantic code vectors 218 are pooled in order to produce a new valid semantic code vector 218′ in place of the selected invalid semantic code vector 602. This results in the creation of an updated collection 702′ of semantic code vectors.”), [comprises:
computing a similarity between the first feature and each second feature; and
regenerating the features of the first cell based on the similarity].
However, Li and Dhamo both fail to teach computing a similarity between the first feature and each second feature; and regenerating the features of the first cell based on the similarity.
Weng teaches computing a similarity between the first feature and each second feature (h[j] and e[i] on Page 3 Section 2 first paragraph equate to the first and second feature) (Page 3 Section 2 first paragraph “where s[j,i] is the similarity between h[j] and e[i] which is computed by the dot product”); and
regenerating the features of the first cell based on the similarity (Page 4 first paragraph “the words’ features are aggregated by their relevance to each region on the image plane
c
[
j
]
=
∑
i
=
1
N
β
[
j
,
i
]
ϕ
(
e
[
i
]
, where c[j] is the aggregated word features (or named context feature vector) for the j-th region.”).
Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Li in view of Dhamo reference to include computing a similarity between the first feature and each second feature; and regenerating the features of the first cell based on the similarity taught by Weng’s reference. The motivation for doing so would have been to enforce the generated content at each region to reflect its relevant words (i.e. features) as suggested by Weng (see Weng, Page 4 first paragraph).
Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Weng with Li and Dhamo to obtain the invention specified in claim 9.
Regarding claim 10, which claim 9 is incorporated, Li and Dhamo both fail to teach wherein regenerating the features corresponding to the first cell based on the similarity, comprises: determining a weight corresponding to each second feature based on the similarity, and computing a weighted sum of the second features; and regenerating the features of the first cell according to the weighted sum.
Weng teaches teach wherein regenerating the features corresponding to the first cell based on the similarity (Page 4 first paragraph “the words’ features are aggregated by their relevance to each region on the image plane:
c
[
j
]
=
∑
i
=
1
N
β
[
j
,
i
]
ϕ
(
e
[
i
]
, where c[j] is the aggregated word features (or named context feature vector) for the j-th region.”), comprises:
determining a weight corresponding to each second feature based on the similarity, and computing a weighted sum of the second features (Page 3 Section 2 first paragraph “Their relevance is estimated as an attention weight β[j,i]:
PNG
media_image1.png
72
546
media_image1.png
Greyscale
”; Page 4 first paragraph “the words’ features are aggregated by their relevance to each region on the image plane:
c
[
j
]
=
∑
i
=
1
N
β
[
j
,
i
]
ϕ
(
e
[
i
]
, where c[j] is the aggregated word features (or named context feature vector) for the j-th region.”); and
regenerating the features of the first cell according to the weighted sum (Page 4 first paragraph “For each region, c[j] is concatenated with h[j], so as to enforce the generated content at each region to reflect its relevant words”.).
Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Li in view of Dhamo reference to include wherein regenerating the features corresponding to the first cell based on the similarity, comprises: determining a weight corresponding to each second feature based on the similarity, and computing a weighted sum of the second features; and regenerating the features of the first cell according to the weighted sum taught by Weng’s reference. The motivation for doing so would have been to enforce the generated content at each region to reflect its relevant words (i.e. features) as suggested by Weng (see Weng, Page 4 first paragraph).
Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Weng with Li and Dhamo to obtain the invention specified in claim 10.
Regarding claim 11, which claim 10 is incorporated, Li and Dhamo both fail to teach wherein regenerating the features of the first cell according to the weighted sum, comprises: performing a stacking processing on the weighted sum and the first feature to obtain the features corresponding to the first cell.
Weng teach wherein regenerating the features of the first cell according to the weighted sum, comprises: performing a stacking processing on the weighted sum and the first feature to obtain the features corresponding to the first cell (Page 4 first paragraph “the words' features are aggregated by their relevance to each region on the image plane:
c
[
j
]
=
∑
i
=
1
N
β
[
j
,
i
]
ϕ
(
e
[
i
]
, where c[j] is the aggregated word features (or named context feature vector) for the j-th region. For each region, c[j] is concatenated with h[j], so as to enforce the generated content at each region to reect its relevant words.”).
Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Li in view of Dhamo reference to include wherein regenerating the features of the first cell according to the weighted sum, comprises: performing a stacking processing on the weighted sum and the first feature to obtain the features corresponding to the first cell taught by Weng’s reference. The motivation for doing so would have been to enforce the generated content at each region to reflect its relevant words (i.e. features) as suggested by Weng (see Weng, Page 4 first paragraph).
Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Weng with Li and Dhamo to obtain the invention specified in claim 11.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Zheng et al. (US 2023/0360180 A1) discloses a method for generating an inpainted digital image by using a first feature map and a second feature map.
El-Khamy et al. (US 2021/0217145 A1) discloses a method for inpainting an image by identifying context frames with respect to reference frames and producing a refined frame by processing the reference frames based on the context.
Bouhnik et al. (US 10,540,757 B1) discloses a method for determining a mixed image by combining pixel values from a first and second image, according to generated mask data, for generating missing data for any undefined pixel regions.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to UROOJ FATIMA whose telephone number is (571)272-2096. The examiner can normally be reached M-F 8:00-5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Henok Shiferaw can be reached at (571) 272-4637. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/UROOJ FATIMA/Examiner, Art Unit 2676
/Henok Shiferaw/Supervisory Patent Examiner, Art Unit 2676