DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgment is made of applicant's claim for foreign priority based on an application filed in Korea on 15 December 2023. It is noted, however, that applicant has not filed a certified copy of the KR10-2023-0183426 application as required by 37 CFR 1.55. The effective filing date of the instant application is therefore 12 December 2024.
Specification
The title of the invention, “Method and Device with Image Enhancement”, does not accurately describe the claimed invention and the embodiments described in the specification and drawings. The phrase “image enhancement” is too broad and covers many disparate areas of image processing including tone mapping, metadata generation and sharpening, for example. A new title is required that is clearly indicative of the invention to which the claims are directed.
The following title is suggested: “Method and Device for Image Enhancement with Style Transfer Based on Local and Global Features”.
Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
Claims 4, 11 and 15 are rejected under 35 U.S.C. 112(b) as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor regards as the invention.
Claims 4 and 15 recite, “wherein the neural harmonic decoder is configured to, based on harmony between the local feature representation and the global feature representation, determine the adjustment parameter set” (emphasis added). It is unclear how the decoder determines the adjustment parameter set because “based on harmony” is vague, subjective, and amounts to describing an intended and result rather than a description of how that result is achieved. A “harmony” between feature representations depends upon the criteria used to distinguish that which is harmonious from that which is not harmonious. Image features can be harmonious based on texture, gradient values, pixel-related statistics, and so forth. It is unclear how the “harmony” is quantified and usable by a processor. The independent claims already establish that the adjustment parameter set is used to retouch the input image. Image “retouching” is performed to improve the subjective appearance of an image. Claims 4 and 15 describe that subjective appearance as a “harmony”, which at best, means that a goal of the retouching is to generate a retouched image that makes some aspect of the local region or other part(s) of the input image more alike than different. However, the claims do not sufficiently explain what about the local and global feature representations is made more harmonious and by what metric such harmony is measured. The broadest reasonable interpretation of a “harmony” includes a “pleasing arrangement of parts” and is synonymous with “agreement, accord”.1 Thus, for purposes of applying prior art, the adjustment parameter set is interpreted as being based on an agreement of style between the local and global feature representations, i.e., the style of one or both of the image regions is adjusted so that the retouched image includes fewer differences in style as compared to the original input image.
Claim 11 recites “instructions that, when executed by one or more processors, configure the one or more processors to perform the method of claim 1” (emphasis added). Claim 1 is a “processor-implemented method”, but does not explicitly recite a processor in relation to any particular step in the body of the claim. Presumably, there is at least one processor involved in performing the method of claim 1. While the antecedent basis of “the one or more processors” would seem to be “one or more processors” recited earlier in claim 11, it is unclear if one of those processors is the same processor performing the method of claim 1. Thus, the antecedent basis of “the one or more processors” is unclear in the context of claim 11 viewed as a whole with claim 1. For purposes of applying prior art, “the one or more processors” is interpreted as referring to the same one or more processors that implement the method of claim 1.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception without significantly more.
The claims recite:
[claim 1] A processor-implemented method with image enhancement, the method comprising:
[limitation (a):] based on an input image, region information about a local region of the input image, and style information about a target style to be applied to the local region, generating a local feature representation;
[limitation (b):] generating a global feature representation based on the input image;
[limitation (c):] based on the local feature representation, the global feature representation, and the region information, determining an adjustment parameter set; and
[limitation (d):] based on the adjustment parameter set, generating a retouch result by adjusting the input image.
[claim 2] The method of claim 1, wherein the local feature representation is generated by inputting the input image, the region information, and the style information into a neural local encoder, and the global feature representation is generated by inputting the input image into a neural global encoder.
[claim 3] The method of claim 1, wherein the adjustment parameter set is generated by inputting the local feature representation, the global feature representation, and the region information into a neural harmonic decoder.
[claim 4] The method of claim 3, wherein the neural harmonic decoder is configured to, based on harmony between the local feature representation and the global feature representation, determine the adjustment parameter set.
[claim 5] The method of claim 1, wherein the region information comprises first region information about a first local region and second region information about a second local region, and the style information comprises first style information to be applied to the first local region and second style information to be applied to the second local region.
[claim 6] The method of claim 5, wherein the generating of the local feature representation comprises, based on the input image, the first region information, the second region information, the first style information, and the second style information, generating a first local feature representation corresponding to the first region information and the first style information and a second local feature representation corresponding to the second region information and the second style information.
[claim 7] The method of claim 6, wherein the determining of the adjustment parameter set comprises, based on the first local feature representation, the second local feature representation, the global feature representation, the first region information, and the second region information, determining the adjustment parameter set.
[claim 8] The method of claim 1, further comprising, based on semantic segmentation, determining the local region from the input image.
[claim 9] The method of claim 1, further comprising, based on the input image, generating a region mask corresponding to the region information.
[claim 10] The method of claim 1, wherein the style information comprises either one or both of a target style image of the target style and a target style vector of the target style.
[claim 11] A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, configure the one or more processors to perform the method of [limitations (a) - (d)].
[claim 12] An electronic device comprising: one or more processors configured to: [limitations (a) - (d)].
[claims 13-20] The electronic device [performing the respective method of each of claims 2-7, 9 and 10].
Claim interpretation
Under the broadest reasonable interpretation, the terms of the claims are presumed to have their plain meaning consistent with the specification as it would be interpreted by one of ordinary skill in the art. See MPEP 2111.
Based on the plain meaning of the words in the claim, the broadest reasonable interpretation of claim 1 is method implemented by a processor (which is equivalent to a method performed by a general purpose computer) for retouching an image based on (a) local region features and a target style, (b) global image features to (c) determine an adjustment parameter set the input image therefrom, and (d) retouch the image using the adjustment parameter set. The preambles of claims 1 and 12 do not positively add limitations to the claimed processor(s) or further modify limitations recited in the body of each claim in relation to the processor(s) specifically.
Based on the plain meaning of the words in the claim, the broadest reasonable interpretation of claim 11 is a non-transitory computer-readable storage medium (CRM) having instructions to carry out the method of claim 1.
Based on the plain meaning of the words in the claim, the broadest reasonable interpretation of claim 12 is a generic computer that performs the method of claim 1.
Claims 1, 11 and 12 put no limitation on the various types of “information” other than relating the information to style and local and global regions of the input image. For example, these claims do not recite any pixels of the input image or any pixel-level processing. These claims also do not put any limitation upon the implicitly or explicitly-recited processor(s) with respect to automatic versus manual processing. While a generic computer is certainly involved in claims 1, 11 and 12, the extent to which the computer merely translates a human operator’s commands as compared to performing operations without human involvement, is unspecified. The processor (computer) is recited at a high level of generality, i.e., as a generic computer performing generic computer functions. In other words, a processor may be involved, but the extent to which it is responsible for carrying out the recited operations in the bodies of the claims is open-ended.
An “input image” reasonably connotes a digital image and not a physical image like a painting. In the context of retouching or enhancing a digital images, e.g., Poisson editing (a well-known type of image blending), claims 1, 11 and 12 describe image style transfer based on local and global image features using a generic computer to modify the input image using a set of parameters derived from local information, global information and style information.
Claims 2-4 and 13-15 refer to various “neural” encoders and decoders, which amounts to a generic description of all neural network-based encoders and decoders without being limited to any particular architecture or sub-field of machine learning.
Claims 5-7 and 16-18 further describe the various types of “information” at a high-level of generality.
Claim 8 refers to “semantic segmentation” which is a general description of certain image segmentation techniques that can distinguish one type of object from another within the same image and the claim does not put any limitation on the particular type of semantic segmentation being performed.
Claims 9 and 19 refer to generating region masks without putting any limitation on the method of generating such masks or what type of region they are designed to segment from the rest of the input image.
Claims 10 and 20 further specify the style information as including one or both of an image or vector representing the target style without substantially differentiating an image of a style as opposed to a vector of a style because an image can also be interpreted as a vector or array of values.
Step 1: do the claims fall within any statutory category?
Claim 1 recites a series of steps and therefore, is a process. Claim 11 recites a non-transitory computer readable storage medium. The disclosure gives magnetic media such as hard disks, optical disks, floppy disks as examples of memory that is a non-transitory computer-readable recording medium. The broadest reasonable interpretation of claim 11 covers only statutory embodiments of a computer-readable storage medium and not a transitory signal. A non-transitory computer readable storage medium falls within the “manufacture” category of invention. Claim 12 recites a device comprising one or more processors and therefore, is a machine. See MPEP 2106.03. (Step 1: YES).
Step 2A, Prong One: do the claims recite a judicial exception?
As explained in MPEP 2106.04, subsection II, a claim “recites” a judicial exception when the judicial exception is “set forth” or “described” in the claim. Limitations (a)-(d) describe retouching an image based on local features, global features and style information using a generic computer as explained above. Under its broadest reasonable interpretation when read in light of the specification, the “generating” and/or “determining” in claims 1-20 encompass mental observations or evaluations that are practically performed in the human mind. For example, the claimed generating of local and global features in claims 1 and 12 can refer to a person deciding which parts of the image to retain and which parts to modify. Claims 1, 11 and 12 therefore respectively describe a processor-implemented method, a non-transitory computer-readable medium, and a device having a processor at a high-level of generality that, under the broadest reasonable interpretation amounts to implementing the abstract idea using a generic computer.
Some dependent claims describe features that typically accompany image retouching processing like masks and vectors. However, a “region mask” as recited in claims 9 and 19, and a “target style vector” as recited in claims 10 and 20, in the context of a digital image, are merely sets of data that distinguish one image area from another. For example, a person can easily distinguish which part of an image region corresponds to an object and which part corresponds to the background. Furthermore, a “vector” is merely an N-dimensional array of data. In the context of an image, a “vector” can be the corresponding pixel locations and/or RGB values that correspond to a particular object or even a bespoke set of values related to a specific compression and/or reconstruction process. Under the broadest reasonable interpretation, determining or generating a “mask” or a “vector” amounts to performing evaluation, judgment, and opinion to make a determination about what each image region actually depicts.
Claims 2-4 and 13-15 describe generating the adjustment parameter set or the local and global feature representations using neural network-based encoders or decoders described at a high-level of generality and therefor encompass performing evaluation, judgment, and opinion to make determinations about what certain image regions represent in relation to the target style.
Claims 5, 16 and 20 describe multiple pieces of information for different local regions and multiple pieces of style information. Such “information” is described at a high-level of generality and therefor encompass performing evaluation, judgment, and opinion to make determinations about what certain image regions represent in relation to the target style.
Under the broadest reasonable interpretation, “semantic segmentation” in claim 8 amounts to performing evaluation, judgment, and opinion to make a determination about what each image region actually depicts. For example, a person may look at an image and mentally decide that the upper portion is the sky, the lower portion is the ground, and the foreground object is a person standing on the ground.
Under the broadest reasonable interpretation, the claims describe the observations, evaluations, judgments, and opinions of a person reviewing an input image, deciding how to manually adjust some portion(s) of the image to transfer a target style, and then executing the desired edit to the input image, which falls within the mental process grouping of abstract ideas. See MPEP 2106.04(a)(2), subsection III. (Step 2A, Prong One: YES).
Step 2A, Prong Two: Do the claims as a whole integrate the recited judicial exception into a practical application of the exception?
Claims 1-4, 8-15, 19 and 20 collectively recite ten additional elements: (I) “processor-implemented”, (II) “neural local encoder”, (III) “neural global encoder”, (IV) “neural harmonic decoder”, (V) “semantic segmentation”, (VI) “region mask”, (VII) “target style vector”, (VIII) “non-transitory computer-readable storage medium”, (IX) “one or more processors” and (X) “electronic device”.
The additional elements (I), (VIII), (IX) and (X) describe generic computer components and/or operations thereof (e.g., using a pre-built PC, image editing using Microsoft Paint to make a composite image by pasting one image into another and editing the composite image to blend the regions together) recited at a high level of generality and do not amount to any of the relevant considerations for evaluating whether additional limitations integrate a judicial exception into a practical application provided in MPEP 2106.04(d), subsection I. These additional elements amount to merely including instructions to implement the abstract idea on a computer or merely using a computer as a tool to perform the abstract idea. See MPEP 2106.05(f). These additional elements invoke computers merely as a tool to perform an existing process: image editing/enhancement. See MPEP 2106.05(f)(2).
The additional elements (II)-(IV) describe image compression and reconstruction components at a high-level of generality and merely confine the use of the abstract idea to a particular technological environment, i.e., neural networks for image enhancement, and thus fail to add an inventive concept to the claims. See MPEP 2106.05(h).
The additional element (V) describes semantic image segmentation to determine the local region at a high-level of generality and merely confines the use of the abstract idea to a particular technological environment, i.e., semantic segmentation, and thus fails to add an inventive concept to the claims. See MPEP 2106.05(h).
The additional elements (VI) and (VII) describe simple image objects, i.e., masks and vectors, at a high-level of generality. These additional elements provide nothing more than mere instructions to implement the abstract idea on a generic computer and using the computer to perform an existing process: image segmentation. See MPEP 2106.05(f)(2).
None of the additional elements improve the functioning of a computer. See MPEP 2106.04(d)(1). The specification’s background section already sets forth that neural networks have been used in the prior art for the particular purpose of image restoration and are trainable to perform other tasks. Such training may be performed using feedback provided by “an image expert”. See par. 43. The claims are not directed to such training. Rather, the claims are directed to implementing pre-trained neural networks. The specification provides that in “contrast to the typical decoder” the “neural harmonic decoder”, which is featured in claims 4 and 15, is different in that it uses both of a “global feature representation” and a “local feature representation”. See pars. 53-54. The extent of detail provided with respect to the “semantic segmentation”, featured in claim 8, is that it is performed “based on deep learning”. See par. 56. Thus, it is not evident from the claims or the specification how the functioning of a computer is achieved.
The specification does not appear to explicitly set forth a technical problem. Rather, it is implied that one problem being solved is how to incorporate global information into the conventional decoder that only uses local information. See par. 53. It is unclear from reading the specification what would amount to a specific and unconventional technical solution to such a technical problem that could be considered to integrate the judicial exception into a practical application.
Even when viewed in combination, the additional elements do not integrate the recited judicial exception into a practical application (Step 2A, Prong Two: NO), and the claims are directed to the judicial exception. (Step 2A: YES).
Step 2B: do the claims as a whole amount to significantly more than the judicial exception?
As explained with respect to Step 2A Prong Two, the additional elements of the pending claims amount to performing the abstract idea using a computer as a tool to perform an existing process (image editing), which cannot provide an inventive concept. See MPEP 2106.05(f).
Based on the high-level of specify of the technical aspects of the independent and dependent claims as compared to subject matter in the specification, including the drawings, that is in certain aspects more specific and rooted in the technical problem(s) being solved in the additional elements, the additional elements do not constitute an improvement to the functioning of a computer or to another technology because they represent what is well-understood, routine, conventional activity. See MPEP 2106.04(d)(1). For example, the specification purports to describe an unconventional decoder that differs from encoders of the prior art in that it uses both global and local feature representations. However, Semantic Context-Aware Image Style Transfer to Liao et al. (published 10 February 2022) and Style Image Harmonization via Global-Local Style Mutual Grid to Yan et al. (published 11 March 2023) demonstrate that neural-network based global and local encoding/decoding was well-known in the field of image enhancement before the instant application was effectively filed. Thus, the mere use of a combined implementation of global and local encoding/decoding cannot elevate the claims to amount to significantly more than the judicial exception. The independent claims do not require any neural-network based component, only dependent claims like claims 2-4 and 13-15 require the use of neural networks, and their description in the claims is so broadly-recited that they amount to merely using known processes to enhance the target image.
Even considering each claim as a whole, the claims do not amount to significantly more than the recited judicial exception and fail to encompass an inventive concept (Step 2B: NO). Claims 1-20, therefore, are not eligible.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-7 and 9-20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Style Image Harmonization via Global-Local Style Mutual Grid to Yan et al. (hereinafter “Yan”).
Regarding claim 1, Yan teaches a processor-implemented method with image enhancement (A computer is implicit from the disclosure of section 4.1, which discloses, for example, using the Adam optimizer, which is a computer-implemented optimizer.), the method comprising (Yan, abstract, “we learn to extract global and local information from the Vision Transformer and Convolutional Neural Networks, and adaptively fuse the two kinds of information under a multi-scale fusion structure to ameliorate disharmony between foreground and background styles. Then we train the blending network GradGAN to smooth the image gradient. Finally, we take both style and gradient into consideration to solve the sudden change in the blended boundary gradient.”, section 3.1, “The goal is to blend the foreground of x into the entire image, maintain its texture and semantics while transferring style, and smoothly transitioning the paste boundary gradient with the surrounding gradient. We learn from a generator with an encoder-decoder structure to turn the source image into the target image.”):
based on an input image (The content image containing both the style to be transferred to the local patch, i.e., the middle-left image in each of the green, pink and orange boxes in Figure 2.), region information about a local region (Yan, pg. 5, “pasted area”. A local region is defined by its boundary via a mask and the pixel values contained within the mask, which is region information) of the input image (Yan, Fig. 4, “mask and mask_dilated corresponding to the foreground”. The mask segments the background region, a first local region, from the foreground region, a second local region.), and style information (background pattern to be transferred to the local region in the blended output image) about a target style to be applied to the local region (Yan, section 3.1, “y is the corresponding background image, and m is the mask of the foreground. The goal is to blend the foreground of x into the entire image, maintain its texture and semantics while transferring style”. Global style information is to be transferred to a local region.), generating a local feature representation (Yan, Fig. 2, “local style transfer based on CNN (pink part)”. In Fig. 2, the top-left box is the “green part” or “global style transfer based on the transformer”, the middle-left box is the “pink part” or “local style transfer based on CNN”, the lower-left box is the “orange part” or “gradient smoothing module”, the top-right box is the “blue part” or “global and local blending module”, and the lower-right box is the “yellow part” or “style-gradient fusion module”. See Yan at pg. 5. The pink box generates a local embedding.);
generating a global feature representation based on the input image (Yan, Fig. 2, “Global style transfer based on the transformer (green part)”. The top-left or green box generates a global embedding.);
based on the local feature representation, the global feature representation, and the region information (Yan, Fig. 2, “The global and local blending module Blending Decoder (blue part) decodes the latent code to get an image of the stylized foreground.”. The outputs of the global style transfer and local style transfer are inputs of the blending decoder.), determining an adjustment parameter set (The Blending Decoder blends global and local styles through multi-scale fusion, which comprises a set of parameters that yield the blended result. See Yan, section 2.2, “The task of image blending is to paste an area of the cropped source image onto the target image and make the image looks harmonious as a whole. Traditional blending methods use low-level appearance statistics [27–30] to adjust the foreground image. ... Recent methods [2, 4] combined neural networks and Poisson blending to generate realistic images. Jiang et al. [39] used cropping perturbed images to handle the stylistic images blending, which uses 3D color lookup tables (LUTs) to find information such as hue, brightness, and contrast. All of these approaches distort the foreground and lose its semantics. Our algorithm blends global and local styles, then smooths the blended boundary, and improves the deficiencies in the existing methods.”, Fig. 3, “We propose an adaptive multi-scale Blending Decoder. Using a multi-scale fusion structure to connect the equivalent feature maps of transformers and CNNs, bridge transformer decoders and CNN decoders, which blends global and local styles.”, pg. 7, “Spatial fusion weights at each scale are obtained through back propagation. The middle output y is decoded as a fusion map after the style transfer.”); and
based on the adjustment parameter set, generating a retouch result by adjusting the input image (Yan, pg. 12, Fig. 6, “Our model shows the best results, balancing global and local styles, maintaining the original information of the object, and smoother with the surrounding gradients.”).
Regarding claim 2, Yan teaches the method of claim 1, wherein the local feature representation is generated by inputting the input image (background source image), the region information (mask), and the style information (background pattern to be transferred to the local region) into a neural local encoder (Yan, Fig. 2, “local style transfer based on CNN (pink part)”), and the global feature representation is generated by inputting the input image into a neural global encoder (Yan, section 3.2, pg. 6, “We use transformer encoder [9] as another encoder, which contains a style encoder and a content encoder.”. Per paragraph 47 of the instant specification, “the neural global encoder 210 and/or the neural local encoder 220 may correspond to a convolutional neural network or a transformer encoder. The neural harmonic decoder 230 may correspond to a transformer decoder.” A transformer-based encoder is a neural encoder.).
Regarding claim 3, Yan teaches the method of claim 1, wherein the adjustment parameter set (Yan, pg. 7, “Spatial fusion weights at each scale are obtained through back propagation. The middle output y is decoded as a fusion map after the style transfer.”) is generated by inputting the local feature representation, the global feature representation, and the region information into a neural harmonic decoder (Yan, Fig. 2, “Blending Decoder”, pg. 3, “We propose a novel Blending Decoder that learns to extract global and local information from the ViT and CNNs, and blends this information to make the pasted foreground have a more reasonable style.”, pg. 6, “CNN decoder path” and “transformer decoder layer contains multi-head attention and a Factorization Machine supported Neural Network (FNN).” The term “harmonic” in the phrase “harmonic decoder” refers to the type of “harmony” which is subsequently described in claim 4. The term “harmonic” is not interpreted as referring to harmonics or frequencies, but rather the aforementioned “harmony”.).
Regarding claim 4, Yan teaches the method of claim 3, wherein the neural harmonic decoder is configured to, based on harmony (style similarity) between the local feature representation and the global feature representation (The stylized image produced by the Blending Decoder, as shown in Figure 2 of Yan, is based on an agreement of style between the local and global feature representations, i.e., style transfer.), determine the adjustment parameter set (The Blending Decoder determines how the foreground object should look stylistically to be a closer match with the background and GradGAN determines how the foreground-background transition should behave in the gradient domain, where the final fusion tries to satisfy both goals. See Yan at sections 3.2 and 3.3. The “style information of the stylized image” produced by the Blending Decoder “and the gradient information of the fusion image” determined by GradGAN are combined in a fusion operation “to make the final output more harmonious”. See Yan at section 3.3. The Blending Decoder is an adaptive global-local feature harmonizer where “harmony” is achieved via the learned balance between global style coherence and local style near the foreground object’s insertion location.).
Regarding claim 5, Yan teaches the method of claim 1, wherein the region information (The binary mask divides the global image area into two local regions: an object region and a background region. See Yan at pg. 5, Figure 2.) comprises first region information about a first local region and second region information about a second local region (The mask segregates first and second regions. See Yan at Fig. 4, “mask”), and the style information comprises first style information to be applied to the first local region and second style information to be applied to the second local region (Yan, pg. 3, “global style transfer” and “local style transfer”. The pasted object and the background image are blended, which uses features of both regions to blend the local style with the global style in a harmonious manner such that the pasted object appears more naturally within the context of the background. See Yan at section 2.2. The final output image is a hybrid of both styles. Therefore, the first region influences the second region and vice versa.).
Regarding claim 6, Yan teaches the method of claim 5, wherein the generating of the local feature representation comprises, based on the input image, the first region information, the second region information, the first style information, and the second style information (Yan’s model takes as inputs, an input composite image having first/second region information (object and background boundaries and corresponding image data) , first and second style information (object and background styles)) See Yan at Figure 2. Local and global style transfer (first and second style information) are determined and then provided to the blending decoder. See Yan at Figure 2.), generating a first local feature representation corresponding to the first region information and the first style information (Background style derived from the “green part”. See Yan at Figure 2) and a second local feature representation corresponding to the second region information and the second style information (Local/object style derived from the “pink part”. See Yan at Figure 2).
Regarding claim 7, Yan teaches the method of claim 6, wherein the determining of the adjustment parameter set comprises, based on the first local feature representation, the second local feature representation, the global feature representation, the first region information, and the second region information, determining the adjustment parameter set (The final output image is a result of all the preceding components of the model. See Yan at Figure 2.).
Regarding claim 9, Yan teaches the method of claim 1, further comprising, based on the input image, generating a region mask corresponding to the region information (Yan, Fig. 4, “mask and mask_dilated corresponding to the foreground”. The mask segments the background region from the foreground region.).
Regarding claim 10, Yan teaches the method of claim 1, wherein the style information comprises a target style image of the target style (Yan, pg. 2, Fig. 1, “transfer the style of the background”. The composite image contains the target style. See Yan at Figures 1 and 2.)
Claim 11 substantially corresponds to claim 1 by reciting a non-transitory computer-readable storage medium (Yan’s model comprises trained transformer and convolutional networks, which require their parameters and the corresponding algorithmic instructions to be stored in a non-transitory computer-readable storage medium in order to obtain the results discussed in section 4.2.) storing instructions that, when executed by one or more processors (The Adam optimizer requires a processor. See Yan at section 4.1), configure the one or more processors to perform the method of claim 1. Therefore, claim 11 is rejected for the same reasons as claim 1.
Claims 12-20 substantially correspond to claims 1-7, 9 and 10 by reciting an electronic device comprising: one or more processors (The Adam optimizer requires a processor. See Yan at section 4.1) configured to perform the methods of claims 1-7, 9 and 10. Therefore, claims 12-20 are rejected for the same reasons as claims 1-7, 9 and 10.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Yan in view of Semantic Context-Aware Image Style Transfer to Liao et al. (hereinafter “Liao”).
Regarding claim 8, Yan teaches the method of claim 1, but does not teach that which is explicitly taught by Liao.
Liao teaches based on semantic segmentation (Liao, Figure 1, “Semantic Context Matching”), determining a local region (object category/label) from an input image (Liao, abstract, “The semantic context matching aims to obtain the corresponding regions between the content and style images by using context correlations of different object categories. Based on the matching results, we retrieve semantic context pairs where each pair is composed of two semantically matched regions from the content and style images.” Figure 1 of Liao shows an input image being semantically segmented into sky, ground, and object regions before local and global networks process the input image.).
Yan discloses an image style transfer method that encodes global and local features to harmonize the style of an image object pasted into another image having a different style than the background. Thus, Yan shows that it was known in the art before the effective filing date of the claimed invention to transfer the style of an image into a specific local object of a different style, which is analogous to the claimed invention in that it is pertinent to the problem being solved by the claimed invention, accurately enhancing images through a global-local style transfer. Liao discloses an image style transfer method that encodes global and local features to harmonize the style of two different images, where semantic context matching is used as an initial step to segment/label an input image according to different semantic categories. Thus, Liao shows that it was known in the art before the effective filing date of the claimed invention to determine local regions in style transfer methods based on semantic segmentation, which is analogous to the claimed invention in that it is pertinent to the problem being solved by the claimed invention, accurately enhancing images through a global-local style transfer..
A person of ordinary skill in the art would have been motivated to combine Yan’s object mask processing with Liao’s semantic object segmentation, to thereby generate a semantic mask according to a specific object category before deriving the style of respective image regions. Based on the foregoing, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have made such modification according to known methods to yield the predictable results to have the benefit of giving a user more flexibility in specifying the particular object to serve as the target object.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to RYAN P POTTS whose telephone number is (571)272-6351. The examiner can normally be reached M-F, 9am-5pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Sumati Lefkowitz can be reached at 571-272-3638. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/RYAN P POTTS/ Examiner, Art Unit 2672
1 See https://web.archive.org/web/20190831112309/https://www.merriam-webster.com/dictionary/harmony.