DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
1. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
2. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
3. Claims 1, 2, 7, 8, 11, 13, 14, 17, 19, 20, 23, 25, 26, 29, and 31-35 are rejected under 35 U.S.C. 103 as being unpatentable over Li et al., US 2019/0244329 A1, and in view of Park et al. US 2022/0148241 A1, and further in view of Weissenborn et al. US 2021/0383199 A1.
4. As per claim 1, Li discloses: One or more processors, comprising: circuitry to use one or more neural networks to generate one or more images depicting an object by modifying one or more features of the object based, at least in part, on: (Li, [0027]:” The photo style transfer neural network model 110 receives a photorealistic content image I.sub.C and a photorealistic style image I.sub.S and generates a stylized photorealistic image Y that includes the content of the content image modified according to the style image.”, and [0029]:”FIG. 1B illustrates a style image, content image, and stylized photorealistic image, in accordance with an embodiment. The photorealistic style image I.sub.S is input to the encoder 165 and the photorealistic content image I.sub.C is input to the second encoder 165. The photorealistic style image and the photorealistic content image I.sub.C are processed by the photo style transfer neural network model 110 to produce the stylized photorealistic image Y. The cloud pattern in the photorealistic content image is retained in the stylized photorealistic image while a blue color of the sky and the green color of the landscape in the photorealistic style image appear in the stylized photorealistic image—the color of the sky and the landscape areas is changed compared with the photorealistic content image. The shape of the road in the stylized photorealistic image is consistent with the road in the photorealistic content image, while the color is changed to be similar to the road in the photorealistic style image. In addition to transferring color, the photo style transfer neural network model 110 may also be configured to synthesize patterns contained in the photorealistic content image, such as a cloud, snow, rain, and the like, in the stylized photorealistic image to be consistent with the content of the photorealistic content image.”)
5. Li doesn’t expressly disclose:
one or more appearance features, extracted by a first encoder of the one or more neural networks, of one or more first objects from one or more first images and
one or more structural features, extracted by a second encoder of the one or more neural networks, of one or more second objects from one or more second images; and
a decoding of the one or more appearance features of the first object and the one or more structural features of the second object, that have been encoded into one or more latent spaces of the one or more neural networks.
6. Park discloses:
one or more appearance features, extracted by a first encoder of the one or more neural networks, of one or more first objects from one or more first images (Park, [0063], “In particular, the deep image manipulation system 102 utilizes the encoder neural network 306 to extract a structure code 308 and a texture code 310 from the first digital image 302. Indeed, the deep image manipulation system 102 applies the encoder neural network 306 to the first digital image 302 to extract structural features for the structure code 308 and textural features for the texture code 310.”, and [0074], “In some embodiments, the encoder neural network 306 (E) includes or represents two different encoders: a structural encoder neural network E.sub.s and a textural encoder neural network E.sub.t that extract structure codes and texture codes, respectively.”,and [0016], “In other embodiments, the encoder manager 1002 extracts a structure code from a first digital image and extracts a texture code from a second digital image.”) and
one or more structural features, extracted by a second encoder of the one or more neural networks, of one or more second objects from one or more second images; (Park ,[0064],” In a similar fashion, the deep image manipulation system 102 utilizes the encoder neural network 306 to extract the structure code 312 and the texture code 314 from the second digital image 304. More specifically, the deep image manipulation system 102 extracts structural features from the second digital image 304 for the structure code 312. In addition, the deep image manipulation system 102 extract textural features from the second digital image 304 for the texture code 314.”, and [0074], “In some embodiments, the encoder neural network 306 (E) includes or represents two different encoders: a structural encoder neural network E.sub.s and a textural encoder neural network E.sub.t that extract structure codes and texture codes, respectively.” and [0016], “In other embodiments, the encoder manager 1002 extracts a structure code from a first digital image and extracts a texture code from a second digital image.”) and
7. Park is analogous art with respect to Li because they are from the same field of endeavor, namely image processing. At the time the application was filed, it would have been obvious to a person of ordinary skill in the art to include the process:” one or more appearance features, extracted by a first encoder of the one or more neural networks, of one or more first objects from one or more first images and one or more structural features, extracted by a second encoder of the one or more neural networks, of one or more second objects from one or more second images; and a decoding of the one or more appearance features of the first object and the one or more structural features of the second object, that have been encoded into one or more latent spaces of the one or more neural networks.”, as taught by Park into the teaching of Li. The suggestion for doing so would provide a flexibility when applied to editing real images. Therefore, it would have been obvious to combine Park with Li.
8. Li in view of Park doesn’t disclose:
a slot attention transformer that generates one or more transformed feature vectors based, at least in part, on the one or more appearance features and the one or more structural features, wherein the one or more appearance features and the one or more structural features are received as inputs to the slot attention transformer.
9. Weissenborn discloses:
a slot attention transformer that generates one or more transformed feature vectors based, at least in part, on the one or more appearance features and the one or more structural features, wherein the one or more appearance features and the one or more structural features are received as inputs to the slot attention transformer. (Weissenborn, [0050], “FIG. 3 illustrates a block diagram of slot attention module 300. Slot attention module 300 may include value function 308, key function 310, query function 312, slot attention calculator 314, slot update calculator 316, slot vector initializer 318, and neural network memory unit 320. Slot attention module 300 may be configured to receive as input perceptual representation 302, which may include feature vectors 304-306. Slot attention module 300 may be configured to generate slot vectors 322-324 based on perceptual representation 302.”, and [0021] “Accordingly, the slot attention module may be configured to generate a plurality of entity-centric representations, referred to herein as slot vectors, based on a plurality of distributed representations, referred to herein as feature vectors. Each slot vector may be an entity-specific semantic embedding that represents the attributes or properties of one or more corresponding entities. The slot attention module may thus be considered to be an interface between perceptual representations and a structured set of variables represented by the slot vectors.” and [0053], “Encoding the position associated with each respective feature vector of feature vectors 304-306 as part of the respective feature vector, rather than by way of the order in which the respective feature vector is provided to slot attention module 300, allows feature vectors 304-306 to be provided to slot attention module 300 in a plurality of different orders.”)
10. Weissenborn is analogous art with respect to Li in view of Park because they are from the same field of endeavor, namely image processing. At the time the application was filed, it would have been obvious to a person of ordinary skill in the art to include the process of that the one or more structural features of the one or more second objects 2and the one or more appearance features are transformed into one or more transformed 3feature vectors using a slot attention transformer, as taught by Weissenborn into the teaching of Li in view of Park. The suggestion for doing so would carry out the processing of data faster and/or utilize fewer computing resources for the processing. Therefore, it would have been obvious to combine Weissenborn with Li in view of Park.
11. As per claim 2, Li in view of Park and in view of Weissenborn discloses: The one or more processors of claim 1, wherein individual appearance features, of 2the one or more appearance features, are associated with respective semantic regions of the one or first mages. (Li, [0029]:”FIG. 1B illustrates a style image, content image, and stylized photorealistic image, in accordance with an embodiment. The photorealistic style image I.sub.S is input to the encoder 165 and the photorealistic content image I.sub.C is input to the second encoder 165. The photorealistic style image and the photorealistic content image I.sub.C are processed by the photo style transfer neural network model 110 to produce the stylized photorealistic image Y. The cloud pattern in the photorealistic content image is retained in the stylized photorealistic image while a blue color of the sky and the green color of the landscape in the photorealistic style image appear in the stylized photorealistic image—the color of the sky and the landscape areas is changed compared with the photorealistic content image. The shape of the road in the stylized photorealistic image is consistent with the road in the photorealistic content image, while the color is changed to be similar to the road in the photorealistic style image. In addition to transferring color, the photo style transfer neural network model 110 may also be configured to synthesize patterns contained in the photorealistic content image, such as a cloud, snow, rain, and the like, in the stylized photorealistic image to be consistent with the content of the photorealistic content image.”, and [0058])
12. As per claim 5, Li in view of Park and in view of Weissenborn discloses: The processor of claim 1, The one or more processors of claim 1, wherein the one or more neural networks include one or more first convolutional neural networks (CNNs) to extract the one or more appearance features from the one or more second objects, using a second encoder of the one or more second CNNs, and one or more second first convolutional neural networks (CNNs) to extract the one or more structural features from the one or more second objects, using a second encoder. (Li, [0029]. “FIG. 1B illustrates a style image, content image, and stylized photorealistic image, in accordance with an embodiment. The photorealistic style image I.sub.S is input to the encoder 165 and the photorealistic content image I.sub.C is input to the second encoder 165. The photorealistic style image and the photorealistic content image I.sub.C are processed by the photo style transfer neural network model 110 to produce the stylized photorealistic image Y.”, [0030], and [0035])
13. Claims 7, 13, 19, and 25 which are similar in scope to claim 1, thus rejected under the same rationale.
14. Claims 8, 14, 20, and 26, which are similar in scope to claim 2, thus rejected under the same rationale.
15. Claims 11, 17, 23, and 29 which are similar in scope to claim 5, thus rejected under the same rationale.
16. As per claim 31, Li in view of Park, and in view of Weissenborn discloses: The one or more processors of claim 1, wherein the one or more images are generated based, at least in part, on a decoding of the one or more appearance features of the first object and the one or more structural features of the second object, wherein the one or more appearance features and the one or more structural features have been encoded into one or more latent spaces of the one or more neural networks. (Park, [0070], "where x.sup.1 represents a latent code representation of the first digital image 402, x.sup.2 represents a latent code representation of the second digital image 404, z.sub.s.sup.1 represents the structure code 406 from the first digital image 402, z.sub.t.sup.2 represents the texture code 412 from the second digital image 404, z.sub.l.sup.1 represents the scene layout of x.sup.1, and the other terms are defined above.", and [0074], "where X represents the first digital image 402, H represents the height of the image, W represents the width of the image, and 3 is the number of channels in an RGB image (i.e. red, green, and blue). For example, the encoder neural network 306 maps the first digital image 402 to a latent space Z, and the generator neural network 318 generates the reconstructed digital image 418 from the encoding in the latent space Z.")
17. Claims 32-35 which are similar in scope to claim 31, thus rejected under the same rationale.
18. Claims 4, 6, 10, 12, 16, 18, 22, 24, 28 and 30 are rejected under 35 U.S.C. 103 as being unpatentable over Li et al., US 2019/0244329 A1, in view of Park et al. US 2022/0148241 A1, and further in view of Weissenborn et al. US 2021/0383199 A1,
and further in view of Barzelay et al. US 2021/0048931 et al.
19. As per claim 4, Li in view of Park, and in view of Weissenborn discloses: Theo or more processors of claim 1, (See rejection of claim 1 above.)
20. Li in view of Park, and in view of Weissenborn doesn’ t expressly discloses: the one or more neural networks 2include a generative adversarial network (GAN) for receiving the one or more transformed 3feature vectors and generate one or more representations.
21. Barzelay discloses: the one or more neural networks 2include a generative adversarial network (GAN) for receiving the one or more transformed 3feature vectors and generate one or more representations. (Barzelay, [0035], “In accordance with at least one embodiment a generative adversarial network (GAN) may be used to reconstruct images given a first image. As will be apparent to one of skill in the art, a generative adversarial network has an aspect that encodes objects (e.g., images) as feature vectors and an aspect that decodes objects from feature vectors (e.g., generates images from feature vector.”)
22. Barzelay is analogous art with respect to Li in view of Park, and in view of Weissenborn because they are from the same field of endeavor, namely image processing. At the time the application was filed, it would have been obvious to a person of ordinary skill in the art to include the process of that the one or more neural networks 2include a generative adversarial network (GAN) for receiving the one or more transformed 3feature vectors and generate one or more representations, as taught by Barzelay into the teaching of Li in view of Park, and in view of Sekharan. The suggestion for doing so would provide a more accurate interpretation of an image of an object. Therefore, it would have been obvious to combine Barzelay with Li in view of Park, and in view of Weissenborn.
23120. As per claim 6, Li in view of Park, and in view of Weissenborn, and in view of Barzelay discloses: The one or more processors of claim 1, wherein the structural features are sampled from 2one or more feature distributions for one or more object types of the one or more second images. (Barzelay, [0030],” In embodiments, the neural network may classify a set of images as according to their shape, color, and texture, as a vector of continuous valued numbers (feature vectors). However, the shape, color, and texture of the image are not explicitly represented by the vector of continuous valued numbers. Instead, the feature vectors are used to identify similar images based on their respective vectors or values having a small Euclidean distance from the source image.”). The proposed combination as well as the motivation for combining the references presented in the rejection of the claim 4 apply to this claim and are incorporated herein by reference.
24. Claims 10, 16, 22, and 28 which are similar in scope to claim 4, thus rejected under the same rationale.
25. Claims 12, 18, 24, and 30, which are similar in scope to claim 6, thus rejected under the same rationale.
Response to Arguments
26. Applicant’s arguments with respect to claims filed 04/22/2026 have been considered but are moot because Applicant submitted new amended claims. Accordingly, new grounds of rejection are set forth above. The new grounds of rejection conclusion have been necessitated by Applicant's amendments to the claims.
Conclusion
27. Applicants amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ABDERRAHIM MEROUAN whose telephone number is (571)270-5254. The examiner can normally be reached on Monday to Friday 7:30 AM to 5:00 PM.
The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ABDERRAHIM MEROUAN/Supervisory Patent Examiner, Art Unit 2683