Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1-30 are pending.
Response to Arguments
Arguments presented with the RCE of 6/18/2026 have been considered but they are moot in view of a new ground of rejection.
Claim Rejections - 35 USC § 112
.The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112:
The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention.
Claims 1-30 is/are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention.
Independent claim 1 recites, in part, “cause one or more neural networks to use one or more classes of one or more objects depicted with a first pose in one or more first images and one or more other objects in one or more second images to determine a second pose of the one or more objects compatible with one or more appearances of the one or more other objects”
The limitations underlined are either misleading, inconsistent, and otherwise unsupported by the specification.
The claim, as originally filed, uses the poses of the objects in the second images to determine new pose of the objects in the first input image(s). The amendment now shift gears, stating “use one or more classes of one or more objects depicted with a first pose in one or more first images and one or more other objects in one or more second images to determine a second pose”. The examiner asserts it is an unreasonable stretch to say the determining new poses uses the classes as input. At nowhere in the Specification specifically states or disclose “use one or more neural networks to use one or more classes of one or more objects depicted with a first pose in one or more first images and one or more other objects in one or more second images to determine a second pose of the one or more objects “. The Specification discloses classes of the first objects are used to determine the appropriate VAE that specialized in encoding/decoding features of the particular class. The class information is never used as part of a function to determine second pose. Fig. 4-5 of the Specification does not mention classes as input for determining new poses. And even if class information is somehow related to the process, the Specification fails to disclose a full, clear, specific, and concise description of how class information comes to play to determine new poses of the first objects. Any implicit inference of disclosures in attempt to justify the amendment would immediately fail the “full, clear, and concise” requirement.
The claim goes on to state “determine a second pose of the one or more objects compatible with one or more appearances of the one or more other objects”.
Left alone the fact that the language “compatible” is subjective, both the claims and the Specification has no explicit support and guidance as to specify criteria/threshold to determine a compatibility. ¶0048-0049 of the original Specification briefly mention pose compatibility and “appearance guided pose suggestion technique” without explaining any manner to determine said compatibility. The Specification merely offer a soft, high level language about compatibility, yet simply does not offer any sufficient description full and clear to support the specific claim language “compatible with one or more appearances of the one or more other object”, thus such language is clear stretch beyond reasonable doubt. Claims 7, 13, 19, and 25 are directed to similar language and is rejected by the same reasoning. All respective dependent claims fail to remedy the issues above and fall together with the base claims.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 1, 7, 13, 19, 25 is/are rejected under 35 U.S.C. 102(a)(2) as being unpatentable over Remine et al. (US 2020/0192389).
As to claim 1:
Remine discloses:
One or more processors, comprising: circuitry to (Abstract, ¶0014, 0054 apparatus with processor having circuitries): cause one or more neural networks to use one or more classes of one or more objects depicted with a first pose in one or more first images and one or more other objects in one or more second images to determine a second pose of the one or more objects compatible with one or more appearances of the one or more other objects, the one or more neural networks using, as input, the one or more first images and the one or more second images to generate the second pose for the one or more objects to be added to the one or more second images; (See at least ¶0031, the system to insert the objects of the first image with known class of “tree” into the real world scene, this information is to ensure said objects is to be place in coherent manner with existing objects of the real world scene, namely “segments of pixels assigned to sidewalks, driveways and buildings in the images of the real-world street. In these examples, the image generator 103 is configured to insert the image of the simulated object into at least one of the segments based on the object classes to which the segments are assigned. For example, the image generator 103 can insert the image of the simulated tree into the segments assigned to sidewalks as commonly seen in reality” ¶0032, 0035-0036, 0045-0048, 0009 using one or more GANs, the one or more processor to receive one or more images of the real world scene, receive also one or more input images of objects to be incorporated to the real world scene image, generate one or more images of said real world scene including said objects to be incorporated. In particular, pose(s) of objects (tree, street, surface) in the real-world scene image is/are determined by a pose estimator component. The first objects received in the one or more input images is/are incorporated to the real-world scene such that their new poses are based on real world scene’s poses, specifically to appear coherently/realistically match (i.e. compatible with) the pose/orientation of the objects/structures of the real world scene images per example described in at least ¶0030-0031)
and use the one or more neural networks to add the one or more objects to the one or more second images, the added one or more objects in the one or more second images having the second pose, the second pose being different from the first pose. (¶0048, 0035, generating the output image with the real world scene now with the object inserted with a photo-realistic poses based on the context of the real-world scene. Per ¶0032, 0035, 0048, the pose of the inserted object must be determined in consideration of other objects in the scene so that it makes contextual sense, thus the pose is altered in the face of such variables)
As to claim 7:
Remine discloses:
A system comprising: one or more processors to: cause one or more neural networks to use one or more classes of one or more objects depicted with a first pose in one or more first images and one or more other objects in one or more second images to determine a second pose of the one or more objects compatible with one or more appearances of the one or more other objects, the one or more neural networks using, as input, the one or more first images and the one or more second images to generate the second pose for the one or more objects to be added to the one or more second images; (See at least ¶0031, the system to insert the objects of the first image with known class of “tree” into the real world scene, this information is to ensure said objects is to be place in coherent manner with existing objects of the real world scene, namely “segments of pixels assigned to sidewalks, driveways and buildings in the images of the real-world street. In these examples, the image generator 103 is configured to insert the image of the simulated object into at least one of the segments based on the object classes to which the segments are assigned. For example, the image generator 103 can insert the image of the simulated tree into the segments assigned to sidewalks as commonly seen in reality” ¶0032, 0035-0036, 0045-0048, 0009 using one or more GANs, the one or more processor to receive one or more images of the real world scene, receive also one or more input images of objects to be incorporated to the real world scene image, generate one or more images of said real world scene including said objects to be incorporated. In particular, pose(s) of objects (tree, street, surface) in the real-world scene image is/are determined by a pose estimator component. The first objects received in the one or more input images is/are incorporated to the real-world scene such that their new poses are based on real world scene’s poses, specifically to appear coherently/realistically match (i.e. compatible with) the pose/orientation of the objects/structures of the real world scene images per example described in at least ¶0030-0031)
and use the one or more neural networks to add the one or more objects to the one or more second images, the added one or more objects in the one or more second images having the second pose, the second pose being different from the first pose; (¶0048, 0035, generating the output image with the real world scene now with the object inserted with a photo-realistic poses based on the context of the real-world scene. Per ¶0032, 0035, 0048, the pose of the inserted object must be determined in consideration of other objects in the scene so that it makes contextual sense, thus the pose is altered in the face of such variables)
Claim 13 is directed to a method with step(s) similar to those in claim 1 and is rejected by the same reasoning.
Claim 19 is directed to a non-transitory CRM with instructions when performed by a one or more processor to perform a method with step(s) similar to those in claim 1 and is rejected by the same reasoning.
As to claim 25:
Remine discloses:
cause one or more neural networks to use one or more classes of one or more objects depicted with a first pose in one or more first images and one or more other objects in one or more second images to determine a second pose of the one or more objects compatible with one or more appearances of the one or more other objects, the one or more neural networks using, as input, the one or more first images and the one or more second images to generate the second pose for the one or more objects to be added to the one or more second images; (See at least ¶0031, the system to insert the objects of the first image with known class of “tree” into the real world scene, this information is to ensure said objects is to be place in coherent manner with existing objects of the real world scene, namely “segments of pixels assigned to sidewalks, driveways and buildings in the images of the real-world street. In these examples, the image generator 103 is configured to insert the image of the simulated object into at least one of the segments based on the object classes to which the segments are assigned. For example, the image generator 103 can insert the image of the simulated tree into the segments assigned to sidewalks as commonly seen in reality” ¶0032, 0035-0036, 0045-0048, 0009 using one or more GANs, the one or more processor to receive one or more images of the real world scene, receive also one or more input images of objects to be incorporated to the real world scene image, generate one or more images of said real world scene including said objects to be incorporated. In particular, pose(s) of objects (tree, street, surface) in the real-world scene image is/are determined by a pose estimator component. The first objects received in the one or more input images is/are incorporated to the real-world scene such that their new poses are based on real world scene’s poses, specifically to appear coherently/realistically match (i.e. compatible with) the pose/orientation of the objects/structures of the real world scene images per example described in at least ¶0030-0031)
and use the one or more neural networks to add the one or more objects to the one or more second images, the added one or more objects in the one or more second images having the second pose, the second pose being different from the first pose; (¶0048, 0035, generating the output image with the real world scene now with the object inserted with a photo-realistic poses based on the context of the real-world scene. Per ¶0032, 0035, 0048, the pose of the inserted object must be determined in consideration of other objects in the scene so that it makes contextual sense, thus the pose is altered in the face of such variables)
memory for storing network parameters for the one or more neural networks (¶0059, Fig. 3, ¶0004, memory to store instructions and model components)
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 2, 4-6, 8, 10-12, 14, 16-18, 20, 22-24, 26, 28-30 is/are rejected under 35 U.S.C. 103 as being unpatentable over Remine et al. (US 2020/0192389) in view of Lee et al. (US 2020/0074707).
As to claims 2, 8, 14, 20 and 26:
Remine discloses all limitations of claim 1/7/13/19 and 25, but silent on the one or more neural networks 2include one or more variational autoencoders (VAEs) to determine features for the one or more other objects within the one or more second image and encode those features to a latent space to act as a 4constraint in adding the one or more first objects to the image.
Lee discloses a system/method for inserting objects into an existing image in which the one or more neural networks 2include one or more variational autoencoders (VAEs) to determine features for the first 3objects and the second objects and encode those features to a latent space to act as a 4constraint in adding the one or more first objects to the image. (See at least ¶0018, 0019, also, 0026-0028 using at least a VAE, features of the object and background are analyzed, to generate a vector in a latent space that is used to generate location/scale of the object to be added in the scene).
It would have been obvious to one of ordinary skill in the art before the effective filing time of the invention that the GAN system of Remine to include one or more variational autoencoders (VAEs) to determine features for the first 3objects and the second objects and encode those features to a latent space to act as a 4constraint in adding the one or more first objects to the image. Given that Remine uses GAN that generates/render object per abstract, as analogous with the GAN of Lee. A VAE allows for advantage of accurate injection by providing specific location/scale of an object in a scene (Lee, ¶0026).
As to claims 4, 10, 16, 22 and 28:
Remine in view of Lee discloses all limitations of claim 2/8/14/20/26, wherein the one or more neural networks 2include a generative network to determine one or more potential poses for the added one or more 3 objects based at least in part upon object types of the one or more other objects and with 4respect to features of the one or more second objects, wherein information for the potential 5poses is to be encoded into the latent space. (Remine, 0032, 0035-0036, determine potential poses by accessing also classes of objects using GAN. Lee, as discussed in above, discloses determining potential placements that maintain contextual coherence with the scene’s features per ¶0018-0019 , which is encoded in latent space, and See at least ¶0018, 0019, also, 0026-0028 using at least a VAE, features of the object and background are analyzed, to generate a vector in a latent space that is used to generate placement/scale of the object to be added in the scene.)
As to claims 5, 11, 17, 23 and 29:
Remine in view of Lee discloses all limitations of claim 4/10/16/22/28, wherein the one or more neural networks 2include a neural network to determine one or more potential positions for the added one or more 3 objects based at least in part upon the object types of the one or more other objects and potential poses of the one or more 4first objects, and with respect to the features of the one or more other objects, wherein 5information for the potential positions is to be encoded into the latent space(Remine, ¶0035, “the pose estimator 1031 can estimate or determine the position and/or rotational orientation of the real-world tree. The image generator 103 is configured to insert the second image into the images of the real-world scene based on the pose of the real-world object in the second image. In this way, the image generator produces the images of the real-world scene including the simulated object and further including the real-world object. For example, the image generator can generate images of the real-world street including the simulated tree and the real-world tree”, ¶0031, “ insert the image of the simulated object into at least one of the segments based on the object classes to which the segments are assigned”. Lee discloses determining potential placements, which include position and orientation, that maintain contextual coherence with the scene’s features per ¶0018-0019, which is encoded in latent space - See at least ¶0018, 0019, also, 0026-0028 using at least a VAE, features of the object and background are analyzed, to generate a vector in a latent space that is used to generate placement/scale of the object to be added in the scene).
As to claims 6, 12, 18, 24 and 30:
Remine in view of Lee discloses all limitations of claims 5/11/17/23/29, wherein the one or more neural networks 2include a generative adversarial network (GAN) to generate one or more output images including 137 \\NORTHCA - 1R2674/010501 - 2773047 vlthe added one or more objects added to the image, wherein the added one or more objects have different 4poses or positions in the output images, the poses and positions to be selected from the potential 5poses and the potential positions determined from the latent space. (Lee, See at least ¶0018-0019, using neural network (Generative adversarial network) model add a desired object into a desired position of a captured real world scene image. See also Remine, ¶0035, “the pose estimator 1031 can estimate or determine the position and/or rotational orientation of the real-world tree. The image generator 103 is configured to insert the second image into the images of the real-world scene based on the pose of the real-world object in the second image. In this way, the image generator produces the images of the real-world scene including the simulated object and further including the real-world object. For example, the image generator can generate images of the real-world street including the simulated tree and the real-world tree”)
Claim(s) 3, 9, 15, 21 and 27 is/are rejected under 35 U.S.C. 103 as being unpatentable over Remine et al. (US 2020/0192389) in view of Lee et al. (US 2020/0074707) in view of Kopf (Mixture of Expert Variational Autoencoder for Clustering and Generating from Similarity-based Representation (01-2020) (IDS entry) and in further view of Irsoy et al. (Unsupervised feature extraction with autoencoder trees” (2017) – IDS entry.
As to claims 3, 9, 15, 21 and 27:
Remine in view of Lee discloses all limitations of claims 2/8/14/20/26, however is silent on the one or more neural networks 2include a gating network to select the one or more VAEs from a set of VAEs each trained 3for a different class of object, the gating network to select the one or more VAEs using a 4hierarchical mixture-of-experts approach.
Kopf discloses a gating network to select the one or more VAEs from a set of VAEs each trained 3for a different class of object (See Abstract, see page 3, a cluster I is gated to a corresponding expert (VAE), note that an expert has sole expertise in a particular class of object).
It would have been obvious to one of ordinary skill in the art before the effective filing time of the invention that the system/method of Lee to incorporate the feature of gating network to select VAEs as such implementation show superior clustering performance of the model on real world data (See page 2 of Kopf).
None of the above further discloses using hierarchical mixture of expert approach.
Irsoy, however, in a related field of endeavor discloses in Abstract, page 64, Section 3 through page 65, which discusses a soft decision node to direct instance to its branches according to different probability as given a gating function (gating network) in a hierarchical mixture of expert approach. Also Fig. 1, left column of page 64 discusses the gating function.
It would have been obvious to one of ordinary skill in the art before the effective filing time of the invention that the system/method of Lee to incorporate the feature of using hierarchical mixture of expert approach to select VAEs as such implementation improved operational accuracy (Irsoy page 71 - Conclusion)
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US 2019/0304076 - Techniques related to synthesizing an image of a person in an unseen pose are discussed. Such techniques include detecting a body part occlusion for a body part in a representation of the person in a first image and, in response to the detected occlusion, projecting a representation of the body part from a second image having a different view into the first image. A geometric transformation based on a source pose of the person and a target pose is then applied to the merged image to generate a synthesized image comprising a representation of the person in the target pose.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to QUAN M HUA whose telephone number is (571)270-7232. The examiner can normally be reached 10:30-6:30.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Anthony Addy can be reached on 571-272-7795. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/QUAN M HUA/Primary Examiner, Art Unit 2645