Prosecution Insights
Last updated: August 06, 2026
Application No. 19/032,713

SYSTEMS AND METHODS FOR CONSTRUCTING CUSTOM PHOTOS BASED ON SKELETAL FEATURE POINT INPUT

Non-Final OA §102§103§112
Filed
Jan 21, 2025
Priority
Jan 26, 2024 — provisional 63/625,453
Examiner
BADER, ROBERT N.
Art Unit
Tech Center
Assignee
Perfect Mobile Corp.
OA Round
1 (Non-Final)
44%
Grant Probability
Moderate
1-2
OA Rounds
1y 10m
Est. Remaining
70%
With Interview

Examiner Intelligence

Grants 44% of resolved cases
44%
Career Allowance Rate
177 granted / 399 resolved
-15.6% vs TC avg
Strong +26% interview lift
Without
With
+25.7%
Interview Lift
resolved cases with interview
Typical timeline
3y 4m
Avg Prosecution
26 currently pending
Career history
429
Total Applications
across all art units

Statute-Specific Performance

§101
11.9%
-28.1% vs TC avg
§103
48.1%
+8.1% vs TC avg
§102
12.9%
-27.1% vs TC avg
§112
20.3%
-19.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 399 resolved cases

Office Action

§102 §103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 5, 6, 11, 12, 17, and 18 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claims 5, 12, and 18 recite the limitation “training a diffusion model using the target image each depicting the target individual”. The target image is singular, i.e. the independent claims establish that “a target image” is obtained, whereas the word “each” suggests there may be more than one target image, conflicting with the independent claim limitation establishing that the target image is singular. That is, claims 5, 12, and 18 are indefinite because “each” suggests there should be a plurality of target images depicting the target individual rather than only one target image. Depending claim 6 does not clarify this issue and is similarly rejected. For purposes of applying prior art, the claims will be read as though they did not recite “each”, i.e. “training a diffusion model using the target image Claims 11 and 17 recites the limitation "the at least one of the attributes of the at least one individual depicted in the input image being mirrored". There is insufficient antecedent basis for this limitation in the claim. The “mirror at least one of the attributes depicted in the input image” limitation is introduced in claims 12 and 18, such that although claim 6, depending from claim 5, has proper antecedent basis, claims 11 and 17, depending from claims 10 and 16, respectively, lack sufficient antecedent basis. For purposes of applying prior art, claims 11 and 17 will be treated as depending on claims 12 and 18, respectively, corresponding to the scope of similar claim 6. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 1-4, 7-10, and 13-16 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by U.S. Patent 11,854,203 B1 (hereinafter Gafni). Regarding claim 1, the limitation “a method implemented in a computing device, comprising: obtaining an input image depicting at least one individual; obtaining a target image depicting a target individual” are taught by Gafni (Gafni, e.g. abstract, cols 4-12, discloses a system for performing context aware human generation for adding a target person to an existing captured photo of one or more people. Gafni, e.g. col 4, line 10 - col 5, line 10, teaches that each output image is generated using a first input image depicting at least one person and a second input image depicting a target person to be inserted into the first input image, corresponding to the claimed input image depicting at least one individual and target image depicting a target individual. Finally, Gafni’s system is implemented using a computing device, which comprises a processor executing instructions stored on a non-transitory medium/memory as further required by similar independent claims 7 and 13, e.g. Gafni, col 13, line 5 - col 15, line 17, describes the computing device used to implement the system.) The limitation “obtaining a location within the input image for inserting the target individual” is taught by Gafni (Gafni, e.g. col 5, lines 12-65, teaches that the computing device generates a source segmentation mask for the one or more persons in the first input image, and uses the source segmentation mask to generate a target segmentation mask for the target person in the context of the first image, wherein each segmentation mask comprises a semantic pose map channel and a face channel, each pixel thereof indicating a semantic pose label of background or a part of a person. That is, the target segmentation mask identifies the location within the input image, i.e. the non-background labeled pixels, for inserting the target individual. Further, Gafni, e.g. col 5, lines 31-34, col 7, lines 1-15, indicates that the user may provide a bounding box indicating the location for inserting the target person, but that the computing device may determine the bounding box when the user does not provide the bounding box. That is, both the bounding box and the generated target segmentation mask correspond to the claimed obtained location within the input image for inserting the target individual.) The limitation “determining attributes of the at least one individual depicted in the input image” are taught by Gafni (As noted above, Gafni, e.g. col 5, lines 12-65, teaches that the computing device generates a source segmentation mask for the one or more persons in the first input image, wherein the segmentation mask comprises a semantic pose map channel and a face channel, each pixel thereof indicating a semantic pose label of background or a part of a person. Further, Gafni, e.g. col 8, line 45 - col 9, line 37, describes rendering the target person into the first image with not only a new pose, but a new facial expression, indicating that the face channel corresponds to a facial expression. That is, the source segmentation map semantic pose map channel and face channel correspond to the claimed attributes of the at least one individual depicted in the input image, wherein with respect to depending claims 3 and 4, the pose map channel includes a pose of each individual, the shape of each segment group making up the body, upper and lower body clothing, shoes, and contextual details, and the face channel includes an expression of the individual.) The limitations “synthesizing a replicate of the target individual based on the attributes of the at least one individual; and inserting the synthesized replicate of the target individual into the input image based on the location to generate a modified image” are taught by Gafni (Gafni, e.g. col 7, line 20 - col 9, line 37, figure 3, describes using a second machine learning model to generate a third image comprising the target person with a new pose and optionally new facial expression using the target segmentation mask, where as discussed above, the target segmentation mask is generated based on the source segmentation mask comprising the attributes of the individual(s) in the first image, i.e. the third image generated by the second machine learning model, 330 in figure 3, is the claimed synthesized replicate of the target individual based on the attributes of the at least one individual. Further, Gafni, e.g. col 9, lines 4-37, describes generating the output image by compositing the third image into the first image using a blending mask which is based on the target segmentation mask, e.g. col 7, lines 57 - 66, where, as discussed above, the target segmentation mask is defined at the obtained location for inserting the target individual, i.e. the output image generated by the compositing operation corresponds to the claimed modified image generated by inserting the synthesized replicate of the target individual into the input image based on the location.) Regarding claim 2, the limitation “wherein obtaining the location within the input image for inserting the target individual further comprises automatically determining whether to insert the target individual to the right or to the left of the at least one individual in the input image” is taught by Gafni (As discussed in the claim 1 rejection above, Gafni, e.g. col 5, lines 31-34, col 7, lines 1-15, indicates that the user may provide a bounding box indicating the location for inserting the target person, but that the computing device may determine the bounding box when the user does not provide the bounding box, i.e. the bounding box may be automatically determined, and by extension, the target segmentation mask generated based on the automatically determined bounding box would also be automatically determined. Further, as discussed in the claim 1 rejection above, both the bounding box and the generated target segmentation mask correspond to the claimed obtained location within the input image for inserting the target individual, i.e. Gafni teaches automatic determination for both of the mappings. Finally, Gafni, e.g. figures 1-2 shows an example where the bounding box/target segmentation mask is to the left of the at least one individual, but Gafni does not indicate any limitation on the relative placement of the target individual relative to the at least one individual, i.e. one of ordinary skill in the art would understand that the system could determine the bounding box/target segmentation mask is to the right of the at least one individual, analogous to the training example of figure 5.) Regarding claims 3 and 4, the limitations “wherein the attributes of the at least one individual depicted in the input image comprise at least one of: an expression of the at least one individual, a gaze of the at least one individual, a pose of the at least one individual, a body shape of the at least one individual, clothing worn by the at least one individual, accessories worn by the at least one individual, or environmental lighting depicted in the input image” and “wherein the attributes of the at least one individual depicted in the input image further comprise at least one of: contextual details depicted in the input image and emotions depicted in the input image” are taught by Gafni (As discussed in the claim 1 rejection above, Gafni, e.g. col 5, lines 12-65, teaches that the computing device generates a source segmentation mask for the one or more persons in the first input image, wherein the segmentation mask comprises a semantic pose map channel and a face channel, each pixel thereof indicating a semantic pose label of background or a part of a person. Further, Gafni, e.g. col 8, line 45 - col 9, line 37, describes rendering the target person into the first image with not only a new pose, but a new facial expression, indicating that the face channel corresponds to a facial expression. That is, the source segmentation map semantic pose map channel and face channel correspond to the claimed attributes of the at least one individual depicted in the input image, wherein with respect to the limitations of claims 3 and 4, the pose map channel includes a pose of each individual, the shape of each segment group making up the body, upper and lower body clothing, shoes, and contextual details, and the face channel includes an expression of the individual. It is noted that the face expressions in the face channel correspond to both the expression of the at least one individual attribute in claim 3, and the attribute of the emotions depicted in the input image as in claim 4, as different facial expressions correspond to different emotions. Further, the semantic pose segment groups indicate semantic context, as noted by Gafni, col 4, lines 10-14, i.e. the segment group labels correspond to attributes comprising contextual details depicted in the input image as in claim 4.) Regarding claims 7 and 13, the limitations are similar to those treated in the above rejection(s) and are met by the references as discussed in claim 1 above. Regarding claims 8 and 14, the limitations are similar to those treated in the above rejection(s) and are met by the references as discussed in claim 2 above. Regarding claims 9, 10, 15, and 16, the limitations are similar to those treated in the above rejection(s) and are met by the references as discussed in claims 3 and 4 above. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 5, 6, 11, 12, 17, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over U.S. Patent 11,854,203 B1 (hereinafter Gafni) as applied to claims 1, 7, and 13 above, and further in view of “Wish You Were Here: Context-Aware Human Generation” by Oran Gafni and Lior Wolf (hereinafter Wolf). Regarding claim 5, the limitation “training a diffusion model using the target image line 31, Gafni does not explicitly teach that the second machine learning model is trained using the input target image and target segmentation mask generated based on the claimed attributes.) However, this limitation is taught by Wolf (Wolf, e.g. abstract, sections 1, 3-6, describes the same system disclosed by Gafni, i.e. the references are authored by the same two inventors and describe the same system for inserting a target person into another photograph using semantic context, although each reference focuses on different details. Wolf, sections 3, 3.2, 4, describe the multi-condition rendering network (MCRN) corresponding to Gafni’s second machine learning network, e.g. Gafni figure 3 is the same as Wolf figure 3(a). Further, Wolf, e.g. section 4, paragraph 1, figure 3(b), teaches that the MCRN is trained for each target person separately using the sub-images from the target image and the semantic map p, i.e. as claimed, Wolf’s MCRN is a diffusion model which is trained using the target image depicting the target individual and based on the attributes of the at least one individual depicted in the input image, i.e. the semantic map p, corresponding to Gafni’s target segmentation mask. That is, while Gafni indicates that a variety of training techniques are possible, Gafni does not describe details of training the second machine learning model, thereby motivating one of ordinary skill in the art to use Wolf’s separate/per-person MCRN training technique for Gafni’s second machine learning model.) Therefore it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to implement Gafni’s system using Wolf’s separate/per-person MCRN training technique for Gafni’s second machine learning model because Gafni does not describe details of training the second machine learning model and Wolf does describe details of training a corresponding machine learning model, thereby motivating one of ordinary skill in the art to use Wolf’s disclosed training details. As noted above, Wolf, e.g. section 4, paragraph 1, figure 3(b), teaches that the MCRN is trained for each target person separately using the sub-images from the target image and the semantic map p, corresponding to Gafni’s target segmentation mask, where the target segmentation mask is generated based on the source segmentation mask comprising the attributes of the individual(s) in the first image, i.e. in the modified system, as claimed, Gafni’s second machine learning model is a diffusion model which is trained using the target image depicting the target individual and based on the attributes of the at least one individual depicted in the input image, i.e. Gafni’s target segmentation mask. The limitation “performing inpainting on a region in the input image using the diffusion model based on the location to fill in the region in the input image with the replicate of the target individual and to mirror at least one of the attributes of the at least one individual depicted in the input image in the replicate of the target individual” is taught by Gafni (As discussed in the claim 1 rejection above, Gafni, e.g. col 7, line 20 - col 9, line 37, figure 3, describes using the second machine learning model to generate a third image comprising the target person with a new pose and optionally new facial expression using the target segmentation mask, and further, Gafni, e.g. col 9, lines 4-37, describes generating the output image by compositing the third image into the first image using a blending mask which is based on the target segmentation mask, e.g. col 7, lines 57 - 66. That is, the target segmentation mask comprises the claimed region(s) based on the obtained location, where the region(s) are inpainted with the third image of the target person rendered by the second machine learning/diffusion model, corresponding to the claimed inpainting. Further, the third image of the target person rendered by the second machine learning/diffusion model mirrors the attributes of the individuals in the input image, i.e. the new pose and expression are defined to match the attributes/semantic context of the individuals in the first image, as in Gafni, cols 5-6, e.g. as in the example of figures 1-2, the target person is rendered with a pose aligned with the individuals in the first input image.) Regarding claim 6, the limitation “wherein the at least one of the attributes of the at least one individual depicted in the input image being mirrored is specified using textual input or is selected by artificial intelligence” is taught by Gafni (Gafni, e.g. col 5, line 12 - col 6, line 67, describes generating the source and target segmentation masks, where, as discussed in the claim 1 rejection above, the source segmentation map semantic pose map channel and face channel correspond to the claimed attributes of the at least one individual depicted in the input image. Gafni, e.g. col 5, lines 21-25, col 6, lines 3-6, 23-33, teaches that the source segmentation mask may be generated using one or more pre-trained machine learning models, i.e. the pose attributes of the individual(s) in the input image being mirrored are specified using artificial intelligence.) Regarding claims 11 and 17, the limitations are similar to those treated in the above rejection(s) and are met by the references as discussed in claim 6 above. Regarding claims 12 and 18, the limitations are similar to those treated in the above rejection(s) and are met by the references as discussed in claim 5 above. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to ROBERT BADER whose telephone number is (571)270-3335. The examiner can normally be reached 11-7 m-f. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tammy Goddard can be reached at 571-272-7773. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ROBERT BADER/Primary Examiner, Art Unit 2611
Read full office action

Prosecution Timeline

Jan 21, 2025
Application Filed
Jul 29, 2026
Non-Final Rejection mailed — §102, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12700148
ELECTRONIC DEVICE FOR GENERATING USER-PREFERRED CONTENT, AND OPERATING METHOD THEREFOR
1y 11m to grant Granted Aug 04, 2026
Patent 12682547
3D MODEL RENDERING USING IMPORTANCE SAMPLING
2y 8m to grant Granted Jul 14, 2026
Patent 12682548
APPARATUS AND METHOD USING TRIANGLE PAIRS AND SHARED TRANSFORMATION CIRCUITRY TO IMPROVE RAY TRACING PERFORMANCE
2y 4m to grant Granted Jul 14, 2026
Patent 12651399
TRAINING DATA SAMPLING FOR NEURAL NETWORKS
2y 8m to grant Granted Jun 09, 2026
Patent 12646247
SPATIOTEMPORAL RESAMPLING WITH DECOUPLED SHADING AND REUSE
2y 2m to grant Granted Jun 02, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
44%
Grant Probability
70%
With Interview (+25.7%)
3y 4m (~1y 10m remaining)
Median Time to Grant
Low
PTA Risk
Based on 399 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month