Prosecution Insights
Last updated: October 04, 2026
Application No. 18/712,268

INFORMATION PROCESSING SYSTEM, INFORMATION PROCESSING METHOD, AND PROGRAM

Non-Final OA §103§112
Filed
Dec 02, 2024
Priority
Dec 29, 2021 — IT 102021000032969 +1 more
Examiner
RUSH, ERIC
Art Unit
Tech Center
Assignee
Fondazione Istituto Italiano di Tecnologia
OA Round
1 (Non-Final)
61%
Grant Probability
Moderate
1-2
OA Rounds
1y 7m
Est. Remaining
97%
With Interview

Examiner Intelligence

Grants 61% of resolved cases
61%
Career Allowance Rate
392 granted / 645 resolved
+0.8% vs TC avg
Strong +36% interview lift
Without
With
+36.1%
Interview Lift
resolved cases with interview
Typical timeline
3y 5m
Avg Prosecution
22 currently pending
Career history
670
Total Applications
across all art units

Statute-Specific Performance

§101
9.3%
-30.7% vs TC avg
§103
46.9%
+6.9% vs TC avg
§102
12.3%
-27.7% vs TC avg
§112
24.1%
-15.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 645 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Preliminary Amendment This action is responsive to the preliminary amendments and remarks received 12 September 2025. Claims 13 - 32 are currently pending. Specification The abstract of the disclosure is objected to because it exceeds 150 words. A corrected abstract of the disclosure is required and must be presented on a separate sheet, apart from any other text. See MPEP § 608.01(b). The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed. The lengthy specification has not been checked to the extent necessary to determine the presence of all possible minor errors. Applicant’s cooperation is requested in correcting any errors of which applicant may become aware in the specification. Claim Objections Claim 13 is objected to because of the following informalities: Line 7 of claim 13 recites, in part, “each training image rendering a representation” which appears to contain a minor informality. The Examiner suggests amending the claim to --each training image of the plurality of training images rendering a representation-- in order to improve the clarity and precision of the claim. Appropriate correction is required. Claim 13 is objected to because of the following informalities: Lines 10 - 11 of claim 13 recite, in part, “positions of the key points on the training images;” which appears to contain inconsistent claim terminology and/or minor informalities. The Examiner suggests amending the claim to --positions of the plurality of key points on the plurality of training images;-- in order to maintain consistency with line 6 and line 9 of claim 13 and to improve the clarity and precision of the claim. Appropriate correction is required. Claim 13 is objected to because of the following informalities: Lines 12 - 13 of claim 13 recite, in part, “that includes the training images, the ground truth data, and the initial images” which appears to contain inconsistent claim terminology and/or minor informalities. The Examiner suggests amending the claim to --that includes the plurality of training images, the ground truth data, and the plurality of initial images-- in order to maintain consistency with line 3 and line 6 of claim 13 and to improve the clarity and precision of the claim. Appropriate correction is required. Claim 14 is objected to because of the following informalities: Lines 2 - 3 of claim 14 recite, in part, “that each indicates positions of the key points in a respective training image, each positional image being generated” which appears to contain inconsistent claim terminology and/or minor informalities. The Examiner suggests amending the claim to --that each indicates positions of the plurality of key points in a respective training image, each positional image of the plurality of positional images being generated-- in order to maintain consistency with line 9 of claim 13 and to improve the clarity and precision of the claim. Appropriate correction is required. Claim 16 is objected to because of the following informalities: Line 2 of claim 16 recites, in part, “for each of the initial images, determining” which appears to contain inconsistent claim terminology and/or a minor informality. The Examiner suggests amending the claim to --for each of the plurality of initial images, determining-- in order to maintain consistency with line 3 of claim 13 and to improve the clarity and precision of the claim. Appropriate correction is required. Claim 16 is objected to because of the following informalities: Lines 3 - 4 of claim 16 recite, in part, “direction of the initial image; generate a plurality of positional images from the initial images” which appears to contain grammatical errors, inconsistent claim terminology and/or minor informalities. The Examiner suggests amending the claim to --direction of the initial image; and generating a plurality of positional images from the plurality of initial images-- in order to maintain consistency with line 3 of claim 13 and to improve the clarity and precision of the claim. Appropriate correction is required. Claim 16 is objected to because of the following informalities: Line 5 of claim 16 recites, in part, “the initial images” which appears to contain inconsistent claim terminology and/or a minor informality. The Examiner suggests amending the claim to --the plurality of initial images-- in order to maintain consistency with line 3 of claim 13 and to improve the clarity and precision of the claim. Appropriate correction is required. Claim 16 is objected to because of the following informalities: Line 6 of claim 16 recites, in part, “the positional images” which appears to contain inconsistent claim terminology and/or a minor informality. The Examiner suggests amending the claim to --the plurality of positional images-- in order to maintain consistency with line 4 of claim 16 and to improve the clarity and precision of the claim. Appropriate correction is required. Claim 17 is objected to because of the following informalities: Line 4 of claim 17 recites, in part, “and the key point” which appears to contain inconsistent claim terminology and/or a minor informality. The Examiner suggests amending the claim to --and the at least one key point-- in order to maintain consistency with line 1 of claim 17 and to improve the clarity and precision of the claim. Appropriate correction is required. Claim 20 is objected to because of the following informalities: Line 11 of claim 20 recites, in part, “each training image rendering a representation” which appears to contain a minor informality. The Examiner suggests amending the claim to --each training image of the plurality of training images rendering a representation-- in order to improve the clarity and precision of the claim. Appropriate correction is required. Claim 20 is objected to because of the following informalities: Lines 14 - 15 of claim 20 recite, in part, “positions of the key points on the training images, and” which appears to contain inconsistent claim terminology and/or minor informalities. The Examiner suggests amending the claim to --positions of the plurality of key points on the plurality of training images, and-- in order to maintain consistency with line 10 and line 13 of claim 20 and to improve the clarity and precision of the claim. Appropriate correction is required. Claim 20 is objected to because of the following informalities: Lines 16 - 17 of claim 20 recite, in part, “that includes the training images, the ground truth data, and the initial images” which appears to contain inconsistent claim terminology and/or minor informalities. The Examiner suggests amending the claim to --that includes the plurality of training images, the ground truth data, and the plurality of initial images-- in order to maintain consistency with line 7 and line 10 of claim 20 and to improve the clarity and precision of the claim. Appropriate correction is required. Claim 25 is objected to because of the following informalities: Lines 2 - 3 of claim 25 recite, in part, “that each indicates the positions of the key points in a respective training image, each positional image being generated” which appears to contain inconsistent claim terminology and/or minor informalities. The Examiner suggests amending the claim to --that each indicates the positions of the plurality of key points in a respective training image, each positional image of the plurality of positional images being generated-- in order to maintain consistency with line 13 of claim 20 and to improve the clarity and precision of the claim. Appropriate correction is required. Claim 27 is objected to because of the following informalities: Line 2 of claim 27 recites, in part, “for each of the initial images, determining” which appears to contain inconsistent claim terminology and/or a minor informality. The Examiner suggests amending the claim to --for each of the plurality of initial images, determining-- in order to maintain consistency with line 7 of claim 20 and to improve the clarity and precision of the claim. Appropriate correction is required. Claim 27 is objected to because of the following informalities: Lines 3 - 4 of claim 27 recite, in part, “direction of the initial image; generate a plurality of positional images from the initial images” which appears to contain grammatical errors, inconsistent claim terminology and/or minor informalities. The Examiner suggests amending the claim to --direction of the initial image; and generating a plurality of positional images from the plurality of initial images-- in order to maintain consistency with line 7 of claim 20 and to improve the clarity and precision of the claim. Appropriate correction is required. Claim 27 is objected to because of the following informalities: Line 5 of claim 27 recites, in part, “the initial images” which appears to contain inconsistent claim terminology and/or a minor informality. The Examiner suggests amending the claim to --the plurality of initial images-- in order to maintain consistency with line 7 of claim 20 and to improve the clarity and precision of the claim. Appropriate correction is required. Claim 27 is objected to because of the following informalities: Line 6 of claim 27 recites, in part, “the positional images” which appears to contain inconsistent claim terminology and/or a minor informality. The Examiner suggests amending the claim to --the plurality of positional images-- in order to maintain consistency with line 4 of claim 27 and to improve the clarity and precision of the claim. Appropriate correction is required. Claim 28 is objected to because of the following informalities: Line 4 of claim 28 recites, in part, “and the key point” which appears to contain inconsistent claim terminology and/or a minor informality. The Examiner suggests amending the claim to --and the at least one key point-- in order to maintain consistency with lines 1 - 2 of claim 28 and to improve the clarity and precision of the claim. Appropriate correction is required. Claim 29 is objected to because of the following informalities: Lines 7 - 8 of claim 29 recite, in part, “each training image rendering a representation” which appears to contain a minor informality. The Examiner suggests amending the claim to --each training image of the plurality of training images rendering a representation-- in order to improve the clarity and precision of the claim. Appropriate correction is required. Claim 29 is objected to because of the following informalities: Line 11 of claim 29 recites, in part, “positions of the key points on the training images;” which appears to contain inconsistent claim terminology and/or minor informalities. The Examiner suggests amending the claim to --positions of the plurality of key points on the plurality of training images;-- in order to maintain consistency with line 7 and line 10 of claim 29 and to improve the clarity and precision of the claim. Appropriate correction is required. Claim 29 is objected to because of the following informalities: Lines 13 - 14 of claim 29 recite, in part, “that includes the training images, the ground truth data, and the initial images” which appears to contain inconsistent claim terminology and/or minor informalities. The Examiner suggests amending the claim to --that includes the plurality of training images, the ground truth data, and the plurality of initial images-- in order to maintain consistency with line 4 and line 7 of claim 29 and to improve the clarity and precision of the claim. Appropriate correction is required. Claim 30 is objected to because of the following informalities: Lines 2 - 3 of claim 30 recite, in part, “that each indicates the positions of the key points in a respective training image, each positional image being generated” which appears to contain inconsistent claim terminology and/or minor informalities. The Examiner suggests amending the claim to --that each indicates the positions of the plurality of key points in a respective training image, each positional image of the plurality of positional images being generated-- in order to maintain consistency with line 10 of claim 29 and to improve the clarity and precision of the claim. Appropriate correction is required. Claim 32 is objected to because of the following informalities: Line 3 of claim 32 recites, in part, “for each of the initial images, determining” which appears to contain inconsistent claim terminology and/or a minor informality. The Examiner suggests amending the claim to --for each of the plurality of initial images, determining-- in order to maintain consistency with line 4 of claim 29 and to improve the clarity and precision of the claim. Appropriate correction is required. Claim 32 is objected to because of the following informalities: Lines 4 - 3 of claim 32 recite, in part, “direction of the initial image; generate a plurality of positional images from the initial images” which appears to contain grammatical errors, inconsistent claim terminology and/or minor informalities. The Examiner suggests amending the claim to --direction of the initial image; and generating a plurality of positional images from the plurality of initial images-- in order to maintain consistency with line 4 of claim 29 and to improve the clarity and precision of the claim. Appropriate correction is required. Claim 32 is objected to because of the following informalities: Line 6 of claim 32 recites, in part, “the initial images,” which appears to contain inconsistent claim terminology and/or a minor informality. The Examiner suggests amending the claim to --the plurality of initial images,-- in order to maintain consistency with line 4 of claim 29 and to improve the clarity and precision of the claim. Appropriate correction is required. Claim 32 is objected to because of the following informalities: Line 7 of claim 27 recites, in part, “the positional images” which appears to contain inconsistent claim terminology and/or a minor informality. The Examiner suggests amending the claim to --the plurality of positional images-- in order to maintain consistency with line 5 of claim 32 and to improve the clarity and precision of the claim. Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 13 - 19, 26 and 31 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 13 recites the limitation "the positions of the key points" in line 10. There is insufficient antecedent basis for this limitation in the claim. Claim 15 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention because it is unclear as to which positional image “the positional image” recited on line 3 is referencing at least because lines 1 - 2 of claim 14 clearly indicates that there are a plurality of positional images. Thus, the Examiner asserts that it is unclear as to which positional image of the plurality of positional images “the positional image” recited on line 3 of claim 15 is referencing, if any. Clarification and appropriate correction are required. For purposes of examination, the Examiner will treat “the positional image” recited on line 3 of claim 15 as referencing any of the plurality of positional images. Claim 19 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention because it is unclear as to which initial image “the initial image” recited on line 1 is referencing at least because line 3 of claim 13 clearly indicates that there are a plurality of initial images. Thus, the Examiner asserts that it is unclear as to which initial image of the plurality of initial images “the initial image” recited on line 1 of claim 19 is referencing, if any. Clarification and appropriate correction are required. For purposes of examination, the Examiner will treat “the initial image” recited on line 1 of claim 19 as referencing any of the plurality of initial images. Claim 26 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention because it is unclear as to which positional image “the positional image” recited on line 4 is referencing at least because lines 1 - 2 of claim 25 clearly indicates that there are a plurality of positional images. Thus, the Examiner asserts that it is unclear as to which positional image of the plurality of positional images “the positional image” recited on line 4 of claim 26 is referencing, if any. Clarification and appropriate correction are required. For purposes of examination, the Examiner will treat “the positional image” recited on line 4 of claim 26 as referencing any of the plurality of positional images. Claim 31 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention because it is unclear as to which positional image “the positional image” recited on line 4 is referencing at least because line 2 of claim 30 clearly indicates that there are a plurality of positional images. Thus, the Examiner asserts that it is unclear as to which positional image of the plurality of positional images “the positional image” recited on line 4 of claim 31 is referencing, if any. Clarification and appropriate correction are required. For purposes of examination, the Examiner will treat “the positional image” recited on line 4 of claim 31 as referencing any of the plurality of positional images. Claims 14 and 16 - 18 are also rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, due to being dependent upon a rejected base claim(s) but would be withdrawn from the rejection if their base claim(s) overcome the rejection. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 13, 16, 19, 20 - 23, 27, 29 and 32 are rejected under 35 U.S.C. 103 as being unpatentable over Atanasoaei et al. U.S. Publication No. 2024/0320304 A1 in view of Li et al. U.S. Publication No. 2022/0114746 A1. - With regards to claim 13, Atanasoaei et al. disclose a method of training a machine learning model configured to detect a pose of an object depicted in an image, (Atanasoaei et al., Abstract, Figs. 3A & 3B, Pg. 1 ¶ 0006 - 0007 and 0009 - 0010, Pg. 4 ¶ 0036 - 0044) the method comprising: receiving a plurality of initial images that each depicts the object; (Atanasoaei et al., Abstract, Figs. 1, 3A & 3B, Pg. 1 ¶ 0006 - 0008, Pg. 2 ¶ 0018 - 0024) generating a three-dimensional shape model of the object; (Atanasoaei et al., Figs. 1 - 3B, Pg. 3 ¶ 0032 - 0034) generating a plurality of training images from the three-dimensional shape model, (Atanasoaei et al., Abstract, Figs. 3A & 3B, Pg. 1 ¶ 0006 - 0007 and 0009 - 0010, Pg. 2 ¶ 0022, Pg. 3 ¶ 0032 - 0035, Pg. 4 ¶ 0039 - 0040) each training image rendering a representation of the three-dimensional shape model from a respective view direction; (Atanasoaei et al., Figs. 3A & 3B, Pg. 1 ¶ 0006 - 0007 and 0009, Pg. 2 ¶ 0019 - 0022, Pg. 3 ¶ 0032 - 0035, Pg. 4 ¶ 0039 - 0040) determining a plurality of key points on the three-dimensional shape model; (Atanasoaei et al., Figs. 1 & 2, Pg. 1 ¶ 0006 - 0007, Pg. 2 ¶ 0020 - 0022, Pg. 3 ¶ 0034, Pg. 4 ¶ 0039 [“3D model 330 can include a number of landmark annotations on the object. As with the landmarks used by pose estimator 315, the nature of the landmarks will depend on the nature of the object. However, since 3D model 330 embodies precise information about the object, 3D model 330 also embodies precise information about the position of the landmarks on the three-dimensional object”]) generating ground truth data indicating the positions of the key points on the training images; (Atanasoaei et al., Pg. 1 ¶ 0006 - 0007 and 0010, Pg. 2 ¶ 0021 - 0022, Pg. 3 ¶ 0034, Pg. 4 ¶ 0039 - 0044) and training the machine learning model based on training data that includes the training images, the ground truth data, and the initial images. (Atanasoaei et al., Figs. 3A - 4, Pg. 1 ¶ 0007 - 0010, Pg. 2 ¶ 0021 - 0022, Pg. 3 ¶ 0033 - Pg. 4 ¶ 0036, Pg. 4 ¶ 0038 - 0043) Atanasoaei et al. fail to disclose expressly generating a three-dimensional shape model of the object based on the received plurality of initial images. Pertaining to analogous art, Li et al. disclose a method of training a machine learning model configured to detect a pose of an object depicted in an image, (Li et al., Figs. 1 - 3, Pg. 9 ¶ 0139 - Pg. 10 ¶ 0143, Pg. 10 ¶ 0147, Pg. 11 ¶ 0153 and 0157 - 0148, Pg. 16 ¶ 0223 - 0225, Pg. 17 ¶ 0228 - 0229 and 0232 - 0234) the method comprising: receiving a plurality of initial images that each depicts the object; (Li et al., Figs. 1 & 5, Pg. 9 ¶ 0142 - Pg. 10 ¶ 0143, Pg. 10 ¶ 0147, Pg. 11 ¶ 0154 - 0157, Pg. 12 ¶ 0175, Pg. 16 ¶ 0223 - 0225, Pg. 17 ¶ 0237) generating a three-dimensional shape model of the object based on the received plurality of initial images; (Li et al., Pg. 9 ¶ 0142 - Pg. 10 ¶ 0143, Pg. 10 ¶ 0147, Pg. 11 ¶ 0154 - 0157, Pg. 12 ¶ 0172 - 0175 [“the 3D model of the target object may be generated based on a plurality of images obtained by photographing the target object”]) and generating a plurality of training images from the three-dimensional shape model, (Li et al., Fig. 5, Pg. 1 ¶ 0012 - 0013, Pg. 12 ¶ 0176 - Pg. 13 ¶ 0180, Pg. 16 ¶ 0223 - 0224, Pg. 18 ¶ 0245, Pg. 19 ¶ 0258 and 0288) each training image rendering a representation of the three-dimensional shape model from a respective view direction. (Li et al., Pg. 1 ¶ 0012 - 0013, Pg. 13 ¶ 0177 - 0180, Pg. 16 ¶ 0223 - 0224) Atanasoaei et al. and Li et al. are combinable because they are both directed towards image processing systems that utilize machine learning models to estimate poses of objects in images. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Atanasoaei et al. with the teachings of Li et al. This modification would have been prompted in order to substitute the 3D model creation process of Atanasoaei et al. for the 3D model generation technique of Li et al. The 3D model generation technique of Li et al. could be substituted in place of the 3D model creation process of Atanasoaei et al. using well-known techniques in the art and would likely yield predictable results, in that, in the combination, the 3D model generation technique of Li et al., which uses a plurality of photographed images of the object, would be utilized to obtain the 3D model of the object. Furthermore, this modification would have been prompted by the teachings and suggestions of Atanasoaei et al. that a three-dimensional shape model of an object may be created by scanning real objects, see at least page 3 paragraph 0032 of Atanasoaei et al. This combination could be completed according to well-known techniques in the art and would likely yield predictable results, in that the base device of Atanasoaei et al. would utilize the 3D model generation technique of Li et al., which uses a plurality of photographed images of the object, to obtain the 3D model of the object. Therefore, it would have been obvious to combine Atanasoaei et al. with Li et al. to obtain the invention as specified in claim 13. - With regards to claim 16, Atanasoaei et al. in view of Li et al. disclose the method of claim 13, further comprising: for each of the initial images, determining a respective pose of the object in the initial image based on a logic for calculating photography direction of the initial image; (Atanasoaei et al., Abstract, Figs. 1, 3A & 3B, Pg. 1 ¶ 0006 - 0007, Pg. 2 ¶ 0019 - 0022, Pg. 3 ¶ 0025 - 0033, Pg. 4 ¶ 0037 - 0041 and 0047) generate a plurality of positional images from the initial images and based on the respective poses of the object in the initial images, (Atanasoaei et al., Abstract, Figs. 3A & 3B, Pg. 1 ¶ 0006 - 0007 and 0009, Pg. 3 ¶ 0031 - Pg. 4 0036) wherein the training data further includes the positional images. (Atanasoaei et al., Figs. 3A & 3B, Pg. 1 ¶ 0007 and 0009 - 0010, Pg. 3 ¶ 0032 - Pg. 4 ¶ 0040) - With regards to claim 19, Atanasoaei et al. in view of Li et al. disclose the method of claim 13, wherein each of the initial image depicts the object from a respective angle of view. (Atanasoaei et al., Fig. 1, Pg. 1 ¶ 0006 - 0009, Pg. 2 ¶ 0019 - 0021, Pg. 3 ¶ 0025 - 0034, Pg. 4 ¶ 0043 - 0048) - With regards to claim 20, Atanasoaei et al. disclose a system (Atanasoaei et al., Abstract, Figs. 1 - 3B, Pg. 1 ¶ 0006 and 0009, Pg. 2 ¶ 0011 and 0022, Pg. 5 ¶ 0049 - 0054) comprising: one or more processors; (Atanasoaei et al., Pg. 2 ¶ 0011 and 0022, Pg. 5 ¶ 0049 - 0054) and a computer-readable storage device coupled to the one or more processors and having instructions stored thereon which, when executed by the one or more processors, cause the one or more processors to perform operations (Atanasoaei et al., Pg. 2 ¶ 0011 and 0022, Pg. 5 ¶ 0049 - 0054) for training a machine learning model, (Atanasoaei et al., Abstract, Figs. 3A & 3B, Pg. 1 ¶ 0006 - 0007 and 0009 - 0010, Pg. 4 ¶ 0036 - 0044) the operations comprising: receiving a plurality of initial images that each depicts an object, (Atanasoaei et al., Abstract, Figs. 1, 3A & 3B, Pg. 1 ¶ 0006 - 0008, Pg. 2 ¶ 0018 - 0024) generating a three-dimensional shape model of the object, (Atanasoaei et al., Figs. 1 - 3B, Pg. 3 ¶ 0032 - 0034) generating a plurality of training images from the three-dimensional shape model, (Atanasoaei et al., Abstract, Figs. 3A & 3B, Pg. 1 ¶ 0006 - 0007 and 0009 - 0010, Pg. 2 ¶ 0022, Pg. 3 ¶ 0032 - 0035, Pg. 4 ¶ 0039 - 0040) each training image rendering a representation of the three-dimensional shape model from a respective view direction, (Atanasoaei et al., Figs. 3A & 3B, Pg. 1 ¶ 0006 - 0007 and 0009, Pg. 2 ¶ 0019 - 0022, Pg. 3 ¶ 0032 - 0035, Pg. 4 ¶ 0039 - 0040) determining a plurality of key points on the three-dimensional shape model, (Atanasoaei et al., Figs. 1 & 2, Pg. 1 ¶ 0006 - 0007, Pg. 2 ¶ 0020 - 0022, Pg. 3 ¶ 0034, Pg. 4 ¶ 0039 [“3D model 330 can include a number of landmark annotations on the object. As with the landmarks used by pose estimator 315, the nature of the landmarks will depend on the nature of the object. However, since 3D model 330 embodies precise information about the object, 3D model 330 also embodies precise information about the position of the landmarks on the three-dimensional object”]) generating ground truth data indicating positions of the key points on the training images, (Atanasoaei et al., Pg. 1 ¶ 0006 - 0007 and 0010, Pg. 2 ¶ 0021 - 0022, Pg. 3 ¶ 0034, Pg. 4 ¶ 0039 - 0044) and training the machine learning model based on training data that includes the training images, the ground truth data, and the initial images. (Atanasoaei et al., Figs. 3A - 4, Pg. 1 ¶ 0007 - 0010, Pg. 2 ¶ 0021 - 0022, Pg. 3 ¶ 0033 - Pg. 4 ¶ 0036, Pg. 4 ¶ 0038 - 0043) Atanasoaei et al. fail to disclose expressly generating a three-dimensional shape model of the object based on the received plurality of initial images. Pertaining to analogous art, Li et al. disclose a system (Li et al., Abstract, Figs. 1 - 4, 10 & 11, Pg. 9 ¶ 0141, Pg. 11 ¶ 0156 - Pg. 12 ¶ 0169, Pg. 20 ¶ 0308, Pg. 22 ¶ 0348) comprising: one or more processors; (Li et al., Fig. 4, Pg. 11 ¶ 0160 - Pg. 12 ¶ 0169, Pg. 20 ¶ 0308, Pg. 22 ¶ 0348) and a computer-readable storage device coupled to the one or more processors and having instructions stored thereon which, when executed by the one or more processors, cause the one or more processors to perform operations (Li et al., Fig. 4, Pg. 11 ¶ 0160 - Pg. 12 ¶ 0169, Pg. 20 ¶ 0308, Pg. 22 ¶ 0348) for training a machine learning model, (Li et al., Figs. 1 - 3, Pg. 9 ¶ 0139 - Pg. 10 ¶ 0143, Pg. 10 ¶ 0147, Pg. 11 ¶ 0153 and 0157 - 0148, Pg. 16 ¶ 0223 - 0225, Pg. 17 ¶ 0228 - 0229 and 0232 - 0234) the operations comprising: receiving a plurality of initial images that each depicts an object, (Li et al., Figs. 1 & 5, Pg. 9 ¶ 0142 - Pg. 10 ¶ 0143, Pg. 10 ¶ 0147, Pg. 11 ¶ 0154 - 0157, Pg. 12 ¶ 0175, Pg. 16 ¶ 0223 - 0225, Pg. 17 ¶ 0237) generating a three-dimensional shape model of the object based on the received plurality of initial images, (Li et al., Pg. 9 ¶ 0142 - Pg. 10 ¶ 0143, Pg. 10 ¶ 0147, Pg. 11 ¶ 0154 - 0157, Pg. 12 ¶ 0172 - 0175 [“the 3D model of the target object may be generated based on a plurality of images obtained by photographing the target object”]) and generating a plurality of training images from the three-dimensional shape model, (Li et al., Fig. 5, Pg. 1 ¶ 0012 - 0013, Pg. 12 ¶ 0176 - Pg. 13 ¶ 0180, Pg. 16 ¶ 0223 - 0224, Pg. 18 ¶ 0245, Pg. 19 ¶ 0258 and 0288) each training image rendering a representation of the three-dimensional shape model from a respective view direction. (Li et al., Pg. 1 ¶ 0012 - 0013, Pg. 13 ¶ 0177 - 0180, Pg. 16 ¶ 0223 - 0224) Atanasoaei et al. and Li et al. are combinable because they are both directed towards image processing systems that utilize machine learning models to estimate poses of objects in images. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Atanasoaei et al. with the teachings of Li et al. This modification would have been prompted in order to substitute the 3D model creation process of Atanasoaei et al. for the 3D model generation technique of Li et al. The 3D model generation technique of Li et al. could be substituted in place of the 3D model creation process of Atanasoaei et al. using well-known techniques in the art and would likely yield predictable results, in that, in the combination, the 3D model generation technique of Li et al., which uses a plurality of photographed images of the object, would be utilized to obtain the 3D model of the object. Furthermore, this modification would have been prompted by the teachings and suggestions of Atanasoaei et al. that a three-dimensional shape model of an object may be created by scanning real objects, see at least page 3 paragraph 0032 of Atanasoaei et al. This combination could be completed according to well-known techniques in the art and would likely yield predictable results, in that the base device of Atanasoaei et al. would utilize the 3D model generation technique of Li et al., which uses a plurality of photographed images of the object, to obtain the 3D model of the object. Therefore, it would have been obvious to combine Atanasoaei et al. with Li et al. to obtain the invention as specified in claim 20. - With regards to claim 21, Atanasoaei et al. in view of Li et al. disclose the system of claim 20, wherein the system is part of a game console. (Atanasoaei et al., Pg. 5 ¶ 0049 - 0054 [“a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name just a few.”]) - With regards to claim 22, Atanasoaei et al. in view of Li et al. disclose the system of claim 20, wherein the computer-readable storage device further includes instructions stored thereon which, when executed by the one or more processors, cause the one or more processors to (Atanasoaei et al., Pg. 2 ¶ 0011 and 0022, Pg. 5 ¶ 0049 - 0054) use the trained machine learning model to determine a pose of a photographed object. (Atanasoaei et al., Abstract, Figs. 1 & 3A - 4, Pg. 1 ¶ 0006 - 0009, Pg. 2 ¶ 0019 - 0023, Pg. 3 ¶ 0025 - 0031, Pg. 4 ¶ 0037 - 0045 and 0047) In addition, analogous art Li et al. disclose wherein the computer-readable storage device further includes instructions stored thereon which, when executed by the one or more processors, cause the one or more processors to (Li et al., Fig. 4, Pg. 11 ¶ 0160 - Pg. 12 ¶ 0169, Pg. 20 ¶ 0308, Pg. 22 ¶ 0348) use the trained machine learning model to determine a pose of a photographed object. (Li et al., Fig. 1, Pg. 10 ¶ 0143 and 0147, Pg. 11 ¶ 0155, Pg. 17 ¶ 0232 - 0237, Pg. 19 ¶ 0259 - 0260, Pg. 20 ¶ 0292 - 0295) - With regards to claim 23, Atanasoaei et al. in view of Li et al. disclose the system of claim 22, further comprising a camera (Atanasoaei et al., Fig. 1, Pg. 2 ¶ 0019 - 0021) configured to take one or more images including the photographed object. (Atanasoaei et al., Figs. 1, 3A & 3B, Pg. 1 ¶ 0006 - 0009, Pg. 2 ¶ 0019 - 0022, Pg. 3 ¶ 0025 - 0032, Pg. 4 ¶ 0037 - 0045) In addition, analogous art Li et al. disclose a camera (Li et al., Pg. 10 ¶ 0143 and 0147, Pg. 11 ¶ 0154 - 0155, Pg. 12 ¶ 0175, Pg. 16 ¶ 0224 - 0225, Pg. 17 ¶ 0232 - 0237) configured to take one or more images including the photographed object. (Li et al., Pg. 10 ¶ 0143 and 0147, Pg. 11 ¶ 0154 - 0155, Pg. 12 ¶ 0175, Pg. 16 ¶ 0224 - 0225, Pg. 17 ¶ 0232 - 0237) - With regards to claim 27, Atanasoaei et al. in view of Li et al. disclose the system of claim 20, wherein the operations further comprise: for each of the initial images, determining a respective pose of the object in the initial image based on a logic for calculating photography direction of the initial image; (Atanasoaei et al., Abstract, Figs. 1, 3A & 3B, Pg. 1 ¶ 0006 - 0007, Pg. 2 ¶ 0019 - 0022, Pg. 3 ¶ 0025 - 0033, Pg. 4 ¶ 0037 - 0041 and 0047) generate a plurality of positional images from the initial images and based on the respective poses of the object in the initial images, (Atanasoaei et al., Abstract, Figs. 3A & 3B, Pg. 1 ¶ 0006 - 0007 and 0009, Pg. 3 ¶ 0031 - Pg. 4 0036) wherein the training data further includes the positional images. (Atanasoaei et al., Figs. 3A & 3B, Pg. 1 ¶ 0007 and 0009 - 0010, Pg. 3 ¶ 0032 - Pg. 4 ¶ 0040) - With regards to claim 29, Atanasoaei et al. disclose a non-transitory, computer-readable medium storing one or more instructions executable by a computer system to perform operations (Atanasoaei et al., Pg. 2 ¶ 0011 and 0022, Pg. 5 ¶ 0049 - 0054) for training a machine learning model, (Atanasoaei et al., Abstract, Figs. 3A & 3B, Pg. 1 ¶ 0006 - 0007 and 0009 - 0010, Pg. 4 ¶ 0036 - 0044) the operations comprising: receiving a plurality of initial images that each depicts an object; (Atanasoaei et al., Abstract, Figs. 1, 3A & 3B, Pg. 1 ¶ 0006 - 0008, Pg. 2 ¶ 0018 - 0024) generating a three-dimensional shape model of the object; (Atanasoaei et al., Figs. 1 - 3B, Pg. 3 ¶ 0032 - 0034) generating a plurality of training images from the three-dimensional shape model, (Atanasoaei et al., Abstract, Figs. 3A & 3B, Pg. 1 ¶ 0006 - 0007 and 0009 - 0010, Pg. 2 ¶ 0022, Pg. 3 ¶ 0032 - 0035, Pg. 4 ¶ 0039 - 0040) each training image rendering a representation of the three-dimensional shape model from a respective view direction; (Atanasoaei et al., Figs. 3A & 3B, Pg. 1 ¶ 0006 - 0007 and 0009, Pg. 2 ¶ 0019 - 0022, Pg. 3 ¶ 0032 - 0035, Pg. 4 ¶ 0039 - 0040) determining a plurality of key points on the three-dimensional shape model; (Atanasoaei et al., Figs. 1 & 2, Pg. 1 ¶ 0006 - 0007, Pg. 2 ¶ 0020 - 0022, Pg. 3 ¶ 0034, Pg. 4 ¶ 0039 [“3D model 330 can include a number of landmark annotations on the object. As with the landmarks used by pose estimator 315, the nature of the landmarks will depend on the nature of the object. However, since 3D model 330 embodies precise information about the object, 3D model 330 also embodies precise information about the position of the landmarks on the three-dimensional object”]) generating ground truth data indicating positions of the key points on the training images; (Atanasoaei et al., Pg. 1 ¶ 0006 - 0007 and 0010, Pg. 2 ¶ 0021 - 0022, Pg. 3 ¶ 0034, Pg. 4 ¶ 0039 - 0044) and training the machine learning model based on training data that includes the training images, the ground truth data, and the initial images. (Atanasoaei et al., Figs. 3A - 4, Pg. 1 ¶ 0007 - 0010, Pg. 2 ¶ 0021 - 0022, Pg. 3 ¶ 0033 - Pg. 4 ¶ 0036, Pg. 4 ¶ 0038 - 0043) Atanasoaei et al. fail to disclose expressly generating a three-dimensional shape model of the object based on the received plurality of initial images. Pertaining to analogous art, Li et al. disclose a non-transitory, computer-readable medium storing one or more instructions executable by a computer system to perform operations (Li et al., Fig. 4, Pg. 11 ¶ 0160 - Pg. 12 ¶ 0169, Pg. 20 ¶ 0308, Pg. 22 ¶ 0348) for training a machine learning model, (Li et al., Figs. 1 - 3, Pg. 9 ¶ 0139 - Pg. 10 ¶ 0143, Pg. 10 ¶ 0147, Pg. 11 ¶ 0153 and 0157 - 0148, Pg. 16 ¶ 0223 - 0225, Pg. 17 ¶ 0228 - 0229 and 0232 - 0234) the operations comprising: receiving a plurality of initial images that each depicts an object; (Li et al., Figs. 1 & 5, Pg. 9 ¶ 0142 - Pg. 10 ¶ 0143, Pg. 10 ¶ 0147, Pg. 11 ¶ 0154 - 0157, Pg. 12 ¶ 0175, Pg. 16 ¶ 0223 - 0225, Pg. 17 ¶ 0237) generating a three-dimensional shape model of the object based on the received plurality of initial images; (Li et al., Pg. 9 ¶ 0142 - Pg. 10 ¶ 0143, Pg. 10 ¶ 0147, Pg. 11 ¶ 0154 - 0157, Pg. 12 ¶ 0172 - 0175 [“the 3D model of the target object may be generated based on a plurality of images obtained by photographing the target object”]) and generating a plurality of training images from the three-dimensional shape model, (Li et al., Fig. 5, Pg. 1 ¶ 0012 - 0013, Pg. 12 ¶ 0176 - Pg. 13 ¶ 0180, Pg. 16 ¶ 0223 - 0224, Pg. 18 ¶ 0245, Pg. 19 ¶ 0258 and 0288) each training image rendering a representation of the three-dimensional shape model from a respective view direction. (Li et al., Pg. 1 ¶ 0012 - 0013, Pg. 13 ¶ 0177 - 0180, Pg. 16 ¶ 0223 - 0224) Atanasoaei et al. and Li et al. are combinable because they are both directed towards image processing systems that utilize machine learning models to estimate poses of objects in images. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Atanasoaei et al. with the teachings of Li et al. This modification would have been prompted in order to substitute the 3D model creation process of Atanasoaei et al. for the 3D model generation technique of Li et al. The 3D model generation technique of Li et al. could be substituted in place of the 3D model creation process of Atanasoaei et al. using well-known techniques in the art and would likely yield predictable results, in that, in the combination, the 3D model generation technique of Li et al., which uses a plurality of photographed images of the object, would be utilized to obtain the 3D model of the object. Furthermore, this modification would have been prompted by the teachings and suggestions of Atanasoaei et al. that a three-dimensional shape model of an object may be created by scanning real objects, see at least page 3 paragraph 0032 of Atanasoaei et al. This combination could be completed according to well-known techniques in the art and would likely yield predictable results, in that the base device of Atanasoaei et al. would utilize the 3D model generation technique of Li et al., which uses a plurality of photographed images of the object, to obtain the 3D model of the object. Therefore, it would have been obvious to combine Atanasoaei et al. with Li et al. to obtain the invention as specified in claim 29. - With regards to claim 32, Atanasoaei et al. in view of Li et al. disclose the computer-readable medium of claim 29, wherein the operations further comprise: for each of the initial images, determining a respective pose of the object in the initial image based on a logic for calculating photography direction of the initial image; (Atanasoaei et al., Abstract, Figs. 1, 3A & 3B, Pg. 1 ¶ 0006 - 0007, Pg. 2 ¶ 0019 - 0022, Pg. 3 ¶ 0025 - 0033, Pg. 4 ¶ 0037 - 0041 and 0047) generate a plurality of positional images from the initial images and based on the respective poses of the object in the initial images, (Atanasoaei et al., Abstract, Figs. 3A & 3B, Pg. 1 ¶ 0006 - 0007 and 0009, Pg. 3 ¶ 0031 - Pg. 4 0036) wherein the training data further includes the positional images. (Atanasoaei et al., Figs. 3A & 3B, Pg. 1 ¶ 0007 and 0009 - 0010, Pg. 3 ¶ 0032 - Pg. 4 ¶ 0040) Claims 14, 15, 17, 25, 26, 28, 30 and 31 are rejected under 35 U.S.C. 103 as being unpatentable over Atanasoaei et al. U.S. Publication No. 2024/0320304 A1 in view of Li et al. U.S. Publication No. 2022/0114746 A1 as applied to claims 13, 16, 20, 27 and 29 above, and further in view of Sida Peng, Yuan Liu, Qixing Huang, Xiaowei Zhou and Hujun Bao, “PVNet: Pixel-wise Voting Network for 6DoF Pose Estimation”, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pages 4556 - 4565, herein referred to as “Peng et al.”. - With regards to claim 14, Atanasoaei et al. in view of Li et al. disclose the method of claim 13, wherein the ground truth data includes a plurality of positional images that each indicates positions of the key points in a respective training image. (Atanasoaei et al., Pg. 1 ¶ 0006 - 0007 and 0010, Pg. 2 ¶ 0021 - 0022, Pg. 3 ¶ 0034, Pg. 4 ¶ 0039 - 0044) Atanasoaei et al. fail to disclose explicitly each positional image being generated based on a relative position between a projected position of a respective key point and different pixels in the respective training image. Pertaining to analogous art, Peng et al. disclose each positional image being generated based on a relative position between a projected position of a respective key point and different pixels in the respective training image. (Peng et al., Pg. 4556 Abstract, Pg. 4556 Fig. 1, Pg. 4557 Left-Hand Column First-Full Paragraph - Second-Full Paragraph, Pg. 4558 § 3 - Pg. 4559 § 3.1 ¶ 3, Pg. 4558 Fig. 2, Pg. 4560 § 3.2 and § 4.1, Pg. 4561 § 5.2) Atanasoaei et al. in view of Li et al. and Peng et al. are combinable because they are all directed towards image processing systems that utilize machine learning models to estimate poses of objects in images. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combined teachings of Atanasoaei et al. in view of Li et al. with the teachings of Peng et al. This modification would have been prompted in order to enhance the combined base device of Atanasoaei et al. in view of Li et al. with the well-known and applicable technique Peng et al. applied to a comparable device. Generating each positional image based on a relative position between a projected position of a respective key point and different pixels in the respective training image, as taught by Peng et al., would enhance the combined base device by enabling it to represent object keypoints that are occluded or outside of the images and by improving its ability to identify consistent keypoint correspondences to predict object poses thereby enhancing its ability to accurately and reliably detect poses of the object in images, as taught and suggested by Peng et al., see at least page 4557 left-hand column second-full paragraph - third-full paragraph and page 4559 section 3.1 second-full paragraph of Peng et al. Furthermore, this modification would have been prompted by the teachings and suggestions of Atanasoaei et al. that the pose of an object can be estimated based on landmarks that are detected, the distance that separates the landmarks, and the relative positions of the landmarks in the image of the object, see at least page 3 paragraph 0025 of Atanasoaei et al. This combination could be completed according to well-known techniques in the art and would likely yield predictable results, in that each positional image would be generated based on a relative position between a projected position of a respective key point and different pixels in the respective training image so as to enhance the ability of the combined base device to accurately and reliably detect poses of the object in images by enabling it to represent object keypoints that are occluded or outside of the images and by improving its ability to identify consistent keypoint correspondences to predict object poses. Therefore, it would have been obvious to combine Atanasoaei et al. in view of Li et al. with Peng et al. to obtain the invention as specified in claim 14. - With regards to claim 15, Atanasoaei et al. in view of Li et al. in view of Peng et al. disclose the method of claim 14, further comprising determining a pose of the three-dimensional shape model based on the positions of the plurality of the key points, (Atanasoaei et al., Pg. 1 ¶ 0006 - 0007 and 0010, Pg. 2 ¶ 0020 - 0022, Pg. 3 ¶ 0032 - 0035) wherein the positional image further includes the pose of the three-dimensional shape model. (Atanasoaei et al., Figs. 3A & 3B, Pg. 1 ¶ 0006 - 0007 and 0009 - 0010, Pg. 3 ¶ 0032 - 0035) - With regards to claim 17, Atanasoaei et al. in view of Li et al. disclose the method of claim 16, further comprising identifying at least one key point on an initial image. (Atanasoaei et al., Pg. 1 ¶ 0007 and 0010, Pg. 2 ¶ 0020 - 0022, Pg. 3 ¶ 0025 and 0034 [“labeling landmarks or refining positions of landmarks on the real images in the proper subset based on the three-dimensional model, and including the real images in the proper subset in the training set.”]) Atanasoaei et al. disclose fail to disclose explicitly wherein each pixel in a positional image generated from the initial image, indicates a positional relationship between a respective pixel in the initial image and the key point. Pertaining to analogous art, Peng et al. disclose wherein each pixel in a positional image generated from the initial image, indicates a positional relationship between a respective pixel in the initial image and the key point. (Peng et al., Pg. 4556 Abstract, Pg. 4556 Fig. 1, Pg. 4557 Left-Hand Column First-Full Paragraph - Second-Full Paragraph, Pg. 4558 § 3 - Pg. 4559 § 3.1 ¶ 3, Pg. 4558 Fig. 2, Pg. 4560 § 3.2 and § 4.1, Pg. 4561 § 5.2) Atanasoaei et al. in view of Li et al. and Peng et al. are combinable because they are all directed towards image processing systems that utilize machine learning models to estimate poses of objects in images. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combined teachings of Atanasoaei et al. in view of Li et al. with the teachings of Peng et al. This modification would have been prompted in order to enhance the combined base device of Atanasoaei et al. in view of Li et al. with the well-known and applicable technique Peng et al. applied to a comparable device. Generating a positional image from the initial image wherein each pixel in the positional image indicates a positional relationship between a respective pixel in the initial image and the key point, as taught by Peng et al., would enhance the combined base device by enabling it to represent object keypoints that are occluded or outside of the images and by improving its ability to identify consistent keypoint correspondences to predict object poses thereby enhancing its ability to accurately and reliably detect poses of the object in images, as taught and suggested by Peng et al., see at least page 4557 left-hand column second-full paragraph - third-full paragraph and page 4559 section 3.1 second-full paragraph of Peng et al. Furthermore, this modification would have been prompted by the teachings and suggestions of Atanasoaei et al. that the pose of an object can be estimated based on landmarks that are detected, the distance that separates the landmarks, and the relative positions of the landmarks in the image of the object, see at least page 3 paragraph 0025 of Atanasoaei et al. This combination could be completed according to well-known techniques in the art and would likely yield predictable results, in that each pixel in the positional image would indicate a positional relationship between a respective pixel in the initial image and the key point so as to enhance the ability of the combined base device to accurately and reliably detect poses of the object in images by enabling it to represent object keypoints that are occluded or outside of the images and by improving its ability to identify consistent keypoint correspondences to predict object poses. Therefore, it would have been obvious to combine Atanasoaei et al. in view of Li et al. with Peng et al. to obtain the invention as specified in claim 17. - With regards to claim 25, Atanasoaei et al. in view of Li et al. disclose the system of claim 20, wherein the ground truth data includes a plurality of positional images that each indicates the positions of the key points in a respective training image. (Atanasoaei et al., Pg. 1 ¶ 0006 - 0007 and 0010, Pg. 2 ¶ 0021 - 0022, Pg. 3 ¶ 0034, Pg. 4 ¶ 0039 - 0044) Atanasoaei et al. fail to disclose explicitly each positional image being generated based on a relative position between a projected position of a respective key point and different pixels in the respective training image. Pertaining to analogous art, Peng et al. disclose each positional image being generated based on a relative position between a projected position of a respective key point and different pixels in the respective training image. (Peng et al., Pg. 4556 Abstract, Pg. 4556 Fig. 1, Pg. 4557 Left-Hand Column First-Full Paragraph - Second-Full Paragraph, Pg. 4558 § 3 - Pg. 4559 § 3.1 ¶ 3, Pg. 4558 Fig. 2, Pg. 4560 § 3.2 and § 4.1, Pg. 4561 § 5.2) Atanasoaei et al. in view of Li et al. and Peng et al. are combinable because they are all directed towards image processing systems that utilize machine learning models to estimate poses of objects in images. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combined teachings of Atanasoaei et al. in view of Li et al. with the teachings of Peng et al. This modification would have been prompted in order to enhance the combined base device of Atanasoaei et al. in view of Li et al. with the well-known and applicable technique Peng et al. applied to a comparable device. Generating each positional image based on a relative position between a projected position of a respective key point and different pixels in the respective training image, as taught by Peng et al., would enhance the combined base device by enabling it to represent object keypoints that are occluded or outside of the images and by improving its ability to identify consistent keypoint correspondences to predict object poses thereby enhancing its ability to accurately and reliably detect poses of the object in images, as taught and suggested by Peng et al., see at least page 4557 left-hand column second-full paragraph - third-full paragraph and page 4559 section 3.1 second-full paragraph of Peng et al. Furthermore, this modification would have been prompted by the teachings and suggestions of Atanasoaei et al. that the pose of an object can be estimated based on landmarks that are detected, the distance that separates the landmarks, and the relative positions of the landmarks in the image of the object, see at least page 3 paragraph 0025 of Atanasoaei et al. This combination could be completed according to well-known techniques in the art and would likely yield predictable results, in that each positional image would be generated based on a relative position between a projected position of a respective key point and different pixels in the respective training image so as to enhance the ability of the combined base device to accurately and reliably detect poses of the object in images by enabling it to represent object keypoints that are occluded or outside of the images and by improving its ability to identify consistent keypoint correspondences to predict object poses. Therefore, it would have been obvious to combine Atanasoaei et al. in view of Li et al. with Peng et al. to obtain the invention as specified in claim 25. - With regards to claim 26, Atanasoaei et al. in view of Li et al. in view of Peng et al. disclose the system of claim 25, wherein the operations further comprise determining a pose of the three-dimensional shape model based on the positions of the plurality of the key points, (Atanasoaei et al., Pg. 1 ¶ 0006 - 0007 and 0010, Pg. 2 ¶ 0020 - 0022, Pg. 3 ¶ 0032 - 0035) wherein the positional image further includes the pose of the three-dimensional shape model. (Atanasoaei et al., Figs. 3A & 3B, Pg. 1 ¶ 0006 - 0007 and 0009 - 0010, Pg. 3 ¶ 0032 - 0035) - With regards to claim 28, Atanasoaei et al. in view of Li et al. disclose the system of claim 27, wherein the operations further comprise identifying at least one key point on an initial image. (Atanasoaei et al., Pg. 1 ¶ 0007 and 0010, Pg. 2 ¶ 0020 - 0022, Pg. 3 ¶ 0025 and 0034 [“labeling landmarks or refining positions of landmarks on the real images in the proper subset based on the three-dimensional model, and including the real images in the proper subset in the training set.”]) Atanasoaei et al. fail to disclose explicitly wherein each pixel in a positional image generated from the initial image, indicates a positional relationship between a respective pixel in the initial image and the key point. Pertaining to analogous art, Peng et al. disclose wherein each pixel in a positional image generated from the initial image, indicates a positional relationship between a respective pixel in the initial image and the key point. (Peng et al., Pg. 4556 Abstract, Pg. 4556 Fig. 1, Pg. 4557 Left-Hand Column First-Full Paragraph - Second-Full Paragraph, Pg. 4558 § 3 - Pg. 4559 § 3.1 ¶ 3, Pg. 4558 Fig. 2, Pg. 4560 § 3.2 and § 4.1, Pg. 4561 § 5.2) Atanasoaei et al. in view of Li et al. and Peng et al. are combinable because they are all directed towards image processing systems that utilize machine learning models to estimate poses of objects in images. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combined teachings of Atanasoaei et al. in view of Li et al. with the teachings of Peng et al. This modification would have been prompted in order to enhance the combined base device of Atanasoaei et al. in view of Li et al. with the well-known and applicable technique Peng et al. applied to a comparable device. Generating a positional image from the initial image wherein each pixel in the positional image indicates a positional relationship between a respective pixel in the initial image and the key point, as taught by Peng et al., would enhance the combined base device by enabling it to represent object keypoints that are occluded or outside of the images and by improving its ability to identify consistent keypoint correspondences to predict object poses thereby enhancing its ability to accurately and reliably detect poses of the object in images, as taught and suggested by Peng et al., see at least page 4557 left-hand column second-full paragraph - third-full paragraph and page 4559 section 3.1 second-full paragraph of Peng et al. Furthermore, this modification would have been prompted by the teachings and suggestions of Atanasoaei et al. that the pose of an object can be estimated based on landmarks that are detected, the distance that separates the landmarks, and the relative positions of the landmarks in the image of the object, see at least page 3 paragraph 0025 of Atanasoaei et al. This combination could be completed according to well-known techniques in the art and would likely yield predictable results, in that each pixel in the positional image would indicate a positional relationship between a respective pixel in the initial image and the key point so as to enhance the ability of the combined base device to accurately and reliably detect poses of the object in images by enabling it to represent object keypoints that are occluded or outside of the images and by improving its ability to identify consistent keypoint correspondences to predict object poses. Therefore, it would have been obvious to combine Atanasoaei et al. in view of Li et al. with Peng et al. to obtain the invention as specified in claim 28. - With regards to claim 30, Atanasoaei et al. in view of Li et al. disclose the computer-readable medium of claim 29, wherein the ground truth data includes a plurality of positional images that each indicates the positions of the key points in a respective training image. (Atanasoaei et al., Pg. 1 ¶ 0006 - 0007 and 0010, Pg. 2 ¶ 0021 - 0022, Pg. 3 ¶ 0034, Pg. 4 ¶ 0039 - 0044) Atanasoaei et al. fail to disclose explicitly each positional image being generated based on a relative position between a projected position of a respective key point and different pixels in the respective training image. Pertaining to analogous art, Peng et al. disclose each positional image being generated based on a relative position between a projected position of a respective key point and different pixels in the respective training image. (Peng et al., Pg. 4556 Abstract, Pg. 4556 Fig. 1, Pg. 4557 Left-Hand Column First-Full Paragraph - Second-Full Paragraph, Pg. 4558 § 3 - Pg. 4559 § 3.1 ¶ 3, Pg. 4558 Fig. 2, Pg. 4560 § 3.2 and § 4.1, Pg. 4561 § 5.2) Atanasoaei et al. in view of Li et al. and Peng et al. are combinable because they are all directed towards image processing systems that utilize machine learning models to estimate poses of objects in images. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combined teachings of Atanasoaei et al. in view of Li et al. with the teachings of Peng et al. This modification would have been prompted in order to enhance the combined base device of Atanasoaei et al. in view of Li et al. with the well-known and applicable technique Peng et al. applied to a comparable device. Generating each positional image based on a relative position between a projected position of a respective key point and different pixels in the respective training image, as taught by Peng et al., would enhance the combined base device by enabling it to represent object keypoints that are occluded or outside of the images and by improving its ability to identify consistent keypoint correspondences to predict object poses thereby enhancing its ability to accurately and reliably detect poses of the object in images, as taught and suggested by Peng et al., see at least page 4557 left-hand column second-full paragraph - third-full paragraph and page 4559 section 3.1 second-full paragraph of Peng et al. Furthermore, this modification would have been prompted by the teachings and suggestions of Atanasoaei et al. that the pose of an object can be estimated based on landmarks that are detected, the distance that separates the landmarks, and the relative positions of the landmarks in the image of the object, see at least page 3 paragraph 0025 of Atanasoaei et al. This combination could be completed according to well-known techniques in the art and would likely yield predictable results, in that each positional image would be generated based on a relative position between a projected position of a respective key point and different pixels in the respective training image so as to enhance the ability of the combined base device to accurately and reliably detect poses of the object in images by enabling it to represent object keypoints that are occluded or outside of the images and by improving its ability to identify consistent keypoint correspondences to predict object poses. Therefore, it would have been obvious to combine Atanasoaei et al. in view of Li et al. with Peng et al. to obtain the invention as specified in claim 30. - With regards to claim 31, Atanasoaei et al. in view of Li et al. in view of Peng et al. disclose the computer-readable medium of claim 30, wherein the operations further comprise determining a pose of the three-dimensional shape model based on the positions of the plurality of the key points, (Atanasoaei et al., Pg. 1 ¶ 0006 - 0007 and 0010, Pg. 2 ¶ 0020 - 0022, Pg. 3 ¶ 0032 - 0035) wherein the positional image further includes the pose of the three-dimensional shape model. (Atanasoaei et al., Figs. 3A & 3B, Pg. 1 ¶ 0006 - 0007 and 0009 - 0010, Pg. 3 ¶ 0032 - 0035) Claim 18 is rejected under 35 U.S.C. 103 as being unpatentable over Atanasoaei et al. U.S. Publication No. 2024/0320304 A1 in view of Li et al. U.S. Publication No. 2022/0114746 A1 as applied to claim 13 above, and further in view of Josip Josifovski, Matthias Kerzel, Christoph Pregizer, Lukas Posniak and Stefan Wermter, “Object Detection and Pose Estimation based on Convolutional Neural Networks Trained with Synthetic Data”, IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2018, pages 6269 - 6276, herein referred to as “Josifovski et al.”. - With regards to claim 18, Atanasoaei et al. in view of Li et al. disclose the method of claim 13. Atanasoaei et al. fail to disclose explicitly masking one or more initial images to filter out respective background portions; and replacing each initial image with a respective masked initial image. Pertaining to analogous art, Josifovski et al. disclose masking one or more initial images to filter out respective background portions; (Josifovski et al., Pg. 6271 § III - Subsection A, Pg. 6271 Fig. 2, Pg. 6274 Subsection B ¶ 1, Pg. 6274 Fig. 6) and replacing each initial image with a respective masked initial image. (Josifovski et al., Pg. 6271 § III - Subsection A, Pg. 6271 Fig. 2, Pg. 6274 Subsection B ¶ 1, Pg. 6274 Fig. 6) Atanasoaei et al. in view of Li et al. and Josifovski et al. are combinable because they are all directed towards image processing systems that utilize machine learning models to estimate poses of objects in images. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combined teachings of Atanasoaei et al. in view of Li et al. with the teachings of Josifovski et al. This modification would have been prompted in order to enhance the combined base device of Atanasoaei et al. in view of Li et al. with the well-known and applicable technique Josifovski et al. applied to a comparable device. Masking one or more initial images to filter out respective background portions and replacing each initial image with a respective masked initial image, as taught by Josifovski et al., would enhance the combined base device by reducing the amount of superfluous image data that is processed and analyzed by the combined base device thereby allowing for the machine learning model to focus solely on image data pertaining to the object and improving the ability of the combined base device to accurately and reliably detect poses of the object in images. This combination could be completed according to well-known techniques in the art and would likely yield predictable results, in that each initial image would be replaced with a masked initial image in which the respective background portions are filtered out so as to reduce the amount of superfluous image data that is processed and analyzed by the combined base device in order to help the machine learning model focus solely on image data pertaining to the object and improve the ability of the combined base device to accurately and reliably detect poses of the object in images. Therefore, it would have been obvious to combine Atanasoaei et al. in view of Li et al. with Josifovski et al. to obtain the invention as specified in claim 18. Claim 24 is rejected under 35 U.S.C. 103 as being unpatentable over Atanasoaei et al. U.S. Publication No. 2024/0320304 A1 in view of Li et al. U.S. Publication No. 2022/0114746 A1 as applied to claim 20 above, and further in view of Ning et al. U.S. Publication No. 2020/0074678 A1. - With regards to claim 24, Atanasoaei et al. in view of Li et al. disclose the system of claim 20, wherein the computer-readable storage device further includes instructions stored thereon which, when executed by the one or more processors, cause the one or more processors to (Atanasoaei et al., Pg. 2 ¶ 0011 and 0022, Pg. 5 ¶ 0049 - 0054) use the trained machine learning model. (Atanasoaei et al., Abstract, Figs. 1 & 3A - 4, Pg. 1 ¶ 0006 - 0009, Pg. 2 ¶ 0019 - 0023, Pg. 3 ¶ 0025 - 0031, Pg. 4 ¶ 0037 - 0045 and 0047) Atanasoaei et al. fail to disclose explicitly tracking a moving object. Pertaining to analogous art, Ning et al. disclose using the trained machine learning model to track a moving object. (Ning et al., Abstract, Figs. 1A, 4 & 5, Pg. 1 ¶ 0009 - Pg. 2 ¶ 0013, Pg. 2 ¶ 0017 - 0021, Pg. 4 ¶ 0057 - 0058, Pg. 5 ¶ 0063 - 0065, Pg. 6 ¶ 0068, Pg. 7 ¶ 0074 - 0077, Pg. 8 ¶ 0081, Pg. 9 ¶ 0098 - 0101, Pg. 10 ¶ 0106 - 0112, Pg. 11 ¶ 0117 - 0120 [“the human pose estimation module 130 which, when loaded into the memory 114 and executed by the processor 112, may cause the processor 112 to, in response to receiving the detected result from the human candidate detection module 120, estimate a pose of each of the objects within each of the plurality of consecutive frames using the detection result, and send the estimated poses to the executed human pose tracking module 140” and “the human pose tracking module 140 which, when loaded into the memory 114 and executed by the processor 112, may cause the processor 112 to, in response to receiving the estimated poses of the objects, track the poses of each of the objects across the plurality of consecutive frames.”]) Atanasoaei et al. in view of Li et al. and Ning et al. are combinable because they are all directed towards image processing systems that utilize machine learning models to estimate poses of objects in images. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combined teachings of Atanasoaei et al. in view of Li et al. with the teachings of Ning et al. This modification would have been prompted in order to enhance the combined base device of Atanasoaei et al. in view of Li et al. with the well-known and applicable technique Ning et al. applied to a comparable device. Using the trained machine learning model to track a moving object, as taught by Ning et al., would enhance the combined base device by enabling it to detect the poses of additional target objects, such as moving objects, and perform additional related image processing tasks with respect to input images thereby allowing for the combined base device to be utilized in various additional image processing applications and increasing its overall appeal and usefulness to potential end-users. Furthermore, this modification would have been prompted by the teachings and suggestions of Li et al. that their method helps reduce difficulties in pose recognition for an object during pose detection and tracking, and improves pose recognition accuracy for objects, see at least page 1 paragraph 0014, page 4 paragraph 0039 and page 18 paragraph 0245 of Li et al. This combination could be completed according to well-known techniques in the art and would likely yield predictable results, in that the trained machine learning model would be used to track a moving object so as to enable the combined base device to detect the poses of additional target objects, such as moving objects, and be utilized in various additional image processing applications, in order to increase its overall appeal and usefulness to potential end-users. Therefore, it would have been obvious to combine Atanasoaei et al. in view of Li et al. with Ning et al. to obtain the invention as specified in claim 24. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Birchfield et al. U.S. Publication No. 2022/0277472 A1; which is directed towards systems and techniques to determine a pose of an object from an image, wherein one or more neural networks are trained to predict a pose of an object in an image. Ilic et al. U.S. Publication No. 2023/0169677 A1; which is directed towards a method and apparatus to estimate poses of objects in a scene, wherein a neural network determines keypoints of an object in an image and the pose of the object in the image is estimated based on the determined keypoints. Wang et al. U.S. Publication No. 2023/0298204 A1; which is directed towards systems and methods for three-dimensional pose estimation, wherein one or more neural networks are trained to detect keypoints of a subject in images and estimate a pose of the subject. Any inquiry concerning this communication or earlier communications from the examiner should be directed to ERIC RUSH whose telephone number is (571) 270-3017. The examiner can normally be reached 9am - 5pm Monday - Friday. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Bee can be reached at (571) 270 - 5183. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ERIC RUSH/Primary Examiner, Art Unit 2677
Read full office action

Prosecution Timeline

Dec 02, 2024
Application Filed
Aug 12, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12734404
METHOD, DEVICE, AND NON-TRANSITORY COMPUTER-READABLE RECORDING MEDIUM FOR ESTIMATING INFORMATION ON GOLF SWING
4y 5m to grant Granted Sep 15, 2026
Patent 12738066
METHOD AND SYSTEM FOR IDENTIFYING EMERGING THREATS IN REAL-TIME
2y 5m to grant Granted Sep 15, 2026
Patent 12738042
ARTIFICIAL INTELLIGENCE SYSTEM BASED ON SPATIAL-TEMPORAL INFORMATION PAIRS
2y 1m to grant Granted Sep 15, 2026
Patent 12725363
TAGGING VIRTUALIZED CONTENT
2y 2m to grant Granted Sep 01, 2026
Patent 12711388
ALIGNING SEQUENCES BY GENERATING ENCODED REPRESENTATIONS OF DATA ITEMS
5y 3m to grant Granted Aug 18, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
61%
Grant Probability
97%
With Interview (+36.1%)
3y 5m (~1y 7m remaining)
Median Time to Grant
Low
PTA Risk
Based on 645 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month