Prosecution Insights
Last updated: October 04, 2026
Application No. 18/414,409

METHODS, SYSTEMS, AND COMPUTER-READABLE STORAGE MEDIUMS FOR POSITIONING TARGET OBJECT

Final Rejection §103§112
Filed
Jan 16, 2024
Priority
Aug 09, 2021 — CN 202110905411.5 +1 more
Examiner
BALI, VIKKRAM
Art Unit
2663
Tech Center
2600 — Communications
Assignee
Zhejiang Huaray Technology Co. Ltd.
OA Round
2 (Final)
82%
Grant Probability
Favorable
3-4
OA Rounds
1m
Est. Remaining
93%
With Interview

Examiner Intelligence

Grants 82% — above average
82%
Career Allowance Rate
527 granted / 647 resolved
+19.5% vs TC avg
Moderate +12% lift
Without
With
+11.9%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
33 currently pending
Career history
676
Total Applications
across all art units

Statute-Specific Performance

§101
16.7%
-23.3% vs TC avg
§103
52.7%
+12.7% vs TC avg
§102
6.2%
-33.8% vs TC avg
§112
18.5%
-21.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 647 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . All amendments filed on 4/21/2026 have been entered and action follows: Response to Arguments Applicant’s arguments with respect to claims have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim 28 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 28 recites the limitation "wherein different categories" in line 1. There is insufficient antecedent basis for this limitation in the claim. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-5, 10, 20-21 and 24 are rejected under 35 U.S.C. 103 as being unpatentable over Visual reconstruction and localization based robust robotic 6-Dof grasping in the wild, by Liang in view of Hsieh et al (US Pub. 2023/0007960). With respect to claim 1, Liang discloses A method for positioning a target object, comprising: determining an identification result by processing an image based on an identification model, wherein the image determining, from the image, a target image of each of determining, based on a first reference image and the target image of each of determine an operating order in which an operating device works on Hwoever, Liang fails to explicitly disclose determining an identification result by processing an image based on an identification model, wherein the image includes a plurality of target objects, (emphasis added), as claimed. Hsieh in the same field teaches determining an identification result by processing an image based on an identification model, wherein the image includes a plurality of target objects, (emphasis added, see figure 3, numerical 600, and paragraph 0030,wherein …In the embodiment of present disclosure, the point cloud computing module 300 is used to obtain the point cloud information of the objects 600 according to the three-dimensional image of the objects 600 captured by the three-dimensional camera 200…), as claimed. It would have been obvious to one ordinary skilled in the art at the effective date of invention to combine the two references as they are analogous because they are solving similar problem of grasping objects using robotic arm using image analysis. The teaching of Hsieh to work the mechanical arm with multiple objects can be incorporated into Liang as suggested (see page 72452 figure 1 workspace), for suggestion, and modifying the system yields identifying object in plurality of objects using a robotic arm (see Hsieh paragraph 0001), for motivation. With respect to claim 2, Liang and Hsieh further discloses determining a first feature of the target image of each of the plurality of target objects; and determining, based on a similarity between the first feature of the target image of each of the plurality of target objects and a second feature, the operating order in which the operating device works on the plurality of target objects, wherein the second feature corresponds to a second reference image, (see Algorithm 1 in section II.E, step 7 for computing feature matching for the object, and section II.E ends with, …operate the manipulator to complete the grasping task), as claimed. With respect to claim 3, Liang and Hsieh further discloses wherein the first feature is obtained based on the target image through a feature extraction model, and the feature extraction model is a machine learning model, (see Algorithm 1, and section II E, Page 72458 right hand column, wherein …In order to select the feature point detection scheme adaptively, we first use the pre-trained YOLO v4 [21] target detector to get the position of the target in the image, and then extract the features by SIFT method), as claimed. With respect to claim 4, Liang and Hsieh further discloses wherein for each of the plurality of target objects, a representation parameter of the first position of the target object includes a direction parameter of an object frame where the target object is located, (see Algorithm 1, step 4 applying YOLO v4, to detect the bounding box for the objects; and section II E, Page 72458 right hand column, wherein …In order to select the feature point detection scheme adaptively, we first use the pre-trained YOLO v4 [21] target detector to get the position of the target in the image, and then extract the features by SIFT method), as claimed. With respect to claim 5, Liang and Hsieh further discloses wherein the representation parameters includes: a plurality of position parameters of a plurality of key points of the object frame, (and section II E, Page 72458 right hand column, wherein …In order to select the feature point detection scheme adaptively, we first use the pre-trained YOLO v4 [21] target detector to get the position of the target in the image, and then extract the features by SIFT method), as claimed. Claims 10 is rejected for the same reasons as set forth in the rejection for claim 1, because claim 10 is claiming subject matter of similar scope as claimed in claim 1. Claims 20 and 21 are rejected for the same reasons as set forth in the rejection for claims 1 and 2, because claims 20 and 21 are claiming subject matter of similar scope as claimed in claim 1 and 2. With respect to claim 24, Liang and Hsieh further discloses wherein the determining the operating order in which the operating device works on the plurality of target objects based on the second position of each of the plurality of target objects includes: adjusting, based on a positional relationship between each of the plurality of target objects and a reference plane in the target image, the second reference image; and determining the operating order in which the operating device works on the plurality of target objects based on an adjusted second reference image, (see Hsieh paragraph 0031, wherein … The depth image module 400 sets the reference point as the origin of the coordinate axis according to the reference point information and readjusts the coordinates in the point cloud information to form a depth image; and paragraph 0034, wherein … The mechanical arm 100 may grasp the correct object based on a sorted order…), as claimed. Claims 6-9, 16 and 26-28 are rejected under 35 U.S.C. 103 as being unpatentable over Visual reconstruction and localization based robust robotic 6-Dof grasping in the wild, by Liang in view of Hsieh et al (US Pub. 2023/0007960) as applied to claim 5 above, and further in view of Taamazyan et al (US Pub. 2022/0405506). With respect to claim 6, Liang and Hsieh disclose all the elements as claimed and rejected in claim 5 above. However, they fail to explicitly disclose wherein the identification model is obtained by a training process, labels in the training process include a sample direction parameter of a sample object frame where each of a plurality of sample objects are located and a plurality of sample position parameters of a plurality of sample key points of the sample object frame; and a loss function includes a first loss item and a second loss item, wherein the first loss item is constructed based on the sample direction parameter, and the second loss item is constructed based on the plurality of sample position parameters by a Wing Loss function, as claimed. Taamazyan in the same field teaches the identification model is obtained by a training process, labels in the training process include a sample direction parameter of a sample object frame where each of a plurality of sample objects are located and a plurality of sample position parameters of a plurality of sample key points of the sample object frame; and a loss function includes a first loss item and a second loss item, wherein the first loss item is constructed based on the sample direction parameter, and the second loss item is constructed based on the plurality of sample position parameters by a Wing Loss function, (see figure 1, numerical 2a-2d as sample objects; paragraph 0117, wherein … a processing pipeline may include receiving images captured by sensor devices (e.g., master cameras 10 and support cameras 30) and outputting control commands for controlling a robot arm, where the processing pipeline is trained, in an end-to-end manner, based on training data that includes sensor data as input and commands for controlling the robot arm (e.g., a destination pose for the end effector 26 of the robotic arm 24) as the labels for the input training data…; and paragraph 0191, wherein … where the modification may include flattening the output of the neural network before supplying the output to the loss function used to train the disparity neural network, such that the loss function accounts identifies and detects disparities along both the x-axis and the y-axis. In some embodiments, an optical flow neural network is trained and/or retrained to operate on the given types of input data…), as claimed. It would have been obvious to one ordinary skilled in the art at the effective date of invention to combine the two references as they are analogous because they are solving similar problem of grasping objects using robotic arm using image analysis. The teaching of Taamazyan to train a model in order to get the position of the object can be incorporated in to Liang and Hsieh as suggested (see Liang page 72454 right hand column wherein …Inspiration from the strategy of neural network training), for suggestion, and modifying the system yields picking an object from plurality of objects using a robotic arm (see Taamazyan paragraph 0004), for motivation. With respect to claim 7, Liang, Hsieh and Taamazyan for the same reasons of combination discloses wherein the identification model includes a feature extraction layer, a feature fusion layer, and an output layer; wherein the feature extraction layer includes a plurality of convolutional layers connected in series, and the plurality of convolutional layers output a plurality of graph features; the feature fusion layer fuses the plurality of graph features to determine a third feature of the image; and the output layer processes the third feature to determine the identification result, (see Taamazyan figure 2, numerical 15 feature extractor; and paragraph 0062, wherein …FIG. 2 is a more detailed block diagram of the vision module 7 according to one embodiment. The vision module 7 may include a feature extractor 15 and a predictor 19 (e.g., a classical computer vision prediction algorithm or a trained statistical model) configured to compute a prediction output 21 (e.g., a statistical prediction) regarding one or more objects 2 in the scene based on the output of the feature extractor…; and paragprah 0166, wherein …one embodiment, the deep learning network 412 is configured to generate feature maps based on the input images 410, and employ a region proposal network (RPN) to propose regions of interest from the generated feature maps. The proposals by the CNN backbone may be provided to a box head 414 for performing classification and bounding box regression. In one embodiment, the classification outputs a class label 416 for each of the object instances in the input images 410, and the bounding box regression predicts bounding boxes 418 for the classified objects…), as claimed. With respect to claim 8, Liang, Hsien and Taamazyan for the same reasons of combination discloses wherein the determining, based on the first reference image and the target image of each of the plurality of target objects, the second position of each of the plurality of target objects in the second coordinate system includes: for each of the plurality of target objects, determining a transformation parameter by processing, based on a transformation model, the first reference image and the target image of the target object; and converting, based on the transformation parameter, a third position of the target object in a third coordinate system into the second position, wherein the third coordinate system is determined based on the target image of the target object, (see Taamazyan paragraph 0100, wherein … A pose estimator 100 according to various embodiments of the present disclosure is configured to compute or estimate poses of the objects 22 based on information captured by the main camera 10 and the support cameras 30 …Types of electronic circuits may include a central processing unit (CPU), a graphics processing unit (GPU), an artificial intelligence (AI) accelerator (e.g., a vector processor, which may include vector arithmetic logic units configured efficiently perform operations common to neural networks, such dot products and softmax); and see Taamazyan paragraph 0169, wherein …block 430, the matching algorithm identifies features of a first object instance in a first segmentation mask. The identified features for the first object instance may include a shape of the region of the object instance, a feature vector in the region, and/or keypoint predictions in the region. The shape of the region for the first object instance may be represented via a set of points sampled along the contours of the region. Where a feature vector in the region is used as the feature descriptor, the feature vector may be an average deep learning feature vector extracted via a convolutional neural network..), as claimed. With respect to claim 9, Liang, Hsieh and Taamazyan for the same reasons of combination discloses wherein the transformation model includes an encoding layer and a conversion layer, wherein the encoding layer processes the target image to determine a first encoding vector, and processes the first reference image to determine a second coding vector; and the conversion layer processes the first encoding vector and the second encoding vector to determine the transformation parameter, (see Taamazyan paragraph 0169, wherein …block 430, the matching algorithm identifies features of a first object instance in a first segmentation mask. The identified features for the first object instance may include a shape of the region of the object instance, a feature vector in the region, and/or keypoint predictions in the region. The shape of the region for the first object instance may be represented via a set of points sampled along the contours of the region. Where a feature vector in the region is used as the feature descriptor, the feature vector may be an average deep learning feature vector extracted via a convolutional neural network), as claimed. Claim 16 is rejected for the same reasons as set forth in the rejection for claim 7, because claim 16 is claiming subject matter of similar scope as claimed in claim 7. With respect to claim 26, Liang, Hsieh and Taamazyan for the same reasons of combination discloses wherein the determining, from the image, the target image of each of the plurality of target objects includes: for each of the plurality of target objects, when the target object is shielded by another target object in the image, obtaining an updated target image of the target object by processing a shielded area in the target image of the target object, wherein the shielded area refers to an area in the target image of the target object that is shielded by the another target object, (see Taamazyan paragraph 0209, wherein … Additionally, using optical flow to compute correspondences results in a correspondence map for every visible pixel depicting the object and therefore the PnP algorithm has more than enough information to solve for a refined pose), as claimed. With respect to claim 27, Liang, Hsieh and Taamazyan for the same reasons of combination discloses wherein the determining, from the image, the target image of each of the plurality of target objects includes: for each of the plurality of target objects, adjusting an object frame in which the target object is located based on a category of the target object; and segmenting the image based on an adjusted object frame of the target object to obtain the target image of the target object, (see Taamazyan paragraph 0150, wherein … Generally, the approach described in the above-referenced international patent application relates to computing a 6-DoF pose of an object in a scene by determining a class or type of the object (e.g., a known or expected object) and aligning a corresponding 3-D model of the object (e.g., a canonical or ideal version of the object based on known design specifications of the object and/or based on the combination of a collection of samples of the object) with the various views of the object, as captured from different viewpoints around the object…; and paragraph 0154, wherein …A process of instance segmentation identifies the pixels in each of the images that depict the five objects, in addition to labeling them separately based on the type or class of object…), as claimed. With respect to claim 28, Liang, Hsieh and Taamazyan for the same reasons of combination discloses wherein different categories of the plurality of target objects correspond to different first reference images, and the plurality of target objects in the first reference image are complete, un-shielded, and unstacked, (see Taamazyan paragraph 0161, wherein …block 402 the pose estimator 100 performs instance segmentation and mask generation based on the captured images. In this regard, the pose estimator 100 classifies various regions (e.g. pixels) of an image captured by a particular camera 10, 30 as belonging to particular classes of objects…), as claimed. Allowable Subject Matter Claims 23, 23 and 25 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to VIKKRAM BALI whose telephone number is (571)272-7415. The examiner can normally be reached Monday-Friday 7:00AM-3:00PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Gregory Morse can be reached at 571-272-3838. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /VIKKRAM BALI/Primary Examiner, Art Unit 2663
Read full office action

Prosecution Timeline

Jan 16, 2024
Application Filed
Feb 06, 2026
Non-Final Rejection mailed — §103, §112
Apr 21, 2026
Response after Non-Final Action
Apr 21, 2026
Response Filed
Jun 25, 2026
Response Filed
Sep 02, 2026
Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749201
PEDESTRIAN TRAJECTORY PREDICTION METHOD
3y 7m to grant Granted Sep 29, 2026
Patent 12749581
DEPTH NETWORK DETECTION METHOD FOR DIABETIC RETINOPATHY BASED ON GENETIC FUZZY TREE
2y 0m to grant Granted Sep 29, 2026
Patent 12743763
SYSTEM FOR TRAINING A DEEP-LEARNING ALGORITHM AND ASSOCIATED METHOD
2y 4m to grant Granted Sep 22, 2026
Patent 12731267
SPATIAL REGIME CHANGE DETECTION FOR VIDEO ANALYTICS
3y 10m to grant Granted Sep 08, 2026
Patent 12725298
MODEL TRAINING METHOD AND APPARATUS, DEVICE, MEDIUM, AND PROGRAM PRODUCT
3y 3m to grant Granted Sep 01, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
82%
Grant Probability
93%
With Interview (+11.9%)
2y 10m (~1m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 647 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month