Prosecution Insights
Last updated: October 04, 2026
Application No. 18/757,104

OBJECT IDENTIFICATION BASED ON MACHINE VISION AND 3D MODELS

Final Rejection §103
Filed
Jun 27, 2024
Priority
Jan 14, 2022 — provisional 63/266,811 +1 more
Examiner
GARCIA, SANTIAGO
Art Unit
2673
Tech Center
2600 — Communications
Assignee
Materialise N V
OA Round
2 (Final)
88%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 88% — above average
88%
Career Allowance Rate
907 granted / 1032 resolved
+25.9% vs TC avg
Moderate +14% lift
Without
With
+13.6%
Interview Lift
resolved cases with interview
Typical timeline
2y 3m
Avg Prosecution
20 currently pending
Career history
1046
Total Applications
across all art units

Statute-Specific Performance

§101
8.0%
-32.0% vs TC avg
§103
62.8%
+22.8% vs TC avg
§102
17.8%
-22.2% vs TC avg
§112
1.5%
-38.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1032 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant's arguments filed 06/11/2026 have been fully considered but they are not persuasive. First applicant argues, neither Yang nor Yao teaches or suggests the recited "image code."… “"generating, for each of the one or more objects, a plurality of image codes,". The Examiner respectfully disagrees. The claim language is too general as to a what is actually meant by the “image code” and so then Yang reads on this limitation. An image code is far too general and Yang teaches code means. And inherently any image to be displayed has some form of code as well, hence why Yang teaches, “image code” there is nothing specific other than moving this code around which Yang does as well in the current claim limitations. And in ¶[0036] and [0083] teach the “code means” and also deals ¶[0083] The object detection pipeline uses machine learning or deep learning algorithms to detect those objects them being neural networks. And the 2D image this machine learning model clearly. And there is nothing in the claim language to make this different. The applicant argues that, Yang does not disclose generating such encoded representations, nor does Yang disclose receiving image codes as outputs of a machine-learning model. However, by having those code means with the detection as in ¶[0083] would teach the claimed limitations. There are no details on what the code is improving and it is not clear from the claims how would this be different, as anything in the process would contain some form of “code”. Second, applicant argues that neither Yang nor Yao teaches or suggests training a second machine-learning model on image codes labeled with identifiers of corresponding objects. Claims 1, 10, and 19 each recite "using supervised learning to train a second machine-learning model on a dataset comprising, for each of the plurality of image codes of each of the one or more objects, the image code labelled with an identifier of a corresponding object." The examiner respectfully disagrees. Yao teaches, in ¶[0026] the use of multiple neural networks, and one neural network connected to the other. Therefore, “training a second machine-learning model” is already taught by Yao and not innovative. There are no details in the claim language about how this neural network would be faster by the specific code of the image code provided or what this image code actually is or differs from any image that contains some form of code. Third, the applicant argues that the Office Action improperly equates Yang's received 3D image with the recited image code. Claims 1, 10, and 19 each recite "receiving a first image code in response to the inputting the one or more captured images" and further recite "inputting the first image code into the second machine-learning model." Thus, the claims recite that the image code is an encoded output generated by the first machine-learning model and subsequently provided as input to the second machine-learning model. The Examiner respectfully disagrees and still points to the fact that there is nothing extra being brought up in the claim limitations to distinguish this code image or what this image code actually is. The interpretation giving is that any image data being used would have some form of code to operate with the particular processors, and at the basic level code of the image being processed. Fourth, applicant argues, that Yao does not remedy these deficiencies. Yao is directed to incremental 2D-to-3D pose lifting for human pose estimation. Yao employs multiple neural-network components that cooperate to estimate and iteratively refine a 3D pose from a detected 2D pose. However, Yao does not disclose generating image codes from images, training a second machine-learning model on image codes labeled with identifiers of corresponding objects, or identifying an object based on such image codes. The Examiner respectfully disagrees. The applicant is correct in that image codes from images is not what Yao was relied upon, however the claims are disclosed in such general terms that Yao can also read on the limitations. And the Office Action did bring in Yao to show the implementation of more than one neural network for training. Further the claims are written in such general terms that any type of image analysis using a neural network having 2D images reads on the claimed limitations. For these reasons the same rejection applies. The Examiner suggests adding more detail on what is meant by the image code and how the second training actually improved the image analysis. Because currently the way they are written is that any image has code in it, and using a neural network to train such image which has image code. And the secondary reference to show using more than one neural network. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Yang (US 2022/0378395) in view of Yao (US 2023/0386072). As per claims 1, 10, and 19 Yang teaches, a method, computer vision system and a non-transitory computer-readable medium for identifying one or more objects by a computer vision system, comprising: obtaining, for each of the one or more objects, a plurality of 2D representations of the object (Yang, fig.8, 803-804 this represents a plurality of 2D representations of the object, by generating 2D feature maps and information about that object with these 2D feature maps ); generating, for each of the one or more objects, a plurality of image codes, the generating comprising inputting, for each of the plurality of 2D representations of the object, the 2D representation into a first machine-leaning model and receiving as output a corresponding image code (Yang, ¶[0036] “In another aspect of the invention, the object of the invention is also realized by computer program comprising code means for implementing any herein described method when said program is run on a processing system.” This represents the image codes, and ¶[0083] “However, the object-detection pipeline may be performed using any machine learning or deep learning algorithms to detect the objects, such as artificial neural networks.” Represent the machine-learning model); using supervised learning to train a second machine-learning model on a dataset comprising (Yang, “This process can repeated until the error converges, and the predicted output data entries are sufficiently similar (e.g. ±1%) to the training output data entries. This is commonly known as a supervised learning technique.” This represents having supervised learning), for each of the plurality of image codes of each of the one or more objects, the image code labelled with an identifier of a corresponding object (Yang, ¶[0135] Thus, the corresponding training input data entries comprise example 3D images and the corresponding training output data entries comprise, for each example 3D image, a first ground truth 2D image that indicates a location, size and shape of the target object with respect to a first view/projection of the 3D image along the first volumetric axis, and a second ground truth 2D image that indicates location, size and shape of the target object with respect to a second view/projection of the 3D image along the second volumetric axis.” This would represent the identifier, location size and shape as an example and ¶ [0137] To constrain the neural network to improve identification of the position of the object in the predicted 2D image, a multi-level loss function can be used.”); inputting one or more captured images of a physical representation of a first object of the one or more objects into the first machine-learning model (Yang, ¶[0132] “An initialized machine-learning algorithm is applied to each input data entry to generate predicted output data entries.” This represents inputting captured images of a physical representation of a first object of the one or more objects into the first machine-learning model); receiving a first image code in response to the inputting the one or more captured images (Yang, 801 receive 3D image represent first image code); inputting the first image code into the second machine-learning model (Yang, fig.8, into 804 feature maps being used to get information 2D, see ¶[0135]); and determining a first identifier corresponding to the first object in response to the inputting the first image code (Yang, 801 first image code and then identifying 804 different things about the object in 2D). Yang doesn’t clearly teach, having a second machine-learning model on a dataset. Yao teaches, having a second machine-learning model on a dataset (Yao, ¶[0026] “In some embodiments, initial 2D to 3D lifting network 112 implements pretrained fully connected networks (FCN). In some embodiments, initial 2D to 3D lifting network 112 implements pretrained graph convolutional networks (GCN). In some embodiments, initial 2D to 3D lifting network 112 implements pretrained locally connected networks (LCN).” teaching multiple networks, and ¶[0048] teaching the insides network as detailed in the current specification). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Yang with Yao’s ability to have multiple machine learning models or neural networks. The motivation would have been to improve quality and efficiency in being able to tell 3D images with 2D graphics as thought by Yao in ¶[002]. As per claims 2, 11 and 20, Yang in view of Yao teaches, the method of claim 1, wherein, for each of the one or more objects, each of the plurality of image codes comprises a corresponding feature vector indicating a probability of the corresponding 2D representation matching a category for each of a plurality of categories (Yang, ¶[0083] “Other machine-learning algorithms such as logistic regression, support vector machines or Naïve Bayesian model are suitable alternatives for forming the object-detection pipeline.” vector machines represents feature vector indicating a probability). As per claims 3 and 12, Yang in view of Yao teaches, the method of claim 2, wherein the first machine-learning model comprises a convolutional neural network having a plurality of computation layers (Yang, ¶[0130] “As previously explained, FIGS. 2 to 4 illustrate an implementation of the underlying inventive concept that uses a hybrid deep convolutional neural network in order to predict whether an object is present (and optionally a location of the object) within a 3D image of a region of interest.” machine-learning model comprises a convolutional neural network having a plurality of computation layers and ¶ 0131] “The structure of an artificial neural network (or, simply, neural network) is inspired by the human brain. Neural networks are comprised of layers, each layer comprising a plurality of neurons.” And this represents a hidden layer of the plurality of computation layers of the first machine-learning model being the feature vector ), an output of a hidden layer of the plurality of computation layers of the first machine-learning model being the feature vector (Yao, ¶[0048] As shown in FIG. 5, neural network 500 includes a fully connected layer 511 and a fully connected layer 512 separated by a residual block 513. Fully connected layer 511 receives feature set 501 and increases the dimensionality of input feature set 501. In some embodiments, fully increases the dimensionality of input feature set 501 to 256 or a vector length of 256 (e.g., from J*2 to 256, where J is the number of body joints). In some embodiments, fully connected layer 511 includes or is followed by (prior to residual block 513) a batch normalization layer, a ReLU layer, and a dropout layer to apply a particular drop out ratio such as 0.25. Residual block 513 includes hidden layers 514, 515, a residual connection 516, and a residual adder 517. In some embodiments, each of hidden layers 514, 515 has a number of nodes equal to the dimensionality of the output of fully connected layer 511 (e.g., 256 nodes). In some embodiments, each of hidden layers 514, 515 is followed by a dropout with a particular drop out ratio such as 0.25. Fully connected layer 511 receives the output from adder 517 (e.g., a sum of the output from hidden layer 515 and features carried forward from fully connected layer 511). Fully connected layer 511 predicts 3D pose increment 502 for use as discussed herein.” Represents the hidden layers ). As per claims 4 and 13, Yang in Yao teaches, teaches, the method of claim 1, wherein the second machine-learning model comprises a network having hidden layers densely connected (Yao, ¶[0048] “As shown in FIG. 5, neural network 500 includes a fully connected layer 511 and a fully connected layer 512 separated by a residual block 513. Fully connected layer 511 receives feature set 501 and increases the dimensionality of input feature set 501. In some embodiments, fully increases the dimensionality of input feature set 501 to 256 or a vector length of 256 (e.g., from J*2 to 256, where J is the number of body joints). In some embodiments, fully connected layer 511 includes or is followed by (prior to residual block 513) a batch normalization layer, a ReLU layer, and a dropout layer to apply a particular drop out ratio such as 0.25. Residual block 513 includes hidden layers 514, 515, a residual connection 516, and a residual adder 517. In some embodiments, each of hidden layers 514, 515 has a number of nodes equal to the dimensionality of the output of fully connected layer 511 (e.g., 256 nodes). In some embodiments, each of hidden layers 514, 515 is followed by a dropout with a particular drop out ratio such as 0.25. Fully connected layer 511 receives the output from adder 517 (e.g., a sum of the output from hidden layer 515 and features carried forward from fully connected layer 511). Fully connected layer 511 predicts 3D pose increment 502 for use as discussed herein.” This represents a network having hidden layers densely connected). As per claim 5, Yang in view of Yao teaches, the method of claim 1, wherein the second machine-learning model is trained only on the dataset specific to the one or more objects and the corresponding 2D digital representations (Yang, ¶[0134] “and the desired output comprises a first predicted 2D image that indicates a predicted location, size and shape of the target object with respect to a first view/projection of the 3D image along a first volumetric axis and a second predicted 2D image that indicates a predicted location, size and shape of the target object with respect to a second view/projection of the 3D image along a second volumetric axis. The first and second volumetric axes are preferably perpendicular.” the second machine-learning model is trained only on the dataset specific to the one or more objects and the corresponding 2D digital representations). As per claims 6 and 15, Yang in view of Yao teaches, the method of claim 1, wherein the one or more captured images comprise a plurality of captured images that are captured from at least two different viewpoints or orientations (Yang, ¶[0062] “Thus, in the context of ultrasound 3D data, a reduction of the 3D data along a depth axis provides an axial projection or “axial view” of the 3D data, whereas a reduction of the 3D data along a width or height axis provides a side projection or “side view” of the 3D data.” Represents 2 views ). As per claims 7 and 16, Yang in view of Yao teaches, the method of claim 1, wherein the one or more objects comprise a plurality of objects, and further comprising: inputting captured images of physical representations of the plurality of objects into the first machine-learning model; receiving a plurality of image codes in response to inputting the captured images; inputting the plurality of image codes into the second machine-learning model; receiving a plurality of identifiers corresponding to the plurality of objects in response to the inputting the plurality of image codes; and outputting an identification report comprising: a corresponding identifier of each of the plurality of objects identified in the captured images, and an indication of at least one object of the one or more objects not identified in the captured images (Yang, ¶[0080] “Thus, in some embodiments, no more than two 2D feature representations (each 2D feature representation associated with a different projection direction) need to be generated to enable identification of a location of a target object within a 3D image or region of interest.” This represents corresponding identifier, and the rest of the claim is repetitive from claim 1 and ¶[0091] In other examples, this process may comprise processing the 2D feature representations to determine whether or not an object is present within the 2D feature representations (and therefore within the overall 3D image).” Represents not identified in the captured images). As per claims 8, 14 and 17, Yang in view of Yao teaches, the method of claim 1, wherein the plurality of 2D representations of the object comprise a plurality of 2D rendered images, and wherein obtaining, for each of the one or more objects, the plurality of 2D representations of the object comprises: receiving, for each of the one or more objects, a 3D digital representation of the object; and generating, for each of the one or more objects, the plurality of 2D rendered images based on the 3D digital representation of the object (Yang, fig.8, 801 receive 3D images, and then 2D feature maps. Fig.1 from 3D to 2D processing). As per claims 9 and 18, Yang in view of Yao teaches, the method of claim 1, further comprising selecting the one or more captured images from a plurality of captured images corresponding to a video stream ( Yang, ¶[008] “he main limitation is that the complex feature processing, i.e. extraction, selection, compression and classification, for a complete image is very time consuming.” This represents selection and ¶[009] “[0009] The current processing pipelines for object detection in ultrasound images, such as machine learning-based or deep learning-based, mainly focus on the accuracy of object detection, which therefore sometimes introduces a complex algorithm design in full 3D space.” This would then create the video). Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SANTIAGO GARCIA whose telephone number is (571)270-5182. The examiner can normally be reached Monday-Friday 9:30am-5:30pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chineyere Wills-Burns can be reached at (571) 272-9752. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SANTIAGO GARCIA/Primary Examiner, Art Unit 2673 /SG/
Read full office action

Prosecution Timeline

Jun 27, 2024
Application Filed
May 13, 2026
Non-Final Rejection mailed — §103
Jun 11, 2026
Response Filed
Aug 18, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749566
SYSTEMS AND METHODS FOR PROCESSING ELECTRONIC IMAGES
3y 1m to grant Granted Sep 29, 2026
Patent 12743804
DETERIORATION DETERMINATION DEVICE, DETERIORATION DETERMINATION METHOD, AND PROGRAM
2y 9m to grant Granted Sep 22, 2026
Patent 12738061
DYNAMICALLY COMPOSABLE OBJECT TRACKER CONFIGURATION FOR INTELLIGENT VIDEO ANALYTICS SYSTEMS
2y 4m to grant Granted Sep 15, 2026
Patent 12731423
CIRCUIT BOARD PROCESSING SYSTEM USING FRAGMENTED SEARCH REGION IMAGE ANALYSIS
2y 8m to grant Granted Sep 08, 2026
Patent 12718530
MODULAR MACHINE LEARNING SYSTEMS, DATABASES, METHODS, AND COMPUTER PROGRAM PRODUCTS FOR DEVELOPING AND DEPLOYING AUTOMATED IMAGE SEGMENTATION PROGRAMS
2y 4m to grant Granted Aug 25, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
88%
Grant Probability
99%
With Interview (+13.6%)
2y 3m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 1032 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month