Prosecution Insights
Last updated: August 17, 2026
Application No. 19/004,417

3D Shape Part Segmentation by Vision-Language Model Distillation

Non-Final OA §103§112
Filed
Dec 29, 2024
Priority
Dec 29, 2023 — provisional 63/615,818
Examiner
RENZE, GEORGE NICHOLAS
Art Unit
Tech Center
Assignee
MediaTek Inc.
OA Round
1 (Non-Final)
72%
Grant Probability
Favorable
1-2
OA Rounds
11m
Est. Remaining
91%
With Interview

Examiner Intelligence

Grants 72% — above average
72%
Career Allowance Rate
23 granted / 32 resolved
+11.9% vs TC avg
Strong +19% interview lift
Without
With
+18.8%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
20 currently pending
Career history
62
Total Applications
across all art units

Statute-Specific Performance

§101
2.5%
-37.5% vs TC avg
§103
74.2%
+34.2% vs TC avg
§102
15.7%
-24.3% vs TC avg
§112
7.6%
-32.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 32 resolved cases

Office Action

§103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 2-7 and 13-18 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claims 2 and 13 recite the limitation "the vision-language model" in line 3 of each of the claims. There is insufficient antecedent basis for this limitation in these claims because “the vision-language model” is lacking antecedent basis and should instead be referred to as “[[the]] a vision-language model”. Claims 4 and 15 recite the limitation "the distillation head" in line 2 of each of the claims. There is insufficient antecedent basis for this limitation in these claims because “the distillation head” is lacking antecedent basis and should instead be referred to as “[[the]] a distillation head”. Claims 3, 5-7, 14 and 16-18 are also rejected due to their dependence on a previously 112(b) rejected claim. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 9-12 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Li et al. (Pub. No.: US 2020/0027215 A1), hereinafter Li, in view of Krishna et al. (US 2025/0103642 A1), hereinafter Krishna. Regarding claim 1, Li discloses a method for three-dimensional (3D) shape part segmentation (FIG. 1 and paragraph 20 teach that FIG. 1 is a flow schematic diagram of one embodiment of a three-dimensional shape expression method of the present disclosure.), performed by a processor (Paragraph 60 teaches that persons of ordinary skills in the art are capable to understand that achieving all or part of the processes in the method of the above embodiments are completed by a computer program instructing relevant hardwares, the computer program is stored in a computer readable storage medium, when the computer program is executed, processes of the embodiments of above methods are included.), comprising: obtaining two-dimensional (2D) predictions for part segmentation of a 3D shape (Paragraph 33 teaches that step 101: extracting a hybrid type framework of a three-dimensional shape and FIG. 2 and paragraph 37 teach that the hybrid type framework in step 101 includes: a one-dimensional curve and a two-dimensional slice. And FIG. 2 is a flow schematic diagram of one embodiment of step 101.); lifting the 2D predictions onto the 3D shape to obtain initial 3D part segmentation knowledge (Paragraph 34 teaches that step 102: obtaining a segmentation of the three-dimensional shape by segmenting the hybrid type framework and paragraph 35 teaches that step 103: obtaining a sub-structure of the three-dimensional shape according to the segmentation of the three-dimensional shape.); processing the 3D shape to extract geometric features (Paragraph 40 teaches that in step 403, the sub-graph with the number of connecting nodes as n=1, . . . , 5 in the connecting graph is extracted, so that the sub-structure corresponding to the three-dimensional shape is obtained, the sub-structure corresponding to the three-dimensional shape is represented by a series of geometrical features, and a value of n is not limited to the above example.). However, Li fails to disclose processing the 3D shape using a 3D encoder. Krishna discloses processing the 3D shape using a 3D encoder (Paragraph 206 teaches that the one or more generative models 90 may include a vision language model. The vision language model can be trained, tuned, and/or configured to process image data and/or text data to generate a natural language output. The vision language model may leverage a pre-trained large language model (e.g., a large autoregressive language model) with one or more encoders (e.g., one or more image encoders and/or one or more text encoders) to provide detailed natural language outputs that emulate natural language composed by a human. Additionally, paragraph 201 teaches that rendering dataset generation may include training one or more neural radiance field models to learn a three-dimensional representation for one or more objects.). Since Li teaches the initial method steps for obtaining two-dimensional data in relation to three-dimensional part segmentation information and the ability to extract geometric features from three-dimensional shapes and Krishna teaches using an encoder to help in encoding data related to three-dimensional objects, it would have been obvious to a person having ordinary skill in the art to combine the functions together so that any three-dimensional object data extracted could be obtained by using an encoder to help in extracting additional geometrical feature data related to a three-dimensional shaped object. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Li to incorporate the teachings of Krishna, so that the combined functions together would provide the capabilities of using a 3D encoder to assist in extracting three-dimensional shape geometry, which could potentially help improve overall accuracy and predictions of a student network within a vision language model. Furthermore, Li in view of Krishna disclose performing a distillation process to refine the initial 3D part segmentation knowledge (Paragraph 209 of Krishna teaches that the one or more generative models 90 may include one or more compact vision language models that may include less parameters than a vision language model stored and operated by the server computing system. The compact vision language model may be trained via distillation training.); and generating a final 3D shape part segmentation according to the refined 3D part segmentation knowledge and the geometric features (Paragraph 41 of Li teaches that finally, normalizing the term vectors to obtain the expression of the three-dimensional shape, that is, step 503. Additionally, paragraph 92 of Krishna teaches that pixels descriptive of the object may be segmented and searched. In some implementations, the image segment can be processed with a generative model (e.g., a vision language model and/or a large language model) to generate a model-generated response 418 to the query. The model-generated response 418 can include a natural language response that summarizes one or more web resources determined to be associated with the segmented object. Additionally, and/or alternatively, the segmented image can be processed to determine one or more visual search results 420. The one or more visual search results 420 may be determined based on classification label matching, embedding search, feature matching, clustering, and/or image matching. The one or more visual search results 420 may include product listings, articles, and/or other web resources. The one or more visual search results 420 may be provided with visual matches 422 that include images that depict objects that match the segmented object. The search results interface can include search results of a plurality of different types and may be displayed in a plurality of different formats in a plurality of different panels.). Regarding claim 9, Li in view of Krishna disclose everything claimed as applied above (see claim 1), in addition, Li in view of Krishna disclose further comprising generating a mask indicating which points of the 3D shape are covered by the 2D predictions (Paragraph 58 of Li teaches that the framework extracting module 111 is specifically configured to obtain the sampling points by sampling the surfaces of the three-dimensional shape, and re-express the sampling points to obtain the hybrid type framework including the one-dimensional curve and the two-dimensional slice. The segmentation module 112 is specifically configured to segment the hybrid type framework, and obtain the segmentation of the three-dimensional shape by segmenting the hybrid type framework, according to the corresponding relationships between the hybrid type framework and the sampling points. Additionally, paragraph 122 of Krishna teaches that the one or more bounding boxes and the display data may be processed with a segmentation model to generate masks for each of the detected objects to segment the objects from the one or more images of the display data and/or generate detailed outlines of the objects that indicate object boundaries. In some implementations, the segmented objects may be processed with a search engine and/or one or more additional machine-learned models to generate the visual search data. The search engine may determine one or more visual search results based on detected features in the image segments, an embedding search (e.g., embedding neighbor determination), one or more object classifications, one or more image classifications, application classification, and/or multimodal search (e.g., search based on the image segment and text data (e.g., input text, metadata, text labels, etc.)). Lastly, paragraph 74 of Krishna teaches that the semantic analysis model 242 can be utilized to process the display data to generate a semantic output descriptive of an understanding of the display data with regards to topic understanding, scene understanding, a focal point, pattern recognition, application understanding, and/or one or more other semantic outputs.). Regarding claim 10, Li in view of Krishna disclose everything claimed as applied above (see claim 1), in addition, Li in view of Krishna disclose wherein data of the 3D shape is stored in a memory (Paragraph 60 of Li teaches that persons of ordinary skills in the art are capable to understand that achieving all or part of the processes in the method of the above embodiments are completed by a computer program instructing relevant hardwares, the computer program is stored in a computer readable storage medium, when the computer program is executed, processes of the embodiments of above methods are included. The storage medium is a magnetic disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM) and so on.). Regarding claim 11, the apparatus steps correspond to and are rejected similarly to the method steps of claim 1 (see claim 1 above). In addition, Li in view of Krishna disclose an apparatus for three-dimensional (3D) shape part segmentation (FIG. 11 and paragraph 30 of Li teach that FIG. 11 is a structural schematic diagram of one embodiment of a three-dimensional expression device of the present disclosure.), comprising: a memory configured to store instructions and 3D shape data (Paragraph 60 of Li teaches that persons of ordinary skills in the art are capable to understand that achieving all or part of the processes in the method of the above embodiments are completed by a computer program instructing relevant hardwares, the computer program is stored in a computer readable storage medium, when the computer program is executed, processes of the embodiments of above methods are included. The storage medium is a magnetic disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM) and so on.); and a processor coupled to the memory, configured to execute the instructions (Paragraph 67 of Krishna teaches that the hardware 218 can include physical parts of the user computing device 210, which can include a central processing unit, a graphics processing unit, random access memory, speakers, a sound card, computer data storage, input components, physical display components (e.g., a visual display), and/or other hardware components.). Regarding claim 12, Li in view of Krishna disclose everything claimed as applied above (see claim 11), in addition, Li in view of Krishna disclose a graphics processing unit (GPU) coupled to the processor, configured to accelerate rendering of the 2D images and processing of the 3D shape (Paragraph 67 of Krishna teaches that the hardware 218 can include physical parts of the user computing device 210, which can include a central processing unit, a graphics processing unit, random access memory, speakers, a sound card, computer data storage, input components, physical display components (e.g., a visual display), and/or other hardware components.); a network interface coupled to the processor, configured to receive 3D shape data and transmit segmentation results (Paragraph 73 of Krishna teaches that the visual search interface 216 can communicate over a network with a server computing system 230 to provide a plurality of additional processing services. The server computing system 230 can include one or more generative models 232, one or more object detection models 234, one or more segmentation models 236, one or more classification models 238, one or more embedding models 240, one or more semantic analysis models 242, and/or one or more search engines 244.); and a display device coupled to the processor, configured to display the final 3D shape part segmentation (Paragraph 53 of Krishna teaches that the user computing device 210 can include a visual display. The visual display can display a plurality of pixels. The plurality of pixels can be configured to display content associated with one or more applications 212.). Regarding claim 20, the apparatus steps correspond to and are rejected similarly to the method steps of claim 9 (see claim 9 above). Claims 2-3, 8, 13-14 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Li in view of Krishna as applied to claims 1 and 11 above, and further in view of Liu et al. (Pub. No.: US 2024/0144589 A1), hereinafter Liu. Regarding claim 2, Li in view of Krishna disclose everything claimed as applied above (see claim 1), in addition, Li in view of Krishna disclose wherein obtaining the 2D predictions comprises: rendering a plurality of 2D images of the 3D shape (Paragraph 201 of Krishna teaches that the augmented-reality experience may render information associated with an environment into the respective environment. Alternatively, and/or additionally, objects related to the processed dataset(s) may be rendered into the user environment and/or a virtual environment. Rendering dataset generation may include training one or more neural radiance field models to learn a three-dimensional representation for one or more objects.). However, Li in view of Krishna fail to disclose rendering a plurality of 2D images from a plurality of views of the 3D shape. Liu discloses rendering a plurality of 2D images from a plurality of views of the 3D shape (Paragraph 40 teaches that the system 300 can receive as input a 3D capture 302 of an object. The process can include rending multi-view 2D images 304 from the 3D capture. The system 300 can also receive part data 308 (e.g., such as text) associated with the part, such as “chair back” or “chair leg” of a chair.). Since Li in view of Krishna teach the initial rendering and displaying process of 3D objects and shapes and Liu teaches a rendering technique that can use multi-view 2D images from a 3D capture to help in rendering different shapes/parts of an object, it would have been obvious to a person having ordinary skill in the art to combine the functions together so that multiple views of multiple 2D images could be utilized to help render a 3D shape of an object. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Li in view of Krishna to incorporate the teachings of Liu, so that the combined functions together would improve the overall accuracy and detail of rendering a 3D shape by incorporating multiple different 2D images and views of the desired 3D shape to render. Furthermore, Li in view of Krisha and Liu disclose and processing the plurality of 2D images using the vision-language model (VLM) to obtain the 2D predictions for part segmentation (Paragraph 204 of Krishna teaches that the one or more generative models 90 can include language models (e.g., large language models and/or vision language models), image generation models (e.g., text-to-image generation models and/or image augmentation models), audio generation models, video generation models, graph generation models, and/or other data generation models (e.g., other content generation models). Additionally, paragraph 60 of Liu teaches that at block 906, the process 900 can include processing the one or more two-dimensional images of the object to generate at least one two-dimensional bounding box (e.g., 2D bounding box(es) 314/346 of FIG. 3A and/or FIG. 3B) that identifies, based on a vision language pretrained model and the data, the part of the object.). Regarding claim 3, Li in view of Krisha and Liu disclose everything claimed as applied above (see claim 2), in addition, Li in view of Krisha and Liu disclose wherein the VLM generates bounding box predictions and/or pixel-wise predictions (Paragraph 40 of Liu teaches that the system 300 can generate multi-view 2D images 304 from the 3D capture 302 of the object. Using the multi-view 2D images 304 and the part data 308, the machine learning model 306 (e.g., a VLP model) can determine or generate one or more 2D bounding boxes 314 that identify the parts. The one or more 2D bounding boxes 314 are provided to a 3D fusion engine 312 that also receives the 3D capture 302 of the object. The 3D fusion engine 312 can output a 3D point cloud 310 of the object with part segmentation. Additionally, FIG. 5 and paragraph 52 of Liu teach that FIG. 5 illustrates various bounding boxes identifying chair parts 500 and shows how multi-view feature aggregation can occur. The chairs are shown in various views and the machine learning model 306/342 can predict good results on some views but poor results on others. Chair 502 shows one bounding box 504 for the whole chair and another bounding box 506 for the seat. Lastly, paragraph 205 of Krishna teaches that the one or more generative models 90 can be trained to process input data and generate model-generated content items, which may include a plurality of predicted words, pixels, signals, and/or other data.). Regarding claim 8, Li in view of Krishna disclose everything claimed as applied above (see claim 1), however, Li in view of Krishna fail to disclose wherein lifting the 2D predictions onto the 3D shape comprises performing back-projection of the 2D predictions using camera parameters associated with the rendering of the 2D images. Liu discloses wherein lifting the 2D predictions onto the 3D shape comprises performing back-projection of the 2D predictions using camera parameters associated with the rendering of the 2D images (Paragraph 36 of Liu teaches that in some cases, the systems and techniques can apply to any 2D visual prediction task, such as semantic segmentation and depth estimation. Visual prediction tasks can be integral parts to many applications or systems, such as extended reality (XR), vehicle systems (e.g., autonomous or semi-autonomous driving, safety systems, etc.), camera image/video processing, robotics (e.g., as shown in FIG. 1), and/or other applications or systems.). Since Li in view of Krishna teach the initial method steps for lifting 2D predictions onto a 3D shape and Liu teaches a 2D prediction function related to semantic segmentation and depth estimations and can utilize camera parameters that would be associated with camera image/video processing, it would have been obvious to a person having ordinary skill in the art to combine the functions together so that any 2D prediction could also utilize a cameras parameters that is associated with any rendering of a 2D image or 3D shape. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Li in view of Krishna to incorporate the teachings of Liu, so that the combined functions together would provide more accurate 2D predictions, which should also improve upon any geometrical features related to the rendering of the 2D images as well. Regarding claim 13, the apparatus steps correspond to and are rejected similarly to the method steps of claim 2 (see claim 2 above). Regarding claim 14, the apparatus steps correspond to and are rejected similarly to the method steps of claim 3 (see claim 3 above). Regarding claim 19, the apparatus steps correspond to and are rejected similarly to the method steps of claim 8 (see claim 8 above). Claims 4-7 and 15-18 are rejected under 35 U.S.C. 103 as being unpatentable over Li in view of Krishna as applied to claims 1 and 11 above, and further in view of Li et al. (Pub. No.: US 2024/0212374 A1), hereinafter Li 2, and Yu et al. (Pub. No.: US 2022/0261593 A1), hereinafter Yu. Regarding claim 4, Li in view of Krishna disclose everything claimed as applied above (see claim 1), however, Li in view of Krishna fail to disclose wherein the distillation process comprises: performing forward distillation by aligning output of the distillation head with the lifted initial 3D part segmentation knowledge. Li 2 discloses performing forward distillation by aligning output of the distillation head with the lifted initial 3D part segmentation knowledge (Paragraph 148 teaches that in this embodiment, the above fused features are obtained based on multi-scale fusion-single knowledge distillation (MSFSKD). MSFSKD is the key to 2DPASS, which aims to improve the three-dimensional representation of each scale by fusion and distillation using assisted two-dimensional priori. The design of the knowledge distillation (KD) of MSFSKD is partly inspired by XMUDA. However, XMUDA deals with KD in a simple cross-modal way, that is, outputs of two sets of single-modal features (i.e., the two-dimensional features or the three-dimensional features) are simply aligned, which inevitably pushes the two sets of modal features into an overlapping space thereof. Thus, this way actually discards the information of the specific modal, which is the key to multi-sensor segmentation. Although this problem can be mitigated by introducing an additional layer of segmented prediction, it is inherent in cross-modal distillation and thus results in biased predictions. Therefore, an MSFSKD module is provided, as shown in FIG. 5. Firstly, the image and the features of the point cloud are fused using an algorithm, and then the fused features of the point cloud are unidirectionally aligned. In the fusion-before and distillation-after method, the fusion preserves the complete information from the multi-modal data. In addition, unidirectional alignment ensures that the features of the enhanced point cloud after fusion does not discard any modal feature information.). Since Li in view of Krishna teach a distillation process for refining lifted initial 3D part segmentation knowledge and Li 2 teaches a forward (knowledge) distillation process that can fuse and align different three-dimensional features and points with one another, it would have been obvious to a person having ordinary skill in the art to combine the functions together so that any 3D part segmentation knowledge lifted during the distillation process, could be aligned with other points and distillation models/networks (such as a distillation head/teacher model). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Li in view of Krishna to incorporate the teachings of Li 2, so that the combined functions together would provide the distillation process with the capabilities to output data from a distillation head/teacher model to a student model with improved geometric and structural accuracy due to proper alignment of the different 3D part segmentation knowledge and the distillation head. Furthermore, however, Li in view of Krishna and Li2 fail to disclose and performing backward distillation to refine the lifted initial 3D part segmentation knowledge based on the aligned output of the distillation head. Yu discloses and performing backward distillation to refine the lifted initial 3D part segmentation knowledge based on the aligned output of the distillation head (FIG. 41A and paragraph 590 teach that FIG. 41A illustrates a data flow diagram for a process 4100 to train, retrain, or update a machine learning model, in accordance with at least one embodiment. Additionally, paragraph 136 teaches that in at least one embodiment, training framework 904 trains untrained neural network 906 repeatedly while adjust weights to refine an output of untrained neural network 906 using a loss function and adjustment algorithm, such as stochastic gradient descent. In at least one embodiment, training framework 904 trains untrained neural network 906 until untrained neural network 906 achieves a desired accuracy and paragraph 90 teaches that in at least one embodiment, segmentation scores of a student network 206 are aligned with those of a teacher network 204 for all box proposals (e.g., box proposals 206D).). Since Li in view of Krishna and Li 2 teach the capabilities for performing a distillation training process which uses 3D part segmentation knowledge and Yu teaches a distillation training process that allows for retraining/refining data by repeating the distillation training process multiple times until a student network segmentation scores are aligned with that of a teacher’s network segmentation scores, it would have been obvious to a person having ordinary skill in the art to combine the functions together so that any of the distillation alignment data acquired initially and associated with the initial 3D part segmentation knowledge could be used to help retrain and refine the overall distillation process. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Li in view of Krishna and Li 2 to incorporate the teachings of Yu, so that the combined functions together would provide the distillation process with a way to retrain and refine the lifted initial 3D part segmentation knowledge from a teacher model, which would then provide more accurate and aligned output data from a student model. Regarding claim 5, Li in view of Krishna, Li 2 and Yu disclose everything claimed as applied above (see claim 4), in addition, Li in view of Krishna, Li 2 and Yu disclose wherein the distillation head comprises the geometric features (Paragraph 89 of Yu teaches that in at least one embodiment, a teacher network 204 shares a same architecture as a student network 206, with an addition of a conditional random fields (CRF) module. In at least one embodiment, a teacher network 204 is initialized using same parameters as a student network 206. In at least one embodiment, a feature backbone 204A, a feature pyramid 204B, a prediction head 204C, a mask coefficients 204E, a prototype network 204F, mask prototypes 204G, and one or more masks 210 of a teacher network 204 are same or similar as a feature backbone 206A, a feature pyramid 206B, a prediction head 206C, a mask coefficients 206E, a prototype network 206F, mask prototypes 206G, and masks 208, respectively, of a student network 206 and paragraph 475 of Yu teaches that during execution, in at least one embodiment, graphics and media pipelines send thread initiation requests to thread execution logic 3200 via thread spawning and dispatch logic. In at least one embodiment, once a group of geometric objects has been processed and rasterized into pixel data, pixel processor logic (e.g., pixel shader logic, fragment shader logic, etc.) within shader processor 3202 is invoked to further compute output information and cause results to be written to output surfaces (e.g., color buffers, depth buffers, stencil buffers, etc.). ... In at least one embodiment, arithmetic operations on texture data and input geometry data compute pixel color data for each geometric fragment or discards one or more pixels from further processing. Additionally, paragraph 44 of Li teaches that h.sub.i and h.sub.j are respectively formed by a connection of geometric feature histograms of components of the node n.sub.i and the node n.sub.j, geometric features include a shape diameter function and the three local features based on PCA, and a dimension of each feature histogram is sixteen.). Regarding claim 6, Li in view of Krishna, Li 2 and Yu disclose everything claimed as applied above (see claim 4), in addition, Li in view of Krishna, Li 2 and Yu disclose wherein performing backward distillation comprises re-scoring confidence values associated with the lifted initial 3D part segmentation knowledge according to agreement between the initial knowledge and the aligned output of the distillation head (Paragraph 121 of Yu teaches that In at least one embodiment, a system performing at least a part of process 700 includes executable code to update 712 a student network using output of a teacher network and one or more loss functions and paragraph 241 of Yu teaches that in at least one embodiment, a primary computer may be configured to provide a supervisory MCU with a confidence score, indicating that primary computer's confidence in a chosen result. Additionally, paragraph 591 of Yu teaches that in at least one embodiment, model training 3714 may include retraining or updating an initial model 4104 (e.g., a pre-trained model) using new training data (e.g., new input data, such as customer dataset 4106, and/or new ground truth data associated with input data). In at least one embodiment, to retrain, or update, initial model 4104, output or loss layer(s) of initial model 4104 may be reset, or deleted, and/or replaced with an updated or new output or loss layer(s). In at least one embodiment, initial model 4104 may have previously fine-tuned parameters (e.g., weights and/or biases) that remain from prior training, so training or retraining 3714 may not take as long or require as much processing as training a model from scratch. In at least one embodiment, during model training 3714, by having reset or replaced output or loss layer(s) of initial model 4104, parameters may be updated and re-tuned for a new data set based on loss calculations associated with accuracy of output or loss layer(s) at generating predictions on new, customer dataset 4106 (e.g., image data 3708 of FIG. 37).). Regarding claim 7, Li in view of Krishna, Li 2 and Yu disclose everything claimed as applied above (see claim 4), in addition, Li in view of Krishna, Li 2 and Yu disclose wherein performing forward distillation comprises minimizing a masked cross-entropy loss between the output of the distillation head and the lifted initial 3D part segmentation knowledge (Paragraph 87 of Yu teaches that in at least one embodiment, an auxiliary softmax cross-entropy loss is calculated by one or more systems for a student network 206 to learn a rough semantic segmentation determined by regions indicated by boxes of bounding box annotations. In at least one embodiment, auxiliary softmax cross-entropy loss refers to loss that is based at least in part on one or more softmax functions and one or more cross-entropy loss functions.). Regarding claim 15, the apparatus steps correspond to and are rejected similarly to the method steps of claim 4 (see claim 4 above). Regarding claim 16, the apparatus steps correspond to and are rejected similarly to the method steps of claim 5 (see claim 5 above). Regarding claim 17, the apparatus steps correspond to and are rejected similarly to the method steps of claim 7 (see claim 7 above). Regarding claim 18, the apparatus steps correspond to and are rejected similarly to the method steps of claim 6 (see claim 6 above). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Fu et al. (Pub. No.: US 2018/0174325 A1) teaches methods and systems for segmenting objects using information from a 2D image to improve a 3D model within a 3D representation of a scene. Nagao (Pub. No.: US 2024/0320910 A1) teaches a three-dimensional shape generation apparatus that includes circuitry to generate three-dimensional shape information indicating a three-dimensional shape corresponding to a three-dimensional point cloud, using model shape information indicating a three-dimensional model shape. Armeni et al. (U.S. Patent: #10,277,859 B2) teaches systems and methods for obtaining model object data and generating multi-modal images of a synthetic scene. Any inquiry concerning this communication or earlier communications from the examiner should be directed to George Renze whose telephone number is (703)756-5811. The examiner can normally be reached Monday-Friday 9:00am - 6:00pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Xiao Wu can be reached at (571) 272-7761. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /G.R./Examiner, Art Unit 2613 /XIAO M WU/Supervisory Patent Examiner, Art Unit 2613
Read full office action

Prosecution Timeline

Dec 29, 2024
Application Filed
Jul 29, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12694597
DYNAMIC FLUID DISPLAY METHOD AND APPARATUS, ELECTRONIC DEVICE, AND READABLE MEDIUM
3y 1m to grant Granted Jul 28, 2026
Patent 12620166
RENDERING AS A SERVICE PLATFORM WITH INDUSTRIAL AUTOMATION EMULATION FOR METAVERSE PLATFORM EXECUTION
2y 4m to grant Granted May 05, 2026
Patent 12602407
SYSTEMS AND METHODS FOR GENERATING A UNIQUE IDENTITY FOR A GEOSPATIAL OBJECT CODE BY PROCESSING GEOSPATIAL DATA
2y 7m to grant Granted Apr 14, 2026
Patent 12573147
LANDMARK DATA COLLECTION METHOD AND LANDMARK BUILDING MODELING METHOD
2y 10m to grant Granted Mar 10, 2026
Patent 12555315
HEURISTIC-BASED VARIABLE RATE SHADING FOR MOBILE GAMES
2y 7m to grant Granted Feb 17, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
72%
Grant Probability
91%
With Interview (+18.8%)
2y 7m (~11m remaining)
Median Time to Grant
Low
PTA Risk
Based on 32 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month