Prosecution Insights
Last updated: October 02, 2026
Application No. 18/774,313

APPARATUS FOR BUILDING A DEEP LEARNING MODEL FOR IMAGE LEARNING, AND A METHOD THEREOF

Non-Final OA §103
Filed
Jul 16, 2024
Priority
Nov 06, 2023 — RE 10-2023-0151997
Examiner
SHERRILLO, DYLAN JOSEPH
Art Unit
2665
Tech Center
2600 — Communications
Assignee
Kia Corporation
OA Round
1 (Non-Final)
90%
Grant Probability
Favorable
1-2
OA Rounds
7m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 90% — above average
90%
Career Allowance Rate
44 granted / 49 resolved
+27.8% vs TC avg
Moderate +13% lift
Without
With
+13.2%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
6 currently pending
Career history
63
Total Applications
across all art units

Statute-Specific Performance

§101
4.7%
-35.3% vs TC avg
§103
45.3%
+5.3% vs TC avg
§102
44.7%
+4.7% vs TC avg
§112
3.3%
-36.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 49 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55. Information Disclosure Statement The information disclosure statement (IDS) submitted on 07/16/2024 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Status of Claim(s) Claim(s) 1-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Vora (US 20220383640 A1) in view of Liang (NPL: Compressing the Multiobject Tracking Model via Knowledge Distillation). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Vora (US 20220383640 A1) in view of Liang (NPL: Compressing the Multiobject Tracking Model via Knowledge Distillation) Regarding Claim 1: Vora teaches: An apparatus, comprising (Paragraph 6, “These and other aspects, features, and implementations can be expressed as methods, apparatuses, systems, components, program products, means or steps for performing a function, and in other ways.”): a memory configured to store program instructions (Paragraph 66, “As illustrated, device 300 includes processor 304, memory 306, storage component 308, input interface 310, output interface 312, communication interface 314, and bus 302.”); and Vora does not explicitly teach the following; however, in related Liang teaches: a processor configured to execute the program instructions to implement a first deep learning model and a second deep learning model (Page 2718, 2nd column, “We conduct the training and evaluation on one Intel 10980xe CPU and four Nvidia RTX 3090 GPUs); wherein the first deep learning model configured to (Figure 2. Top half of deep learning model consists of a teacher model and bottom half has a student model) obtain a first spatial feature map by learning an image, and obtain a first heatmap including center information of an object belonging to the image by learning the image (Figure 4. Image a denotes a first map that has both spatial and heatmap features wherein a center is identified by areas in red by the teacher model); and wherein the second deep learning model configured to obtain a second spatial feature map by learning the image (Figure 2. Bottom half has the second deep learning model being a “student model” that generates a feature map and spatial map), perform learning such that the second spatial feature map imitates the first spatial feature map (Figure 4. Image c/d denotes another map that has both spatial and heatmap features wherein a center is identified by areas in red by a student model), obtain a second heatmap including a center of the object (Figure 4. Image b-d Student model maps show heatmap with red areas that denote center of an object), and perform learning such that the second heatmap imitates the first heatmap (Figure 4 paragraph, “Visualization of the student’s feature map distilled with different attention mechanisms. The student’s feature map distilled with our solution (spatial attention + spatial and channel difference attention) is more similar to the teacher’s than other methods.”). Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Vora’s objection detection system that generates heatmaps for 3D spaces with Liang’s multi-object tracking system for identifying difference between centers of objects between a first and second neural network model. Regarding Claim 2: Vora and Liang teach the embodiments of claim 1. Liang further teaches: wherein the second deep learning model is configured to: output the second spatial feature map through a backbone (Page 2716, Table 1. “COMPARISONS OF TEACHER MODEL AND STUDENT MODEL. ALL PARTS (BACKBONE, NECK, AND HEADS) IN THE MODEL ARE CONSIDERED”) Regarding Claim 3: Vora and Liang teach the embodiments of claim 2. Liang further teaches: wherein: the first deep learning model is configured to obtain a first representative feature value of feature values arranged in a straight direction in the first spatial feature map (Figure 3. Output map of features in a straight line. Applied to both the teacher model and student models), and obtain a first feature vector of a single row by sorting the first representative feature value (Page 2716, Column 2, Both teacher and student model features are put into a distribution vector); and the second deep learning model is configured to obtain a second representative feature value of feature values arranged in a straight direction in the second spatial feature map (Figure 3. Output map of features in a straight line. Applied to both the teacher model and student models), obtain a second feature vector of a single row by sorting the second representative feature value (Page 2716, Column 2, Both teacher and student model features are put into a distribution vector), and perform imitation learning to reduce a difference between the second feature vector and the first feature vector (Page 2716-2717, B. Attention-Guided Feature Distillation, Differences between teacher and student attention maps are identified between spatial and channel masks for both). Regarding Claim 4: Vora and Liang teach the embodiments of claim 3. Liang further teaches: the first deep learning model is configured to obtain a first height feature value based on the feature values having x-axis coordinate values same as each other (Page 2716, Column 1, Feature maps of both teacher and student models have height values set to the same as one another), obtain a first height feature vector based on the first height feature value (Page 2716, Column 1, backbone network outputs a feature map with a channel, height, and width value), obtain a first width feature value based on the feature values having y-axis coordinate values same as each other (Page 2716, Column 1, backbone network outputs a feature map with a channel, height, and width value), and obtain a first width feature vector based on the first width feature value (Page 2716, Column 1, backbone network outputs a feature map with a channel, height, and width value); and the second deep learning model is configured to obtain a second height feature value based on the feature values having the x-axis coordinate values same as each other (Page 2716, Column 1, Feature maps of both teacher and student models have height values set to the same as one another), obtain a second height feature vector based on the second height feature value (Page 2716, Column 1, backbone network outputs a feature map with a channel, height, and width value), obtain a second width feature value based on the feature values having the y-axis coordinate values same as each other (Page 2716, Column 1, backbone network outputs a feature map with a channel, height, and width value), and obtain a second width feature vector based on the second width feature value (Page 2716, Column 1, backbone network outputs a feature map with a channel, height, and width value). Regarding Claim 5: Vora and Liang teach the embodiments of claim 4. Liang further teaches: perform learning to reduce a difference between the second height feature vector and the first height feature vector (Page 2717, Column 1, spatial and channel attention maps are calculated with width and height dimensions where a difference mask is then identified including the height and width); and perform learning to reduce a difference between the second width feature vector and the first width feature vector (Page 2717, Column 1, spatial and channel attention maps are calculated with width and height dimensions where a difference mask is then identified including the height and width). Regarding Claim 6: Vora and Liang teach the embodiments of claim 5. Liang further teaches: the first deep learning model is configured to obtain a first height distribution by normalizing the first height feature vector (Page 2716, “[0,1]1×(H/r)×(W/r),where the local peaks are regarded as the object centers and the rest are the background. The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”), and obtain a first width distribution by normalizing the first width feature vector (Page 2716, “[0,1]1×(H/r)×(W/r),where the local peaks are regarded as the object centers and the rest are the background. The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”); and the second deep learning model is configured to obtain a second height distribution by normalizing the second height feature vector (Page 2716, “[0,1]1×(H/r)×(W/r),where the local peaks are regarded as the object centers and the rest are the background. The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”), obtain a second width distribution by normalizing the second width feature vector (Page 2716, “[0,1]1×(H/r)×(W/r),where the local peaks are regarded as the object centers and the rest are the background. The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”), and perform learning such that the second height distribution imitates the first height distribution and the second width distribution imitates the first width distribution (Page 2716, “[0,1]1×(H/r)×(W/r),where the local peaks are regarded as the object centers and the rest are the background. The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”). Regarding Claim 7: Vora and Liang teach the embodiments of claim 1. Liang further teaches: wherein the second deep learning model is configured to: output the second heatmap through a head (Page 2716, “We list the comparisons of teacher and student models in Table I. Given an input image I ∈R3×H×W, the backbone network outputs a feature map F∈RC×(H/r)×(W/r), where C, H, and W is the channel, height, and width, respectively, and r is the down-sample stride; here we set C=64, r=4 for both teacher and student models. Upon the backbone, there are four separate heads for classification, box size regression, center offset, and ReID. In particular, the classification head predicts a heatmap ˆH∈ [0,1]1×(H/r)×(W/r), where the local peaks are regarded as the object centers and the rest are the background.”). Regarding Claim 8: Vora and Liang teach the embodiments of claim 7. Vora further teaches: the first deep learning model is configured to obtain a first representative center value of center values arranged in a straight direction in the first heatmap (Paragraph 169-170, “where g and q are trained together with a main network (e.g., network 1000 of FIG. 10), and during inference w.sub.k′ and b.sub.k′ are fixed for each location p.sub.k so it does not need extra runtime for g and q. In some embodiments, feature undistortion is applied in center heatmap prediction. Range stratified convolution and normalization 1030 is applied to a center offset. In some cases, the center offset of polar pillars is dependent on a range and azimuth of data points within the pillar. Accordingly, the pillar has different statistics (e.g., center offset, orientation, etc.) at different regions. For example, suppose a heatmap center is at (r.sub.c,θ.sub.c) and the target is at (r.sub.t,θ.sub.t).”), and obtain a first center vector of a single row by sorting the first representative center value (Paragraph 136, “That is, for each feature vector of size C in the dense (C, P) tensor, the image creating component 714 looks up the 2D coordinates of the feature vector using the P dimensional pillar index vector 709, and places the feature vector into the pseudo-image 716 at the 2D coordinates. As a result, each location on the pseudo-image 716 corresponds to one of the pillars and represents features of the data points in the pillar.”); and the second deep learning model is configured to obtain a second representative center value of center values arranged in a straight direction in the second heatmap (Paragraph 169-170, “where g and q are trained together with a main network (e.g., network 1000 of FIG. 10), and during inference w.sub.k′ and b.sub.k′ are fixed for each location p.sub.k so it does not need extra runtime for g and q. In some embodiments, feature undistortion is applied in center heatmap prediction. Range stratified convolution and normalization 1030 is applied to a center offset. In some cases, the center offset of polar pillars is dependent on a range and azimuth of data points within the pillar. Accordingly, the pillar has different statistics (e.g., center offset, orientation, etc.) at different regions. For example, suppose a heatmap center is at (r.sub.c,θ.sub.c) and the target is at (r.sub.t,θ.sub.t).”), obtain a second center vector of a single row by sorting the second representative center value (Paragraph 136, “That is, for each feature vector of size C in the dense (C, P) tensor, the image creating component 714 looks up the 2D coordinates of the feature vector using the P dimensional pillar index vector 709, and places the feature vector into the pseudo-image 716 at the 2D coordinates. As a result, each location on the pseudo-image 716 corresponds to one of the pillars and represents features of the data points in the pillar.”), and Vora does not explicitly teach the following; however, in related Liang teaches: perform imitation learning to reduce a difference between the second center vector and the first center vector (Page 2717, Column 1, calculation of difference masks between different masks relating to height, width, and centers). Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Vora’s objection detection system that generates heatmaps for 3D spaces with Liang’s multi-object tracking system for identifying difference between centers of objects between a first and second neural network model. Regarding Claim 9: Vora and Liang teach the embodiments of claim 8. Liang teaches: the first deep learning model is configured to obtain a first height center value based on the center values having x-axis coordinate values same as each other (Page 2716, “[0,1]1×(H/r)×(W/r),where the local peaks are regarded as the object centers and the rest are the background. The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”), obtain a first height center vector based on the first height center value (Page 2716, “[0,1]1×(H/r)×(W/r),where the local peaks are regarded as the object centers and the rest are the background. The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”), obtain a first width center value based on the center values having y-axis coordinate values same as each other (Page 2716, “[0,1]1×(H/r)×(W/r),where the local peaks are regarded as the object centers and the rest are the background. The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”),, and obtain a first width center vector based on the first width center value (Page 2716, “[0,1]1×(H/r)×(W/r),where the local peaks are regarded as the object centers and the rest are the background. The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”); and the second deep learning model is configured to obtain a second height center value based on the center values having the x-axis coordinate values same as each other (Page 2716, “[0,1]1×(H/r)×(W/r),where the local peaks are regarded as the object centers and the rest are the background. The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”),, obtain a second height center vector based on the second height center value (Page 2716, “[0,1]1×(H/r)×(W/r),where the local peaks are regarded as the object centers and the rest are the background. The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”), obtain a second width center value based on the center values having the y-axis coordinate values same as each other (Page 2716, “[0,1]1×(H/r)×(W/r),where the local peaks are regarded as the object centers and the rest are the background. The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”), and obtain a second width center vector based on the second width center value (Page 2716, “[0,1]1×(H/r)×(W/r),where the local peaks are regarded as the object centers and the rest are the background. The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”). Regarding Claim 10: Vora and Liang teach the embodiments of claim 9. Liang further teaches: perform learning to reduce a difference between the second height center vector and the first height center vector (Page 2716, “The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”); and perform learning to reduce a difference between the second width center vector and the first width center vector (Page 2716, As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆH and input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”). Regarding Claim 11: Vora and Liang teach the embodiments of claim 10. Liang teaches: the first deep learning model is configured to obtain a first height distribution by normalizing the first height center vector (Page 2716, “The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”), and obtain a first width distribution by normalizing the first width center vector (Page 2716, “The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”); and the second deep learning model is configured to obtain a second height distribution by normalizing the second height center vector (Page 2716, “The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”), obtain a second width distribution by normalizing the second width center vector (Page 2716, “The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”), and perform learning such that the second height distribution imitates the first height distribution and the second width distribution imitates the first width distribution (Page 2717, Column 1, calculation of difference masks between different masks relating to height, width, and centers). Regarding Claim 12: Vora teaches: A method comprising (Paragraph 6, “These and other aspects, features, and implementations can be expressed as methods, apparatuses, systems, components, program products, means or steps for performing a function, and in other ways.”): Vora does not explicitly teach the following; however, in related Liang teaches: obtaining a first spatial feature map and a first heatmap including center information of an object belonging to an image by learning the image based on a first deep learning model (Figure 4. Image a denotes a first map that has both spatial and heatmap features wherein a center is identified by areas in red by the teacher model); obtaining a second spatial feature map by learning the image based on a second deep learning model (Figure 2. Bottom half has the second deep learning model being a “student model” that generates a feature map and spatial map); performing learning of the second deep learning model such that the second spatial feature map imitates the first spatial feature map (Figure 4. Image c/d denotes a another map that has both spatial and heatmap features wherein a center is identified by areas in red by a student model); obtaining a second heatmap including a center of the object based on the second deep learning model (Figure 4. Image b-d Student model maps show heatmap with red areas that denote center of an object); and performing learning of the second deep learning model such that the second heatmap imitates the first heatmap (Figure 4 paragraph, “Visualization of the student’s feature map distilled with different attention mechanisms. The student’s feature map distilled with our solution (spatial attention + spatial and channel difference attention) is more similar to the teacher’s than other methods.”). Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Vora’s objection detection system that generates heatmaps for 3D spaces with Liang’s multi-object tracking system for identifying difference between centers of objects between a first and second neural network model. Regarding Claim 13: Vora and Liang teach the embodiments of claim 12. Liang further teaches: wherein performing learning of the second deep learning model such that the second spatial feature map imitates the first spatial feature map includes: obtaining a first representative feature value of feature values arranged in a straight direction in the first spatial feature map (Figure 3. Output map of features in a straight line. Applied to both the teacher model and student models); obtaining a first feature vector of a single row by sorting the first representative feature value (Page 2716, Column 2, Both teacher and student model features are put into a distribution vector); obtaining a second representative feature value of feature values arranged in a straight direction in the second spatial feature map (Figure 3. Output map of features in a straight line. Applied to both the teacher model and student models); obtaining a second feature vector of a single row by sorting the second representative feature value (Page 2716, Column 2, Both teacher and student model features are put into a distribution vector); and learning the second deep learning model to reduce a difference between the second feature vector and the first feature vector (Page 2716-2717, B. Attention-Guided Feature Distillation, Differences between teacher and student attention maps are identified between spatial and channel masks for both). Regarding Claim 14: Vora and Liang teach the embodiments of claim 13. Liang further teaches: obtaining the first feature vector includes obtaining a first height feature value based on the feature values having x-axis coordinate values same as each other (Page 2716, Column 1, Feature maps of both teacher and student models have height values set to the same as one another), obtaining a first height feature vector based on the first height feature value (Page 2716, Column 1, backbone network outputs a feature map with a channel, height, and width value), obtaining a first width feature value based on the feature values having y-axis coordinate values same as each other (Page 2716, Column 1, backbone network outputs a feature map with a channel, height, and width value), and obtaining a first width feature vector based on the first width feature value (Page 2716, Column 1, backbone network outputs a feature map with a channel, height, and width value); and obtaining of the second feature vector includes obtaining a second height feature value based on the feature values having the x-axis coordinate values same as each other (Page 2716, Column 1, Feature maps of both teacher and student models have height values set to the same as one another), obtaining a second height feature vector based on the second height feature value (Page 2716, Column 1, backbone network outputs a feature map with a channel, height, and width value), obtaining a second width feature value based on the feature values having the y-axis coordinate values same as each other (Page 2716, Column 1, backbone network outputs a feature map with a channel, height, and width value), and obtaining a second width feature vector based on the second width feature value (Page 2716, Column 1, backbone network outputs a feature map with a channel, height, and width value). Regarding Claim 15: Vora and Liang teach the embodiments of claim 14. Liang further teaches: performing learning to reduce a difference between the second height feature vector and the first height feature vector (Page 2716, “The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”); and performing learning to reduce a difference between the second width feature vector and the first width feature vector (Page 2716, As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆH and input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”). Regarding Claim 16: Vora and Liang teach the embodiments of claim 15. Liang further teaches: obtaining a first height distribution by normalizing the first height feature vector (Page 2716, “The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”); obtaining a first width distribution by normalizing the first width feature vector (Page 2716, “The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”); obtaining a second height distribution by normalizing the second height feature vector (Page 2716, “The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”); obtaining a second width distribution by normalizing the second width feature vector (Page 2716, “The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”); and performing learning of the second deep learning model such that the second height distribution imitates the first height distribution and the second width distribution imitates the first width distribution (Page 2717, Column 1, calculation of difference masks between different masks relating to height, width, and centers). Regarding Claim 17: Vora and Liang teach the embodiments of claim 12. Liang further teaches: obtaining, by the first deep learning model, a first representative center value of center values arranged in a straight direction in the first heatmap (Page 2716, “[0,1]1×(H/r)×(W/r),where the local peaks are regarded as the object centers and the rest are the background. The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”); obtaining, by the first deep learning model, a first center vector of a single row by sorting the first representative center value (Page 2716, “[0,1]1×(H/r)×(W/r),where the local peaks are regarded as the object centers and the rest are the background. The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”); obtaining, by the second deep learning model, a second representative center value of center values arranged in a straight direction in the second heatmap (Page 2716, “[0,1]1×(H/r)×(W/r),where the local peaks are regarded as the object centers and the rest are the background. The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”); obtaining, by the second deep learning model, a second center vector of a single row by sorting the second representative center value (Page 2716, “[0,1]1×(H/r)×(W/r),where the local peaks are regarded as the object centers and the rest are the background. The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”); and performing, by the second deep learning model, imitation learning to reduce a difference between the second center vector and the first center vector (Page 2716, “[0,1]1×(H/r)×(W/r),where the local peaks are regarded as the object centers and the rest are the background. The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”). Regarding Claim 18: Vora and Liang teach the embodiments of claim 17. Liang further teaches: obtaining, by the first deep learning model, a first height center value based on the center values having x-axis coordinate values same as each other (Page 2716, “[0,1]1×(H/r)×(W/r),where the local peaks are regarded as the object centers and the rest are the background. The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”); obtaining, by the first deep learning model, a first height center vector based on the first height center value (Page 2716, “[0,1]1×(H/r)×(W/r),where the local peaks are regarded as the object centers and the rest are the background. The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”); obtaining, by the first deep learning model, a first width center value based on the center values having y-axis coordinate values same as each other (Page 2716, “[0,1]1×(H/r)×(W/r),where the local peaks are regarded as the object centers and the rest are the background. The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”), and obtaining a first width center vector based on the first width center value (Page 2716, “[0,1]1×(H/r)×(W/r),where the local peaks are regarded as the object centers and the rest are the background. The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”); obtaining, by the second deep learning model, a second height center value based on the center values having the x-axis coordinate values same as each other (Page 2716, “[0,1]1×(H/r)×(W/r),where the local peaks are regarded as the object centers and the rest are the background. The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”); obtaining, by the second deep learning model, a second height center vector based on the second height center value (Page 2716, “[0,1]1×(H/r)×(W/r),where the local peaks are regarded as the object centers and the rest are the background. The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”); obtaining, by the second deep learning model, a second width center value based on the center values having the y-axis coordinate values same as each other (Page 2716, “[0,1]1×(H/r)×(W/r),where the local peaks are regarded as the object centers and the rest are the background. The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”); and obtaining, by the second deep learning model, a second width center vector based on the second width center value (Page 2716, “[0,1]1×(H/r)×(W/r),where the local peaks are regarded as the object centers and the rest are the background. The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”). Regarding Claim 19: Vora and Liang teach the embodiments of claim 18. Liang further teaches: learning the second deep learning model to reduce a difference between the second height center vector and the first height center vector (Page 2716, “The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”); and learning the second deep learning model to reduce a difference between the second width center vector and the first width center vector (Page 2716, “The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”). Regarding Claim 20: Vora and Liang teach the embodiments of claim 19. Liang further teaches: obtaining, by the first deep learning model, a first height distribution by normalizing the first height center vector (Page 2716, “The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”); obtaining, by the first deep learning model, a first width distribution by normalizing the first width center vector (Page 2716, “The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”); obtaining, by the second deep learning model, a second height distribution by normalizing the second height center vector (Page 2716, “The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”); obtaining, by the second deep learning model, a second width distribution by normalizing the second width center vector (Page 2716, “The box size head generates a size map ˆ S∈R2×(H/r)×(W/r), which regresses the height and width of objects from the estimated centers in ˆH.As the r-stride downsample can introduce discretization errors between coordinates in heatmap ˆHand input image I, we use the center offset head to predict a continuous offset ˆO∈R2×(H/r)×(W/r) for each estimated center, therefore localizing objects more precisely.”); and performing, by the second deep learning model, learning such that the second height distribution imitates the first height distribution and the second width distribution imitates the first width distribution (Page 2717, Column 1, calculation of difference masks between different masks relating to height, width, and centers). Relevant Prior Art Directed to State of Art Birchfield (US 20220277472 A1) Shen (US 20220222477 A1) Akbas (US 11521373 B1) Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to DYLAN J SHERRILLO whose telephone number is (703)756-5605. The examiner can normally be reached 1st week of bi-week: Mon-Wed 7am-5:30pm PST, Thurs: 7am-4:30pm PST, Fri off / 2nd week of bi-week: Mon-Wed 7am-5:30pm PST, Thurs-Fri: 7am-4:30pm PST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Stephen R Koziol can be reached at (408) 918-7630. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /D.J.S./Examiner, Art Unit 2665 /Stephen R Koziol/Supervisory Patent Examiner, Art Unit 2665
Read full office action

Prosecution Timeline

Jul 16, 2024
Application Filed
Sep 10, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749171
IMAGE ANALYSIS-BASED BUILDING INSPECTION
2y 11m to grant Granted Sep 29, 2026
Patent 12718348
MONITORING SYSTEM, MONITORING METHOD, PROGRAM, AND COMPUTER-READABLE RECORDING MEDIUM IN WHICH COMPUTER PROGRAM IS STORED
2y 11m to grant Granted Aug 25, 2026
Patent 12705747
METHOD AND SYSTEM FOR DETERMINING THE BPE IN A CONTRAST MEDIUM-ENHANCED X-RAY EXAMINATION OF A BREAST
2y 11m to grant Granted Aug 11, 2026
Patent 12694531
DATA PROCESSING APPARATUS, DATA PROCESSING METHOD, AND DATA PROCESSING SYSTEM
2y 9m to grant Granted Jul 28, 2026
Patent 12688666
Systems and Methods for Entropy-Based Treatment
3y 10m to grant Granted Jul 21, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
90%
Grant Probability
99%
With Interview (+13.2%)
2y 10m (~7m remaining)
Median Time to Grant
Low
PTA Risk
Based on 49 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month