Prosecution Insights
Last updated: October 02, 2026
Application No. 18/936,477

SCENE RECONSTRUCTION IN THREE-DIMENSIONS FROM TWO-DIMENSIONAL IMAGES

Non-Final OA §101§103
Filed
Nov 04, 2024
Priority
Jun 17, 2019 — provisional 62/862,139 +2 more
Examiner
KOPPOLU, VAISALI RAO
Art Unit
Tech Center
Assignee
Snap Inc.
OA Round
1 (Non-Final)
79%
Grant Probability
Favorable
1-2
OA Rounds
10m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 79% — above average
79%
Career Allowance Rate
107 granted / 135 resolved
+19.3% vs TC avg
Strong +27% interview lift
Without
With
+26.6%
Interview Lift
resolved cases with interview
Typical timeline
2y 9m
Avg Prosecution
22 currently pending
Career history
146
Total Applications
across all art units

Statute-Specific Performance

§101
10.1%
-29.9% vs TC avg
§103
55.8%
+15.8% vs TC avg
§102
13.5%
-26.5% vs TC avg
§112
20.2%
-19.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 135 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) submitted is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claim 20 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claim does/do not fall within at least one of the four categories of patent eligible subject matter because the claim is directed to a “computer readable medium”. However, the claim is not limited to non-transitory embodiments, and the specification does not provide a definition limiting the meaning of this term to only non-transitory embodiments (see [0127]). The specification recites that according to one embodiment computer readable medium may include removable and non-removable storage devices including, but not limited to, Read Only Memory (ROM), Random Access Memory (RAM), compact discs (CDs), digital versatile discs (DVD), etc. but does not exclude transitory signals. The claim therefore can be reasonably interpreted as encompassing transitory signal embodiments, which are non-statutory (In re Nuijten, 500 F.3d 1346, 84 USPQ2d 1495 (Fed. Cir. 2007)). If the specification includes written description support, this rejection can be overcome by including the term “non-transitory” in the claim (see USPTO Official Gazette notice 1351 OG 212.). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 1, 9 – 10, 14 – 20 are rejected under 35 U.S.C. 103 as being unpatentable Gong et al. (US 20220318621 A1; hereafter referred to as Gong) in view of Li et al. (Li, D., Chen, X., Zhang, Z., & Huang, K. (2017, July). Learning deep context-aware features over body and latent parts for person re-identification. In 2017 IEEE conference on computer vision and pattern recognition (CVPR) (pp. 7398-7407). IEEE; hereafter referred to as Li). Regarding Claim 1, Gong teaches: A method comprising: extracting object/person re-identification (REID) embeddings from object/human crops by a teacher-student network (Gong, [0097] FIG. 3 provides an overview of this knowledge distillation teacher model construction”; Gong, [0115] “the Re-ID features may be extracted via the CNN network, where ns is a pre-defined number of the gallery candidates”; Gong, [0146] “The final FC layer output feature vector (2,048-D) is extracted as the re-id feature vector in the present model by resizing all of the training images as 256×128”; Gong, [0162] “The teacher model provides a constant uniform target distribution. For the offline competitor KD, we used a large network ResNet-110 as the teacher and a small network ResNet-32 as the student”; Gong, [0143] “CUHK01 is a remarkable small-scale re-id dataset, which consists of 971 identities from two camera views, where each identity has two images per camera view and thus includes 3884 images which are manually cropped”); using the REID embeddings as a supervision signal to train a REID branch for a neural network (Gong, [0101] Knowledge transfer may be attempted between varying-capacity network models…Hinton et al. [28] distilled knowledge from a large pre-trained teacher model to improve a small target net. The rationale behind this is in taking advantage of extra supervision provided by the teacher model during training the target model, beyond a conventional supervised learning objective such as the cross-entropy loss subject to the training data labels. Extra supervision may be extracted from a pre-trained powerful teacher model in form of class posterior probabilities, feature representations”); and processing the REID embeddings generated by the teacher-student network to track object/person identities across multiple images in a scene (Gong, [0081] The following examples describe image and video data sets where individual people with such images are targets. The aim is to identify the same people in different locations obtained by separate video and image feeds”; Gong, [0143] “The Market-1501 is a widely adopted large-scale re-id dataset that contains 1,501 identities obtained by Deformable Part Model pedestrian detector. It includes 32,668 images obtain from 6 non-overlapping camera views on a campus. CUHK01 is a remarkable small-scale re-id dataset, which consists of 971 identities from two camera views, where each identity has two images per camera view and thus includes 3884 images which are manually cropped. Duke is one of the most popular large scale re-id dataset which consists 36411 pedestrian images captured from 8 different camera views. Among them, 16522 images (702 identities) are adopted for training, 2228 (702 identities) images are taken as query to be retrieved from the remaining 17661 images”). However, Gong does not explicitly recite: processing the REID embeddings generated by the teacher-student network to track object/person identities across multiple images in a scene having object overlaps and occlusions. In the same field of endeavor, Li teaches: processing the REID embeddings generated by the teacher-student network to track object/person identities across multiple images in a scene having object overlaps and occlusions (Li, page 7402, 4.1 Datasets and Protocols – page 7405, col. 1, para 1, Li teaches that the proposed model learns features of persons across multiple images and can detect persons with occlusions and overlap). Gong and Li are considered analogous art as they are reasonably pertinent to the same field of endeavor of computer vision and image data. Therefore, it would have been obvious to one of the ordinary skill the art before the effective filing date of the claimed invention to modify the invention of Gong with the method of processing embeddings as taught by Li to make the invention that track object identities across multiple images in a scene having object overlaps and occlusions; doing so can result in effectively and accurately identifying the person in a scene with overlaps and occlusions (Li , 7405, Effectiveness of localization loss); thus one of the ordinary skill in the art would have been motivated to combine the reference. Regarding Claim 9, Gong in view of Li teaches the method of claim 1, wherein the teacher-student network is trained using a pre-trained REID network to generate high-dimensional embeddings (Gong, [0105] “a generic deep Convolutional Neural Network (CNN) architecture may be provided as the base network with ImageNet pre-training, e.g. either Resnet-50 [26] or ResNet-110 [26]. It may be straightforward to apply any other network architectures as alternatives. To effectively learn the ID discriminative feature embedding”; Li, 3.3 Feature Extraction and Fusion, “For the body-based representation, we use MSCAN to extract the global feature maps and then learn a 128-dimension feature embedding”; Li, 4.2 Implementation details, optimization, “network training, we initialize the network using pretrained body-based and part-based model”). Regarding Claim 10, Gong in view of Li teaches the method of claim 1, further comprising using the REID embeddings to maintain object/person identity across a sequence of video frames (Gong, [0081] The following examples describe image and video data sets where individual people with such images are targets. The aim is to identify the same people in different locations obtained by separate video and image feeds”; Gong, [0098] “A person Re-ID task may be used to search for the same people among multiple camera views”). Regarding Claim 14, Gong in view of Li teaches the method of claim 1, further comprising using the REID embeddings to enhance accuracy of object/person-specific parameter smoothing over time (Gong; [0149] accuracy for different date sets have shown improved performance; Li, page 7402, 4.1 datasets, “we use 1,260 person identities for training and the rest 100 identities for testing. Experiments are conducted 20 times and the mean result is reported. MARS: It is the largest sequence-based person ReID dataset”; Li, page 7403, 4.3. “Compared with metric learning methods, such as the state-of-the-art approach DNS, the proposed fusion model improves the Rank-1 identification rate by 11.66% and 13.29% on the labeled and detected datasets respectively. Compared with the similar multi-class person identification network DGD, the Rank-1 identification rate improves by 1.63% using our fusion model on the labeled dataset”). Regarding Claim 15, Gong in view of Li teaches the method of claim 1, wherein the REID embeddings are used to generate a discriminative signature for each object/person, invariant to changes in camera position (Gong, [0143] invariant to camera positions), (Li, 7400, 3. Proposed method, “a multiscale context-aware network for efficient feature learning (Section 3.1), the latent parts learning and localization for better local part-based feature representation (Section 3.2), the fusion of global full-body and local body-part features for person ReID (Section 3.3)”; Li, page 7399, Lol. 2, “a multi-scale context-aware network to enhance the visual context information for better feature representation of fine-grained visual cues. (b) Instead of using rigid parts, we propose to learn and localize pedestrian parts using spatial transformer networks with novel prior spatial constraints”; Li, page 7405, 5. Conclusions, “the fusion of full-body and body-part identity discriminative features for powerful pedestrian representation). Regarding Claim 16, Gong in view of Li teaches the method of claim 1, further comprising employing a memory-based online learning approach to update the REID embeddings as new video frames are processed (Gong, [0030] “updating the model parameters of the initial reinforcement learning model to form an updated reinforcement learning model using the further new training data set”; Gong, [0094] “A reinforcement learning policy enables active selection of new training data from a large pool of un-labelled test data using human feedback. A Convolutional Neural Network (CNN) model introduces both active learning (AL) and reinforcement learning (RL) in a single human-in-the-loop model learning framework”). Regarding Claim 17, Gong teaches: A system comprising: a memory (Gong, [0069] “The computer system may include a memory including volatile and non-volatile storage medium”); and at least one processor, wherein the at least one processor is configured to perform operations (Gong, [0069] The computer system may include a processor or processors (e.g. local, virtual or cloud-based) such as a Central Processing unit (CPU), and/or a single or a collection of Graphics Processing Units (GPUs). The processor may execute logic in the form of a software program”) comprising: extracting object/person re-identification (REID) embeddings from object/human crops by a teacher-student network (Gong, [0097] FIG. 3 provides an overview of this knowledge distillation teacher model construction”; Gong, [0115] “the Re-ID features may be extracted via the CNN network, where ns is a pre-defined number of the gallery candidates”; Gong, [0146] “The final FC layer output feature vector (2,048-D) is extracted as the re-id feature vector in the present model by resizing all of the training images as 256×128”; Gong, [0162] “The teacher model provides a constant uniform target distribution. For the offline competitor KD, we used a large network ResNet-110 as the teacher and a small network ResNet-32 as the student”; Gong, [0143] “CUHK01 is a remarkable small-scale re-id dataset, which consists of 971 identities from two camera views, where each identity has two images per camera view and thus includes 3884 images which are manually cropped”); using the REID embeddings as a supervision signal to train a REID branch for a neural network (Gong, [0101] Knowledge transfer may be attempted between varying-capacity network models…Hinton et al. [28] distilled knowledge from a large pre-trained teacher model to improve a small target net. The rationale behind this is in taking advantage of extra supervision provided by the teacher model during training the target model, beyond a conventional supervised learning objective such as the cross-entropy loss subject to the training data labels. Extra supervision may be extracted from a pre-trained powerful teacher model in form of class posterior probabilities, feature representations”); and processing the REID embeddings generated by the teacher-student network to track object/person identities across multiple images in a scene (Gong, [0081] The following examples describe image and video data sets where individual people with such images are targets. The aim is to identify the same people in different locations obtained by separate video and image feeds”; Gong, [0143] “The Market-1501 is a widely adopted large-scale re-id dataset that contains 1,501 identities obtained by Deformable Part Model pedestrian detector. It includes 32,668 images obtain from 6 non-overlapping camera views on a campus. CUHK01 is a remarkable small-scale re-id dataset, which consists of 971 identities from two camera views, where each identity has two images per camera view and thus includes 3884 images which are manually cropped. Duke is one of the most popular large scale re-id dataset which consists 36411 pedestrian images captured from 8 different camera views. Among them, 16522 images (702 identities) are adopted for training, 2228 (702 identities) images are taken as query to be retrieved from the remaining 17661 images.)”). However, Gong does not explicitly recite: processing the REID embeddings generated by the teacher-student network to track object/person identities across multiple images in a scene having object overlaps and occlusions. In the same field of endeavor, Li teaches: processing the REID embeddings generated by the teacher-student network to track object/person identities across multiple images in a scene having object overlaps and occlusions (Li, page 7402, 4.1 Datasets and Protocols – page 7405, col. 1, para 1, Li teaches that the proposed model learns features of persons across multiple images and can detect persons with occlusions and overlap). Gong and Li are considered analogous art as they are reasonably pertinent to the same field of endeavor of computer vision and image data. Therefore, it would have been obvious to one of the ordinary skill the art before the effective filing date of the claimed invention to modify the invention of Gong with the method of processing embeddings as taught by Li to make the invention that track object identities across multiple images in a scene having object overlaps and occlusions; doing so can result in effectively and accurately identifying the person in a scene with overlaps and occlusions (Li , 7405, Effectiveness of localization loss); thus one of the ordinary skill in the art would have been motivated to combine the reference. Regarding Claim 18, Gong in view of Li teaches the system of claim 17, further comprising employing a memory-based online learning approach to update the REID embeddings as new video frames are processed (Gong, [0030] “updating the model parameters of the initial reinforcement learning model to form an updated reinforcement learning model using the further new training data set”; Gong, [0094] “A reinforcement learning policy enables active selection of new training data from a large pool of un-labelled test data using human feedback. A Convolutional Neural Network (CNN) model introduces both active learning (AL) and reinforcement learning (RL) in a single human-in-the-loop model learning framework”). Regarding Claim 19, Gong in view of Li teaches the system of claim 17, wherein the REID embeddings are used to generate a discriminative signature for each object/person, invariant to changes in camera position (Gong, [0143] invariant to camera positions), (Li, 7400, 3. Proposed method, “a multiscale context-aware network for efficient feature learning (Section 3.1), the latent parts learning and localization for better local part-based feature representation (Section 3.2), the fusion of global full-body and local body-part features for person ReID (Section 3.3)”; Li, page 7399, Lol. 2, “a multi-scale context-aware network to enhance the visual context information for better feature representation of fine-grained visual cues. (b) Instead of using rigid parts, we propose to learn and localize pedestrian parts using spatial transformer networks with novel prior spatial constraints”; Li, page 7405, 5. Conclusions, “the fusion of full-body and body-part identity discriminative features for powerful pedestrian representation”). Regarding Claim 20, Gong teaches: A computer readable medium that stores a set of instructions that is executable by at least one processor to cause the at least one processor to perform operations (Gong, [0069] “The processor may execute logic in the form of a software program. The computer system may include a memory including volatile and non-volatile storage medium. A computer-readable medium may be included to store the logic or program instructions”) comprising: extracting object/person re-identification (REID) embeddings from object/human crops by a teacher-student network (Gong, [0097] FIG. 3 provides an overview of this knowledge distillation teacher model construction”; Gong, [0115] “the Re-ID features may be extracted via the CNN network, where ns is a pre-defined number of the gallery candidates”; Gong, [0146] “The final FC layer output feature vector (2,048-D) is extracted as the re-id feature vector in the present model by resizing all of the training images as 256×128”; Gong, [0162] “The teacher model provides a constant uniform target distribution. For the offline competitor KD, we used a large network ResNet-110 as the teacher and a small network ResNet-32 as the student”; Gong, [0143] “CUHK01 is a remarkable small-scale re-id dataset, which consists of 971 identities from two camera views, where each identity has two images per camera view and thus includes 3884 images which are manually cropped”); using the REID embeddings as a supervision signal to train a REID branch for a neural network (Gong, [0101] Knowledge transfer may be attempted between varying-capacity network models…Hinton et al. [28] distilled knowledge from a large pre-trained teacher model to improve a small target net. The rationale behind this is in taking advantage of extra supervision provided by the teacher model during training the target model, beyond a conventional supervised learning objective such as the cross-entropy loss subject to the training data labels. Extra supervision may be extracted from a pre-trained powerful teacher model in form of class posterior probabilities, feature representations”); and processing the REID embeddings generated by the teacher-student network to track object/person identities across multiple images in a scene (Gong, [0081] The following examples describe image and video data sets where individual people with such images are targets. The aim is to identify the same people in different locations obtained by separate video and image feeds”; Gong, [0143] “The Market-1501 is a widely adopted large-scale re-id dataset that contains 1,501 identities obtained by Deformable Part Model pedestrian detector. It includes 32,668 images obtain from 6 non-overlapping camera views on a campus. CUHK01 is a remarkable small-scale re-id dataset, which consists of 971 identities from two camera views, where each identity has two images per camera view and thus includes 3884 images which are manually cropped. Duke is one of the most popular large scale re-id dataset which consists 36411 pedestrian images captured from 8 different camera views. Among them, 16522 images (702 identities) are adopted for training, 2228 (702 identities) images are taken as query to be retrieved from the remaining 17661 images.)”). However, Gong does not explicitly recite: processing the REID embeddings generated by the teacher-student network to track object/person identities across multiple images in a scene having object overlaps and occlusions. In the same field of endeavor, Li teaches: processing the REID embeddings generated by the teacher-student network to track object/person identities across multiple images in a scene having object overlaps and occlusions (Li, page 7402, 4.1 Datasets and Protocols – page 7405, col. 1, para 1, Li teaches that the proposed model learns features of persons across multiple images and can detect persons with occlusions and overlap). Gong and Li are considered analogous art as they are reasonably pertinent to the same field of endeavor of computer vision and image data. Therefore, it would have been obvious to one of the ordinary skill the art before the effective filing date of the claimed invention to modify the invention of Gong with the method of processing embeddings as taught by Li to make the invention that track object identities across multiple images in a scene having object overlaps and occlusions; doing so can result in effectively and accurately identifying the person in a scene with overlaps and occlusions (Li , 7405, Effectiveness of localization loss); thus one of the ordinary skill in the art would have been motivated to combine the reference. Claims 11 and 13 are rejected under 35 U.S.C. 103 as being unpatentable Gong et al. (US 20220318621 A1; hereafter referred to as Gong) in view of Li et al. (Li, D., Chen, X., Zhang, Z., & Huang, K. (2017, July). Learning deep context-aware features over body and latent parts for person re-identification. In 2017 IEEE conference on computer vision and pattern recognition (CVPR) (pp. 7398-7407). IEEE; hereafter referred to as Li) further in view of Choi et al. (US 20180268292 A1; hereafter referred to as Choi). Regarding Claim 11, Gong in view of Li teaches the method of claim 1, wherein the REID branch of the neural network is fully convolutional (Gong, [0147] “baseline results are computed by directly employing the pre-trained CNN model, and the upper bound result indicates that the model is fine-tuned on the dataset with fully supervised training data”). However, Gong in view of Li does not explicitly recite: allowing for real-time processing independent of a number of objects in the scene. In the same field of endeavor, Choi teaches: allowing for real-time processing independent of a number of objects in the scene (Choi, [0061] “regarding teacher and student networks 110, 120, in the exemplary embodiments of the present invention, Faster R-CNN can be employed as the model for real-time object detection. The detection includes shared convolutional layers, a Region Proposal Network (RPN) and a Region Classification Network (RCN)”). Gong, Li and Choi are considered analogous art as they are reasonably pertinent to the same field of endeavor of computer vision and image data. Therefore, it would have been obvious to one of the ordinary skill the art before the effective filing date of the claimed invention to modify the invention of Gong in view of Li with the method of real-time processing as taught by Choi to make the invention that processes real-time objects in the scene; doing so can improve speed and accuracy of object detection (Choi , [0003] – [0005]); thus one of the ordinary skill in the art would have been motivated to combine the reference. Regarding Claim 13, Gong in view of Li teaches the method of claim 1, but does not explicitly recite: wherein the REID embeddings are used to associate virtual objects or skins with tracked objects/persons across video frames. In the same field of endeavor, Choi teaches: wherein the REID embeddings are used to associate virtual objects or skins with tracked objects/persons across video frames (Choi, [0054] “Regarding hint learning, distillation only transfers knowledge from the last layer. In conventional works, it has been indicated that employing the intermediate representation of the teacher model 110 as hints can improve the training process and final performance of the student model 120”; Choi, [0060] “learning from hints could help the student model 120 converge faster. An adaption layer to map from layer Ls in student network 120 to layer Lt in teacher network 110 can be employed”; Choi [0020] “A plurality of images 105 are input into the teacher model 110 and the student model 120. Hint learning module 130 can be employed to aid the student model 120. The teacher model 110 interacts with a detection module 112 and a prediction module 114, and the student model 120 interacts with a detection module 122 and a prediction module 124. Bounding box regression module 140 can also be used to adjust a location and a size of the bounding box. The prediction modules 114, 116 communicate with soft label module 150 and ground truth module 160”, student network can learn hints in the form of virtual objects and identify objects). Gong, Li and Choi are considered analogous art as they are reasonably pertinent to the same field of endeavor of computer vision and image data. Therefore, it would have been obvious to one of the ordinary skill the art before the effective filing date of the claimed invention to modify the invention of Gong in view of Li with the method of associating objects as taught by Choi to make the invention that associate virtual objects with tracked objects/persons across video frames; doing so can improve speed and accuracy of object detection (Choi , [0003] – [0005]); thus one of the ordinary skill in the art would have been motivated to combine the reference. Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable Gong et al. (US 20220318621 A1; hereafter referred to as Gong) in view of Li et al. (Li, D., Chen, X., Zhang, Z., & Huang, K. (2017, July). Learning deep context-aware features over body and latent parts for person re-identification. In 2017 IEEE conference on computer vision and pattern recognition (CVPR) (pp. 7398-7407). IEEE; hereafter referred to as Li) further in view of Gou et al. (Gou, M., Wu, Z., Rates-Borras, A., Camps, O., & Radke, R. J. (2018). A systematic evaluation and benchmark for person re-identification: Features, metrics, and datasets. IEEE transactions on pattern analysis and machine intelligence, 41(3), 523-536”; hereafter referred to as Gou). Regarding Claim 12, Gong in view of Li teaches the method of claim 1, but does not explicitly recite: further comprising integrating the REID embeddings with a tracking algorithm to handle occlusions and reappearances of objects/persons in video frames. In the same field of endeavor Gou teaches: further comprising integrating the REID embeddings with a tracking algorithm to handle occlusions and reappearances of objects/persons in video frames (Gou, page 525, 3. Datasets, “Based on difficult examples, we also annotate each dataset with challenging attributes from the following list: view point variations(VV),illumination variations(IV),detection errors(DE), occlusions(OCC)”; Gou, page 526, col. 2, “Each video clip was then run through a prototype end-to end re-id system comprised of automatic person detection and tracking algorithms… describe how this dataset can be used to validate detection and tracking algorithms typically used in an end-to-end re-id system”). Gong, Li and Gou are considered analogous art as they are reasonably pertinent to the same field of endeavor of computer vision and image data. Therefore, it would have been obvious to one of the ordinary skill the art before the effective filing date of the claimed invention to modify the invention of Gong in view of Li with the method of using a tracking algorithm as taught by Gou to make the invention that uses a tracking algorithm to handle occlusions and reappearances of objects/persons in video frames; doing so can efficiently evaluate datasets that closely mimic real-world problem setting in tracking objects and help in metric learning and ranking the re-id algorithms (Gou , abstract); thus one of the ordinary skill in the art would have been motivated to combine the reference. Allowable Subject Matter Claims 2 – 8 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Regarding Claim 2, the closest prior arts are Kuntsevich et al. (US 10282898 B1; hereafter referred to as Kuntsevich), Reisner-Kollmann et al. (US 20150062120 A1; hereafter referred to as Reisner) and Sinha et al. (US 20100315412 A1; hereafter referred to as Sinha). However, when looking at all the available prior arts none teaches the limitations: “measuring an error value based on comparing the three-dimensional plane with a two-dimensional plane shown in the single two-dimensional image; and adjusting the three-dimensional plane based on the error value”. Regarding Claims 3 – 8, they are objected to for being dependent on rejected base claim. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. US 20170124415 A1 SUBCATEGORY-AWARE CONVOLUTIONAL NEURAL NETWORKS FOR OBJECT DETECTION A computer-implemented method for detecting objects by using subcategory-aware convolutional neural networks (CNNs) is presented. The method includes generating object region proposals from an image by a region proposal network (RPN) which utilizes subcategory information, and classifying and refining the object region proposals by an object detection network (ODN) that simultaneously performs object category classification, subcategory classification, and bounding box regression. US 20200111250 A1 METHOD FOR RECONSTRUCTING THREE-DIMENSIONAL SPACE SCENE BASED ON PHOTOGRAPHING A method for reconstructing a three-dimensional space scene based on photographing is disclosed, comprising the following steps: S1, importing photos of all spaces, and making the photos correspond to a three-dimensional space according to directions and viewing angles during capture, so that a viewing direction of each pixel, when viewed from the camera position of the three-dimensional space, is in line with that during capture; S2, regarding a room as a set of multiple planes, determining a first plane, and then determining all the planes one by one according to relationships and intersections between the planes; S3, marking a spatial structure of the room by a marking system and obtaining dimension information; and S4, establishing a three-dimensional space model of the room by point coordinate information collected in the step S3. Contact Information Any inquiry concerning this communication or earlier communications from the examiner should be directed to VAISALI RAO KOPPOLU whose telephone number is (571)270-0273. The examiner can normally be reached Monday - Friday 8:30 - 5. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Mehmood can be reached at (571) 272-2976. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. VAISALI RAO. KOPPOLU Examiner Art Unit 2664 /VAISALI RAO KOPPOLU/Examiner of Art Unit 2664
Read full office action

Prosecution Timeline

Nov 04, 2024
Application Filed
Aug 11, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743894
COMPUTING APPARATUS AND METHOD FOR INSPECTING LEARNING DATA
2y 6m to grant Granted Sep 22, 2026
Patent 12731266
METHOD AND COMPUTING DEVICE FOR ENHANCED DEPTH SENSOR COVERAGE
3y 7m to grant Granted Sep 08, 2026
Patent 12731287
VEHICLE LOCATION CALCULATION APPARATUS AND VEHICLE LOCATION CALCULATION METHOD
2y 10m to grant Granted Sep 08, 2026
Patent 12731376
Methods for Automatically Generating a Training Dataset for Training an Optical Recognition Model for Reading Street Signs
2y 6m to grant Granted Sep 08, 2026
Patent 12725293
POSITIONING DEVICE, MOUNTING DEVICE, POSITIONING METHOD, AND METHOD FOR MANUFACTURING ELECTRONIC COMPONENT
2y 4m to grant Granted Sep 01, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
79%
Grant Probability
99%
With Interview (+26.6%)
2y 9m (~10m remaining)
Median Time to Grant
Low
PTA Risk
Based on 135 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month