Prosecution Insights
Last updated: August 16, 2026
Application No. 18/960,754

SUBJECT RE-IDENTIFICATION USING SEMANTIC ATTRIBUTE RECOGNITION

Non-Final OA §102§103
Filed
Nov 26, 2024
Examiner
KAUR, JASPREET
Art Unit
2662
Tech Center
2600 — Communications
Assignee
NVIDIA Corporation
OA Round
1 (Non-Final)
78%
Grant Probability
Favorable
1-2
OA Rounds
11m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 78% — above average
78%
Career Allowance Rate
18 granted / 23 resolved
+16.3% vs TC avg
Strong +42% interview lift
Without
With
+41.7%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
23 currently pending
Career history
54
Total Applications
across all art units

Statute-Specific Performance

§101
20.9%
-19.1% vs TC avg
§103
56.4%
+16.4% vs TC avg
§102
6.1%
-33.9% vs TC avg
§112
8.6%
-31.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 23 resolved cases

Office Action

§102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (“IDS”) filed on 02/26/2025 has been reviewed and the listed references have been considered. Drawings The 10-page drawings have been considered and placed on record in the file. Status of Claims Claims 1-20 are pending. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1 and 7-9 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Yiqiang et al. ("Deep and low-level feature based attribute learning for person re-identification" - Published 2018). Regarding claim 1, Yiqiang teaches “A method comprising: generating semantic data corresponding to appearance features of a subject within a first image (Yiqiang page 3 right hand column paragraph 1 "higher level discriminative features by several succeeding convolution and pooling operations that become specific to different body parts at a given stage (P3) in order to account for the possible displacements of pedestrians due to pose variations"); generating, using one or more models and the semantic data, attribute features of the subject (Yiqiang page 3 right hand column paragraph 1 "Another branch extracts the viewpoint-invariant Local Maximal Occurrence (LOMO) features, a robust visual feature representation that has been specifically designed for viewpoint-invariant pedestrain attribute recognition and achieving state-of-the-art results [23] (cf. Section 3.1.3)"); generating, using the one or more models, the semantic data, and the attribute features, an embedding (Yiqiang page 4 right hand column paragraph 3 "Person re-identification consists of matching images of the same individuals across multiple camera views. In order to achieve this, we need to learn a distance function that has large values for images from different people and small values for images from the same person" and page 5 right hand column paragraph 2 "From the re-identification data, the network learns informative features that distinguish individuals, and the semantic attributes that we want to recognise can be considered such as identify features at a higher level"); and identifying, using the one or more models and the embedding (Yiqiang page 3 left hand column paragraph 1 "In summary, two CNN embeddings are learned based on attribute and identity annotation. Then, an improved triplet loss is used to learn the fusion. We will experimentally show the performance improvement brought by this fusion achieving state-of-the-art results on a public person re-identification benchmark"), the subject within a second image (Yiqiang page 1 right hand column paragraph 2 "Person re-identification consists of matching a query person among a large set of people detected in multiple non-overlapping camera views").“ Regarding claim 7, Yiqiang teaches “The method of claim 1, wherein the attribute features are indicative of at least one intrinsic attribute of the subject and at least one extrinsic attribute of the subject (Yiqiang Table 5 and page 7 right hand column paragraph 4 "The results on the APiS dataset are shown in Table 5").” PNG media_image1.png 253 817 media_image1.png Greyscale Yiqiang Table 5 Regarding claim 8, Yiqiang teaches “The method of claim 7, wherein the at least one intrinsic attribute corresponds to at least one of an age, sex, hair style, or body shape of the subject (Yiqiang Table 5 and page 7 right hand column paragraph 4 "The results on the APiS dataset are shown in Table 5").” Regarding claim 9, Yiqiang teaches “The method of claim 7, wherein the at least one extrinsic attribute corresponds to at least one of a clothing article worn by the subject or an object attached to or carried by the subject (Yiqiang Table 5 and page 7 right hand column paragraph 4 "The results on the APiS dataset are shown in Table 5").” Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 2-3 and 6 are rejected under 35 U.S.C. 103 as being unpatentable over Yiqiang in view of Lin et al. ("Improving person re-identification by attribute and identity learning" - from IDS). Regarding claim 2, Yiqiang teaches “The method of claim 1, wherein the one or more models comprises a first network trained to generate training attribute features from training semantic data associated with a second subject in an initial image using a first loss function (Yiqiang page 4 left hand column paragraph 3 and right hand column paragraph 1 "have several properties at the same time, the attribute recognition is a multi-label classification problem. Thus, the multi-label version of the sigmoid cross entropy is used as the overall loss function"), a second network trained to generate a training embedding using the training semantic data and the training attribute features using a second loss function different from the first loss function (Yiqiang page 6 left hand column paragraph 3 "The pre-trained attribute CNN and identification CNN are combined and trained in a triplet architecture similar to the one explained in Section 3.1.4. Here, we propose to use an improved triplet loss with hard example selection to learn the optimal fusion of the two types of features. The fcl layer of the attribute network and the fcl layer of the identification network are normalised and concatenated, and another fully-connected layer which allows to merge attribute and identification features"), and However, Yiqiang is not relied on to teach “a third network trained to identify the second subject within a subsequent image using the training embedding and a third loss function different from the first and second loss functions”. Lin teaches “a third network trained to identify the second subject within a subsequent image using the training embedding and a third loss function different from the first and second loss functions (Lin page 6 left hand column paragraph 2 "Considering both attribute recognition and identity prediction, we define the overall objective function as followings where λ   is a hyper-parameter to balance the identity classification loss and the attribute recognition losses").” It would have been obvious to a person having ordinary skill in the art before effective filing date of the claimed invention of the instant application to combine a person recognition and tracking method as taught by Yiqiang to train a network for identification using loss function as taught by Lin. The suggestion/motivation for doing so would have been “By combining the attribute recognition task and identity classification task, the APR network is capable of learning more discriminative feature representations for pedestrians, including global and local descriptions. Specifically, we take attribute predictions as additional cues for the identity classification. Considering the dependencies among pedestrian attributes, we first re-weight the attribute predictions and then build identification upon these re-weighted attributes descriptions" as noted by the Lin on page 2 right hand column paragraph 2. Therefore, it would have been obvious to combine the disclosure of Yiqiang with the Lin disclosure to obtain the invention as specified in claim 2 as there is a reasonable expectation of success and/or because doing so merely combines prior art elements according to known methods to yield predictable results. Regarding claim 3, the combination of Yiqiang and Lin teaches “The method of claim 2, wherein at least one of the first, second, or third loss function is a triplet loss function (Yiqiang page 6 left hand column paragraph 3 "The pre-trained attribute CNN and identification CNN are combined and trained in a triplet architecture similar to the one explained in Section 3.1.4. Here, we propose to use an improved triplet loss with hard example selection to learn the optimal fusion of the two types of features") or a binary cross-entropy (BCE) loss function (Yiqiang page 4 left hand column paragraph 3 and right hand column paragraph 1 "have several properties at the same time, the attribute recognition is a multi-label classification problem. Thus, the multi-label version of the sigmoid cross entropy is used as the overall loss function").” Regarding claim 6, the combination of Yiqiang and Lin teaches “The method of claim 1, wherein the embedding generated by the one or more models is based at least on weights derived from relationships between the attribute features and the appearance features (Lin Figure 5 and page 5 left hand column paragraph 4 "network firstly computes attribute losses for the M individual attributes. Then the M prediction scores are concatenated and fed into an Attribute Re-weighting Module").” PNG media_image2.png 213 579 media_image2.png Greyscale Lin Figure 5 The proposed combination as well as the motivation for combining Yiqiang and Lin references presented in the rejection of claim 2, applies to claim 6. Finally the method recited in claim 6 is met by Yiqiang and Lin. Claims 4-5 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Yiqiang in view of Chen et al. ("Beyond appearance: a semantic controllable self-supervised learning framework for human-centric visual tasks" -from IDS). Regarding claim 4, Yiqiang teaches the method of claim 1. However, Yiqiang is not relied on to teach “wherein a particular model of the one or more models generates the semantic data using the first image”. Chen teaches “wherein a particular model of the one or more models generates the semantic data using the first image (Chen Figure 2 and page 3 left hand column paragraph 3 "The whole pipeline of the proposed SOLIDER is shown in Fig. 2. In this section, we first explain how to generate pseudo semantic labels from human prior knowledge and use it to supervise a token-level semantic classification pretext task").” PNG media_image3.png 229 785 media_image3.png Greyscale Chen Figure 2 It would have been obvious to a person having ordinary skill in the art before effective filing date of the claimed invention of the instant application to combine a person recognition and tracking method as taught by Yiqiang to include a SOLIDER model to generate semantic data as taught by Chen. The suggestion/motivation for doing so would have been “It takes advantages of prior knowledge in human images to produce pseudo semantic labels, and utilize it to train the human representation with more semantic information " as noted by the Chen on page 2 left hand column paragraph 4. Therefore, it would have been obvious to combine the disclosure of Yiqiang with the Chen disclosure to obtain the invention as specified in claim 4 as there is a reasonable expectation of success and/or because doing so merely combines prior art elements according to known methods to yield predictable results. Regarding claim 5, the combination of Yiqiang and Chen teaches “The method of claim 4, wherein the particular model is a semantic controllable self-supervised learning framework (SOLIDER), and wherein the semantic data comprises a pseudo-semantic label (Chen Figure 2 and page 3 left hand column paragraph 3 "The whole pipeline of the proposed SOLIDER is shown in Fig. 2. In this section, we first explain how to generate pseudo semantic labels from human prior knowledge and use it to supervise a token-level semantic classification pretext task").” The proposed combination as well as the motivation for combining Yiqiang and Chen references presented in the rejection of claim 4, applies to claim 5. Finally the method recited in claim 5 is met by Yiqiang and Chen. Regarding claim 19, the combination of Yiqiang and Chen teaches “A method comprising: determining training sets of semantic data, wherein each training set of semantic data corresponds to one of a plurality of first images that depicts one of a plurality of subjects (Chen page 5 left hand column paragraph 5 "For pretext tasks, LUPerson [25, 26] is used for training, the same as [55, 97]. It contains 4.18M human images without any label"); training a first network using the training sets of semantic data to generate training sets of attribute features, wherein each training set of attribute features corresponds to one of the training sets of semantic data (Yiqiang page 4 left hand column paragraph 3 "To train the parameters of the proposed CNN, the weights are initialized at random and updated using stochastic gradient descent minimizing the global loss function (Eq. (1)) on the given training set […] Thus, the multi-label version of the sigmoid cross entropy is used as the overall loss function: […] where L is the number of labels (attributes), N is the number of training examples, and Yu,xil are respectively the ith label and classifier output for the ith image"); training a second network using the training sets of attribute features and the training sets of semantic data to generate training embeddings, wherein each training embedding corresponds to one of the training sets of attribute features and to one of the training sets of semantic data (Yiqiang page 4 right hand column paragraph 2 "It is beneficial to pre-train the CNN with a (possibly larger) pedestrian re-identification dataset in a triplet architecture on the re-identification task. Since pedestrian attribute recognition and re-identification are two similar tasks, the visual features learned from re-identification can be useful for recognizing attributes. Thus, after this pre-training, we transfer the re-identification knowledge to attribute recognition by fine-tuning the pre-trained convolution layers on the actual small attribute datasets"); and training a third network using the training embeddings to identify the plurality of subject in a plurality of second images (Yiqiang page 5 left hand column paragraph 1 "During training, for a given triplet, the loss function "pushes" the negative example away from the reference in the output feature space and "pulls" the positive example closer to it. Thus, by presenting many different triplet combinations, the network effectively learns a no-linear projection to a feature space that better represents the semantic similarity of pedestrians").“ The proposed combination as well as the motivation for combining Yiqiang and Chen references presented in the rejection of claim 4, applies to claim 19. Finally the method recited in claim 19 is met by Yiqiang and Chen. Claims 10 and 16-18 are rejected under 35 U.S.C. 103 as being unpatentable over Yiqiang in view Miyano (US 2015/0262019 A1). Regarding claim 10, Yiqiang teaches “A device comprising: obtain semantic data corresponding to appearance features of a subject within a first image (Yiqiang page 3 right hand column paragraph 1 "higher level discriminative features by several succeeding convolution and pooling operations that become specific to different body parts at a given stage (P3) in order to account for the possible displacements of pedestrians due to pose variations"); generate, using one or more models and the semantic data, attribute features of the subject (Yiqiang page 3 right hand column paragraph 1 "Another branch extracts the viewpoint-invariant Local Maximal Occurrence (LOMO) features, a robust visual feature representation that has been specifically designed for viewpoint-invariant pedestrain attribute recognition and achieving state-of-the-art results [23] (cf. Section 3.1.3)"); generate, using the one or more models, the semantic data, and the attribute features, an embedding (Yiqiang page 4 right hand column paragraph 3 "Person re-identification consists of matching images of the same individuals across multiple camera views. In order to achieve this, we need to learn a distance function that has large values for images from different people and small values for images from the same person" and page 5 right hand column paragraph 2 "From the re-identification data, the network learns informative features that distinguish individuals, and the semantic attributes that we want to recognise can be considered such as identify features at a higher level"); and identify, using the one or more models and the embedding (Yiqiang page 3 left hand column paragraph 1 "In summary, two CNN embeddings are learned based on attribute and identity annotation. Then, an improved triplet loss is used to learn the fusion. We will experimentally show the performance improvement brought by this fusion achieving state-of-the-art results on a public person re-identification benchmark", the subject within a second image (Yiqiang page 1 right hand column paragraph 2 "Person re-identification consists of matching a query person among a large set of people detected in multiple non-overlapping camera views").” However, Yiqiang is not relied on to teach “one or more processors; and a memory storing instructions that, when executed by the one or more processors, configure the device to”. Miyano teaches “one or more processors; and a memory storing instructions that, when executed by the one or more processors, configure the device to (Miyano paragraph [0096] "The processor 1001 controls the various types of processing in the information processing server 100 by executing the programs stored in the memory 1003. For example, the processing relating to the input unit 110, the similarity calculation unit 120, the person-to-be-tracked registration unit 130, the correspondence relationship estimation unit 140, and the display control unit 150 explained in FIG. 1 can be realized as programs that mainly run on the processor 1001 upon temporarily being stored in the memory 1003")”. It would have been obvious to a person having ordinary skill in the art before effective filing date of the claimed invention of the instant application to combine a method for person recognition and tracking method as taught by Yiqiang to include computer architecture to implement a method as taught by Miyano. The suggestion/motivation for doing so would have been utilizing computer architecture to implement a method is known by one ordinary skilled in the art. One of ordinary skill in the art would recognize that computer hardware is utilized for a device performing a series of steps such as “The information processing server 100 performs various types of processing such as the detection of persons, the registration of the person to be tracked and the tracking of the registered person by analyzing the moving images captured by the video cameras 200” as disclosed by Miyano in paragraph 29, can be performed using a processor and memory. Therefore, it would have been obvious to combine the disclosure of Yiqiang with the Miyano disclosure to obtain the invention as specified in claim 10 as there is a reasonable expectation of success and/or because doing so merely combines prior art elements according to known methods to yield predictable results. Regarding claim 16, the combination of Yiqiang and Miyano teaches “The device of claim 10, wherein the attribute features are indicative of at least one intrinsic attribute of the subject and at least one extrinsic attribute of the subject (Yiqiang Table 5 and page 7 right hand column paragraph 4 "The results on the APiS dataset are shown in Table 5").“ Regarding claim 17, the combination of Yiqiang and Miyano teaches “The device of claim 16, wherein the at least one intrinsic attribute corresponds to at least one of an age, sex, hair style, or body shape of the subject (Yiqiang Table 5 and page 7 right hand column paragraph 4 "The results on the APiS dataset are shown in Table 5").“ Regarding claim 18, the combination of Yiqiang and Miyano teaches “The device of claim 16, wherein the at least one extrinsic attribute corresponds to at least one of a clothing article worn by the subject or an object attached to or carried by the subject (Yiqiang Table 5 and page 7 right hand column paragraph 4 "The results on the APiS dataset are shown in Table 5").“ Claims 11-12 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Yiqiang and Miyano in view of Lin. Regarding claim 11, Yiqiang teaches “The device of claim 10, wherein the one or more models comprises a first network trained to generate training attribute features from training semantic data associated with a second subject in an initial image using a first loss function (Yiqiang page 4 left hand column paragraph 3 and right hand column paragraph 1 "have several properties at the same time, the attribute recognition is a multi-label classification problem. Thus, the multi-label version of the sigmoid cross entropy is used as the overall loss function"), a second network trained to generate a training embedding using the training semantic data and the training attribute features using a second loss function different from the first loss function (Yiqiang page 6 left hand column paragraph 3 "The pre-trained attribute CNN and identification CNN are combined and trained in a triplet architecture similar to the one explained in Section 3.1.4. Here, we propose to use an improved triplet loss with hard example selection to learn the optimal fusion of the two types of features. The fcl layer of the attribute network and the fcl layer of the identification network are normalised and concatenated, and another fully-connected layer which allows to merge attribute and identification features"), and However, the combination of Yiqiang and Miyano is not relied on to teach “a third network trained to identify the second subject within a subsequent image using the training embedding and a third loss function different from the first and second loss functions”. Lin teaches “a third network trained to identify the second subject within a subsequent image using the training embedding and a third loss function different from the first and second loss functions (Lin page 6 left hand column paragraph 2 "Considering both attribute recognition and identity prediction, we define the overall objective function as followings where λ   is a hyper-parameter to balance the identity classification loss and the attribute recognition losses").” It would have been obvious to a person having ordinary skill in the art before effective filing date of the claimed invention of the instant application to combine a person recognition and tracking method as taught by Yiqiang and Miyano to train a network for identification using loss function as taught by Lin. The suggestion/motivation for doing so would have been “By combining the attribute recognition task and identity classification task, the APR network is capable of learning more discriminative feature representations for pedestrians, including global and local descriptions. Specifically, we take attribute predictions as additional cues for the identity classification. Considering the dependencies among pedestrian attributes, we first re-weight the attribute predictions and then build identification upon these re-weighted attributes descriptions" as noted by the Lin on page 2 right hand column paragraph 2. Therefore, it would have been obvious to combine the disclosure of Yiqiang and Miyano with the Lin disclosure to obtain the invention as specified in claim 11 as there is a reasonable expectation of success and/or because doing so merely combines prior art elements according to known methods to yield predictable results. Regarding claim 12, the combination of Yiqiang, Miyano, and Lin teaches “The device of claim 11, wherein at least one of the first, second, or third loss function is a triplet loss function (Yiqiang page 6 left hand column paragraph 3 "The pre-trained attribute CNN and identification CNN are combined and trained in a triplet architecture similar to the one explained in Section 3.1.4. Here, we propose to use an improved triplet loss with hard example selection to learn the optimal fusion of the two types of features") or a binary cross-entropy (BCE) loss function (Yiqiang page 4 left hand column paragraph 3 and right hand column paragraph 1 "have several properties at the same time, the attribute recognition is a multi-label classification problem. Thus, the multi-label version of the sigmoid cross entropy is used as the overall loss function").” Regarding claim 15, the combination of Yiqiang, Miyano, and Lin teaches “The device of claim 10, wherein the embedding generated by the one or more models is based at least on weights derived from relationships between the attribute features and the appearance features (Lin Figure 5 and page 5 left hand column paragraph 4 "network firstly computes attribute losses for the M individual attributes. Then the M prediction scores are concatenated and fed into an Attribute Re-weighting Module").” Claims 13-14 are rejected under 35 U.S.C. 103 as being unpatentable over Yiqiang and Miyano in view of Chen. Regarding claim 13, the combination of Yiqiang and Miyano teaches the device of claim 10. However, the combination of Yiqiang and Miyano is not relied on to teach “wherein a particular model of the one or more models generates the semantic data using the first image”. Chen teaches “wherein a particular model of the one or more models generates the semantic data using the first image (Chen Figure 2 and page 3 left hand column paragraph 3 "The whole pipeline of the proposed SOLIDER is shown in Fig. 2. In this section, we first explain how to generate pseudo semantic labels from human prior knowledge and use it to supervise a token-level semantic classification pretext task").” PNG media_image3.png 229 785 media_image3.png Greyscale Chen Figure 2 It would have been obvious to a person having ordinary skill in the art before effective filing date of the claimed invention of the instant application to combine a person recognition and tracking method as taught by Yiqiang and Miyano to include a SOLIDER model to generate semantic data as taught by Chen. The suggestion/motivation for doing so would have been “It takes advantages of prior knowledge in human images to produce pseudo semantic labels, and utilize it to train the human representation with more semantic information " as noted by the Chen on page 2 left hand column paragraph 4. Therefore, it would have been obvious to combine the disclosure of Yiqiang and Miyano with the Chen disclosure to obtain the invention as specified in claim 13 as there is a reasonable expectation of success and/or because doing so merely combines prior art elements according to known methods to yield predictable results. Regarding claim 14, the combination of Yiqiang, Miyano, and Chen teaches “The device of claim 13, wherein the particular model is a semantic controllable self-supervised learning framework (SOLIDER), and wherein the semantic data comprises a pseudo-semantic label (Chen Figure 2 and page 3 left hand column paragraph 3 "The whole pipeline of the proposed SOLIDER is shown in Fig. 2. In this section, we first explain how to generate pseudo semantic labels from human prior knowledge and use it to supervise a token-level semantic classification pretext task").” The proposed combination as well as the motivation for combining Yiqiang and Chen references presented in the rejection of claim 4, applies to claim 5. Finally the method recited in claim 5 is met by Yiqiang and Chen. Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Yiqiang and Chen in view of Bozzo et al. ("A Multimodal Neural Network with Gradient Blending Improves Predictions of Survival and Metastasis in Sarcoma"). Regarding claim 20, the combination of Yiqiang and Chen teaches the method of claim 19. However, the combination of Yiqiang and Chen is not relied on to teach “wherein during at least the training of the second network and the training of the third network, parameters of the second and third networks are jointly updated based at least on a combined loss gradient algorithm that combines a first loss gradient corresponding to the second network and a second loss gradient corresponding to the third network”. Bozzo teaches “wherein during at least the training of the second network and the training of the third network, parameters of the second and third networks are jointly updated based at least on a combined loss gradient algorithm that combines a first loss gradient corresponding to the second network and a second loss gradient corresponding to the third network (Bozzo page 8 left hand column paragraph 2 "Gradient blending is used to moderate the loss contributions of the different modalities during the training of multimodal neural networks. During training, gradient blending uses the overfitting to generalization ratios of each modality to promote losses from subnetworks that are generalizing well to the validation set, while down weighting losses from subnetworks that are overfitting the training set. This enables the simultaneous, end-to-end training of two separate model architectures that would typically converge and overfit at different rates. Model selection during training is performed considering only the unweighted contribution of the multimodal head to the loss function").” It would have been obvious to a person having ordinary skill in the art before effective filing date of the claimed invention of the instant application to combine a training method for person recognition and tracking method as taught by Yiqiang and Chen to include a multi-network gradient loss as taught by Bozzo. The suggestion/motivation for doing so would have been “the demonstrated advances in prediction capability are due to the multimodal nature of our work, and the implementation of gradient blending which was crucial to properly combining the inputs from different modalities " as noted by the Bozzo on page 2 right hand column paragraph 5. Therefore, it would have been obvious to combine the disclosure of Yiqiang and Chen with the Bozzo disclosure to obtain the invention as specified in claim 20 as there is a reasonable expectation of success and/or because doing so merely combines prior art elements according to known methods to yield predictable results. Reference Cited The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. US Publication 20230351794 A1 to Dou et al. discloses a system and method of tracking a pedestrian using a multi-modal approach for extraction features and identification. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to JASPREET KAUR whose telephone number is (571)272-5534. The examiner can normally be reached Monday - Friday 7:30 am - 4:00 PST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amandeep Saini can be reached at (571)272-3382. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JASPREET KAUR/Examiner, Art Unit 2662 /AMANDEEP SAINI/Supervisory Patent Examiner, Art Unit 2662
Read full office action

Prosecution Timeline

Nov 26, 2024
Application Filed
Jul 29, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705710
UPSAMPLING BLOCKS OF PIXELS
2y 10m to grant Granted Aug 11, 2026
Patent 12682435
ADAPTIVE SHARPENING FOR BLOCKS OF PIXELS
2y 9m to grant Granted Jul 14, 2026
Patent 12676243
QUANTIFYING VARIATION IN SURGICAL APPROACHES
2y 9m to grant Granted Jul 07, 2026
Patent 12675860
COMPUTERIZED IMAGE ANALYSIS FOR AUTOMATICALLY DETERMINING WAIT TIMES FOR A QUEUE AREA
3y 1m to grant Granted Jul 07, 2026
Patent 12670558
APPARATUS FOR EXTRACTING NOISE FROM IMAGE AND METHOD THEREOF
2y 9m to grant Granted Jun 30, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
78%
Grant Probability
99%
With Interview (+41.7%)
2y 7m (~11m remaining)
Median Time to Grant
Low
PTA Risk
Based on 23 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month