DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
The reply filed on 7/23/2026 has been entered. Some of the applicant’s arguments with respect to claims 1, 8, 9, and 10 have been considered but are moot in view of new rejection.
Claims 1, 2, 4, 5, 7, 12, 155, 16, 17, 18, 19, and 20 are pending in this application and have been considered below. Claims 3, 6, 13, and 14 are canceled by the applicant. In response to the Amendment, the previous rejection of claims 2, 4, 5, 7, 12, 15, 16, 17, 18, 19, and 20 are withdrawn.
Information Disclosure Statement
The IDSs dated 8/28/2026 have been considered are placed in the application file.
1st Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1, 8, and 9 are rejected under 35 U.S.C. 103 as obvious over US Patent Publication 2022 0075994 A1, (Shapira et al.) in view of US Patent Publication 2025 0299506 A1, (Miri et al.) and US Patent Publication 2023 0410483 A1, (Chen et al.).
Claim 1
Regarding claim 1, Shapira et al. teach a computer implemented method for self-supervised learning (SSL) training of a network for facial landmark detection for a face in an input image, wherein the network comprises encoder components encoding features of the face and decoder components determining local correspondences between the features for determining landmark estimates, the method comprising: ("By using the attention network, the facial landmark detection model learns to focus on the crucial information and ignore less relevant information, which increases the facial landmark detection model's accuracy. Using the attention network, the high-weighted patch 600 will get higher weighting than the low-weighted patch 605. The attention network may be used in deep learning models for natural language processing and vision," par. 75).
Shapira et al. do not explicitly teach all of training a first network comprising a Masked Image Modeling (MIM) network that processes non-overlapping patches determined from the input image with a SSL objective, wherein the encoder components comprise a portion of the MIM network to encode features of the input image for decoding; and training a second network comprising the encoder components, as trained, in series with the decoder components, the decoder components trained to determine local correspondences comprising respective relationships between the features of the input image provided by the encoder components.
However, Chen et al. teach training a first network comprising a Masked Image Modeling (MIM) network ("masked image modeling (MIM) training process," par. 4) that processes non-overlapping patches determined from the input image with a SSL
PNG
media_image1.png
456
618
media_image1.png
Greyscale
objective, wherein the encoder components comprise a portion of the MIM network to encode features of the input image for decoding; ("Some extant MIM models employ an encoder-decoder design [AltContent: textbox (Figure 2A shows the system layout of the self-supervised MIM.)]followed by a projection head. The encoder aids in the modeling of latent feature representations," par. 33) and training a second network comprising the encoder components, as trained, in series with the decoder components, the decoder components trained to determine local correspondences comprising respective relationships between the features of the input image provided by the encoder components ("two-layer convolutional transpose can be used as a projection head 260 (FIG. 2A) during the self-supervised MIM training process 200 for pre-training the image encoder 150 and the UPerNet decoder," par. 44) .
Therefore, taking the teachings of Shapira et al. and Chen et al. as a whole, it would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the facial landmark detection system as taught by Shapira et al. to use the masked image modeling and masked autoencoder as taught by Chen et al. The suggestion/motivation for doing so would have been that, “self-supervised MIM training as disclosed herein in is especially advantageous for modeling 3D medical images by significantly speeding up training convergence and improves downstream performance. For instance, when compared to naive contrastive learning, training convergence can save up to a 1.40× training cost to reach a same or higher dice score when the pre-trained image encoder 150 is adapted and fine-tuned to perform a downstream vision task. Similarly, the downstream performance of the downstream vision task of image segmentation can achieve over 5-percent (5%) improvements without any hyper parameter tuning. Additionally, downstream applications incorporating the image encoder pre-trained via self-supervised MIM training are faster and more cost-effective then transfer learning to the particular downstream task for prognosis, treatment sensitivity prediction, tissue segmentation, image classification, and digital representations of patients” as noted by the Chen et al. disclosure in paragraph [0041], which also motivates combination because the combination would predictably have a greater efficiency as there is a reasonable expectation that the modified facial landmark detection system would achieve faster training convergence, reduced computational overhead, and enhanced accuracy in landmark placement without requiring extensive parameter tuning; and/or because doing so merely combines prior art elements according to known methods to yield predictable results.
Additionally, Miri et al. teach a network that processes non-overlapping patches determined from the input image ("image 305 may be passed into the patching and embedding block 230 that may convert the image 305 into non-overlapping equal sized 2D grid of patches," par. 51) with a SSL objective ("the key to success of an SSL model may lie in wisely making use of the information derived from the image," par. 39).
Therefore, taking the teachings of Shapira et al., Chen et al., and Miri et al. as a whole, it would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the facial landmark detection system as taught by Shapira et al. and the masked image modeling and masked autoencoder as taught by Chen et al. to use the patchification and vision transformer as taught by Miri et al. The suggestion/motivation for doing so would have been that, “a self-supervised multi-class-token hierarchical ViT is a novel backbone that capture both coarse and fine-grained features. (The ViT may have been tested on one or more frameworks, such as the DINO, MOCO, and/or SimCLR SSL frameworks). Compared to ImageNet pretraining and other state-of-the-art SSL methods, this model presents at least two advantages: yielding features of considerably higher quality compared to other state-of-the-art SSL frameworks and tile retrieval demonstration; learning more precise morphological phenotypes down to pixel level, different from the grid structure attention map extracted from the multi-head attention heads from regular ViT” as noted by the Miri et al. disclosure in paragraph [0034], which also motivates combination because the combination would predictably have a greater efficiency as there is a reasonable expectation that the system would more accurately locate facial landmarks, even when obscured by masks, by using both coarse structural context and precise, pixel-level morphological features; and/or because doing so merely combines prior art elements according to known methods to yield predictable results.
Claim 8
Regarding claim 8, Shapira et al., Chen et al., and Miri et al. teach the method of claim 1 as noted above.
Shapira et al. do not explicitly teach all of wherein the second network comprises a projector network, the decoder components comprising a portion of the projector network.
However, Chen et al. teach wherein the second network comprises a projector network, the decoder components comprising a portion of the projector network ("decoder 152 may include a UPerNet to perform image segmentation tasks based on the encoded features 225 output from the image encoder 150. That is, a two-layer convolutional transpose can be used as a projection head 260 (FIG. 2A)," par. 44).
Shapira et al., Chen et al., and Miri et al. are combined as per claim 1.
Claim 9
Regarding claim 9, Shapira et al., Chen et al., and Miri et al. teach the method of claim 1 as noted above.
Shapira et al. do not explicitly teach all of wherein training the decoder components trains the projector network using a locality constrained repellence (LCR) loss.
However, Chen et al. teach wherein training the decoder components trains the projector network using a locality constrained repellence (LCR) loss ("The training loss may be based on a distance in a voxel space between the recovered/estimated raw voxel values 270 and the original voxels from the corresponding sets of raw voxel values that represent the masked image patches. The training loss may include either an l.sub.1 or l.sub.2 loss function. Notably, the training loss may only be computed for the masked matches 210M to prevent the encoder 150 from engaging in self-reconstruction and potentially dominate the learning process and ultimately impeded knowledge learning. Thereafter, the training process 200 updates parameters of the image encoder 150 (and optionally the decoder 250) based on the training loss," par. 58).
Shapira et al., Chen et al., and Miri et al. are combined as per claim 1.
2nd Claim Rejections - 35 USC § 103
Claim 10 is rejected under 35 U.S.C. 103 as obvious over US Patent Publication 2022 0075994 A1, (Shapira et al.), US Patent Publication 2025 0299506 A1, (Miri et al.), and US Patent Publication 2023 0410483 A1, (Chen et al.) in view of US Patent Publication 2023 0029608 A1, (Chafni et al.).
Claim 10
Regarding claim 10, Shapira et al., Chen et al., Miri et al. teach the method of claim 9 as noted above.
Shapira et al. teach wherein the LCR operates on features of facial landmark regions and combined information from non-landmark regions that reduces processing to achieve selective correspondence processing for the local correspondences ("The loss function is a weighted average of the normalized Euclidean distance between the regressed landmarks and the ground truth and the absolute difference between the predicted error and the actual error," par. 79).
Shapira et al. do not explicitly teach all of wherein the non-landmark regions include cheek regions and forehead regions of the input image.
Additionally, Chafni et al. teach wherein the non-landmark regions include cheek regions and forehead regions of the input image ("Additionally or alternatively, the second compression template 222 may include one or more elements 226 corresponding to elements of the second image 220 including insignificant information, such as the hair, forehead, ears, jaw, etc. of the human face," par. 37).
Therefore, taking the teachings of Shapira et al., Chen et al., Miri et al., and Chafni et al. as a whole, it would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the facial landmark detection system as taught by Shapira et al., the masked image modeling and masked autoencoder as taught by Chen et al., and the patchification and vision transformer as taught by Miri et al. to use the non-landmark regions of the cheek and forehead as taught by Chafni et al. The suggestion/motivation for doing so would have been that, “the second compression template 222 may include one or more elements 226 corresponding to elements of the second image 220 including insignificant information, such as the hair, forehead, ears, jaw, etc. of the human face” as noted by the Chafni et al. disclosure in paragraph [0037], which also motivates combination because the combination would predictably have a higher accuracy as there is a reasonable expectation that identifying and selectively compressing or masking these insignificant, non-landmark regions allows the system to focus computational resources on enhancing and processing the significant landmark regions without losing critical facial information; and/or because doing so merely combines prior art elements according to known methods to yield predictable results.
Allowable Subject Matter
Claims 2, 4, 5, 7, and 11 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Claims 12, 15, 16, 17, 18, 19, and 20 are allowed.
Reasons for Indicating Allowable Subject Matter
The following is an examiner’s statement of reasons for allowance: The prior art of record does not teach certain distinguishing features as described below in reference to independent claim 1.
Regarding claim 12, the prior art Shapira et al. teaches to provide a network for facial landmark detection for faces in input images; process, using the network, an input image comprising a face to determine and provide facial landmarks therefor; and wherein the network comprises: configured for encoding features of the face, and decoder components configured for determining local correspondences between the features for determining estimates for the facial landmarks, the decoder components trained to determine local correspondences comprising respective relationships between the features of the input image.
Chen et al. teach a system comprising at least one processor, a non-transient storage device coupled to the at least one processor, the storage device storing instructions executable by the at least one processor to cause the system to, and encoder components comprising trained components of a Masked Image Modeling (MIM) network.
Additionally, Miri et al. teach configured to process non-overlapping patches determined from the input image, the MIM network trained with a SSL objective.
None teaches: wherein the MIM network, as trained, is configured to provide respective tokens for the patches for processing by the decoder components, including approximating inattentive tokens related to non-landmark regions including cheek regions and forehead regions of the input image..
Further, none of the reference teaches or fairly suggests the combination of claimed elements. The Examiner finds no reason or motivation to combine the above references in an obviousness rejection thus placing the claims in condition for allowance.
Claim 2 is objected to by analogy.
Claims 15-20 are allowable as depending on an allowable independent claim.
Reference Cited
The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure.
US Patent Publication 2024 0096072 A1 to He et al. discloses a masked autoencoder system wherein the method consists of dividing the input image into a set a patches, selecting a first subset of the patches to be visible and a second subset of the patches to be masked during the pre-training, processing, using the encoder, the first subset of patches to generate corresponding first latent representations, processing, using the decoder, the first latent representations corresponding to the first subset of patches.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KARSTEN F LANTZ whose telephone number is (571) 272-4564. The examiner can normally be reached Monday-Friday 8:00-4:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ms. Jennifer Mehmood can be reached on 571-272-2976. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Karsten F. Lantz/Examiner, Art Unit 2664
Date: 9/3/2026
/JENNIFER MEHMOOD/Supervisory Patent Examiner, Art Unit 2664