Prosecution Insights
Last updated: August 14, 2026
Application No. 18/891,650

METHOD AND SYSTEM OF IMAGE PROCESSING FOR DETERMINING LIVENESS OF A SUBJECT

Non-Final OA §103§112
Filed
Sep 20, 2024
Priority
Sep 22, 2023 — provisional 63/584,825
Examiner
BOYAR, NOAH WILLIAM
Art Unit
Tech Center
Assignee
Hyperverge Inc.
OA Round
1 (Non-Final)
100%
Grant Probability
Favorable
1-2
OA Rounds
3m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 100% — above average
100%
Career Allowance Rate
2 granted / 2 resolved
+40.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Fast prosecutor
2y 2m
Avg Prosecution
20 currently pending
Career history
17
Total Applications
across all art units

Statute-Specific Performance

§101
14.6%
-25.4% vs TC avg
§103
50.0%
+10.0% vs TC avg
§102
18.3%
-21.7% vs TC avg
§112
13.4%
-26.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 2 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f): (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f). The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f). The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) except as otherwise indicated in an Office action. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f), because the claim limitations use a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitations are: “a capturing unit”; “a face detector unit”; “a monocular depth estimation unit”; “a creator unit”; “an input unit”; “a determination unit”; “an aggregator unit”. Because the claim limitations are being interpreted under 35 U.S.C. 112(f), they are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have these limitations interpreted under 35 U.S.C. 112(f) applicant may: (1) amend the claim limitations to avoid them being interpreted under 35 U.S.C. 112(f) (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitations recite sufficient structure to perform the claimed function so as to avoid them being interpreted under 35 U.S.C. 112(f). Claim Objections Claims 13 and 14 are objected to for the following informalities: “The method comprises” is improper. As a dependent claim, it is proper under 35 U.S.C. § 112(d) to recite that “The method further comprises (the further limitations)”. Claims 12 and 27 are objected to for the following informalities: “further based on” is improper, when the claims in question are expanding on the “determining” step of claim 1 in more detail, rather than presenting an additional determination criteria. Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. Claims 1-30 are rejected under 35 U.S.C. 112(b) as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint invention, regards as the invention. With respect to claims 1 and 16, an issue of antecedent basis is present. First, a “depth map corresponding to each target image from the plurality of target images” is generated (i.e., a plurality of depth maps). However, these depth maps are then referred back to in the singular, “based on an addition of the depth map and a set of color models associated with the plurality of target images”. Accordingly, it is unclear which of the depth maps is being referred to. For the purposes of compact prosecution, the examiner will interpret any depth map as qualifying. With respect to claims 1 and 16, the recitation of “an addition of the depth map and a set of color models” is technically improper. Where Applicant acts as his or her own lexicographer to specifically define a term of a claim contrary to its ordinary meaning, the written description must clearly redefine the claim term and set forth the uncommon definition so as to put one reasonably skilled in the art on notice that the applicant intended to so redefine that claim term. Process Control Corp. v. HydReclaim Corp., 190 F.3d 1350, 1357, 52 USPQ2d 1029, 1033 (Fed. Cir. 1999). A depth map comprises pixel data, whereas a color model (i.e., RGB, HSV, YCbCr in accordance with the disclosure) is an abstract, higher-level classification label of colors. It is semantically impossible to “add” a depth map to a color model accordingly. For the purposes of compact prosecution, the examiner will understand the general intention of the invention to involve adding respective pixel data (for example, red, green, and blue pixels), in accordance with page 13 of the claimed invention’s specification (“Therefore, a final model input (i.e., the plurality of modified images) is created by adding the depth map, HSV and YCbCr images to the RGB image as a subsequent channel.”). With respect to claim 16, a nesting issue is present. Following the recitation of “wherein:”, the plurality of multi-branch models are elaborated upon. Then a semicolon is used, and the determination unit is introduced. However, grammatically, the determination unit still exists under the “wherein:” colon. It is thus unclear whether the determination unit is part of the previously recited input unit, or exists as its own separate component. For the purposes of compact prosecution, the examiner will analyze both interpretations as reading. With respect to claim 16, it is also unclear whether the “plurality of multi-branch image liveness models” exist external to the system or internally, as the system only comprises “an input unit configured to provide…to a plurality of multi-branch image liveness models”. If the claim were to include external multi-branch image liveness models, it would go against the “single entity” requirement for 35 U.S.C. § 271 infringement. As the claim is currently written, it is not sufficiently clear whether these models are part of the “system comprising:” or could merely exist alongside it. For the purposes of compact prosecution, the examiner will analyze both interpretations as reading. With respect to claims 12 and 27, the claims recite “by each corresponding multi-branch image liveness model”. There is no antecedent basis for a “corresponding multi-branch image liveness model”, neither is it made clear how specifically they are to “correspond[]” to the various modified images and non-live attacks. For the purposes of compact prosecution, the examiner will interpret the word “corresponding” as merely indicating a model which is responsible for detection of any given non-live attack within any given modified image. With respect to claims 14 and 29, the claims recite “a weighted combination of the image liveness score”, however, only a single “image liveness score” was referred to in claim 13. As you cannot have a combination of a single score, the claim is indefinite. The examiner understands the intended meaning to be the combination of a plurality of scores, however, an interpretation of claim 13 as producing a single score still stands as reasonable (for the reasons outlined in the prior art rejection below). With respect to claims 2-15 and 17-30, they are rejected by virtue of their dependency on base and intervening claims discussed above. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 3-5, 7-13, 16, 18-20, and 22-28 are rejected under 35 U.S.C. 103 as being unpatentable over Wang et. al Multi-modal Face Presentation Attack Detection via Spatial and Channel Attentions (Hereinafter, “Wang”) in view of Ye et. al Deep Joint Depth Estimation and Color Correction from Monocular Underwater Images Based on Unsupervised Adaptation Networks) (Hereinafter, “Ye”) With respect to claim 1, Wang teaches: A method of image processing for determining liveness of a subject, the method comprising: capturing, by a capturing unit, an image of the subject ([4.1] “These RGB, Depth and Infrared (IR) videos simultaneously captured using the Intel RealSense SR300 camera”) processing, by a face detector unit, the image to create a plurality of target images ([3.3] image preprocessing; [3.3] “We resize the cropped face region”, wherein the broad task of facial detection reads on cropping to a facial region; [4.1] “The background image area of the face was removed from original videos to make the face PAD task more challenging) generating, Fig. 2; [4.1] “These RGB, Depth and Infrared (IR) videos simultaneously captured using the Intel RealSense SR300 camera”) creating, by a creator unit, a plurality of modified images based on an addition of the depth map and a set of color models associated with the plurality of target images (Fig. 2, “Fusion”; [3.1] “We build a multi-stream architecture and use the feature-level fusion category which is to fuse the features extracted from RGB, Depth, IR and fused modality subnetworks and then fed into the shared layers to learn joint representations”; “plurality of” being satisfied as the model is trained over numerous epochs with different photographs, [3.3] “our model is trained with end-to-end style for 200 epochs”) providing, by an input unit, the plurality of modified images to a plurality of multi-branch image liveness models (Fig. 2, “res5”, “res6” (with additional GAP/fully connected architecture) satisfying the “plurality”, multi-branch in the sense that they receive the multiple branches of Fig. 2) detecting, by the plurality of multi-branch image liveness models, one of a presence of a set of non-live attacks and an absence of the set of non-live attacks in the plurality of modified images (Fig. 2, “which are shared to learn more discriminative features”) determining, by a determination unit, the liveness of the subject based on detection of the absence of the set of non-live attacks in the plurality of modified images (Fig. 3, final output; [5] “Then the extracted features from four branches are concatenated and fed into the shared layers to classify supervised with the joint of the center loss and softmax loss”) Wang does not explicitly teach: by a monocular depth estimation unit However, Ye in the same field of endeavor of depth and color estimation, teaches: A method of image processing capturing, by a capturing unit, an image of the subject (Fig. 1) processing, Fig. 1) generating, by a monocular depth estimation unit, a depth map corresponding to each target image from the plurality of target images (Fig. 1; [1] “Unfortunately, compared to the task of monocular depth estimation in the air, almost no research focuses on depth estimation in the underwater environment due to the lack of effective training data”) creating, by a creator unit, a plurality of modified images based on an addition of the depth map and a set of color models associated with the plurality of target images (Fig. 1; [5B] “Jointly training depth estimation and color correction modules (both without DA) will improve the performance, because the stacked mode aims to learn a shared embedding across tasks by simply aggregating the supervisions from each individual task.”) PNG media_image1.png 495 1432 media_image1.png Greyscale providing, by an input unit, the plurality of modified images to a plurality of multi-branch Fig. 2, loss term of SAN, additional processing of TN) It would have been obvious to one of ordinary skill in the art as of the effective filing date of the claimed invention, to modify Wang to include the limitations of monocular depth estimation, as taught by Ye. Doing so would have the advantage of allowing for depth estimation using simpler hardware. The systems readily integrate, as an estimated depth map could otherwise be used for similar processing as a reference depth map, under the pipeline of Wang. To the extent Applicant may argue that Wang does not teach modified “images” (since Wang uses mathematical tensor image representations instead), Ye further evidences the possibility of generating decoded image outputs and then taking further representations once again (Fig. 1, Gs and Gd). Ye also evidences the usage of two distinct networks for processing “modified images” (Fig. 1, SAN and TN respectively), if the Applicant argues the narrower interpretation that Wang’s resnets are part of the same “model”. A person of ordinary skill in the art could reach such modifications to a reasonable expectation of predictable success, for the advantage of providing additional control over the tuning of the model’s intermediate outputs. With respect to claim 3, Wang/Ye teaches: The method as claimed in claim 1, wherein the plurality of target images comprises at least a first image, a second image and a third image (Wang, Fig. 2; Wang, [4.1] “These RGB, Depth and Infrared (IR) videos simultaneously captured using the Intel RealSense SR300 camera”) With respect to claim 4, Wang/Ye teaches: The method as claimed in claim 3, wherein the plurality of target images comprises a face of the subject, the second image comprises the face and a first pre-defined percentage of a background detected in the image, and third image comprises the face and a second pre-defined percentage of the background (Wang, Fig. 4, noting that differing portions of the background are visible; Wang [3.3] “We resize the cropped face region to the size 56 x 56, and use an open source imgaug library to do data augmentation, i.e., random flipping, rotation, resizing, cropping and color distortion”, wherein the cropping is “pre-defined” in the sense that mathematically the amount of cropping must be determined (even if determined randomly) prior to the background being fixed to its new value) With respect to claim 5, Wang/Ye teaches: The method as claimed in claim 1, wherein prior to generating the depth map the method comprises resizing the plurality of target images in a pre-determined size (Wang, [3.3]) With respect to claim 7, Wang/Ye teaches: The method as claimed in claim 1, wherein each modified image from the plurality of modified images comprises a set of channels (Wang, Fig. 2; Ye, Fig. 1; modified image outputs necessarily possess channels as digital color images) With respect to claim 8, Wang/Ye teaches: The method as claimed in claim 1, wherein each multi-branch image liveness model from the plurality of multi-branch image liveness models receives the plurality of modified images in Wang, Fig. 2, understood that res5 and res6 process in sequence) With respect to claim 9, Wang/Ye teaches: The method as claimed in claim 1, wherein each multi-branch image liveness model from the plurality of multi-branch image liveness models is a neural network based model (Wang, Fig. 2, “res5”, “res6”), and wherein said each multi-branch image liveness model is trained for detecting a specific type of non-live attack ([4.1] “Specifically, this dataset contains 1,000 Chinese people and each person has 1 live video clip and 6 fake video clips (6 different attack manners) for each modality”) With respect to claim 10, Wang/Ye teaches: The method as claimed in claim 9, wherein the specific type of non-live attack is one of a display attack type, a print attack type, and a mask-based attack type (Wang, Fig. 1; Wang, Fig. 4) PNG media_image2.png 569 377 media_image2.png Greyscale With respect to claim 11, Wang/Ye teaches: The method as claimed in claim 1, wherein the set of non-live attacks comprises at least one of one or more display attacks, one or more print attacks, and one or more mask-based attacks (Wang, Fig. 1; Wang, Fig. 4) With respect to claim 12, Wang/Ye teaches: The method as claimed in claim 1, wherein the determining, by the determining unit, the liveness of the subject is further based on a detection of an absence of each non-live attack from the set of non-live attacks in each modified image from the plurality of modified image, by each corresponding multi-branch image liveness model from the plurality of multi-branch image liveness models (Wang, Fig. 2 “Center Loss”; [1] “In order to enhance the discriminative power of the deeply learned features, the network is using SGRD strategy to update the parameters and optimized with the joint supervision of softmax loss and center loss [32], aiming to minimize the intra-class variations while keep the features of different classed separable”; Wang, Fig. 4, “‘G, G’ (or ‘S, S’) denotes a genuine (spoof) face image is correctly classified as genuine (spoof)”; [1] “The main contributions of this work are…(ii) SGRD solver to update network parameters and joint supervision of softmax loss and center loss to obtain more discriminative feature representation for live and spoof faces”) In other words, the center loss strives to remove all “spoof” features from the live category, such that it can be said that with optimization, the decision that a subject is live is based on the absence of the learned spoof features rather than mere presence of live features. Such that an image with even a single spoof feature could no longer be considered “live”, and would be grouped in the other category. With respect to claim 13, Wang/Ye teaches: The method as claimed in claim 1, the method comprises generating an image liveness score by each multi-branch image liveness model from the plurality of multi-branch image liveness models (Wang, Fig. 2; unlike claim 14, this language does not explicitly require that the multi-branch image liveness models generate scores one by one, but rather can mutually participate in the generation of a final score, such that it could be said that an image liveness score is generated “by” (i.e., through the contribution of) each multi-branch image liveness model, even if only one such model produces the actual final output score) based on one of the presence of the set of non-live attacks and the absence of the set of non-live attacks in the plurality of modified images. ([Abstract] “Finally, we get the classification confidence scores w.r.t. PAD or not”) With respect to claim 16, the claim is functionally parallel to claim 1. Claim 16 describes a system configured to perform the method of claim 1. As the method of Wang/Ye is generally understood to take place on a computer capable of performing the method (Wang, [3.3]), the claim language is met in line with the citations above. Accordingly, the claim is rejected. With respect to claims 18-20 and 22-28, they are rejected in line with the rejections of claims 16, 3-5, and 7-13 above. Claims 2 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Wang and Ye in view of Saban (US 20220398781 A1) (Hereinafter, “Saban”) With respect to claim 2, Wang/Ye teaches the method of claim 1. Wang/Ye does not explicitly teach the further limitations of claim 2. However, Saban, in the same field of endeavor of facial recognition, teaches: wherein for capturing the image the method comprises performing at least one of a set of compliance checks and a set of sanity checks ([0062]-[0064] “The visual feedback module 602 gets the output from module 601 as absolute state or offsets from optimal view and converts it into a visual feedback (e.g., avatar in a form of 3D spheroid or trimmed 2D from 3D projection of sphere on 2D plane) so that the user can change the head position to match the avatar to the graphical target, e.g., can change his gaze, head rotation, head translation, and eyes opening state to get into best possible state for grabbing optimal image.”) It would have been obvious to one of ordinary skill in the art as of the effective filing date of the claimed invention to modify Wang/Ye to include the limitations of compliance and sanity checks as, as taught by Saban. Doing so would ensure the optimal image is taken for processing. The systems readily integrate, as the processing of Wang/Ye requires the isolation of a face. With respect to claim 17, it is rejected in line with the rejections of claims 16 and 2 above. Claims 6 and 21 are rejected under 35 U.S.C. 103 as being unpatentable over Wang and Ye in view of Banerji et. al LBP and Color Descriptors for Image Classification (Hereinafter, “Banerji”) With respect to claim 6, Wang/Ye teach the method of claim 1. Wang/Ye do not explicitly teach the further limitations of claim 6. However, Banerji, in the same field of endeavor of anti-spoofing, teaches: wherein the set of color models comprises a Hue, Saturation, Value (HSV) color model, a Luminance, Chrominance (YCbCr) color model, and a Red, Green, Blue (RGB) color model ([Abstract] “Four novel color Local Binary Pattern (LBP) descriptors are presented in this chapter for scene image and image texture classification with applications to image search and retrieval… the Color LBP Fusion (CLF) descriptor is constructed by integrating the RGB-LBP, the YCbCr-LBP, the HSV-LBP, the rgb-LBP, as well as the oRGB-LBP descriptor”) It would have been obvious to one of ordinary skill in the art as of the effective filing date of the claimed invention to modify Wang/Ye to include the limitations of fused color features, as taught by Banerji. Doing so would have the advantage of providing more color information for analysis. The systems readily integrate as the fused color feature can otherwise be processed in a similar manner to the RGB features of Wang/Ye. With respect to claim 21, it is rejected in line with the rejections of claims 16 and 6 above. Claims 14-15 and 29-30 are rejected under 35 U.S.C. 103 as being unpatentable over Wang and Ye in view of Atoum et. al Face anti-spoofing using patch and depth-based CNNs (Hereinafter, “Atoum”) With respect to claim 14, Wang/Ye teaches the method of claim 13. Wang/Ye does not teach the further limitations of claim 14. However, Atoum, in the same field of endeavor of anti-spoofing, teaches: the method comprises generating by an aggregator unit a final image liveness score based on a weighted combination of the image liveness score generated by said each multi-branch image liveness model ([1] “Hence, we use two CNNs to learn local and holistic features respectively. The first CNN is end-to-end trained, and assign a score to each randomly extracted patch from a face image. We assign the face image with the average of scores. The second CNN estimates the depth map of the face image and provide the face image with a liveness score based on estimated depth map. The fusion of the scores of both CNNs lead to the final estimated class of live vs. spoof”; [4.2] “We use the weighted average of two streams' scores as the final score of our proposed method, where the weights are experimentally determined to be 1 and 0.4 for patch and depth-based streams, respectively”) It would have been obvious to one of ordinary skill as of the effective filing date of the claimed invention, to modify Wang/Ye to include the limitations of a fused image score, as taught by Atoum. Doing so would allow for a more robust score measurement that can potentially consider the advantages of two distinct processing modalities. The systems readily integrate, as Atoum represents an additional parallel processing pipeline which could coexist alongside Wang/Ye prior to final fusion. With respect to claim 15, Wang/Ye/Atoum teaches: The method as claimed in claim 14, wherein the determining, by the determination unit, the liveness of the subject is further based on a comparison of the final image liveness score with a pre-defined threshold score (Atoum, [3] “A face image or video clip is classified as spoof if its spoof-score is above a pre-defined threshold”) With respect to claims 29-30, they are rejected in line with the rejections of claims 16 and 14-15 above. Additional References Additionally cited references (see attached PTO-892) otherwise not relied upon above have been made of record in view of the manner in which they evidence the general state of the art. Inquiry Any inquiry concerning this communication or earlier communications from the examiner should be directed to NOAH WILLIAM BOYAR whose telephone number is (571)272-8392. The examiner can normally be reached 8:30 – 5:00 EST, Monday – Friday. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chan Park can be reached at 571-272-7409. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /NOAH W BOYAR/Examiner, Art Unit 2669 /CHAN S PARK/Supervisory Patent Examiner, Art Unit 2669
Read full office action

Prosecution Timeline

Sep 20, 2024
Application Filed
Jul 22, 2026
Non-Final Rejection mailed — §103, §112 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
100%
Grant Probability
99%
With Interview (+0.0%)
2y 2m (~3m remaining)
Median Time to Grant
Low
PTA Risk
Based on 2 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month