DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f):
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f). The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f). The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f), because the claim limitations use a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitations are: “a receiving unit”; “a decision unit”; “a determination unit” of claim 9.
Because the claim limitations are being interpreted under 35 U.S.C. 112(f), they are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have these limitations interpreted under 35 U.S.C. 112(f) applicant may: (1) amend the claim limitations to avoid them being interpreted under 35 U.S.C. 112(f) (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitations recite sufficient structure to perform the claimed function so as to avoid them being interpreted under 35 U.S.C. 112(f).
Claim Objections
Claims 7 and 15 are objected to for the following informalities: “a set of sensor-based liveness score” reads as a typographical error for “a set of sensor-based liveness scores”.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-16 are rejected under 35 U.S.C. 101 because they are directed to ineligible patent subject matter. The claims are directed to the Abstract Idea groupings of mental processes under MPEP § 2106.04(a)(2)(III) and mathematical calculations under MPEP § 2106.04(a)(2)(I). These are judicial exceptions under Step 2A, Prong One of the framework established by the cases of Alice Corp. Pty. Ltd. v. CLS Bank Int'l, 573 U.S. 208, 216, 110 USPQ2d 1976, 1980 (2014) and Mayo Collaborative Servs. v. Prometheus Labs., Inc., 566 U.S. 66, 71, 101 USPQ2d 1961, 1965 (2012). See MPEP § 2106.04(II).
PNG
media_image1.png
200
400
media_image1.png
Greyscale
Step 1: The claims in question are directed to a method and system (machine) for “determining liveness”. Machines and processes are statutory categories. See MPEP 2106.03(I), “A machine is a "concrete thing, consisting of parts, or of certain devices and combination of devices." Digitech, 758 F.3d at 1348-49, 111 USPQ2d at 1719 (quoting Burr v. Duryee, 68 U.S. 531, 570, 17 L. Ed. 650, 657 (1863)). This category "includes every mechanical device or combination of mechanical powers and devices to perform some function and produce a certain effect or result." Nuijten, 500 F.3d at 1355, 84 USPQ2d at 1501 (quoting Corning v. Burden, 56 U.S. 252, 267, 14 L. Ed. 683, 690 (1854))”; See MPEP 2106.03(I), “NTP, Inc. v. Research in Motion, Ltd., 418 F.3d 1282, 1316, 75 USPQ2d 1763, 1791 (Fed. Cir. 2005) ("[A] process is a series of acts.") (quoting Minton v. Natl. Ass’n. of Securities Dealers, 336 F.3d 1373, 1378, 67 USPQ2d 1614, 1681 (Fed. Cir. 2003)). As defined in 35 U.S.C. 100(b), the term "process" is synonymous with "method."”. (Step 1: Yes).
Step 2A, Prong One: As explained in MPEP 2106.04(II), a claim “recites” a judicial
exception when the judicial exception is “set forth” or “described” in the claim. Here, each
claim recites or depends upon the mental processes of performing sanity checks, analyzing, grouping, identifying, and determining liveness, as well as various acts of data input. (Claim 1, “analyzing…grouping…identifying…determining; Claim 2, “performing a set of sanity checks…”; Claim 4, “comparing…determining…”; Claim 6, “providing…”). The claims additionally recite or depend upon mathematical calculations (Claim 1, “analyzing…”; Claim 2, “normalizing…sampling…generating, by the CNN-LSTM based model…”; Claim 3, “voting-based mechanism, a weighted average based mechanism, and an auxiliary network-based mechanism…”; Claim 4, “determining…the subject as live…in an event…the score…is greater than the corresponding pre-defined threshold score”; Claim 5, “determining…the subject as live…in an event the final liveness score is greater than a preset threshold score…”)
The claims are recited at a high level of generality and lack any specifics precluding such
an analysis from being interpreted under the mental processes grouping of “practically performed
in the mind” (see also MPEP § 2106.04(a)(2) identifying how e.g. a use of pen and paper, a ruler,
or a computer as a tool (to assist in visually/mentally analyzing/observing acquired
images/video) fails to preclude such an interpretation under the mental processes judicial
exception). Activities such as “by the [unit]” therefore may be performed mentally, even if they may require an additional computer tool. Similarly, basic digital data input and output do not elevate these claims past a mental process.
Regarding artificial intelligence, to the extent it is implicated, the recitations are comparable to Claim 2 of Example 47 of the July 2024 PEG regarding subject matter eligibility (https://www.uspto.gov/sites/default/files/documents/2024-AISMEUpdateExamples47-49.pdf). As stated therein, an artificial intelligence’s analyses, detections, and reinforcement learnings may be practically performed in the human mind. To the extent mathematical calculations are required to operate and train the artificial intelligence in image/video/sensor analysis, the separate judicial exception is also implicated.
As such, the usage of a computer to score and determine liveness does not elevate these claims beyond a mental process and/or mathematical calculation. (Step 2A, Prong One: Yes).
Step 2A, Prong Two: If Prong One of Step 2A is met, the examiner must consider (1)
whether there are any ‘additional elements’ recited in the claim beyond the judicial exception,
and (2) evaluate those additional elements individually and in combination to determine whether
the claim as a whole integrates the exception into a practical application. See MPEP §
2106.04(d).
Limitations the courts have found indicative of integration include: an improvement in
the functioning of a computer, or an improvement to other technology or technical field, as
discussed in MPEP §§ 2106.04(d)(1) and 2106.05(a); applying or using a judicial exception to
effect a particular treatment or prophylaxis for a disease or medical condition, as discussed in
MPEP § 2106.04(d)(2); implementing a judicial exception with, or using a judicial exception in
conjunction with, a particular machine or manufacture that is integral to the claim, as discussed
in MPEP § 2106.05(b); effecting a transformation or reduction of a particular article to a
different state or thing, as discussed in MPEP § 2106.05(c); and applying or using the judicial
exception in some other meaningful way beyond generally linking the use of the judicial
exception to a particular technological environment, such that the claim as a whole is more than
a drafting effort designed to monopolize the exception, as discussed in MPEP § 2106.05(e).
Limitations that the courts have found non-indicative of integration include: merely
reciting the words "apply it" (or an equivalent) with the judicial exception, or merely including
instructions to implement an abstract idea on a computer, or merely using a computer as a tool to
perform an abstract idea, as discussed in MPEP § 2106.05(f); adding insignificant extra-solution
activity to the judicial exception, as discussed in MPEP § 2106.05(g); and generally linking the
use of a judicial exception to a particular technological environment or field of use, as discussed
in MPEP § 2106.05(h).
As an additional note, ‘additional elements’ are generally limitations excluded from
interpretation under the Abstract Idea groupings, and may comprise portions of limitations
otherwise identified as falling under those Abstract Idea groupings of the 2019 PEG (e.g. any
‘determination’ that may be made mentally by a user, neural network and/or generic computer hardware is considered under the ‘apply it’ considerations of 2106.05(f)). Any ‘providing’/outputting broadly, and ‘collection/input’ of data (i.e “receiving by a receiving unit”), also fail(s) to integrate at least in view of MPEP 2106.05(g) (extra-solution data gathering/output) and/or 2106.05(h) as ‘generally linking’ the exception to a field of use involving machine learning and/or data so acquired (e.g. the use of a “receiving unit” to acquire score data). The same determination holds for dependent claims that serve to limit the collection/output of data/images (by means of what is collected based on recited conditions) and/or introduce limitations generally linking to a field of use.
None of the instant claims appear to explicitly/clearly capture/recite any disclosed
improvement in technology (see MPEP 2106.05(a), with note that ‘functioning of a computer’
concerns functions integral to the way a computer operates and not ‘functions’ that a generic
computer can be programmed/adapted to perform (see also 2106.05(f))) and any ‘additional
elements’, even when considered in combination, fail to integrate at Prong Two of Step 2A
accordingly. Integration in view of subsection (a) requires an identification of the manner in
which the improvement is achieved, to be explicitly and specifically recited in the claims, as
‘additional elements’ precluded from interpretation under any of the Abstract Idea groupings
(since the improvement cannot be to the exception itself). With reference to MPEP 2106.05(a):
It is important to note, the judicial exception alone cannot provide the improvement. The improvement can be provided by one or more additional elements. See the discussion of Diamond v. Diehr, 450 U.S. 175, 187 and 191-92, 209 USPQ 1, 10 (1981))
As applicable here, additional limitations not directed to a judicial exception fail to
integrate at Prong Two of Step 2A. Claim 1 recites various high-level “unit[s]” and “system[s]”; Claim 2 recites a “Convolutional Neural Network (CNN)-Long Short-Term Memory (LSTM) based model”. The incorporation of conventional computer and machine-learning systems does little more than generally link the judicial exceptions of mental processes to a field-of-use and technological environment. See MPEP §§ 2106.05(h); 2106.05(f).
Claim 1 recites “receiving, by a receiving unit…”; Claim 2 recites “receiving a sensor data…”. These limitations constitute insignificant extra-solution activity under MPEP § 2106.05(g). Specifically, the limitations amount to no more than necessary data inputting, under rationale 3 of MPEP § 2106.05(g).
Even when viewed in combination, any additional elements present do not integrate the
recited judicial exception into a practical application (Step 2A, Prong Two: No), and the claims
are directed to the judicial exception. (Revised Step 2A: Yes → Step 2B).
Step 2B: If Prong Two of Step 2A is not met, the examiner must consider whether the
claim as a whole amounts to ‘significantly more’ than the recited exception, i.e., whether any
‘additional element’, or combination of additional elements, adds an inventive concept to the
claim. The considerations of Step 2A Prong 2 and Step 2B overlap, but differ in that 2B also
requires considering whether the claims feature any “specific limitation(s) other than what is
well-understood, routine, conventional activity in the field” (WURC) (MPEP § 2106.05(d)).
Such a limitation if specifically recited however, must still be excluded from interpretation under
any of the Abstract Idea groupings. Step 2B further requires a re-evaluation of any additional
elements drawn to extra-solution activity in Step 2A (e.g. gathering data, rendering output) – however no limitations appear directed to any novel collection or output generation per se. Limitations not indicative of an inventive concept/‘significantly more’ include those that are not specifically recited (instead recited at a high level of generality), those that are established as
WURC (a plurality of cited references serves to evidence the WURC nature of ‘analysis’ based at least in part on corroborating/additional ground data), and/or those that are not ‘additional
elements’ by nature of their analysis at Prong One of Step 2A (i.e. directed to the exception – see
above re. deciding that a second acquisition may be advantageous/desired). The July 2024 PEG
describes that an improvement/ inventive concept (for ‘significantly more’ determination(s))
cannot be to the judicial exception itself. As additionally recited by Claim 2, machine learning is understood to encompass a plurality of WURC machine-learning models, including “Convolutional Neural Network (CNN)-Long Short-Term Memory (LSTM) based model[s]”.
The claims in question recite little beyond those limitations recited at a high level
of generality and falling under e.g. the mental processes and mathematical calculation Abstract Idea groupings, and would monopolize the exceptions accordingly. The additional limitations of computer processing and machine-learning as recited are WURC, as evidenced by the body of prior art cited by the examiner in this office action. (Step 2B: No).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 3-5, 9, and 11-13 are rejected under 35 U.S.C. 103 as being unpatentable over Suneja (US 20230058259 A1) (Hereinafter, “Suneja”) in view of Atoum et. al Face Anti-Spoofing Using Patch and Depth-Based CNNs (Hereinafter, “Atoum”) and Mittal et. al Static-dynamic features and hybrid deep learning models based spoof detection system for ASV (Hereinafter, “Mittal”)
With respect to claim 1, Suneja teaches:
A method for determining liveness of a subject, the method ([Abstract]) comprising:
receiving, by a receiving unit from an image processing system ([0035]), a final image Fig. 3, 302)
receiving by the receiving unit from a video processing system ([0026] “The user devices may be used for inputting, processing, and displaying information…in other embodiments, the user device may include a digital camera that is integral with the computing device, such as a camera on a smartphone, laptop, or tablet”, establishing that image/video sources may come from multiple sources; [0028]-[0029]; [0035]), a video liveness score (Fig. 3, 302)
receiving by the receiving unit from a sensor-data processing system ([0035]), a sensor-based Fig. 3, 302. Read in line with page 11 of the claimed invention’s specification, “via an accelerometer sensor, a gyroscope sensor and such other sensor(s) associated with the electronic device”)
analyzing by a decision unit the final image Fig. 3, 306), the video liveness score (Fig. 3, 304, 308), and the sensor-based Fig. 3, 310, 312)
grouping by the decision unit the final image Fig. 3, 312; Table 1; [0062]-[0070])
identifying by the decision unit a set of liveness detection mechanisms based on the grouping ([0062]-[0070] discussing voting mechanism logic; Table 1 showing voting mechanism)
determining by a determination unit the liveness of the subject based on the set of liveness mechanisms ([0062] “Multiple scenarios may lead to a final determination that a user’s identity has been authenticated. Likewise, multiple scenarios may lead to the final determination that a user’s identity has not been authenticated; [0063]-[0070] noting additional factors beyond the vote alone that could change outcome, “If the face ID match score is “fail”, but the probability of a match between the video face and the photo ID face is within a predetermined range (e.g., between 80% and 96%), this “fail” for the face ID match score may be offset by other factors/conditions discovered during analysis (e.g., age of user is above 60 years old, or user was not facing the camera through much of the video).”)
Suneja does not explicitly teach:
a final image liveness score
a sensor-based liveness score
(As Suneja describes these scores as being used for authentication generally, rather than determination of a non-live attack specifically)
However, Atoum, in the same field of endeavor of anti-spoofing, teaches:
a final image liveness score (Fig. 1; Fig. 2)
It would have been obvious to one of ordinary skill in the art as of the effective filing date of the claimed invention to modify Suneja to include the limitations of a final image liveness score. Doing so would have the advantage of providing an additional means to prevent a non-live attack. The systems readily integrate, as Suneja is configured to capture a plurality of images from multiple sources. The score of Atoum could further be used as another factor within the broader voting system of Suneja as well.
Suneja/Atoum do not explicitly teach:
a sensor-based liveness score
However, Mittal, in the same field of endeavor of anti-spoofing, teaches:
a sensor-based liveness score ([Introduction] “Replay attacks are the one of the easiest form of attacks in which spoofed speech is the recorded voice signal of targeted user”; [Experimental setups] “It finds out the probability or score for an utterance between zero and one…where FAR is ratio of number of spoofed utterances having score more than or equal to the threshold Ψ to the total number of spoofed utterances and FRR is ration of the number of bonafide utterances having score less than the threshold Ψ value to the total number of bonafide utterances”)
It would have been obvious to one of ordinary skill in the art as of the effective filing date of the claimed invention, to modify Suneja/Atoum to include the limitations of a sensor-based liveness score. Doing so would provide an additional metric by which to prevent unauthenticated attack. The systems readily integrate, as Suneja/Atoum is already configured to collect and analyze voice data. Integrating a further liveness check would increase the robustness of the overall score, and/or otherwise provide for an additional metric in the voting system.
With respect to claim 3, Suneja/Atoum/Mittal teaches:
The method as claimed in claim 1, wherein the set of liveness detection mechanisms comprises at least one of a voting-based mechanism (Suneja, Table 1; Mittal, [Voting protocol based two-level ASV system (System_1)]), a weighted average based mechanism (Atoum, [4.2] “We use the weighted average of two streams' scores as the final score of our proposed method, where the weights are experimentally determined to be 1 and 0.4 for patch and depth-based streams, respectively), and an auxiliary network-based mechanism
With respect to claim 4, Suneja/Atoum/Mittal teaches:
The method as claimed in claim 3, wherein the determining, by the determination unit, the liveness of the subject based on the voting-based mechanism comprises:
comparing, by the determination unit, the final image liveness score, the video liveness score, and the sensor-based liveness score with a corresponding pre-defined threshold score (Suneja, [0061] “In some embodiments, the modules may output a score of either “pass” or “fail”. These scores may be based on thresholds of probabilities. For example, a score of “pass” for the face ID match could be anything above a threshold of 97% match. In other examples, the threshold may be set at 94% match or 96% match”)
determining, by the determination unit, the subject as a live subject in an event each of the final image liveness score, the video liveness score, and the sensor-based liveness score is greater than the corresponding pre-defined threshold (Suneja, [0061]; Suneja, Table 1)
With respect to claim 5, Suneja/Atoum/Mittal teaches:
The method as claimed in claim 3, wherein the determining, by the determination unit, the liveness of the subject based on the weighted average based mechanism comprises:
determining, by the determination unit, a final liveness score based on a weighted average of each of the final image liveness score, the video liveness score, and the sensor-based liveness score (Suneja, [0061]-[0070]. While not explicitly describing a “weighted average”, the pass/fail voting mechanism can be considered broadly to be a type of average, as it takes the mode of a binary determination. Suneja further describes adjusting the thresholds of each metric, which could also be construed as “weighting”; Suneja, [0070] “For example, the threshold for “pass” may be raised for one or more components (e.g., face ID match, liveness, etc)”; Suneja, [0061] “These scores may be based on thresholds of probabilities. For example, a score of “pass” for the face ID match could be anything above a threshold of 97% match. In other examples, the threshold may be set at 94% match or 96% match”; Atoum, [4.2] “We use the weighted average of two streams' scores as the final score of our proposed method, where the weights are experimentally determined to be 1 and 0.4 for patch and depth-based streams, respectively”)
determining, by the determination unit, the subject as a live subject in an event the final liveness score is greater than a preset threshold score (Suneja, [0061]-[0070]; Atoum, [3] “A face image or video clip is classified as spoof if its spoof-score is above a pre-defined threshold”)
To the extent a weighted average is not explicitly taught by Suneja, it would have been obvious to one of ordinary skill in the art as of the effective filing date of the claimed invention, to modify Suneja/Mittal to include the limitations of a weighted average, as taught by Atoum. Doing so would allow a user to adjust the importance of each stream of data to suit individual preferences or risk-specific vectors. Additionally or alternatively, one of ordinary skill in the art could modify each percentage match score generated by Suneja/Mittal to generate a final fused score, as taught by Atoum, and then determine if such a final fused score is greater than a pre-determined threshold. The systems readily integrate, as both Suneja/Mittal and Atoum seek to fuse multiple sources of information.
With respect to claim 9, it is functionally parallel to claim 1, but claimed as a system configured to perform the method of claim 1. Suneja teaches the general high-level hardware configured to perform the same (Suneja, Fig. 1; Suneja, Fig. 2). Accordingly, the claim is rejected in line with the analysis above.
With respect to claims 11-13, they are rejected in line with the rejections of claims 9 and 3-5 above.
Claims 2 and 10 are rejected under 35 U.S.C. 103 as being unpatentable over Suneja/Atoum/Mittal in view of Zhao et. al A lighten CNN-LSTM model for speaker verification on embedded devices (Hereinafter, “Zhao”)
With respect to claim 2, Suneja/Atoum/Mittal teaches:
The method as claimed in claim 1, wherein the sensor-based liveness score is generated by the sensor-data processing system based on:
receiving a sensor data from a set of sensors for a predefined time duration (Suneja, [0029]; Mittal, [Proposed method] “Speech signals taken from the dataset…”)
performing a set of sanity checks to generate a target sensor data (Suneja, [0029] “Voice module 122 may be used to process audio sounds (e.g., user's voice) captured in the live video to determine a voice score indicating whether the voice captured by live video is a sufficient sample for comparing to PEP voice data”)
normalizing the target sensor data to generate a normalized sensor data (Suneja, Fig. 8, “Mel Frequency Cepstral Coefficients”, “Dynamic Time Warping”; Suneja, [0049] “Dynamic time warping (DTW) may include finding an optimal assignment path”. The examiner notes that a decision has yet to be reached at this point in the module as to whether the data is “the target sensor data”, but understands that when this data ends up becoming “the target” sensor data, the processing will have been performed on “the target” sensor data, thus satisfying the language of the claim)
sampling a set of data points from the normalized sensor data (Suneja, [0029]; Suneja, Fig. 8, “Vector Quantizer”; Mittal, [Feature extraction using CQCC features] “resample (): This function converts the geometrically spaced bins provided by CQT into linearly spaced bins. Bins are converted into linear space to make the signal compatible with Discrete Cosine Transformation (DCT).”)
providing the sampled set of data points to a Convolutional Neural Network (CNN)-Long Short-Term Memory (LSTM) based model (Suneja, ([0048] “In some embodiments, a vector quantizer may be used in addition to or in place of a CNN”; Mittal, [Related works])
PNG
media_image2.png
322
775
media_image2.png
Greyscale
generating, by the CNN-LSTM basedSuneja, Fig. 8; Suneja, [0048]), the sensor-based liveness score based on the sampled set of data points for classifying a specific subject and a non-live specific subject based on a set threshold (Suneja, [0061] “The disclosed system and method may include machine learning induction scoring as part of video authentication. As discussed above with respect to the embodiment of FIG. 3, the scores generated by each module may be considered in the determination of whether or not a user's identity is authenticated. In some embodiments, the modules may output a score of either “pass” or “fail”. These scores may be based on thresholds of probabilities. For example, a score of “pass” for the face ID match could be anything above a threshold of 97% match. In other examples, the threshold may be set at 94% match or 96% match.”)
Adopting a narrower view in the interest of compact prosecution, the examiner will further evidence the obvious nature of the following limitation:
a CNN-LSTM based model
(Were Applicant to argue that Mittal does not constitute a CNN-LSTM model, as CNN and LSTM are not explicitly labeled as one single model)
Zhao, in the same field of endeavor of audio authentication, teaches:
providing the sampled set of data points to a Convolutional Neural Network (CNN)-Long Short-Term Memory (LSTM) based model ([1] “We use a 3-stage training method to train the lighten CNN-LSTM network and achieve the best Equal Error Rate (EER)”; [4.5] “our CNN-LSTM model”)
generating by the CNN-LSTM based model, the sensor-based liveness score ([3])
It would have been obvious to one of ordinary skill in the art as of the effective filing date of the claimed invention to modify Suneja/Atoum/Mittal to include the limitation of a CNN-LSTM model. Doing so would have the advantage of analyzing temporal relationships for the purposes of verification. The systems readily integrate, as Suneja/Atoum/Mittal already teach a CNN/LSTM system.
With respect to claim 10, it is rejected in line with the rejections of claims 9 and 2 above.
Claims 6-8 and 14-16 are rejected under 35 U.S.C. 103 as being unpatentable over Suneja/Atoum/Mittal in view of Oh et. al (US 20200134427 A1) (Hereinafter, “Oh”)
With respect to claim 6, Suneja/Atoum/Mittal teaches the method of claim 3.
Suneja/Atoum/Mittal does not explicitly teach the further limitations of claim 6. However, Oh, in the same field of endeavor of human recognition neural networks, teaches:
The method as claimed in claim 3, wherein the determining, by the determination unit, the liveness of the subject based on the auxiliary network-based mechanism ([0054] “FIG. 4 is a view for explaining a knowledge distillation process between a first neural network model and a second neural network model, according to an example embodiment. FIG. 4 shows unlabeled data 401, labeled data 403, a teacher model 410, and a student model 430”; [0061] “The student model 430 may be trained through supervised learning”; [0074]; “Auxiliary neural-network based mechanism” read in line with page 35 of the claimed invention’s specification and in accordance with plain technical meaning) comprises:
providing, by the decision unit to the neural network-based model, at least one of the final image liveness score, the video liveness score, and the sensor-based liveness core ([0070] “Object data refers to data to be recognized based on the second neural network model that has been trained through the above-described process, and may include, for example but not limited to, image data, video data, voice data, time-series data, sensor data, or various combinations thereof”; [0074] “For example, the first neural network model may detect or recognize the face of a person included in the input data from the input data of the image form. In this case, the second neural network model may be trained based on training data generated by refining a result of detecting or recognizing the face of the person by the first neural network model in correspondence with the input data. As another example, the first neural network model may convert voice data into text data”)
determining, by the neural network-based model, the subject Fig. 4; Fig. 6)
PNG
media_image3.png
854
1056
media_image3.png
Greyscale
It would have been obvious to one of ordinary skill in the art as of the effective filing date of the claimed invention, to modify Suneja/Atoum/Mittal to include the limitations of auxiliary neural-network based prediction methods, as taught by Oh. Doing so would have the advantage of providing an additional score for comparison which could then be fused or voted under the methods of Suneja/Atoum, or simply left alone by itself as additional information to be considered for its own sake The examiner notes that under the dependencies of claim 3 and claim 1, claim 6 is not precluded from being interpreted as merely an additional liveness mechanism in the system. Such that these systems could simply coexist, with the individual underlying processing mechanisms of Suneja/Atoum/Mittal and Oh otherwise remaining intact.
Oh is configured to make predictions based on a broad class of data, including image, video, and voice. Suneja/Atoum/Mittal collects the necessary information, which could be processed to predictable success under the disclosure of Oh. While not specifically directed to “liveness” determinations, Oh seeks to recognize humans based on multiple classes of information, which could reasonably include spoof/live data. (Oh, [0074]). Accordingly, the systems readily integrate as complements.
With respect to claim 7, Suneja/Atoum/Mittal/Oh teaches:
The method as claimed in claim 6, wherein the neural network-based model is trained based on a set of final image liveness scores, a set of video liveness scores, and a set of sensor-based liveness score, to determine a target subject in a target video as one of a live target subject and a non-live target subject (Oh, [0070])
With respect to claim 8, Suneja/Atoum/Mittal/Oh teaches:
The method as claimed in claim 6, wherein the neural network-based model is trained using one or more supervised learning techniques (Oh, [0061] “The student model 430 may be trained through supervised learning”. Noting that under the plain language of claim 6, “a/the neural network-based model” does not necessarily need to be the “auxiliary network” itself)
With respect to claims 14-16, they are rejected in line with the rejections of claims 9 and 6-8 above.
Additional References
Additionally cited references (see attached PTO-892) otherwise not relied upon above have been made of record in view of the manner in which they evidence the general state of the art.
Inquiry
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NOAH WILLIAM BOYAR whose telephone number is (571)272-8392. The examiner can normally be reached 8:30 – 5:00 EST, Monday – Friday.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chan Park can be reached at 571-272-7409. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/NOAH W BOYAR/Examiner, Art Unit 2669
/CHAN S PARK/Supervisory Patent Examiner, Art Unit 2669