DETAILED ACTION
Response to Amendment
Claims 1-20 were previously pending. Applicant’s amendment filed August 21, 2026, has been entered in full. Claims 1, 10 and 19 are amended. No claims are added or cancelled. Claims 1-20 remain pending.
Response to Argument
Applicant’s summary of the interview (Remarks filed August 21, 2026, hereinafter Remarks: Page 8) is acknowledged.
Applicant traverses the previous rejections under 35 U.S.C. 103, arguing that the previously-cited Kim reference does not teach features of the amended claims (Remarks: Pages 8-12). Applicant first argues that Kim’s consistency check does not disclose the “determining … there is no associated face detection …” step added to the independent claims (Remarks: Page 9). Applicant second argues that the amended claim “requires generating the face detection based on the detected body location in response to determining that there is no associated face detection” (Remarks: Page 10; emphasis added). Applicant third argues that any modification of Kim to perform the added “determining” step would render Kim unsatisfactory for its intended purpose and thus would not be obvious (Remarks: Pages 10-11). Examiner respectfully disagrees.
Regarding Applicant’s second argument, Examiner notes that it is improper to import a specific ordering of steps from the specification where not required as a matter of logic or grammar in the claims – see, e.g., MPEP 2111.01, Subsection II. Claim 1 does not actually require that the “generating …” step is performed in response to the “determining …” step. For example, the claim does not expressly recite “in response to” (or equivalent) and neither logic nor grammar requires the “generating …” to be performed after the “determining …”.
Furthermore, Examiner notes that claim 3 further defines claim 1 and does specifically recite that “said generating … is performed in response to a determination that none of the one or more face detections are associated with the person detection” (emphasis added). This addition of a requirement that the generating step is “performed in response to” the determining step in the dependent claim is further evidence that the requirement is not present in the independent claim.
Thus, Applicant’s second argument is respectfully non-persuasive because it relies on a non-existent requirement that the “generating …” is performed “in response to” the “determining …”.
Regarding Applicant’s first argument, it focuses on an assertion that the consistency checking described at Sec. 3.2 does not fall within the scope of the claimed “determining … there is no associated face detection …” step. Applicant’s first argument does not mention or address Sec. 3.3 of Kim, which is the only explicit discussion of “association” within the reference. Kim seeks to track a specific person across multiple frames in a video. Kim detects people
p
in each frame. A person
p
is defined as a combination of a human body detection
g
and a face detection
d
. Kim attempts to associate a person
p
from one frame to the next, and so on, with a sequence of associated person detections forming a tracklet. At Page 43, Kim teaches that “Old tracklets are terminated whenever the new detection results cannot match to any existing tracklet for a certain frames, set to 15 in our experiments.” The decision to terminate a tracklet is a determination that there is no associated person detection
p
in the following 15 frames for a person detection
p
received in the current frame. As a person detection
p
includes a body detection
g
and a face detection
d
(see above), this is a determination that there is no associated face detection for the person detection received as an output of the machine-learned person detection model as required by the claimed invention I.e., there are no associated people (and, therefore, no associated faces) for the person detection in the current frame, so the tracklet is terminated. Thus, Kim’s teachings do fall within the scope of the claimed “determining …” step.
Regarding Applicant’s third argument, the hypothetical modification discussed in Applicant’s argument is unnecessary because, as discussed above, Kim does teach the claimed “determining …” step.
For at least these reasons, Applicant’s arguments are respectfully non-persuasive.
Admitted Prior Art
In the Office Action dated May 21, 2026, Examiner took Official Notice of facts in the following instance(s):
At Page 6:
“However, Examiner takes Official Notice that it is old and well-known in the art of image analysis to implement a video processing method as a computing system comprising: one or more processors; and a non-transitory, computer-readable medium comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform operations of the video processing method. Such computer implementation advantageously allows the video processing method to be performed quickly and efficiently over time.”
Regarding Official Notice, MPEP 2144.03(C) includes the following instructions:
“To adequately traverse such a finding, an applicant must specifically point out the supposed errors in the examiner’s action, which would include stating why the noticed fact is not considered to be common knowledge or well-known in the art.”
“A general allegation that the claims define a patentable invention without any reference to the examiner’s assertion of official notice would be inadequate.”
“If applicant does not traverse the examiner’s assertion of official notice or applicant’s traverse is not adequate, the examiner should clearly indicate in the next Office action that the common knowledge or well-known in the art statement is taken to be admitted prior art because applicant either failed to traverse the examiner’s assertion of official notice or that the traverse was inadequate. If the traverse was inadequate, the examiner should include an explanation as to why it was inadequate.”
In the reply filed August 21, 2026, Applicant generally alleges that the claims define a patentable invention without any reference to Examiner’s assertion of Official Notice, which is an inadequate traverse.
Therefore, as required by the MPEP, Examiner clearly indicates that the Official Notice statement(s) noted above is/are taken to be admitted prior art because Applicant either failed to traverse it/them or inadequately traversed it/them.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-2, 5-11, and 14-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over ‘Kim’ (“Face and Body Association for Video-based Face Recognition,” 2018).
Regarding claim 1, Examiner notes that the claim recites a method that is substantially the same as the method performed by the system of claim 10. The system of claim 10 is obvious over Kim (see below). Accordingly, claim 1 is also rejected under 35 U.S.C. 103 as being unpatentable over Kim for substantially the same reasons as claim 10.
Regarding claim 2, Examiner notes that the claim recites limitations that are substantially the same as limitations recited in claim 11. Kim teaches the invention of claim 11 (see below). Accordingly, claim 2 is also rejected under 35 U.S.C. 103 as being unpatentable over Kim for substantially the same reasons as claim 11.
Regarding claim 5, Examiner notes that the claim recites limitations that are substantially the same as limitations recited in claim 14. Kim teaches the invention of claim 14 (see below). Accordingly, claim 5 is also rejected under 35 U.S.C. 103 as being unpatentable over Kim for substantially the same reasons as claim 14.
Regarding claim 6, Examiner notes that the claim recites limitations that are substantially the same as limitations recited in claim 15. Kim teaches the invention of claim 15 (see below). Accordingly, claim 6 is also rejected under 35 U.S.C. 103 as being unpatentable over Kim for substantially the same reasons as claim 15.
Regarding claim 7, Examiner notes that the claim recites limitations that are substantially the same as limitations recited in claim 16. Kim teaches the invention of claim 16 (see below). Accordingly, claim 7 is also rejected under 35 U.S.C. 103 as being unpatentable over Kim for substantially the same reasons as claim 16.
Regarding claim 8, Examiner notes that the claim recites limitations that are substantially the same as limitations recited in claim 17. Kim teaches the invention of claim 17 (see below). Accordingly, claim 8 is also rejected under 35 U.S.C. 103 as being unpatentable over Kim for substantially the same reasons as claim 17.
Regarding claim 9, Examiner notes that the claim recites limitations that are substantially the same as limitations recited in claim 18. Kim teaches the invention of claim 18 (see below). Accordingly, claim 9 is also rejected under 35 U.S.C. 103 as being unpatentable over Kim for substantially the same reasons as claim 18.
Regarding claim 10, Kim teaches a computing system comprising:
one or more processors (see Note Regarding Computer below); and
a non-transitory, computer-readable medium comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform operations (see Note Regarding Computer below), the operations comprising:
obtaining an input image (e.g., Fig. 2, a frame in the probe video);
inputting the input image into a machine-learned person detection model (Section 3.2, OpenPose person detection model is applied to each frame) that is configured to detect human bodies depicted in images (Human body detection examples are shown in Fig. 4 [may be best seen in color]);
receiving a person detection as an output of the machine-learned person detection model (e.g., Fig. 4), wherein the person detection indicates a detected body location of a detected human body in the input image (Sec. 3.2, OpenPose returns locations of 18 body keypoints in the input image, which is then used to extract an upper-body region bounding box);
determining, by the one or more computing devices, there is no associated face detection for the person detection received as an output of the machine-learned person detection model (Examiner notes that it is improper to import a specific ordering of steps from the specification where not required as a matter of logic or grammar in the claims – MPEP 2111.01, Subsection II; Claim 10 does not require any specific order of this determining step and the following generating step; Sec. 3.3 of Kim describes a data association procedure that seeks to associate a person
p
(which includes a face detection
d
and a body detection
g
) with other persons detected in later frames; For clarity, Examiner notes that the claim uses “person detection” to refer to a body detection, while the Kim reference uses “person” to refer to a concatenation of face and body detections; At page 43, Tracklet initialization and termination, Kim discloses “Old tracklets are terminated whenever the new detection results cannot match to any existing tracklet for a certain frames, set to 15 in our experiments”; The decision to terminate a tracklet is a determination that there is no associated person detection
p
in the following 15 frames for a person detection
p
received in the current frame; As a person detection
p
includes a body detection
g
and a face detection
d
(see above), this is a determination that there is no associated face detection for the person detection received as an output of the machine-learned person detection model as required by the claimed invention); and
generating a face detection based at least in part on the detected body location of the detected human body provided by the person detection (e.g., Sec. 3.2, face detections from YOLO are matched with head positions from detected body location; e.g., Sec. 3.3, corresponding face
d
and body
g
are combined to generate a person
p
; The association of a specific face
d
with a specific person
p
and/or body
g
is within the scope of generating a face detection; Additionally, or alternatively, Sec. 3.3 further describes a data association procedure where a person in a current frame
t
+
1
is matched to a person in a previous frame
t
and such an association of a person
p
is also within the scope of generating a face detection at least because a specific person – including their face – from a previous frame is detected in a current frame), wherein the face detection indicates a face location in the input image of a face associated with the detected human body (e.g., Sec. 3.3,
x
,
y
,
w
,
h
in the face vector
d
included within person
p
).
Note Regarding Computer. While Kim’s use of convolutional neural networks and processing of digital videos certainly implies the use of a computer, Kim does not explicitly describe implementing its video processing method using a computer (i.e., computing device). In particular, Kim does not explicitly teach implementing its video processing method as a computing system comprising: one or more processors; and a non-transitory, computer-readable medium comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform operations of the video processing method.
However, it has been taken as admitted prior art that it is old and well-known in the art of image analysis to implement a video processing method as a computing system comprising: one or more processors; and a non-transitory, computer-readable medium comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform operations of the video processing method. Such computer implementation advantageously allows the video processing method to be performed quickly and efficiently over time.
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to implement the video processing method of Kim as a computing system comprising: one or more processors; and a non-transitory, computer-readable medium comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform operations of the video processing method with the reasonable expectation that this would advantageously allow the video processing method to be performed quickly and efficiently over time.
Therefore, it would have been obvious to one of ordinary skill in the art to combine the teachings of Kim to obtain the invention as specified in claim 10.
Regarding claim 11, Kim teaches the computing system of claim 10, and Kim further teaches that:
the person detection comprises one or more body pose landmarks respectively associated with one or more body components of the detected human body (e.g., Sec. 3.2, 18 keypoints indicating 2D joints of body skeleton); and
generating the face detection based at least in part on the detected body location of the detected human body comprises generating, based at least in part on the one or more body pose landmarks (Sec. 3.2, head joint from body pose is matched to head detection in order to associate them; See further explanation regarding generating the face detection in rejection of claim 10), one or more face pose landmarks respectively associated with one or more face components of the face associated with the detected human body (e.g., Sec. 3.3, Scoring function, 1st paragraph, center position
(
x
,
y
)
is included in face vector
d
; The center of a face is within the scope of a face pose landmark at least because it lies at the center of all of the components of the face and is used as a landmark to define the position of the face bounding box).
Regarding claim 14, Kim teaches the computing system of claim 10, and Kim further teaches:
determining whether the face detection satisfies one or more quality criteria (e.g., page 43, Tracklet filtering, face detections with small bounding boxes are considered low quality – i.e., determined not to satisfy one or more quality criteria); and
when it is determined that the face detection does not satisfy the one or more quality criteria, discarding the face detection (e.g., page 43, Tracklet filtering, small face detections are “filtered out” – i.e., discarded).
Regarding claim 15, Kim teaches the computing system of claim 14, and Kim further teaches that the one or more quality criteria comprise one or more of: a tilt angle criterion, a yaw angle criterion, a roll angle criterion, a blur criterion, an eyes open criterion, and a recognizability criterion (e.g., page 43, Tracklet filtering, small, “low resolution faces cannot help in recognition but only add noise”; Therefore, the size criterion applied by Kim falls within the scope of a recognizability criterion).
Regarding claim 16, Kim teaches the computing system of claim 10, the instructions further comprising: generating a whole person detection that associates the face detection with the person detection (Sec. 3.3, person
p
).
Regarding claim 17, Kim teaches the computing system of claim 16, and further teaches: carrying the whole person detection forward to a subsequent image (Sec. 3.3, data association is performed to carry whole person detection
p
forward from current frame/image
t
to subsequent frame/image
t
+
1
) to perform whole person tracking over plural image frames (e.g., Sec. 3.3, person is tracked over whole video via association from frame to frame).
Regarding claim 18, Kim teaches the computing system of claim 10, and Kim further teaches: providing the face detection to one or both of a machine-learned facial recognition model for facial recognition (e.g., Fig. 2, Face Recognition; e.g., Sec. 3, the faces are associated and tracked through the video to obtain a large set of images of the same subject, which are used to create a face representation for use in a machine-learned facial recognition model; e.g., Page 44, Experiment setup, ResFace-101 model) or a machine-learned gaze detection model for gaze detection.
Regarding claim 19, Examiner notes that the claim recites a non-transitory, computer-readable medium that is substantially the same as the non-transitory, computer-readable medium included in the system of claim 10. The system of claim 10 is obvious over Kim (see above). Accordingly, claim 19 is also rejected under 35 U.S.C. 103 as being unpatentable over Kim for substantially the same reasons as claim 10.
Regarding claim 20, Examiner notes that the claim recites limitations that are substantially the same as limitations recited in claim 11. Kim teaches the invention of claim 11 (see above). Accordingly, claim 20 is also rejected under 35 U.S.C. 103 as being unpatentable over Kim for substantially the same reasons as claim 11.
Allowable Subject Matter
Claims 3-4 and 12-13 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to GEOFFREY E SUMMERS whose telephone number is (571)272-9915. The examiner can normally be reached Monday-Friday, 7:00 AM to 3:30 PM ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chan Park can be reached at (571) 272-7409. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/GEOFFREY E SUMMERS/Examiner, Art Unit 2669