DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgement is made of Applicant’s claim of priority from Provisional Application No. 63565987, filed March 15, 2024.
Claim Objections
Claims 10 and 20 are objected to because of the following informalities: “changing in a field of view of the image” should read “changing a field of view of the image”. Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112:
The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.\
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-3, 7, 9, 11-13, 17 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Bradley et al. (US 2024/0161540 A1, filed November 8, 2023) in view of Gernoth et al. (US 2019/0042833 A1).
Regarding claim 1, Bradley teaches an apparatus to estimate a head pose, the apparatus comprising:
interface circuitry (Bradley, Para. [0022], an input/output (I/O) device interface);
instructions (Bradley, Para. [0088], computer program instructions); and
at least one processor circuit to be programmed by the instructions (Bradley, Para. [0088], computer program instructions provided to a processor) to:
identify a plurality of facial landmarks in a plurality of images (Bradley, Para. [0031], landmark detection engine 202 determines one or more landmarks for a given input image. In various embodiments, a landmark is a distinguishing characteristic or point of interest in an image. In various embodiments, landmarks are specified as a 2D coordinate (e.g., an x-y coordinate) on an image. Examples of facial landmarks include the inner or outer corners of the eyes, the inner or outer corners of the mouth, the inner or outer corners of the eyebrows, the tip of the nose, the tips of the ears, the location of the nostrils, the location of the chin, the corners or tips of other facial marks or points, or the like);
identify initial image data based on the plurality of facial landmarks (Bradley, Para. [0039], input image(s) 120 includes images divided into training datasets, testing datasets, or the like. In other embodiments, the training data set is divided into minibatches, which include small, non-overlapping subsets of the dataset. In some embodiments, input image(s) include labeled images, high-definition images (e.g., resolution above 1000×1000 pixels), images with indoor or outdoor footage, images with different lighting and facial expressions, images with variations in poses and facial expressions, images of faces with occlusions, images labelled or re-labelled with a set of landmarks (e.g., 68-point landmarks, 70-point landmarks, or dense landmarks with 50000-point landmarks) (i.e., image data based on the plurality of facial landmarks), video clips with one or more frames annotated with a set of landmarks (e.g., 68 landmarks), images with variations in resolution, videos with archive grayscale footage, or the like);
augment the initial image data with a transformation operation (Bradley, Para. [0039], the landmark detection engine 202 is trained with data augmentations that makes the resulting landmark detection engine 202 more robust landmark detection engine 202 in an end-to-end fashion using a Gaussian negative log likelihood loss function. Para. [0040], a data-augmentation parameter that applies transformations to features inputted into landmark detection engine 202 (e.g., scaling, translating, rotating, shearing, shifting, and/or otherwise transforming an image) (i.e., augment the image data with a transformation operation)); and
train a neural network based on the initial image data and the augmented image data (Bradley, Para. [0039], the landmark detection engine 202 is trained with data augmentations that makes the resulting landmark detection engine 202 more robust landmark detection engine 202 in an end-to-end fashion using a Gaussian negative log likelihood loss function (i.e., train based on augmented image). In each training iteration, landmark detection engine 202 receives an input image (i.e., train a neural network based on the initial image data) from storage 104 and one or more position queries associated with the canonical shape C. Landmark detection engine 202 processes the input image and position queries to generate a set of 2D positions on the input image corresponding to the desired landmarks) to:
infer a confidence metric (Bradley, Para. [0039], In addition to a set of 2D positions on the input image, landmark detection engine 202 also generates a scalar confidence value for each landmark (i.e., infer a confidence metric)).
Although Bradley teaches training a neural network to infer 2D positions and confidence values (Bradley, Para. [0039]), Bradley does not explicitly teach training the neural network to “infer three-dimensional model parameters”. However, in an analogous field of endeavor, Gernoth teaches when the presence of one or more faces are detected in a region, the predictions generated by decoder process includes assessments (e.g., determinations) of one or more properties of the detected face(s) in the region. The assessed properties may include a position of the face relative to a center of the region (e.g., offset of the face from the center of the region), a pose of the face in the region, and a distance between the face in the region and the camera. Pose of the face may include pitch, yaw, and/or roll of the face (i.e., three-dimensional model parameters). The assessed properties may be included in output data along with the decision on the presence of one or more faces in image input (Gernoth, Para. [0056]).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the apparatus of Bradley with the teachings of Gernoth by including training the neural network to infer three-dimensional model parameters (i.e., pose of the face). One having ordinary skill in the art would have been motivated to combine these references because doing so would allow for determining properties of a detected face, as recognized by Gernoth. Thus, the claimed invention would have been obvious to one having ordinary skill in the art before the effective filing date.
Regarding claim 2, Bradley in view of Gernoth teaches the apparatus of claim 1, wherein one or more of the at least one processor circuit is to:
perform an analysis of an input image using the neural network (Bradley, Para. [0039], landmark detection engine 202 processes the input image and position queries to generate a set of 2D positions on the input image corresponding to the desired landmarks. In addition to a set of 2D positions on the input image, landmark detection engine 202 also generates a scalar confidence value for each landmark); and
output, based on the analysis:
three-dimensional model parameters for the input image (Gernoth, Para. [0056], when the presence of one or more faces are detected in a region, the predictions generated by decoder process includes assessments (e.g., determinations) of one or more properties of the detected face(s) in the region. The assessed properties may include a position of the face relative to a center of the region (e.g., offset of the face from the center of the region), a pose of the face in the region, and a distance between the face in the region and the camera. Pose of the face may include pitch, yaw, and/or roll of the face (i.e., three-dimensional model parameters). The assessed properties may be included in output data along with the decision on the presence of one or more faces in image input); and
a confidence metric for the input image (Bradley, Para. [0039], landmark detection engine 202 processes the input image and position queries to generate a set of 2D positions on the input image corresponding to the desired landmarks. In addition to a set of 2D positions on the input image, landmark detection engine 202 also generates a scalar confidence value for each landmark).
The proposed combination as well as the motivation for combining the Bradley and Gernoth references presented in the rejection of Claim 1, apply to Claim 2 and are incorporated herein by reference. Thus, the apparatus recited in Claim 2 is met by Bradley in view of Gernoth.
Regarding claim 3, Bradley in view of Gernoth teaches the apparatus of claim 2, wherein one or more of the at least one processor circuit is to estimate a head pose based on the three-dimensional model parameters for the input image when the confidence metric for the input image satisfies a threshold (Gernoth, Para. [0055], the multiple predictions may be used to determine a confidence that a face, or a portion of a face, is present in each region (e.g., the predictions may be used to rank confidence for the regions). The region(s) with the highest confidence for the detected face may then be selected as the region used in training process 200. Para. [0056], when the presence of one or more faces are detected in a region (i.e., confidence metric satisfies a threshold), the predictions generated by decoder process includes assessments (e.g., determinations) of one or more properties of the detected face(s) in the region. The assessed properties may include a position of the face relative to a center of the region (e.g., offset of the face from the center of the region), a pose of the face in the region, and a distance between the face in the region and the camera. Pose of the face may include pitch, yaw, and/or roll of the face (i.e., three-dimensional model parameters). The assessed properties (i.e., head pose) may be included in output data along with the decision on the presence of one or more faces in image input).
The proposed combination as well as the motivation for combining the Bradley and Gernoth references presented in the rejection of Claim 1, apply to Claim 3 and are incorporated herein by reference. Thus, the apparatus recited in Claim 3 is met by Bradley in view of Gernoth.
Regarding claim 7, Bradley in view of Gernoth teaches the apparatus claim 2, wherein one or more of the at least one processor circuit is to determine at least one of an expression of a face in the input image or an identity of the face based on the model and when the confidence metric of the input image satisfies a threshold (Gernoth, Para. [0036], ISP 110 may process the images independently to determine characteristics of each image separately. SEP 112 may then compare the separate image characteristics with stored template images for each type of image to generate an authentication score (i.e., confidence metric) (e.g., a matching score or other ranking of matching between the user in the captured image and in the stored template images) for each separate image. The authentication scores for the separate images (e.g., the flood IR and depth map images) may be combined to make a decision on the identity of the user (i.e., identity of the face) and, if authenticated, allow the user to use device 100 (e.g., unlock the device)).
The proposed combination as well as the motivation for combining the Bradley and Gernoth references presented in the rejection of Claim 1, apply to Claim 7 and are incorporated herein by reference. Thus, the apparatus recited in Claim 7 is met by Bradley in view of Gernoth.
Regarding claim 9, Bradley in view of Gernoth teaches the apparatus of claim 1, wherein the transformation operation includes one or more of a crop of the image, a change in a field of view of the image, a resizing of the image, a scaling of the image, a rotation of the image, or a shifting of the image (Bradley, Para. [0040], a data-augmentation parameter that applies transformations to features inputted into landmark detection engine 202 (e.g., scaling, translating, rotating, shearing, shifting, and/or otherwise transforming an image)).
Claims 11-13, 17 and 19 recite computer-readable storage mediums storing programs with instructions corresponding to the elements recited in Claims 1-3, 7 and 9, respectively. Therefore, the recited programming instructions of these claims are mapped to the proposed combination in the same manner as the corresponding elements in their corresponding apparatus claims. Additionally, the rationale and motivation to combine the Bradley and Gernoth references, presented in rejection of Claim 1, apply to these claims. Finally, the combination of the Bradley and Gernoth references discloses a computer readable storage medium (Bradley, Para. [0074], one or more non-transitory computer readable media).
Claims 4 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Bradley et al. (US 2024/0161540 A1, filed November 8, 2023) in view of Gernoth et al. (US 2019/0042833 A1), as applied to claims 1-3, 7, 9, 11-13, 17 and 19 above, and further in view of Ludovico Novelli (US 10,928,904 B1).
Regarding claim 4, Bradley in view of Gernoth teaches the apparatus of claim 2, as described above.
Although Bradley in view of Gernoth teaches estimating head pose when a confidence metric satisfies a threshold (Gernoth, Para. [0055]-[0056]), the references do not explicitly teach “wherein one or more of the at least one processor circuit is to track a face in the input image when the confidence metric for the input image satisfies a threshold”. However, in an analogous field of endeavor, Novelli teaches confirming that the detected user is the user and not the bystander based on a detected face of a plurality of detected faces that is closest to the display or facing the display. In some embodiments, tracking a location of the user's face can be based on the observational data of the detected user and calculating and periodically updating a confidence score that indicates a likelihood that the detected user is the user and not the bystander based on the tracked location of the user's face (Novelli, Col. 17, lines 34-46).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the apparatus of Bradley in view of Gernoth with the teachings of Novelli by including confirming that a user is the detected user based on a confidence score and performing tracking of the face based on the confirmation. One having ordinary skill in the art would have been motivated to combine these references because doing so would allow for performing real-time facial tracking, as recognized by Novelli. Thus, the claimed invention would have been obvious to one having ordinary skill in the art before the effective filing date.
Claim 14 recites a computer-readable storage medium storing a program with instructions corresponding to the elements recited in Claim 4. Therefore, the recited programming instructions of this claim are mapped to the proposed combination in the same manner as the corresponding elements in its corresponding apparatus claim. Additionally, the rationale and motivation to combine the Bradley, Gernoth and Novelli references, presented in rejection of Claim 4, apply to this claim. Finally, the combination of the Bradley, Gernoth and Novelli references discloses a computer readable storage medium (Bradley, Para. [0074], one or more non-transitory computer readable media).
Claims 5-6, 8, 15-16 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Bradley et al. (US 2024/0161540 A1, filed November 8, 2023) in view of Gernoth et al. (US 2019/0042833 A1), as applied to claims 1-3, 7, 9, 11-13, 17 and 19 above, and further in view of Rui et al. (US 2003/0103647 A1).
Regarding claim 5, Bradley in view of Gernoth teaches the apparatus of claim 2, as described above.
Although Bradley in view of Gernoth teaches estimating head pose when a confidence metric satisfies a threshold (Gernoth, Para. [0055]-[0056]), the references do not explicitly teach “wherein one or more of the at least one processor circuit is to track the head pose over multiple images when the confidence metric of the input image satisfies a threshold”. However, in an analogous field of endeavor, Rui teaches detection and tracking module 132 analyzes content on a frame by frame basis. For each frame, module 132 activates the auto-initialization module 140 which operates to detect candidates for new face regions. Each such candidate is a region of the video content that potentially includes a new face (that is, a face that is not currently being tracked). Once detected, a candidate region is passed to hierarchical verification module 142, which in turn verifies whether the candidate region does indeed include a face. Hierarchical verification module 142 generates a confidence level for each candidate and determines to keep the candidate as a face region if the confidence level exceeds a threshold value (i.e., track the head pose over multiple images when the confidence metric satisfies a threshold), adding a description of the region to tracking list 146. If the confidence level does not exceed the threshold value, then hierarchical verification module 142 discards the candidate (Rui, Para. [0041]). Two tasks, face detection and pose estimation (i.e., track the head pose), are performed jointly by classifying the input Ip into one of the L+1 classes. If the input is classified into one of the L face classes, a face is detected and the corresponding view is the estimated pose (Rui, Para. [0096]).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date to modify the apparatus of Bradley in view of Gernoth with the teachings of Rui by including tracking head pose over multiple images when the confidence metric satisfies a threshold. One having ordinary skill in the art would have been motivated to combine these references because doing so would allow for tracking faces frame to frame, as recognized by Rui. Thus, the claimed invention would have been obvious to one having ordinary skill in the art before the effective filing date.
Regarding claim 6, Bradley in view of Gernoth further in view of Rui teaches the apparatus of claim 5, wherein one or more of the at least one processor circuit is to set a bounding box in the input image as an expectation for a position of the head in a subsequent image (Rui, Para. [0038], face/candidate tracking list 146 maintains information for each detected region that includes, or potentially includes, a human face. Those regions that potentially include a face but for which the presence of a face has not been verified are referred to as candidate regions. In the illustrated example, each region is described by a center coordinate 148, a bounding box 150, a tracking duration 152, and a time since last verification 154. The regions of video content that include faces or face candidates are defined by a center coordinate and a bounding box. Center coordinate 148 represents the approximate center of the region, while bounding box 150 represents a rectangular region around the center coordinate. This rectangular region is the region that includes a face or face candidate (i.e., expectation for a position of the head) and is tracked by detection and tracking module 132. Tracking duration 152 represents how long the face or face candidate in the region has been tracked, while the time since last verification 154 represents how long ago the face or face candidate in the region was verified (i.e., expectation for position of the head in a subsequent image).
The proposed combination as well as the motivation for combining the Bradley, Gernoth and Rui references presented in the rejection of Claim 5, apply to Claim 6 and are incorporated herein by reference. Thus, the apparatus recited in Claim 6 is met by Bradley in view of Gernoth further in view of Rui.
Regarding claim 8, Bradley in view of Gernoth teaches the apparatus of claim 2, as described above.
Although Bradley in view of Gernoth teaches estimating head pose when a confidence metric satisfies a threshold (Gernoth, Para. [0055]-[0056]), the references do not explicitly teach “wherein one or more of the at least one processor circuit is to preclude tracking an object in the input image when the confidence metric of the input image does not satisfy a threshold”. However, in an analogous field of endeavor, Rui teaches detection and tracking module 132 analyzes content on a frame by frame basis. For each frame, module 132 activates the auto-initialization module 140 which operates to detect candidates for new face regions. Each such candidate is a region of the video content that potentially includes a new face (that is, a face that is not currently being tracked). Once detected, a candidate region is passed to hierarchical verification module 142, which in turn verifies whether the candidate region does indeed include a face. Hierarchical verification module 142 generates a confidence level for each candidate and determines to keep the candidate as a face region if the confidence level exceeds a threshold value, adding a description of the region to tracking list 146. If the confidence level does not exceed the threshold value, then hierarchical verification module 142 discards the candidate (i.e., preclude tracking an object when the confidence metric does not satisfy a threshold) (Rui, Para. [0041]).
The proposed combination as well as the motivation for combining the Bradley, Gernoth and Rui references presented in the rejection of Claim 5, apply to Claim 8 and are incorporated herein by reference. Thus, the apparatus recited in Claim 8 is met by Bradley in view of Gernoth further in view of Rui.
Claims 15-16 and 18 recite computer-readable storage mediums storing programs with instructions corresponding to the elements recited in Claims 5-6 and 8, respectively. Therefore, the recited programming instructions of these claims are mapped to the proposed combination in the same manner as the corresponding elements in their corresponding apparatus claims. Additionally, the rationale and motivation to combine the Bradley, Gernoth and Rui references, presented in rejection of Claim 5, apply to these claims. Finally, the combination of the Bradley, Gernoth and Rui references discloses a computer readable storage medium (Bradley, Para. [0074], one or more non-transitory computer readable media).
Claims 10 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Bradley et al. (US 2024/0161540 A1, filed November 8, 2023) in view of Gernoth et al. (US 2019/0042833 A1), as applied to claims 1-3, 7, 9, 11-13, 17 and 19 above, and further in view of Tal et al. (US 2022/0019829 A1).
Regarding claim 10, Bradley in view of Gernoth teaches the apparatus of claim 1, wherein the instructions program one or more of the at least one processor circuit to implement transformation operations including:
scaling the image (Bradley, Para. [0040], a data-augmentation parameter that applies transformations to features inputted into landmark detection engine 202 (e.g., scaling, translating, rotating, shearing, shifting, and/or otherwise transforming an image)), and
shifting of the image (Bradley, Para. [0040], a data-augmentation parameter that applies transformations to features inputted into landmark detection engine 202 (e.g., scaling, translating, rotating, shearing, shifting, and/or otherwise transforming an image)); and
the at least one processor circuit is to augment the image data with one or more of the transformation operations (Bradley, Para. [0040], a data-augmentation parameter that applies transformations to features inputted into landmark detection engine 202 (e.g., scaling, translating, rotating, shearing, shifting, and/or otherwise transforming an image)).
Although Bradley in view of Gernoth teaches augmenting the image data with transformations (Bradley, Para. [0040]), the references do not explicitly teach that the transformations include “cropping the image”, “changing in a field of view of the image” and “resizing the image”. However, in an analogous field of endeavor, Tal teaches the image(s) may need to undergo image processing operations to optimize their compatibility with the neural networks used and the object(s) of interest which they are trained to identify. Some examples of image processing operations are resizing or adjusting resolution, field of view adjustments, cropping, and/or color space conversion (Tal, Para. [0102]).
Therefore, it would have been obvious to one having ordinary skill in the art to modify the apparatus of Bradley in view of Gernoth with the teachings of Tal by including performing transformations on the image including resizing, cropping and/or field of view adjustments. One having ordinary skill in the art would have been motivated to combine these references because doing so would allow for optimizing identification of objects of interest using a neural network, as recognized by Tal. Thus, the claimed invention would have been obvious to one having ordinary skill in the art before the effective filing date.
Claim 20 recites a computer-readable storage medium storing a program with instructions corresponding to the elements recited in Claim 10. Therefore, the recited programming instructions of this claim are mapped to the proposed combination in the same manner as the corresponding elements in its corresponding apparatus claim. Additionally, the rationale and motivation to combine the Bradley, Gernoth and Tal references, presented in rejection of Claim 10, apply to this claim. Finally, the combination of the Bradley, Gernoth and Tal references discloses a computer readable storage medium (Bradley, Para. [0074], one or more non-transitory computer readable media).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Emma Rose Goebel whose telephone number is (703)756-5582. The examiner can normally be reached Monday - Friday 7:30-5.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amandeep Saini can be reached at (571) 272-3382. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Emma Rose Goebel/Examiner, Art Unit 2662
/AMANDEEP SAINI/Supervisory Patent Examiner, Art Unit 2662