DETAILED ACTION
Notice of Pre-AIA or AIA Status
Claims 1-20 are pending in this application. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1, 13 and 20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by (Bhargava et al. US PGPub US20220405500A1, hereby referred to as “Bhargava”).
Consider Claims 1, 13 and 20.
Bhargava teaches:
1. A computer-implemented method for audio processing based on a head pose of a user, the computer-implemented method comprising: / 13. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform audio processing based on a head pose of a user by performing the steps of: / 20. A system comprising: one or more speakers; a camera that captures one or more images of a user; a memory storing instructions; and one or more processors, that when executing the instructions, are configured to perform audio processing based on a head pose of a user by performing the steps of: (Bhargava: abstract, A computer-implemented method includes receiving a two-dimensional (2-D) side view face image of a person, identifying a bounded portion or area of the 2-D side view face image of the person as an ear region-of-interest (ROI) area showing at least a portion of an ear of the person, and processing the identified ear ROI area of the 2-D side view face image, pixel-by-pixel, through a trained fully convolutional neural network model (FCNN model) to predict a 2-D ear saddle point (ESP) location for the ear shown in the ear ROI area. The FCNN model has an image segmentation architecture. [0004]-[0008] Figure 1-2, [0120]-[0123], Figure 9; [0063] In some example implementations of system 100, a simple machine learning (ML) model or a cross validation (CV) approach (e.g., a convolutional filter) may be used to further refine (if required) the ear ROI area derived using a single landmark point on the ear or on the face before image processing at stage 160 in image processing pipeline 110 to identify ESPs.)
1. acquiring one or more images of a user; / 13. acquiring one or more images of a user; / 20. acquiring one or more images of a user; (Bhargava: [0027] In virtual settings, where the person is remote (e.g., on-line, or on the Internet), a virtual 3-D prototype of the glasses may be constructed after inferring the 3-D features of the person's head from a set of two-dimensional (2-D) images of the person's head. The glasses may be custom fitted by positioning the virtual 3-D prototype on a 3-D head model of the person in a virtual-try-on (VTO) session (simulating an actual physical fitting of the glasses on the person's head). Proper sizing and accurate virtual-try-on (VTO) are important factors for successfully making custom fitted glasses for remote consumers. [0028] In some virtual fitting situations, the ESPs of a remote person can be identified and located on the 2-D images using a sizing application (app) to process 2-D images (e.g., digital photographs or pictures) of the person's head. The sizing app may involve a machine learning model (e.g., a trained neural network model) to process the 2-D images to identify or locate the ESPs. To run such a sizing app, for example, on a mobile phone, to efficiently identify or locate the ESP of the person based on a 2-D image, the processes or algorithms used in the sizing app to process the 2-D images should be fast, and consume little memory and other computational resources. [0029] Previous efforts at using sizing apps (e.g., on mobile phones) to locate the ESPs in the 2-D images have been inefficient and have yielded less than satisfactory results. The previous sizing apps have utilized two detection models (a first model and a second model) to locate the ESPs in the 2-D images.)
1. processing the one or more images to identify a plurality of face landmarks representing locations on a head of the user; / 13. processing the one or more images to identify a plurality of face landmarks representing locations on a head of the user; / 20. processing the one or more images to identify a plurality of face landmarks representing locations on a head of the user; (Bhargava: [0032] The disclosed image processing solutions involve receiving 2-D images (pictures) of the person's head in different orientations, identifying fiducial facial landmark features (landmark points) on the person's face in the 2-D images, and using at least one of the fiducial landmark points as a geometrical reference point or marker to define an area or portion (i.e., an ear region-of-interest (ROI)) in a side view face image of the person for ESP analysis and detection. The defined ear ROI may be a small portion of the side view face image, and may show or include at least a portion of an ear (left ear or right ear) of the person. For a side view face image having a typical size of ˜1000×1000 pixels, the defined ear ROI area may, for example, be less than ˜200×200 pixels. For reference, an average human ear is about 2.5 inches (6.3 centimeters) long. However, there can be large variations in ear shape, size and orientation from individual to individual and even between the left ears and right ears of individuals. [0040] Image processing pipeline 110 may include an input stage 120, a pose estimator stage 130, a fiducial landmarks detection stage 140, an ear ROI extraction stage 150, and an ESP identification stage 160. Processing images through the various stages 120-150 may involve processing the images through the one or more CNN and FCNN models (e.g., CNN 15, ESP-FCNN 16, etc.).)
1. selecting, from the plurality of face landmarks and based on an estimated head pose of the user, a set of one or more landmark pairs; / 13. selecting, from the plurality of face landmarks and based on an estimated head pose of the user, a set of one or more landmark pairs; / 20. selecting, from the plurality of face landmarks and based on an estimated head pose of the user, a set of one or more landmark pairs; (Bhargava: [0034] The disclosed image processing solutions can be used to determine an ESP of a person, for example, for fitting glasses on the person. The fitting of glasses (e.g. sizing of the glasses) may be conducted in a virtual-try-on (VTO) system, in which the fitting is accomplished remotely (e.g., over the Internet) on a 3-D head model of the person. For proper fitting, the 2-D location of the ESP on the 2-D image is projected to a 3-D point at a depth on a side of the ear on the 3-D head model. The projected point may represent a 3-D ESP in 3-D space for fitting glasses on the person. [0041] Input stage 120 may be configured to receive 2-D images of a person's head. The 2-D images may be captured using, for example, a smartphone camera. The received 2-D images (e.g., image 60) may include images (e.g., front and side view face images) taken at different orientations (e.g., neck rotations or tilt) of the person's head. The received 2-D images (e.g., image 60) may be processed through a pose estimator stage (e.g., pose estimator stage 130) and segregated for further processing according to whether the image is a front face view (corresponding, e.g., to a face tilt or head rotation of less than ˜5 degrees), or a side face view (corresponding, e.g., to a face tilt or head rotation of greater than ˜30 degrees). The front view face image may be expected to show little of the person's ears, while the side view face image may be expected to show more of the person's ear (either left ear or right ear).)
1. determine, based on the set of landmark pairs, three-dimensional positions of ears of the user; / 13. determine, based on the set of landmark pairs, three-dimensional positions of ears of the user; / 20. determine, based on the set of landmark pairs, three-dimensional positions of ears of the user; (Bhargava: [0091] In method 700, identifying the ear ROI area on the 2-D side view face image 720 may include receiving a 2-D front view face image of the person corresponding to the 2-D side view face image of the person (received at 710), and processing the 2-D front view face image through a trained fully convolutional neural network model (e.g., a Face-SSD model) to identify the ear ROI area. A shape (e.g., a rectangular shape) and a pixel-size of the bounded area of the ear ROIs may be predefined. In example implementations, the pixel-size of the ear ROI area may be less than about 1000×1000 pixels (e.g., 200×200 pixels, 128×96 pixels, 140×110 pixels, etc.). In example implementations, the size of the bounded area of the ear ROIs may be based on a face size parameter related to the size of the face shown, for example, in the front view face image of the person. [0092] In example implementations, the Face-SSD model may identify one or more facial landmark points on the 2-D front view face image. The identified facial landmark points may for example, include a left ear tragion (LET) point and a right ear tragion (RET) point (disposed on the left ear tragus and the right ear tragus of the person, respectively). The Face-SSD model may define a portion or area of the 2-D side view face image as being bounded, for example, by a rectangle. The position of the bounding rectangle may be determined using one or more of the identified facial landmark points as geometrical fiducial reference points.)
1. and processing one or more audio signals to generate one or more processed audio signals based on the three-dimensional positions of the ears./ 13. and processing one or more audio signals to generate one or more processed audio signals based on the three-dimensional positions of the ears. / 20. and processing one or more audio signals to generate one or more processed audio signals based on the three-dimensional positions of the ears. (Bhargava: [0093] After the ear ROI area is identified (at 720) in method 700, processing the ear ROI area, pixel-by-pixel, through the trained ESP-CNN model 730 may include image segmentation of the ear ROI area and using each pixel for category prediction. The trained ESP-CNN model may, for example, predict a probability or confidence value for each pixel in the ear ROI area that the pixel is an actual or correct 2-D ESP location. The predicted confidence value for a pixel may be a floating point number reflecting an inverse distance from the pixel to the actual or correct ESP location (instead of a binary decision whether or not the pixel is the correct 2-D ESP). In example implementations, processing the ear ROI area, pixel-by-pixel, through the trained ESP-CNN model 730 may be include generating a confidence map (prediction heatmap) in which pixels with high confidence are predicted to be the correct 2-D ESP. [0094] In example implementations of method 700, when the identified ear ROI area may have a size less than 1000×1000 pixels, the trained ESP-CNN model (e.g., a U-Net) may have a size less than 1000 Kb (e.g., 246 Kb). [0095] Method 700 may further include projecting the predicted 2-D ESP located in the ear ROI area on the 2-D side view face image through 3-D space to a 3-D ESP location on a 3-D head model of the person (740), and fitting virtual glasses to the 3-D head model of the person with a temple piece of the glasses resting on the projected 3-D ESP in a virtual-try-on-session (750).)
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent may not be obtained though the invention is not identically disclosed or described as set forth in section 102 of this title, if the differences between the subject matter sought to be patented and the prior art are such that the subject matter as a whole would have been obvious at the time the invention was made to a person having ordinary skill in the art to which said subject matter pertains. Patentability shall not be negatived by the manner in which the invention was made.
Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Kondrashov et al. (US PGPub US20230102851A1, hereby referred to as “Kondrashov”), in view of Bhargava et al. (US PGPub US20220405500A1, hereby referred to as “Bhargava”).
Consider Claims 1, 13 and 20.
Kondrashov teaches:
1. A computer-implemented method for audio processing based on a head pose of a user, the computer-implemented method comprising: / 13. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform audio processing based on a head pose of a user by performing the steps of: / 20. A system comprising: one or more speakers; a camera that captures one or more images of a user; a memory storing instructions; and one or more processors, that when executing the instructions, are configured to perform audio processing based on a head pose of a user by performing the steps of: (Kondrashov: abstract Disclosed herein is an apparatus comprising a camera and a processing unit operatively coupled to the camera, wherein the processing unit is configured to: receive a sequence of images captured by the camera; process a first image of the received sequence of images to compute respective likelihoods of each of a plurality of predetermined facial features being visible in the first image; compute, from the computed likelihoods, a probability that the first image depicts a predetermined first side of a human head; responsive to at least the computed probability exceeding the predetermined detection probability, trigger performance of a predetermined action. [0061], [0068]-[0076], Figures 1-2 [0068] FIG. 1 schematically shows an example of an apparatus 10 for triggering performance of a head-pose dependent action, i.e. an action triggered by a predetermined pose of a human head 30. In the example of FIG. 1 , the apparatus 10 is a portable device, in particular a smartphone, having an integrated digital camera 12. For example, the camera 12 may be a front facing camera of the smartphone. However, in other embodiments, the camera may be a rear-facing camera. It will further be appreciated that, in other embodiments, the portable device 10 may be a tablet computer, a laptop computer or another type of portable processing device having an integrated digital camera. The portable device 10 comprises a processing unit (not explicitly shown in FIG. 1 ), such as a suitable programmed central processing unit. The processing unit is programmed or otherwise configured to: [0069] receive a sequence of images captured by the digital camera 12;)
1. acquiring one or more images of a user; / 13. acquiring one or more images of a user; / 20. acquiring one or more images of a user; (Kondrashov: [0068]-[0076], Figures 1-2 [0068] FIG. 1 schematically shows an example of an apparatus 10 for triggering performance of a head-pose dependent action, i.e. an action triggered by a predetermined pose of a human head 30. In the example of FIG. 1 , the apparatus 10 is a portable device, in particular a smartphone, having an integrated digital camera 12. For example, the camera 12 may be a front facing camera of the smartphone. However, in other embodiments, the camera may be a rear-facing camera. It will further be appreciated that, in other embodiments, the portable device 10 may be a tablet computer, a laptop computer or another type of portable processing device having an integrated digital camera. The portable device 10 comprises a processing unit (not explicitly shown in FIG. 1 ), such as a suitable programmed central processing unit. The processing unit is programmed or otherwise configured to: [0069] receive a sequence of images captured by the digital camera 12;)
1. processing the one or more images to identify a plurality of face landmarks representing locations on a head of the user; / 13. processing the one or more images to identify a plurality of face landmarks representing locations on a head of the user; / 20. processing the one or more images to identify a plurality of face landmarks representing locations on a head of the user; (Kondrashov: [0068]-[0076], Figures 1-2, [0070] process a first image of the received sequence of images to compute respective likelihoods of each of at least three predetermined facial features being visible in the first image; [0071] compute, from the detected likelihoods, a probability that the first image depicts a predetermined first side of a human head 30; [0072] process one or more images of the sequence of images to compute a stability parameter indicative of a stability of the captured images over time; [0076] FIG. 2 schematically shows another example of an apparatus 10 for triggering performance of a head-pose dependent action. The apparatus 10 of FIG. 2 is similar to the apparatus of FIG. 1 in that the apparatus 10 includes a digital camera 12 and a processing unit 112 configured to perform the acts as described in connection with FIG. 1 . The apparatus of FIG. 2 differs from the apparatus of FIG. 1 in that the apparatus 10 of FIG. 2 comprises two distinct and separate devices, each having its own housing. In particular, the apparatus of FIG. 2 includes a data processing device 110 and the digital camera 12 separate from and external to the data processing device 110.)
1. selecting, from the plurality of face landmarks and based on an estimated head pose of the user, a set of one or more landmark pairs; / 13. selecting, from the plurality of face landmarks and based on an estimated head pose of the user, a set of one or more landmark pairs; / 20. selecting, from the plurality of face landmarks and based on an estimated head pose of the user, a set of one or more landmark pairs; (Kondrashov: [0072] process one or more images of the sequence of images to compute a stability parameter indicative of a stability of the captured images over time; [0073] determine whether the computed probability exceeds a predetermined detection probability and whether the computed stability parameter fulfills a predetermined stability condition; [0074] responsive to the computed probability exceeding the predetermined detection probability and the computed stability parameter fulfilling the predetermined stability condition, triggering performance of a predetermined head-pose dependent action, e.g., the capturing of an image or the recording of one of the already captured images as an image of the first side of the human head 30. [0077] The digital camera 12 may be a webcam or another type of digital camera communicatively coupled to the data processing device 110 and operable to capture images of a human head 30. In the example of FIG. 2 , the digital camera 12 is communicatively coupled to the data processing device 110 via a short-range wireless communications link 80. The short-range wireless communications link 80 may be a radio communications link, e.g., Bluetooth communication slink or a wireless communications link using another suitable wireless communications technology. In other examples, the digital camera 12 is communicatively coupled to the data processing device 110 via a different type of wired or wireless communications link. Examples of a different type of wireless communications link include an indirect link, e.g., via a wireless access point, via a local wireless computer network or the like. Yet further, examples of a wired communications link include direct or indirect wired connections, e.g., via a USB cable, a wired local area network, or another suitable wired communication technology.)
1. determine, based on the set of landmark pairs, three-dimensional positions of ears of the user; / 13. determine, based on the set of landmark pairs, three-dimensional positions of ears of the user; / 20. determine, based on the set of landmark pairs, three-dimensional positions of ears of the user; (Kondrashov: [0025] Finally, the third facial feature may be chosen such that it is visible in a side view of the human head seen when facing the third side. The third facial feature may be selected such that it is also visible on a side view of the human head seen when facing the first side and/or on a side view of the human head seen when facing the second side. In particular, the first facial feature may be a feature of a first ear of the human head, e.g., a landmark point on a first earlobe, an attachment point of the first ear, etc. The second facial feature may be a feature of a second ear, opposite the first ear, e.g., a landmark point on a second earlobe, an attachment point of the second ear, etc. The third facial feature may be a landmark point of the nose, e.g., a tip of the nose, which is also visible on a lateral side view of the human head. [0072] process one or more images of the sequence of images to compute a stability parameter indicative of a stability of the captured images over time; [0073] determine whether the computed probability exceeds a predetermined detection probability and whether the computed stability parameter fulfills a predetermined stability condition; [0074] responsive to the computed probability exceeding the predetermined detection probability and the computed stability parameter fulfilling the predetermined stability condition, triggering performance of a predetermined head-pose dependent action, e.g., the capturing of an image or the recording of one of the already captured images as an image of the first side of the human head 30.)
1. and processing one or more signals to generate one or more processed signals based on the three-dimensional positions of the ears./ 13. and processing one or more signals to generate one or more processed signals based on the three-dimensional positions of the ears. / 20. and processing one or more signals to generate one or more processed signals based on the three-dimensional positions of the ears. (Kondrashov: [0080] FIG. 3 schematically shows an example of a method for triggering performance of a head-pose dependent action. The method may e.g., be performed by the apparatus of any of FIGS. 1-2 or by another suitable apparatus. [0081] At initial step S1, the process receives one or more captured images depicting a human head. The subsequent steps S2-S4 may be performed in real-time or quasi-real time, i.e., the individual images, e.g., individual frames of video, may be processed as they are received rather than waiting for an entire plurality of images having been received. In other examples, the process may receive a certain number of images, e.g., a certain number of frames, and then process the number of images by performing steps S2-S4. [0082] In subsequent step S2, the process processes the one or more captured images to detect an orientation of the depicted head relative to the camera having captured the image(s). In particular, this step may detect whether a predetermined lateral side view of a human head is depicted in the image. Accordingly, this step may output a corresponding lateral side view detection condition Cside={true, false}, e.g., a left side detection condition Cright={true, false} and/or a right side detection condition Cleft={true, false}. At least some embodiments of this step may be based on detected facial landmarks. An example of a process for detecting the orientation of the head will be described below in more detail with reference to FIG. 4 .)
Even if Kondrashav does not teach specifically teach “audio signal” from : 1./13./20. and processing one or more audio signals to generate one or more processed audio signals based on the three-dimensional positions of the ears
Bhargava teaches:
1. A computer-implemented method for audio processing based on a head pose of a user, the computer-implemented method comprising: / 13. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform audio processing based on a head pose of a user by performing the steps of: / 20. A system comprising: one or more speakers; a camera that captures one or more images of a user; a memory storing instructions; and one or more processors, that when executing the instructions, are configured to perform audio processing based on a head pose of a user by performing the steps of: (Bhargava: abstract, A computer-implemented method includes receiving a two-dimensional (2-D) side view face image of a person, identifying a bounded portion or area of the 2-D side view face image of the person as an ear region-of-interest (ROI) area showing at least a portion of an ear of the person, and processing the identified ear ROI area of the 2-D side view face image, pixel-by-pixel, through a trained fully convolutional neural network model (FCNN model) to predict a 2-D ear saddle point (ESP) location for the ear shown in the ear ROI area. The FCNN model has an image segmentation architecture. [0004]-[0008] Figure 1-2, [0120]-[0123], Figure 9; [0063] In some example implementations of system 100, a simple machine learning (ML) model or a cross validation (CV) approach (e.g., a convolutional filter) may be used to further refine (if required) the ear ROI area derived using a single landmark point on the ear or on the face before image processing at stage 160 in image processing pipeline 110 to identify ESPs.)
1. acquiring one or more images of a user; / 13. acquiring one or more images of a user; / 20. acquiring one or more images of a user; (Bhargava: [0027] In virtual settings, where the person is remote (e.g., on-line, or on the Internet), a virtual 3-D prototype of the glasses may be constructed after inferring the 3-D features of the person's head from a set of two-dimensional (2-D) images of the person's head. The glasses may be custom fitted by positioning the virtual 3-D prototype on a 3-D head model of the person in a virtual-try-on (VTO) session (simulating an actual physical fitting of the glasses on the person's head). Proper sizing and accurate virtual-try-on (VTO) are important factors for successfully making custom fitted glasses for remote consumers. [0028] In some virtual fitting situations, the ESPs of a remote person can be identified and located on the 2-D images using a sizing application (app) to process 2-D images (e.g., digital photographs or pictures) of the person's head. The sizing app may involve a machine learning model (e.g., a trained neural network model) to process the 2-D images to identify or locate the ESPs. To run such a sizing app, for example, on a mobile phone, to efficiently identify or locate the ESP of the person based on a 2-D image, the processes or algorithms used in the sizing app to process the 2-D images should be fast, and consume little memory and other computational resources. [0029] Previous efforts at using sizing apps (e.g., on mobile phones) to locate the ESPs in the 2-D images have been inefficient and have yielded less than satisfactory results. The previous sizing apps have utilized two detection models (a first model and a second model) to locate the ESPs in the 2-D images.)
1. processing the one or more images to identify a plurality of face landmarks representing locations on a head of the user; / 13. processing the one or more images to identify a plurality of face landmarks representing locations on a head of the user; / 20. processing the one or more images to identify a plurality of face landmarks representing locations on a head of the user; (Bhargava: [0032] The disclosed image processing solutions involve receiving 2-D images (pictures) of the person's head in different orientations, identifying fiducial facial landmark features (landmark points) on the person's face in the 2-D images, and using at least one of the fiducial landmark points as a geometrical reference point or marker to define an area or portion (i.e., an ear region-of-interest (ROI)) in a side view face image of the person for ESP analysis and detection. The defined ear ROI may be a small portion of the side view face image, and may show or include at least a portion of an ear (left ear or right ear) of the person. For a side view face image having a typical size of ˜1000×1000 pixels, the defined ear ROI area may, for example, be less than ˜200×200 pixels. For reference, an average human ear is about 2.5 inches (6.3 centimeters) long. However, there can be large variations in ear shape, size and orientation from individual to individual and even between the left ears and right ears of individuals. [0040] Image processing pipeline 110 may include an input stage 120, a pose estimator stage 130, a fiducial landmarks detection stage 140, an ear ROI extraction stage 150, and an ESP identification stage 160. Processing images through the various stages 120-150 may involve processing the images through the one or more CNN and FCNN models (e.g., CNN 15, ESP-FCNN 16, etc.).)
1. selecting, from the plurality of face landmarks and based on an estimated head pose of the user, a set of one or more landmark pairs; / 13. selecting, from the plurality of face landmarks and based on an estimated head pose of the user, a set of one or more landmark pairs; / 20. selecting, from the plurality of face landmarks and based on an estimated head pose of the user, a set of one or more landmark pairs; (Bhargava: [0034] The disclosed image processing solutions can be used to determine an ESP of a person, for example, for fitting glasses on the person. The fitting of glasses (e.g. sizing of the glasses) may be conducted in a virtual-try-on (VTO) system, in which the fitting is accomplished remotely (e.g., over the Internet) on a 3-D head model of the person. For proper fitting, the 2-D location of the ESP on the 2-D image is projected to a 3-D point at a depth on a side of the ear on the 3-D head model. The projected point may represent a 3-D ESP in 3-D space for fitting glasses on the person. [0041] Input stage 120 may be configured to receive 2-D images of a person's head. The 2-D images may be captured using, for example, a smartphone camera. The received 2-D images (e.g., image 60) may include images (e.g., front and side view face images) taken at different orientations (e.g., neck rotations or tilt) of the person's head. The received 2-D images (e.g., image 60) may be processed through a pose estimator stage (e.g., pose estimator stage 130) and segregated for further processing according to whether the image is a front face view (corresponding, e.g., to a face tilt or head rotation of less than ˜5 degrees), or a side face view (corresponding, e.g., to a face tilt or head rotation of greater than ˜30 degrees). The front view face image may be expected to show little of the person's ears, while the side view face image may be expected to show more of the person's ear (either left ear or right ear).)
1. determine, based on the set of landmark pairs, three-dimensional positions of ears of the user; / 13. determine, based on the set of landmark pairs, three-dimensional positions of ears of the user; / 20. determine, based on the set of landmark pairs, three-dimensional positions of ears of the user; (Bhargava: [0091] In method 700, identifying the ear ROI area on the 2-D side view face image 720 may include receiving a 2-D front view face image of the person corresponding to the 2-D side view face image of the person (received at 710), and processing the 2-D front view face image through a trained fully convolutional neural network model (e.g., a Face-SSD model) to identify the ear ROI area. A shape (e.g., a rectangular shape) and a pixel-size of the bounded area of the ear ROIs may be predefined. In example implementations, the pixel-size of the ear ROI area may be less than about 1000×1000 pixels (e.g., 200×200 pixels, 128×96 pixels, 140×110 pixels, etc.). In example implementations, the size of the bounded area of the ear ROIs may be based on a face size parameter related to the size of the face shown, for example, in the front view face image of the person. [0092] In example implementations, the Face-SSD model may identify one or more facial landmark points on the 2-D front view face image. The identified facial landmark points may for example, include a left ear tragion (LET) point and a right ear tragion (RET) point (disposed on the left ear tragus and the right ear tragus of the person, respectively). The Face-SSD model may define a portion or area of the 2-D side view face image as being bounded, for example, by a rectangle. The position of the bounding rectangle may be determined using one or more of the identified facial landmark points as geometrical fiducial reference points.)
1. and processing one or more audio signals to generate one or more processed audio signals based on the three-dimensional positions of the ears./ 13. and processing one or more audio signals to generate one or more processed audio signals based on the three-dimensional positions of the ears. / 20. and processing one or more audio signals to generate one or more processed audio signals based on the three-dimensional positions of the ears. (Bhargava: [0093] After the ear ROI area is identified (at 720) in method 700, processing the ear ROI area, pixel-by-pixel, through the trained ESP-CNN model 730 may include image segmentation of the ear ROI area and using each pixel for category prediction. The trained ESP-CNN model may, for example, predict a probability or confidence value for each pixel in the ear ROI area that the pixel is an actual or correct 2-D ESP location. The predicted confidence value for a pixel may be a floating point number reflecting an inverse distance from the pixel to the actual or correct ESP location (instead of a binary decision whether or not the pixel is the correct 2-D ESP). In example implementations, processing the ear ROI area, pixel-by-pixel, through the trained ESP-CNN model 730 may be include generating a confidence map (prediction heatmap) in which pixels with high confidence are predicted to be the correct 2-D ESP. [0094] In example implementations of method 700, when the identified ear ROI area may have a size less than 1000×1000 pixels, the trained ESP-CNN model (e.g., a U-Net) may have a size less than 1000 Kb (e.g., 246 Kb). [0095] Method 700 may further include projecting the predicted 2-D ESP located in the ear ROI area on the 2-D side view face image through 3-D space to a 3-D ESP location on a 3-D head model of the person (740), and fitting virtual glasses to the 3-D head model of the person with a temple piece of the glasses resting on the projected 3-D ESP in a virtual-try-on-session (750).)
It would have been obvious before the effective filing date of the claimed invention was made to one of ordinary skill in the art to modify the method and system for determining a head pose triggered action of Kondrashov with the facial detection algorithm of Bhagrava, as they are both directed towards the usage of facial landmark and feature detection for image processing and analysis. The determination of obviousness is predicated upon the following findings: One skilled in the art would have been motivated to modify Kondrashov in order to more accurately leverage a computational efficient and robust ear saddle and landmark feature-based facial detection for determining head-pose dependent actions. Furthermore, the prior art collectively includes each element claimed (though not all in the same reference), and one of ordinary skill in the art could have combined the elements in the manner explained above using known engineering design, interface and/or programming techniques, without changing a “fundamental” operating principle of Kondrashov while the teaching of Bhagrava continues to perform the same function as originally taught prior to being combined, in order to produce the repeatable and predictable result of more accurately detection facial features and poses for downstream image processing algorithms. It is for at least the aforementioned reasons that the examiner has reached a conclusion of obviousness with respect to the claim in question.
Consider Claims 2 and 14.
The combination of Kondrashov and Bhagrava teaches:
2. The computer-implemented method of claim 1, wherein the set of landmark pairs includes at least two landmark pairs./ 14. The one or more non-transitory computer-readable media of claim 13, wherein the set of landmark pairs includes at least two landmark pairs. (Bhargava: [0032]-[0033], [0043] The processing of image 62 at fiducial landmarks detection stage 140 may mark image 62 with the identified facial fiducial landmarks to generate a marked image (e.g., image 62L, FIG. 2A) for output. [0044] FIG. 2A shows an example marked image (e.g., image 62L) with two eye pupils EP and 36 facial landmark points LP marked on the image at fiducial landmarks detection stage 140 (e.g., by a Face-SSD model coupled to a face landmark model). The 36 facial landmark points LP can include landmark points on various facial features (e.g., brows, cheek, chin, lips, etc.) and include two anthropometric landmark tragion points (e.g., a left ear tragion (LET) and a right ear tragion (RET) marked on the left ear tragus and the right ear tragus of the person, respectively). In example implementations, a single landmark tragion point (e.g., the LET point for the left ear, or the RET point for the right ear) may be used as a geometrical reference point or fiducial marker to define a bounded portion or area (e.g., a rectangular area) of the image as an ear ROI area (e.g., ROI 64R) for the ear (left ear or the right ear) shown in the corresponding side view image (e.g., image 64). [0045] In example implementations, of the identified fiducial landmarks (identified at fiducial landmarks detection stage 140) only the LET point or only the RET point may be used as a single geometrical reference point to identify the ear ROI area according to whether the 2-D side view face image shows a left ear or a right ear of the person. Kondrashav: [0097] Most landmark detection algorithms produce numerous facial landmarks related to different features of a human face. At least some embodiments of the process disclosed herein only use information about selected, predetermined facial landmarks as an input for the detection of the orientation of the head depicted in the image. In particular, in order to detect a lateral side view of the human head, the process may utilize three groups of landmark features as schematically illustrated in FIG. 5B, i.e., a first group 501 of facial landmarks associated with the left ear of the human head, a second group 502 of facial landmarks associated with the right ear of the human head and a third group 503 of facial landmarks associated with the nose of the human head. [0098] It will be appreciated that each of the groups of landmarks may include a single landmark or multiple landmarks. The groups may include equal numbers of landmarks or different numbers of landmarks. For each group of landmarks, the process may determine a representative image position, e.g., as a geometric center of the detected individual landmarks of the group or another aggregate position. Similarly, the process may determine an aggregate confidence level of the group of landmarks having been detected, e.g., as a product, average or other combination of the individual landmark confidence levels of the respective landmarks of the group.)
Consider Claims 3 and 15.
The combination of Kondrashov and Bhagrava teaches:
3. The computer-implemented method of claim 2, further comprising averaging the three-dimensional positions of the ears./ 15. The one or more non-transitory computer-readable media claim 14, further comprising averaging the three-dimensional positions of the ears. (Bhargava: [0032]-[0033], [0043] The processing of image 62 at fiducial landmarks detection stage 140 may mark image 62 with the identified facial fiducial landmarks to generate a marked image (e.g., image 62L, FIG. 2A) for output. [0044] FIG. 2A shows an example marked image (e.g., image 62L) with two eye pupils EP and 36 facial landmark points LP marked on the image at fiducial landmarks detection stage 140 (e.g., by a Face-SSD model coupled to a face landmark model). The 36 facial landmark points LP can include landmark points on various facial features (e.g., brows, cheek, chin, lips, etc.) and include two anthropometric landmark tragion points (e.g., a left ear tragion (LET) and a right ear tragion (RET) marked on the left ear tragus and the right ear tragus of the person, respectively). In example implementations, a single landmark tragion point (e.g., the LET point for the left ear, or the RET point for the right ear) may be used as a geometrical reference point or fiducial marker to define a bounded portion or area (e.g., a rectangular area) of the image as an ear ROI area (e.g., ROI 64R) for the ear (left ear or the right ear) shown in the corresponding side view image (e.g., image 64). [0045] In example implementations, of the identified fiducial landmarks (identified at fiducial landmarks detection stage 140) only the LET point or only the RET point may be used as a single geometrical reference point to identify the ear ROI area according to whether the 2-D side view face image shows a left ear or a right ear of the person. Kondrashav: [0097] Most landmark detection algorithms produce numerous facial landmarks related to different features of a human face. At least some embodiments of the process disclosed herein only use information about selected, predetermined facial landmarks as an input for the detection of the orientation of the head depicted in the image. In particular, in order to detect a lateral side view of the human head, the process may utilize three groups of landmark features as schematically illustrated in FIG. 5B, i.e., a first group 501 of facial landmarks associated with the left ear of the human head, a second group 502 of facial landmarks associated with the right ear of the human head and a third group 503 of facial landmarks associated with the nose of the human head. [0098] It will be appreciated that each of the groups of landmarks may include a single landmark or multiple landmarks. The groups may include equal numbers of landmarks or different numbers of landmarks. For each group of landmarks, the process may determine a representative image position, e.g., as a geometric center of the detected individual landmarks of the group or another aggregate position. Similarly, the process may determine an aggregate confidence level of the group of landmarks having been detected, e.g., as a product, average or other combination of the individual landmark confidence levels of the respective landmarks of the group.)
Consider Claims 4 and 16.
The combination of Kondrashov and Bhagrava teaches:
4. The computer-implemented method of claim 1, further comprising: for each face landmark included in the plurality of face landmarks, determining a confidence score associated with a location of the face landmark in the one or more images of the user; ordering, based on the confidence scores, the set of one or more landmark pairs to generate an ordered set of landmark pairs, wherein selecting the one or more landmark pairs is based on the ordered set of landmark pairs. / 16. The one or more non-transitory computer-readable media of claim 13, further comprising: for each face landmark included in the plurality of face landmarks, determining a confidence score associated with a location of the face landmark in the one or more images of the user; ordering, based on the confidence scores, the set of one or more landmark pairs to generate an ordered set of landmark pairs, wherein selecting the one or more landmark pairs is based on the ordered set of landmark pairs. (Bhargava: [0033] A trained neural network model analyzes the ear ROI area, pixel-by-pixel, to predict a pixel-sized 2-D location (or a few pixels-sized location) of the ESP in the ear ROI area of the 2-D side view face image. The model takes as input the ear ROI area image, predicts a probability (i.e., a probability value between 0% and 100% or equivalently a confidence value between 0 and 1) that each pixel is the actual or correct ESP, and outputs a confidence map of the predicted ESP locations. The output confidence map may have the same pixel resolution as the input ear ROI area image. Pixels with high confidence values in the confidence map are designated or deemed to be the actual or correct ESP. 9. The computer-implemented method of claim 1, wherein the plurality of face landmarks include one or more of an eye landmark, an eyebrow landmark, a nose landmark, a glabella landmark, a mouth landmark, a chin landmark, or a jawline landmark. (Bhargava: [0032]-[0033], [0043] The processing of image 62 at fiducial landmarks detection stage 140 may mark image 62 with the identified facial fiducial landmarks to generate a marked image (e.g., image 62L, FIG. 2A) for output. [0044] FIG. 2A shows an example marked image (e.g., image 62L) with two eye pupils EP and 36 facial landmark points LP marked on the image at fiducial landmarks detection stage 140 (e.g., by a Face-SSD model coupled to a face landmark model). The 36 facial landmark points LP can include landmark points on various facial features (e.g., brows, cheek, chin, lips, etc.) and include two anthropometric landmark tragion points (e.g., a left ear tragion (LET) and a right ear tragion (RET) marked on the left ear tragus and the right ear tragus of the person, respectively). In example implementations, a single landmark tragion point (e.g., the LET point for the left ear, or the RET point for the right ear) may be used as a geometrical reference point or fiducial marker to define a bounded portion or area (e.g., a rectangular area) of the image as an ear ROI area (e.g., ROI 64R) for the ear (left ear or the right ear) shown in the corresponding side view image (e.g., image 64). [0045] In example implementations, of the identified fiducial landmarks (identified at fiducial landmarks detection stage 140) only the LET point or only the RET point may be used as a single geometrical reference point to identify the ear ROI area according to whether the 2-D side view face image shows a left ear or a right ear of the person. [0073] In example implementations, for training the U-Net model, the GT ESP locations may be defined, for example, by a Gaussian distribution function: C=exp−(d 2/(2*δ2),
where C is the confidence value, d is the distance to the GT ESP location, and δ is the standard deviation of the Gaussian distribution. A confidence in the model's ESP prediction will be higher for pixels closer to the GT ESP (and equal to 1 for the GT). A small value of the standard deviation δ in the definition of the GT may produce a largely blank confidence map, which can mislead the model in to generating a trivial result predicting zero confidence everywhere. Conversely, a large value of the standard deviation δ in the definition of the GT, may produce an overly diffuse confidence map, which can cause the model to fail to predict a precise location for the EPS. In example implementations, a value of standard deviation δ in the definition of the GT may be selected based on a desired precision in the predicted locations of the EPS. In example implementations, the value of standard deviation δ may be selected to be in a range of about 2 to 10 pixels (e.g., 3 pixels) as a satisfactory or acceptable precision required in the predicted locations of the EPS predicted by the U-Net model. Kondrashav: [0097] Most landmark detection algorithms produce numerous facial landmarks related to different features of a human face. At least some embodiments of the process disclosed herein only use information about selected, predetermined facial landmarks as an input for the detection of the orientation of the head depicted in the image. In particular, in order to detect a lateral side view of the human head, the process may utilize three groups of landmark features as schematically illustrated in FIG. 5B, i.e., a first group 501 of facial landmarks associated with the left ear of the human head, a second group 502 of facial landmarks associated with the right ear of the human head and a third group 503 of facial landmarks associated with the nose of the human head. [0098] It will be appreciated that each of the groups of landmarks may include a single landmark or multiple landmarks. The groups may include equal numbers of landmarks or different numbers of landmarks. For each group of landmarks, the process may determine a representative image position, e.g., as a geometric center of the detected individual landmarks of the group or another aggregate position. Similarly, the process may determine an aggregate confidence level of the group of landmarks having been detected, e.g., as a product, average or other combination of the individual landmark confidence levels of the respective landmarks of the group.)
Consider Claims 5 and 17.
The combination of Kondrashov and Bhagrava teaches:
5. The computer-implemented method of claim 1, wherein determining the three-dimensional positions of ears based on the set of landmark pairs of the user comprises: generating, based on the one or more images, two-dimensional landmark coordinates for the plurality of face landmarks using a face detection model; and generating, based on the estimated head pose, landmark depth estimates for the two-dimensional landmark coordinates; generating, based on the two-dimensional landmark coordinates and the landmark depth estimates, three-dimensional landmark coordinates, wherein the three-dimensional positions of the ears are based on the three-dimensional landmark coordinates. / 17. The one or more non-transitory computer-readable media of claim 13, wherein determining the three-dimensional positions of ears based on the set of landmark pairs of the user comprises: generating, based on the one or more images, two-dimensional landmark coordinates for the plurality of face landmarks using a face detection model; generating, based on the estimated head pose, landmark depth estimates for the two-dimensional landmark coordinates; and generating, based on the two-dimensional landmark coordinates and the landmark depth estimates, three-dimensional landmark coordinates, wherein the three-dimensional positions of the ears are based on the three-dimensional landmark coordinates. (Kondrashov: [0111] In initial step S31, the process receives a sequence of input data sets associated with a corresponding sequence of video frames. The sequence of video frames having been captured at a certain frame rate, measured as frames per second (FSP). The input data may include the actual video frames. In that case, for each video frame, the process performs landmark detection, e.g., as described in connection with step S21 of the side view detection process of FIG. 4 . However, it will be appreciated that the stability detection process may reuse any detected landmark coordinates that have already been detected in a video frame as part of the side view detection. Accordingly, instead of, or in addition to, receiving the video frames, the stability detection process may receive a sequence of landmark data sets as its input. Each landmark data set includes information about the detected landmarks in an image, e.g., in a video frame. For each detected landmark, the information includes the 2D image coordinates of the detected landmark and the associated confidence level or score. [0112] In subsequent step S32, the process computes a weighted sum of the image coordinates of the detected landmarks:
PNG
media_image1.png
45
114
media_image1.png
Greyscale
[0113] where zi is a coordinate vector associated with i-th landmark and z is the weighted sum of landmark positions, each landmark coordinate vector being weighted by its detection confidence level ci. The weighted sum z may be considered as a generalized head center. N is the number of landmarks. It will be appreciated that the stability detection may be performed based on all detected landmarks or only based on a subset of the detected landmarks, e.g., the landmarks selected for the side view detection, as described in connection with step S21 of the side view detection process of FIG. 4 . It will further be appreciated that the process may be based on a different combination of landmark positions of all or of a selected subset of the detected landmarks. Bhargava: [0081] System 600 may include a processor 17, a memory 18, a display 19, and a 3-D head model 610 of the person. 3-D head model 610 of the person's head may include 3-D representations or depictions of the person's facial features (e.g., eyes, ears, nose, etc.). The 3-D head model may be used, for example, as a mannequin or dummy, for fitting glasses to the person in VTO sessions. System 600 may be included in, or coupled to, system 100. [0082] System 600 may receive 2-D coordinates (e.g. (x, y)) of the predicted 2-D ESP (e.g., 500R-ESP, FIG. 5 ) for the person, for example, from system 100. In system 600, processor 17 may execute instructions (stored, e.g., in memory 18) to snap the predicted 2-D ESP having two-dimensional co-ordinates (x, y) on to the model of the person's ear (e.g., to a lobe of the ear), and project it by ray projection through 3-D space to a 3-D ESP point (x, y, z) on a side of the person's ear. A depth search may be carried out in a predefined cuboid region of the 3-D head model to find a depth point (e.g., co-ordinate z) for locating a projected 3-D ESP point on the 3-D head model 610 at the depth z behind or to a side of the person's ear. The (x, y) coordinates of the projected 3-D ESP point may be the same as the (x, y) coordinates of the 2-D ESP point. However, the z coordinate of the projected 3-D ESP point may be set to be the z coordinate of the deepest point found in the depth search of the cuboid region. System 600 may then generate virtual glasses to fit the 3-D head model with temple pieces of the glasses resting on, or passing through, the projected 3-D ESP point for a good fit. [0083] FIG. 6A shows, for example, a portion of 3-D head model 610 of a person processed by system 600 with an original predicted 2-D ESP (e.g., ESP 62 (x, y)) snapped on an outer lobe of the person's ear. [0084] FIG. 6B illustrates cuboid regions (i.e., convex polyhedrons) of 3-D head model 610 that may be searched by system 600 to find a depth point for locating a projected 3-D ESP point (e.g., ESP 64 (x, y, z), FIG. 6C) at a depth z behind, or to a side of, the person's ear. [0085] FIG. 6C shows, for example, the portion of 3-D head model 610 including the person's ear with the original predicted 2-D ESP (e.g., ESP 62 (x, y)) snapped on an outer lobe of the person's ear, and the projected 3-D ESP point (e.g., ESP 64 (x, y, z),) disposed at a depth z behind, or to a side of, the person's ear. [0086] FIG. 6D illustrates another view of 3-D head model 610 with the projected 3-D ESP point (e.g., ESP 64 (x, y, z)) disposed at a depth z behind, and to a side of, the person's ear. [0087] FIG. 6E illustrates the example 3-D head model 610 fitted with a pair of virtual glasses (e.g., glasses 90) having a temple piece (e.g., temple piece 92) passing through or attached to the projected 3-D ESP point (e.g., ESP 64 (x, y, z)) in 3-D space.)
Consider Claims 6 and 18.
The combination of Kondrashov and Bhagrava teaches:
6. The computer-implemented method of claim 5, further comprising: determining a head pose vector based on the two-dimensional landmark coordinates for the plurality of face landmarks; and determining the landmark depth estimates based on the head pose vector. / 18. The one or more non-transitory computer-readable media of claim 17, further comprising: determining a head pose vector based on the two-dimensional landmark coordinates for the plurality of face landmarks; and determining the landmark depth estimates based on the head pose vector. (Kondrashov: [0111] In initial step S31, the process receives a sequence of input data sets associated with a corresponding sequence of video frames. The sequence of video frames having been captured at a certain frame rate, measured as frames per second (FSP). The input data may include the actual video frames. In that case, for each video frame, the process performs landmark detection, e.g., as described in connection with step S21 of the side view detection process of FIG. 4 . However, it will be appreciated that the stability detection process may reuse any detected landmark coordinates that have already been detected in a video frame as part of the side view detection. Accordingly, instead of, or in addition to, receiving the video frames, the stability detection process may receive a sequence of landmark data sets as its input. Each landmark data set includes information about the detected landmarks in an image, e.g., in a video frame. For each detected landmark, the information includes the 2D image coordinates of the detected landmark and the associated confidence level or score. [0112] In subsequent step S32, the process computes a weighted sum of the image coordinates of the detected landmarks:
PNG
media_image1.png
45
114
media_image1.png
Greyscale
[0113] where zi is a coordinate vector associated with i-th landmark and z is the weighted sum of landmark positions, each landmark coordinate vector being weighted by its detection confidence level ci. The weighted sum z may be considered as a generalized head center. N is the number of landmarks. It will be appreciated that the stability detection may be performed based on all detected landmarks or only based on a subset of the detected landmarks, e.g., the landmarks selected for the side view detection, as described in connection with step S21 of the side view detection process of FIG. 4 . It will further be appreciated that the process may be based on a different combination of landmark positions of all or of a selected subset of the detected landmarks. Bhargava: [0081] System 600 may include a processor 17, a memory 18, a display 19, and a 3-D head model 610 of the person. 3-D head model 610 of the person's head may include 3-D representations or depictions of the person's facial features (e.g., eyes, ears, nose, etc.). The 3-D head model may be used, for example, as a mannequin or dummy, for fitting glasses to the person in VTO sessions. System 600 may be included in, or coupled to, system 100. [0082] System 600 may receive 2-D coordinates (e.g. (x, y)) of the predicted 2-D ESP (e.g., 500R-ESP, FIG. 5 ) for the person, for example, from system 100. In system 600, processor 17 may execute instructions (stored, e.g., in memory 18) to snap the predicted 2-D ESP having two-dimensional co-ordinates (x, y) on to the model of the person's ear (e.g., to a lobe of the ear), and project it by ray projection through 3-D space to a 3-D ESP point (x, y, z) on a side of the person's ear. A depth search may be carried out in a predefined cuboid region of the 3-D head model to find a depth point (e.g., co-ordinate z) for locating a projected 3-D ESP point on the 3-D head model 610 at the depth z behind or to a side of the person's ear. The (x, y) coordinates of the projected 3-D ESP point may be the same as the (x, y) coordinates of the 2-D ESP point. However, the z coordinate of the projected 3-D ESP point may be set to be the z coordinate of the deepest point found in the depth search of the cuboid region. System 600 may then generate virtual glasses to fit the 3-D head model with temple pieces of the glasses resting on, or passing through, the projected 3-D ESP point for a good fit. [0083] FIG. 6A shows, for example, a portion of 3-D head model 610 of a person processed by system 600 with an original predicted 2-D ESP (e.g., ESP 62 (x, y)) snapped on an outer lobe of the person's ear. [0084] FIG. 6B illustrates cuboid regions (i.e., convex polyhedrons) of 3-D head model 610 that may be searched by system 600 to find a depth point for locating a projected 3-D ESP point (e.g., ESP 64 (x, y, z), FIG. 6C) at a depth z behind, or to a side of, the person's ear. [0085] FIG. 6C shows, for example, the portion of 3-D head model 610 including the person's ear with the original predicted 2-D ESP (e.g., ESP 62 (x, y)) snapped on an outer lobe of the person's ear, and the projected 3-D ESP point (e.g., ESP 64 (x, y, z),) disposed at a depth z behind, or to a side of, the person's ear. [0086] FIG. 6D illustrates another view of 3-D head model 610 with the projected 3-D ESP point (e.g., ESP 64 (x, y, z)) disposed at a depth z behind, and to a side of, the person's ear. [0087] FIG. 6E illustrates the example 3-D head model 610 fitted with a pair of virtual glasses (e.g., glasses 90) having a temple piece (e.g., temple piece 92) passing through or attached to the projected 3-D ESP point (e.g., ESP 64 (x, y, z)) in 3-D space.)
Consider Claims 7.
The combination of Kondrashov and Bhagrava teaches:
7. The computer-implemented method of claim 5, further comprising: determining a head pose based on the two-dimensional landmark coordinates for the plurality of face landmarks, wherein the head pose is within 45 degrees of a principal point; and determining the landmark depth estimates based on the head pose vector. (Bhargava: [0072] In an example implementation, the U-Net model is trained using only about 200 images of persons taken from only two different camera viewpoints (e.g., ˜90 degrees for front view face images, and ˜45 degrees for side view face (ear) images). The model generalizes well on different lighting and camera angles. FIG. 4A shows, for purposes of illustration, three example ear ROI area images (e.g., ear ROI 64R-c, ear ROI 64R-d. and ear ROI 64R-e) that can be used as training data for the U-Net model. The ear ROI area images in the training data may be annotated with the actual or ground truth (GT) ESP locations of the persons' ears in the images. [0073] In example implementations, for training the U-Net model, the GT ESP locations may be defined, for example, by a Gaussian distribution function: C=exp−(d 2/(2*δ2),
where C is the confidence value, d is the distance to the GT ESP location, and δ is the standard deviation of the Gaussian distribution. A confidence in the model's ESP prediction will be higher for pixels closer to the GT ESP (and equal to 1 for the GT). A small value of the standard deviation δ in the definition of the GT may produce a largely blank confidence map, which can mislead the model in to generating a trivial result predicting zero confidence everywhere. Conversely, a large value of the standard deviation δ in the definition of the GT, may produce an overly diffuse confidence map, which can cause the model to fail to predict a precise location for the EPS. In example implementations, a value of standard deviation δ in the definition of the GT may be selected based on a desired precision in the predicted locations of the EPS. In example implementations, the value of standard deviation δ may be selected to be in a range of about 2 to 10 pixels (e.g., 3 pixels) as a satisfactory or acceptable precision required in the predicted locations of the EPS predicted by the U-Net model.)
Consider Claim 8.
The combination of Kondrashov and Bhagrava teaches:
8. The computer-implemented method of claim 5, wherein the three-dimensional positions of the ears of the user are determined based on one or more relationships in an enrollment head geometry, wherein the one or more relationships relate the three-dimensional landmark coordinates to the three-dimensional positions of the ears./ 19. The one or more non-transitory computer-readable media of claim 17, wherein the three-dimensional positions of the ears of the user are determined based on one or more relationships in an enrollment head geometry, wherein the one or more relationships relate the three-dimensional landmark coordinates to the three-dimensional positions of the ears. (Kondrashov: [0111] In initial step S31, the process receives a sequence of input data sets associated with a corresponding sequence of video frames. The sequence of video frames having been captured at a certain frame rate, measured as frames per second (FSP). The input data may include the actual video frames. In that case, for each video frame, the process performs landmark detection, e.g., as described in connection with step S21 of the side view detection process of FIG. 4 . However, it will be appreciated that the stability detection process may reuse any detected landmark coordinates that have already been detected in a video frame as part of the side view detection. Accordingly, instead of, or in addition to, receiving the video frames, the stability detection process may receive a sequence of landmark data sets as its input. Each landmark data set includes information about the detected landmarks in an image, e.g., in a video frame. For each detected landmark, the information includes the 2D image coordinates of the detected landmark and the associated confidence level or score. [0112] In subsequent step S32, the process computes a weighted sum of the image coordinates of the detected landmarks:
PNG
media_image1.png
45
114
media_image1.png
Greyscale
[0113] where zi is a coordinate vector associated with i-th landmark and z is the weighted sum of landmark positions, each landmark coordinate vector being weighted by its detection confidence level ci. The weighted sum z may be considered as a generalized head center. N is the number of landmarks. It will be appreciated that the stability detection may be performed based on all detected landmarks or only based on a subset of the detected landmarks, e.g., the landmarks selected for the side view detection, as described in connection with step S21 of the side view detection process of FIG. 4 . It will further be appreciated that the process may be based on a different combination of landmark positions of all or of a selected subset of the detected landmarks. Bhargava: [0043] The processing of image 62 at fiducial landmarks detection stage 140 may mark image 62 with the identified facial fiducial landmarks to generate a marked image (e.g., image 62L, FIG. 2A) for output. [0044] FIG. 2A shows an example marked image (e.g., image 62L) with two eye pupils EP and 36 facial landmark points LP marked on the image at fiducial landmarks detection stage 140 (e.g., by a Face-SSD model coupled to a face landmark model). The 36 facial landmark points LP can include landmark points on various facial features (e.g., brows, cheek, chin, lips, etc.) and include two anthropometric landmark tragion points (e.g., a left ear tragion (LET) and a right ear tragion (RET) marked on the left ear tragus and the right ear tragus of the person, respectively). In example implementations, a single landmark tragion point (e.g., the LET point for the left ear, or the RET point for the right ear) may be used as a geometrical reference point or fiducial marker to define a bounded portion or area (e.g., a rectangular area) of the image as an ear ROI area (e.g., ROI 64R) for the ear (left ear or the right ear) shown in the corresponding side view image (e.g., image 64). [0045] In example implementations, of the identified fiducial landmarks (identified at fiducial landmarks detection stage 140) only the LET point or only the RET point may be used as a single geometrical reference point to identify the ear ROI area according to whether the 2-D side view face image shows a left ear or a right ear of the person. [0081] System 600 may include a processor 17, a memory 18, a display 19, and a 3-D head model 610 of the person. 3-D head model 610 of the person's head may include 3-D representations or depictions of the person's facial features (e.g., eyes, ears, nose, etc.). The 3-D head model may be used, for example, as a mannequin or dummy, for fitting glasses to the person in VTO sessions. System 600 may be included in, or coupled to, system 100. [0082] System 600 may receive 2-D coordinates (e.g. (x, y)) of the predicted 2-D ESP (e.g., 500R-ESP, FIG. 5 ) for the person, for example, from system 100. In system 600, processor 17 may execute instructions (stored, e.g., in memory 18) to snap the predicted 2-D ESP having two-dimensional co-ordinates (x, y) on to the model of the person's ear (e.g., to a lobe of the ear), and project it by ray projection through 3-D space to a 3-D ESP point (x, y, z) on a side of the person's ear. A depth search may be carried out in a predefined cuboid region of the 3-D head model to find a depth point (e.g., co-ordinate z) for locating a projected 3-D ESP point on the 3-D head model 610 at the depth z behind or to a side of the person's ear. The (x, y) coordinates of the projected 3-D ESP point may be the same as the (x, y) coordinates of the 2-D ESP point. However, the z coordinate of the projected 3-D ESP point may be set to be the z coordinate of the deepest point found in the depth search of the cuboid region. System 600 may then generate virtual glasses to fit the 3-D head model with temple pieces of the glasses resting on, or passing through, the projected 3-D ESP point for a good fit. [0083] FIG. 6A shows, for example, a portion of 3-D head model 610 of a person processed by system 600 with an original predicted 2-D ESP (e.g., ESP 62 (x, y)) snapped on an outer lobe of the person's ear. [0084] FIG. 6B illustrates cuboid regions (i.e., convex polyhedrons) of 3-D head model 610 that may be searched by system 600 to find a depth point for locating a projected 3-D ESP point (e.g., ESP 64 (x, y, z), FIG. 6C) at a depth z behind, or to a side of, the person's ear. [0085] FIG. 6C shows, for example, the portion of 3-D head model 610 including the person's ear with the original predicted 2-D ESP (e.g., ESP 62 (x, y)) snapped on an outer lobe of the person's ear, and the projected 3-D ESP point (e.g., ESP 64 (x, y, z),) disposed at a depth z behind, or to a side of, the person's ear. [0086] FIG. 6D illustrates another view of 3-D head model 610 with the projected 3-D ESP point (e.g., ESP 64 (x, y, z)) disposed at a depth z behind, and to a side of, the person's ear. [0087] FIG. 6E illustrates the example 3-D head model 610 fitted with a pair of virtual glasses (e.g., glasses 90) having a temple piece (e.g., temple piece 92) passing through or attached to the projected 3-D ESP point (e.g., ESP 64 (x, y, z)) in 3-D space.)
Consider Claim 9.
The combination of Kondrashov and Bhagrava teaches:
9. The computer-implemented method of claim 1, wherein the plurality of face landmarks include one or more of an eye landmark, an eyebrow landmark, a nose landmark, a glabella landmark, a mouth landmark, a chin landmark, or a jawline landmark. (Bhargava: [0032]-[0033], [0043] The processing of image 62 at fiducial landmarks detection stage 140 may mark image 62 with the identified facial fiducial landmarks to generate a marked image (e.g., image 62L, FIG. 2A) for output. [0044] FIG. 2A shows an example marked image (e.g., image 62L) with two eye pupils EP and 36 facial landmark points LP marked on the image at fiducial landmarks detection stage 140 (e.g., by a Face-SSD model coupled to a face landmark model). The 36 facial landmark points LP can include landmark points on various facial features (e.g., brows, cheek, chin, lips, etc.) and include two anthropometric landmark tragion points (e.g., a left ear tragion (LET) and a right ear tragion (RET) marked on the left ear tragus and the right ear tragus of the person, respectively). In example implementations, a single landmark tragion point (e.g., the LET point for the left ear, or the RET point for the right ear) may be used as a geometrical reference point or fiducial marker to define a bounded portion or area (e.g., a rectangular area) of the image as an ear ROI area (e.g., ROI 64R) for the ear (left ear or the right ear) shown in the corresponding side view image (e.g., image 64). [0045] In example implementations, of the identified fiducial landmarks (identified at fiducial landmarks detection stage 140) only the LET point or only the RET point may be used as a single geometrical reference point to identify the ear ROI area according to whether the 2-D side view face image shows a left ear or a right ear of the person. Kondrashav: [0097] Most landmark detection algorithms produce numerous facial landmarks related to different features of a human face. At least some embodiments of the process disclosed herein only use information about selected, predetermined facial landmarks as an input for the detection of the orientation of the head depicted in the image. In particular, in order to detect a lateral side view of the human head, the process may utilize three groups of landmark features as schematically illustrated in FIG. 5B, i.e., a first group 501 of facial landmarks associated with the left ear of the human head, a second group 502 of facial landmarks associated with the right ear of the human head and a third group 503 of facial landmarks associated with the nose of the human head. [0098] It will be appreciated that each of the groups of landmarks may include a single landmark or multiple landmarks. The groups may include equal numbers of landmarks or different numbers of landmarks. For each group of landmarks, the process may determine a representative image position, e.g., as a geometric center of the detected individual landmarks of the group or another aggregate position. Similarly, the process may determine an aggregate confidence level of the group of landmarks having been detected, e.g., as a product, average or other combination of the individual landmark confidence levels of the respective landmarks of the group.)
Consider Claims 10.
The combination of Kondrashov and Bhagrava teaches:
10. The computer-implemented method of claim 1, wherein the one or more landmark pairs include at least one of a bridge-to-chin landmark pair, a bridge-to-jawline landmark pair, or an eye edge-to-jawline landmark pair. (Bhargava: [0032]-[0033], [0043] The processing of image 62 at fiducial landmarks detection stage 140 may mark image 62 with the identified facial fiducial landmarks to generate a marked image (e.g., image 62L, FIG. 2A) for output. [0044] FIG. 2A shows an example marked image (e.g., image 62L) with two eye pupils EP and 36 facial landmark points LP marked on the image at fiducial landmarks detection stage 140 (e.g., by a Face-SSD model coupled to a face landmark model). The 36 facial landmark points LP can include landmark points on various facial features (e.g., brows, cheek, chin, lips, etc.) and include two anthropometric landmark tragion points (e.g., a left ear tragion (LET) and a right ear tragion (RET) marked on the left ear tragus and the right ear tragus of the person, respectively). In example implementations, a single landmark tragion point (e.g., the LET point for the left ear, or the RET point for the right ear) may be used as a geometrical reference point or fiducial marker to define a bounded portion or area (e.g., a rectangular area) of the image as an ear ROI area (e.g., ROI 64R) for the ear (left ear or the right ear) shown in the corresponding side view image (e.g., image 64). [0045] In example implementations, of the identified fiducial landmarks (identified at fiducial landmarks detection stage 140) only the LET point or only the RET point may be used as a single geometrical reference point to identify the ear ROI area according to whether the 2-D side view face image shows a left ear or a right ear of the person. Kondrashav: [0097] Most landmark detection algorithms produce numerous facial landmarks related to different features of a human face. At least some embodiments of the process disclosed herein only use information about selected, predetermined facial landmarks as an input for the detection of the orientation of the head depicted in the image. In particular, in order to detect a lateral side view of the human head, the process may utilize three groups of landmark features as schematically illustrated in FIG. 5B, i.e., a first group 501 of facial landmarks associated with the left ear of the human head, a second group 502 of facial landmarks associated with the right ear of the human head and a third group 503 of facial landmarks associated with the nose of the human head. [0098] It will be appreciated that each of the groups of landmarks may include a single landmark or multiple landmarks. The groups may include equal numbers of landmarks or different numbers of landmarks. For each group of landmarks, the process may determine a representative image position, e.g., as a geometric center of the detected individual landmarks of the group or another aggregate position. Similarly, the process may determine an aggregate confidence level of the group of landmarks having been detected, e.g., as a product, average or other combination of the individual landmark confidence levels of the respective landmarks of the group.)
Consider Claims 11.
The combination of Kondrashov and Bhagrava teaches:
11. The computer-implemented method of claim 1, wherein: one or more speakers generate an audio output from the processed audio signals; and the one or more speakers include one or more of headrest speakers, gaming chair speakers, or sound bar speakers. (Bhargava: [0126]
Device 950 may also communicate audibly using audio codec 960, which may receive spoken information from a user and convert it to usable digital information. Audio codec 960 may likewise generate audible sound for a user, such as through a speaker, e.g., in a handset of device 950. Such sound may include sound from voice telephone calls, may include recorded sound (e.g., voice messages, music files, etc.) and may also include sound generated by applications operating on device 950. [0127] The computing device 950 may be implemented in a number of different forms, as shown in the figure. For example, it may be implemented as a cellular telephone 990. It may also be implemented as part of a smart phone 99892, personal digital assistant, or other similar mobile device. Kondrashov: [0110] FIG. 6 schematically illustrates a process for detecting whether a human head is depicted in a sequence of images, in particular in a sequence of video frames, in a sufficiently stable manner. The process of FIG. 6 is a possible implementation of step S3 of FIG. 3 . [0111] In initial step S31, the process receives a sequence of input data sets associated with a corresponding sequence of video frames. The sequence of video frames having been captured at a certain frame rate, measured as frames per second (FSP). The input data may include the actual video frames. In that case, for each video frame, the process performs landmark detection, e.g., as described in connection with step S21 of the side view detection process of FIG. 4 . [0118] In the evolution equation, x is related to the current video frame while xprev is related to the preceding video frame, i.e., the video frame processed during the preceding iteration of the process. ΔT=1/FPS is the reciprocal of the frame rate FPS (frames per second). The evolution equation further includes a process noise term w, which may be drawn from a zero mean multivariate normal distribution with covariance Q, i.e. w˜N(0,Q). The covariance Q may be a diagonal matrix with suitable predetermined values.)
Consider Claims 12.
The combination of Kondrashov and Bhagrava teaches:
12. The computer-implemented method of claim 1, wherein the one or more processed audio signals apply one or more audio effects to the one or more audio signals, wherein the audio effects include one or more of a spatial audio effect, noise cancellation, or crosstalk cancellation. (Bhargava: [0126] Device 950 may also communicate audibly using audio codec 960, which may receive spoken information from a user and convert it to usable digital information. Audio codec 960 may likewise generate audible sound for a user, such as through a speaker, e.g., in a handset of device 950. Such sound may include sound from voice telephone calls, may include recorded sound (e.g., voice messages, music files, etc.) and may also include sound generated by applications operating on device 950. [0127] The computing device 950 may be implemented in a number of different forms, as shown in the figure. For example, it may be implemented as a cellular telephone 990. It may also be implemented as part of a smart phone 99892, personal digital assistant, or other similar mobile device. Kondrashov: [0110] FIG. 6 schematically illustrates a process for detecting whether a human head is depicted in a sequence of images, in particular in a sequence of video frames, in a sufficiently stable manner. The process of FIG. 6 is a possible implementation of step S3 of FIG. 3 . [0111] In initial step S31, the process receives a sequence of input data sets associated with a corresponding sequence of video frames. The sequence of video frames having been captured at a certain frame rate, measured as frames per second (FSP). The input data may include the actual video frames. In that case, for each video frame, the process performs landmark detection, e.g., as described in connection with step S21 of the side view detection process of FIG. 4 . [0118] In the evolution equation, x is related to the current video frame while xprev is related to the preceding video frame, i.e., the video frame processed during the preceding iteration of the process. ΔT=1/FPS is the reciprocal of the frame rate FPS (frames per second). The evolution equation further includes a process noise term w, which may be drawn from a zero mean multivariate normal distribution with covariance Q, i.e. w˜N(0,Q). The covariance Q may be a diagonal matrix with suitable predetermined values.)
Conclusion
The prior art made of record in form PTO-892 and not relied upon is considered pertinent to applicant's disclosure.
PNG
media_image2.png
288
902
media_image2.png
Greyscale
Any inquiry concerning this communication or earlier communications from the examiner should be directed to TAHMINA ANSARI whose telephone number is 571-270-3379. The examiner can normally be reached on IFP Flex - Monday through Friday 9 to 5.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, O’NEAL MISTRY can be reached on 313-446-4912. The fax phone numbers for the organization where this application or proceeding is assigned are 571-273-8300 for regular communications and 571-273-8300 for After Final communications. TC 2600’s customer service number is 571-272-2600.
Any inquiry of a general nature or relating to the status of this application or proceeding should be directed to the receptionist whose telephone number is 571-272-2600.
/Tahmina Ansari/
July 17, 2026
/TAHMINA N ANSARI/Primary Examiner, Art Unit 2674