Detailed Action
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s arguments with respect to claims 1, 3-10, 12-19, and 28-30 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 9, 10, 18 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Hicks (US 20160066034 A1), in view of Mande et al. (US 2020/0311392 A1), (hereinafter, Mande) and further in view of Rawat et al. (US 2021/0056667 A1), (hereinafter, Rawat) and Mu (CN 110769148 A).
Regarding claim 1, Hicks teaches a meter device configured to measure facial attention to a media device by at least one panelist member of a panelist household and configured to protect the privacy of the at least one panelist member (Hicks, “In some audience measurement systems, people data is collected for a media exposure environment (e.g., a television room, a family room, a living room, a bar, a restaurant, an office space, a cafeteria, etc.) by capturing a series of images of the environment and analyzing the images to determine, for example, an identity of one or more persons present in the media exposure environment, an amount of people present in the media exposure environment during one or more times and/or periods of time, an amount of attention being paid to a media presentation by one or more persons, a gesture made by a person in the media exposure environment, etc.”, pg. 1, paragraph 0012, lines 1-11, “Example methods, apparatus, and articles of manufacture disclosed herein detect people in an environment in an efficient, accurate and private manner that avoids drawbacks of known image data recognition systems”, pg. 1, paragraph 0014, lines 1-4), the meter device comprising:
at least one memory storing machine readable instructions; and a processor configured to at least one of instantiate or execute the machine readable instructions (Hicks, “Flowcharts representative of example machine readable instructions for implementing the example person detector 200 of FIGS. 2 and/or 3 are shown in FIGS. 6 and 7. In this example, the machine readable instructions comprise a program for execution by a processor Such as the processor 812 shown in the example processor platform 800 discussed below in connection with FIG. 8.”, pg. 8, paragraph 0045, lines 1-7) to:
obtain, from a camera embedded within the meter device and installed at a room of the panelist household, input image data comprising image frames of the room of the panelist household, wherein the panelist household provides audience attention data about the at least one panelist member of the panelist household to an audience measurement central facility (Hicks, “FIG. 1 is an illustration of an example media exposure environment 100 including an information presentation device 102, an example meter 104, and an audience 106 including a first person 108 and a second person 110. In the illustrated example of FIG. 1, the information presentation device 102 is a television and the media exposure environment 100 is a room of a household (e.g., a room in a home of a panelist such as the home of a “Nielsen family’) that has been statistically selected to develop television ratings data for population(s)/demographic(s) of interest. In the illustrated example of FIG. 1, one or more persons of the house hold have registered with an audience measurement entity (e.g., by agreeing to be a panelist) and have provided demographic information to the audience measurement entity as part of a registration process to enable associating demographics with viewing activities (e.g., media exposure).”, pg. 2, paragraph 0017, lines 1-16, “As disclosed below in connection with FIGS. 2 and 3, the example meter 104 of FIG. 1 uses light information (e.g., brightness value sequences) representative of the environment 100 to detect people (e.g., the first and second persons 108, 100). In some examples, the example meter 104 of FIG. 1 correlates the people detection information with media identifying information collected from the environment 100. In some examples, the correlation is done by another device Such as a remote data collection server. Thus, the example meter 104 of FIG. 1 generates audience measurement data (e.g., exposure information for particular media) for the environment 100.”, pg. 3, paragraph 0018, lines 46-57, Image frames are collected for a panelist household to detect the number of people present for a given . This information is combined with data corresponding to the media being presented in the environment (e.g., television program) to generate exposure data, representing how many people are paying attention to the presented media.);
provide an attentiveness metric about the at least one panelist member as the audience attention data to the audience measurement facility, wherein the input image data (and the adjusted input image data) is not obtained by the audience measurement facility to protect the privacy of the at least one panelist member of the panelist household (Hicks, “Notably, when an array of light sensors is used to obtain the light information, examples disclosed herein are able to detect a number of people in the environment while maintaining the privacy of the people and without having to perform computationally expensive image recognition algorithms.”, pg. 2, paragraph 0016, lines 47-52, “The audience measurement entity associated with the example data collection facility 214 of FIG. 2 utilizes the people tallies generated by the person detector 200 in conjunction with the media identifying data collected by the media detector 210 to generate exposure information. The information from many panelist locations may be collected and analyzed to generate ratings representative of media exposure by one or more populations of interest.”, pg. 5, paragraph 0029, lines 16-22, “Ones of the example pulse-based person detections 310 of FIG. 3 captured during a same period of time (e.g., according to a time stamp) can be added together to generate a pulse-based person count.”, pg. 7, paragraph 0037, lines 47-50, The detected count of panelists acts as an attentiveness metric, indicating how many individuals are engaging with the presented media. Rather than transmitting full image data, only these derived counts are sent to a data collection facility for further analysis.).
Hicks does not teach attempt to identify facial landmarks of the at least one panelist member from the image frames of the room of the panelist household, the facial landmarks corresponding to coordinates of landmarks of a face of the at least one panelist member detected in the input image data; determine a first distance; determine a second distance; and compare a quotient of the first distance and the second distance to a threshold to determine an attentiveness metric about the at least one panelist member.
However, Mande teaches attempt to identify facial landmarks of the at least one panelist member from the image frames of the room of the panelist household, the facial landmarks corresponding to coordinates of landmarks of a face of the at least one panelist member detected in the input image data; determine a first distance; determine a second distance; and compare a quotient of the first distance and the second distance to a threshold to determine an attentiveness metric about the at least one panelist member (Mande, “As illustrated in FIG. 6, once a face is detected (blocks 520 or 540 of FIG. 5), landmark points are detected on the face (block 610).”, pg. 9, paragraph 0147, lines 1-3, “1) Calculating a sum of distance between points on right eyebrow (e.g. points 18 to 22 illustrated in FIG. 7) to the points on the right border of the face ( e.g. points 1 to 5 illustrated in FIG. 7). Adding to it a sum of distance between points on the nose to the points on the right border of the face. 2) Calculating a total distance for the left side of face in a similar manner to that described in stage (1)”, pg. 9, paragraphs 0148 and 0149, “Calculating a distance ratio i.e. a ratio between total distance of right side to the total distance of left side… Based on the distance ratio, a direction of attention to be front, right or left may be estimated (block 630). For example, if the distance ratio is close to the value of 1, e.g. 0.66 to 1.5, then a person may substantially be looking to the front direction. If a distance ratio is greater than 1.5, then a person may substantially be looking to his left direction.”, pg. 9, paragraphs 0150-0152, see landmark examples in Fig. 7, see resulting attention metric data in Fig. 3)
Hicks teaches detecting a number of panelists for a particular time to privately transmit attention data for media exposure analysis (Hicks, “To generate the people data, some systems attempt to recognize objects as humans in image data representative of the monitored environment. In Such systems, a tally is maintained for each frame of image data to reflect an amount of people in the environment at a particular time.”, pg. 1, paragraph 0013, lines 1-5). Hicks teaches that audience measurement systems determine an amount of attention being paid to a presentation of media (Hicks, “In some audience measurement systems, people data is collected for a media exposure environment (e.g., a television room, a family room, a living room, a bar, a restaurant, an office space, a cafeteria, etc.) by capturing a series of images of the environment and analyzing the images to determine, for example… an amount of attention being paid to a media presentation by one or more persons”, pg. 1, paragraph 0012, lines 1-10). Mande teaches measuring facial attention by detecting facial landmarks and evaluating their distances to generate individual facial attention data (Mande, pg. 9, paragraphs 0148-0152, see Figs. 3 and 7). Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to have modified Hicks to include the facial landmark measurements and individual facial attention data generation methods as taught by Mande (Mande, pg. 9, paragraphs 0148-0152, see Figs. 3 and 7). The motivation for doing so would have been to consider, in addition to the count of panelists, individual facial attention corresponding to particular objects in the presentation of media, thereby improving the overall measurement of panelists interest and attention (as suggested by Mande, “Moreover, obtaining information on the interest of the audience towards specific objects, or even parts of such objects, is even more valuable, as the interest of the watched scene can be measured in a more accurate manner, while evaluating the objects that contributed to the success of the event, in terms of the audience's interest in these objects during the event.”, pg. 1, paragraph 0005). The combination of Hicks in view of Mande would detect individual panelists and track facial attention for each across a set of frames, generate attention metric data per panelist (Mande, see Fig. 3) to be transmitted privately to the data collection facility for media exposure analysis. Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine the teachings of Hicks with Mande to obtain the invention above.
Hicks in view of Mande teaches determining the first distance between a center point and the left border and a second distance between a center point and the right border (Mande, pg. 9, paragraphs 0150-0152), but does not expressly teach a facial landmark which is not the border of the face.
However, Rawat teaches determining a first distance between a first facial landmark of the facial landmarks and a second facial landmark of the facial landmarks and determining a second distance between the first facial landmark and a third facial landmark of the facial landmarks (Rawat, “the process 1000 can include determining a plurality of facial landmarks of the face in the image, and determining the alignment of the face using the plurality of facial landmarks… The plurality of facial landmarks can include one or more left facial landmarks, one or more right facial landmarks, and one or more center facial landmarks (e.g., as shown in FIG. 5). The process 1000 can include determining at least a first distance between a first left facial landmark of the one or more left facial landmarks and a first center facial landmark of the one or more center facial landmarks, and determining at least a second distance between a first right facial landmark of the one or more right facial landmarks and the first center facial landmark.”, pg. 10, paragraph 0086, lines 1-15).
Hicks in view of Mande teaches that alternative methods could be used for face symmetry estimation (Mande, “It should be noted that other stages can be executed for estimating the symmetry of the landmark points.”, pg. 9, paragraph 0151). Rawat describes an alternative method for face symmetry estimation (Rawat, pg. 10, paragraph 0086, lines 1-15). Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to have modified Hicks in view of Mande by using the direct landmark-to-landmark distance computation methods as taught by Rawat. The motivation for doing so would have been to reduce the number of landmarks used, thereby increasing the speed of computation by the system. Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine the teachings of Hicks in view of Mande with Rawat to obtain the invention above.
Hicks in view of Mande and further in view of Rawat does not teach determine, through the attempt, that no facial landmarks were identified due a face in at least one image frame of the image frames being far from the camera; notify control circuitry of the camera embedded within the meter device that no facial landmarks of the face were identified; responsive to a determination by the control circuitry that the image frames of the room of the panelist household are not suitable for measuring the facial attention to the media device, adjust a configuration of the camera; obtain adjusted input image data from the camera, to facilitate identifying the facial landmarks of the at least one panelist member.
However, Mu teaches determine, through the attempt, that no facial landmarks were identified due a face in at least one image frame of the image frames being far from the camera; notify control circuitry of the camera embedded within the meter device that no facial landmarks of the face were identified; responsive to a determination by the control circuitry that the image frames of the room of the panelist household are not suitable for measuring the facial attention to the media device, adjust a configuration of the camera; obtain adjusted input image data from the camera, to facilitate identifying the facial landmarks of the at least one panelist member (Mu, “In order to clearly and completely identify a human face in a human face image shot by a camera, firstly, a camera or equipment with information processing capability except the camera acquires image parameters of the human face image shot by the camera in a first viewing area, processes the image parameters to obtain a processing result which can indicate the image quality of the human face image, and the camera obtains the processing result, wherein the image quality can be the brightness of the human face image, the size of the human face in the human face image…”, pg. 8, lines 24-36, “Further, the image quality includes: the size of the face in the face image; the lens optical parameters include: a focal length; adjusting the optical parameters of the lens of the camera according to the processing result, so that the image quality of the face image shot by the camera meets the preset requirement, comprising the following steps: when the processing result shows that the size of the face in the face image is smaller than a first preset size threshold, the focal length of the camera is increased, so that the size of the face in the face image is larger than or equal to the first preset size threshold; and when the processing result shows that the size of the face in the face image is larger than a second preset size threshold, reducing the focal length of the camera so that the size of the face in the face image is smaller than or equal to the second preset size threshold. Specifically, the difficulty of recognizing the face in the face image is increased when the subject stands at a position other than the preset shot point. Here, the first preset size threshold and the second preset size threshold are minimum and maximum size values of a size of a face that can normally recognize the face in the face image, where the size values may be a distance from a top of a forehead to a bottom of a chin in the face. If the face size interval for normally recognizing the face in the face image is c-d, when the face size in the face image shot by the camera is smaller than c, the face in the face image is smaller, the shot person is far away from the camera, the camera is difficult to recognize the face in the face image, the focal length of the camera needs to be increased, and the face size in the face image shot by the camera is equal to or larger than c… therefore, the face can be easily recognized from the face image shot by the camera after the focal length is adjusted, and the camera can conveniently adjust the viewing area of the camera according to the condition of the recognized face.”, pg. 9 and 10, lines 23-43 and 1-5, respectively, When faces are determined to be under a size threshold, this is equivalent to determining they are too far from the camera. Controls circuitry of the camera is then notified to automatically adjust focal length to facilitate facial recognition. ).
Hicks in view of Mande and further in view of Rawat teaches detecting faces of panelists to determine attention metrics. Mu teaches automatically adjusting focal length of a camera to facilitate facial recognition when faces are determined to be too far from the camera (see above). Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to have modified the facial detection of Hicks in view of Mande and further in view of Rawat to include automatic focal length correction when panelist faces are undetectable, as taught by Mu (Mu, pg. 9 and 10, lines 23-43 and 1-5, respectively). The motivation for doing so would have been to continuously facilitate facial tracking and measurements under varying facial distances, thereby improving facial tracking efficiency (as suggested by Mu, “…the complete face can be recognized in the face image shot by the camera, the effect of shooting the face image by the camera is improved, and the face recognition efficiency is improved.”, pg. 12, lines 1-3). Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine the teachings of Hicks in view of Mande and further in view of Rawat with Mu to obtain the invention according to claim 1.
Hicks in view of Mande and further in view of Rawat and Mu teaches wherein the input image data and the adjusted input image data is not obtained by the audience measurement facility to protect the privacy of the at least one panelists member of the panelist household. The combination would adjust captured images (i.e., focal length of the camera) of panelist faces, when the size of the face is determined to be under a size threshold, to facilitate tracking even under varying facial distances. This adjusted image data would allow the system to measure quantitative facial attention metrics for each panelists (Mande, see Fig. 3), which is then transmitted privately to the data collection facility for media exposure analysis.
Regarding claim 9, Hicks in view of Mande and further in view of Rawat and Mu teach the meter device according to claim 1, wherein the threshold is indicative of a maximum amount of deviation from a neutral face before the face is identified as distracted (Mande, “Based on the distance ratio, a direction of attention to be front, right or left may be estimated (block 630). For example, if the distance ratio is close to the value of 1, e.g. 0.66 to 1.5, then a person may substantially be looking to the front direction.”, pg. 9, paragraph 0152, lines 1-3, the threshold compared to the distance ratio defines the boundary between a face maintaining attention in a neutral or forward direction and faces that are turned)
Claim 10 corresponds to claim 1, additionally reciting a non-transitory machine readable storage medium. Hicks in view of Mande and further in view of Rawat and Mu teaches a non-transitory machine readable storage medium (Hicks, “ …stored on a non-transitory computer and/or machine readable medium such as a hard disk drive, a flash memory, a read-only memory, a compact disk, a digital Versatile disk, a cache, a random-access memory and/or any other storage device or storage disk in which information is stored for any duration…”, pg. 8 and 9, paragraphs 0046, lines 20-25) to perform the functions of claim 1. As indicated in the analysis of claim 1, Hicks in view of Mande and further in view of Rawat and Mu teaches all the limitations of claim 1. Therefore, claim 10 is rejected for the same reasons of obviousness as claim 1.
Claim 18 corresponds to claim 9, additionally reciting a non-transitory computer-readable medium. Hicks in view of Mande and further in view of Rawat and Mu teaches a non-transitory computer readable medium (see analysis of claim 10) to perform the functions of claim 9. As indicated in the analysis of claim 9, Hicks in view of Mande and further in view of Rawat and Mu teaches all the limitations of claim 9. Therefore, claim 18 is rejected for the same reasons of obviousness as claim 9.
Claim 19 corresponds to claim 1, additionally reciting a method. Hicks in view of Mande and further in view of Rawat and Mu teaches a method to perform the functions of claim 1 (see analysis of claim 1). As indicated in the analysis of claim 1, Hicks in view of Mande and further in view of Rawat and Mu teaches all the limitations of claim 1. Therefore, claim 19 is rejected for the same reasons of obviousness as claim 1.
Claims 3, 4, 12 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Hicks (US 20160066034 A1), in view of Mande et al. (US 2020/0311392 A1), (hereinafter, Mande) and further in view of Rawat et al. (US 2021/0056667 A1), (hereinafter, Rawat), Mu (CN 110769148 A), and Wang (2018/0225842 A1).
Regarding claim 3, Hicks in view of Mande and further in view of Rawat and Mu teach the meter device according to claim 1, wherein the quotient is a first quotient, the threshold is a first threshold, and the processor is to:
determine the first quotient satisfies the first threshold (Mande, “if the distance ratio is close to the value of 1, e.g. 0.66 to 1.5, then a person may substantially be looking to the front direction.”, pg. 9, paragraph 0152, lines 3-5).
Hicks in view of Mande and further in view of Rawat and Mu teaches generating an attention metric based on horizontal skew (Mande, “Based on the distance ratio, a direction of attention to be front, right or left may be estimated (block 630).”, pg. 9, paragraph 0152, lines 1-2) but does not expressly teach comparing a second quotient of a third distance and a fourth distance to a second threshold to determine the attentiveness metric, the third distance being between the first facial landmark and a fourth facial landmark of the facial landmarks, the fourth distance being between the first facial landmark and a fifth facial landmark of the facial landmarks.
Thus, Hicks in view of Mande and further in view of Rawat and Mu does not teach determining users’ attentiveness based on horizontal and vertical skew.
However, Wang teaches determining users’ attentiveness based on horizontal and vertical skew (Wang, “the method described herein are used to determine whether the user is paying attention to the display, and/or where on the display the user's attention is focused”, pg. 1, paragraph 0008, lines 1-4 ,“According to a first aspect, the embodiments of the present technology provide a method for determining a facial pose angle ( e.g., including the yaw, pitch, and roll angles ) ”, pg. 1, paragraph 0010).
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to have modified Hicks in view of Mande and further in view of Rawat and Mu to include determining users’ attentiveness based on horizontal and vertical skew, as suggested by Wang (pg. 1, paragraph 0010). The motivation for doing so would have been to track facial poses by considering vertical and horizontal directions, thereby increasing the accuracy of the system. Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine the teachings of Hicks in view of Mande and further in view of Rawat and Mu with Wang to obtain the invention as specified in claim 3.
Hicks in view of Mande and further in view of Rawat and Mu teaches determining the first quotient (Rawat, “a first distance between a first left facial landmark of the one or more left facial landmarks and a first center facial landmark of the one or more center facial landmarks, and determining at least a second distance between a first right facial landmark of the one or more right facial landmarks and the first center facial landmark.”, pg. 10, paragraph 0086, lines 1-15, a first quotient corresponding to the horizontal skew is determined based on a distance from a left facial landmark to a nose center point and a distance from a right facial landmark to a nose center point) satisfies the first threshold (Mande, “if the distance ratio is close to the value of 1, e.g. 0.66 to 1.5, then a person may substantially be looking to the front direction.”, pg. 9, paragraph 0152, lines 3-5). The combination of Hicks in view of Mande and further in view of Rawat and Mu with Wang teaches compare a second quotient of a third distance and a fourth distance to a second threshold to determine the attentiveness metric, the third distance being between the first facial landmark and a fourth facial landmark of the facial landmarks, the fourth distance being between the first facial landmark and a fifth facial landmark of the facial landmarks (Wang, “The first ratio of the length A′N′ of the first line segment to the length B′N′ of the second line segment is calculated, a pre-established correspondence between the first ratio and the face pitch angle is queried according to the first ratio, and the face pitch angle corresponding to the calculated first ratio is queried from the correspondence, and the face pitch angle is determined as the face pitch angle of the to-be-determined face image.”, pg. 5, paragraph 0063, a second quotient corresponding to the pitch of the face is determined based on a distance from a midpoint between the eyebrows to the nose center point and a distance from a midpoint of the lips to the nose center point, the second quotient is compared to a threshold to determine vertical skew).
Regarding claim 4, Hicks in view of Mande and further in view of Rawat, Mu, and Wang teaches the meter device according to claim 3, wherein the processor circuitry is to, in response to the second quotient not satisfying the threshold, identify that the face is distracted (Mande, “In some cases, a direction of attention of an individual can be classified by using general directions indications such 'front', 'right' or 'left'… A person versed in the art would realize that other classifications can be used for indicating the direction of attention of an individual.”, pg. 9, paragraph 0154, in a case where the horizontal skew is determined as looking forward but the vertical skew is determined as looking up or down past the threshold, the face would be classified as ‘up’ or ‘down’).
Claims 12 and 13 corresponds to claims 3 and 4, respectively, additionally reciting a non-transitory computer-readable medium. Hicks in view of Mande and further in view of Rawat, Mu, and Wang teaches a non-transitory computer-readable medium (see analysis of claim 10) to perform the functions of claims 3 and 4, respectively. As indicated in the analysis of claim 3 and 4, Hicks in view of Mande and further in view of Rawat, Mu, and Wang teaches all the limitations of claim 3 and 4, respectively. Therefore, claims 12 and 13 are rejected for the same reasons of obviousness as claims 3 and 4, respectively.
Claims 5 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Hicks (US 20160066034 A1), in view of Mande et al. (US 2020/0311392 A1), (hereinafter, Mande) and further in view of Rawat et al. (US 2021/0056667 A1), (hereinafter, Rawat), Mu (CN 110769148 A), and De Martino et al. (BR 102012016227 A2), (hereinafter, De Martino).
Regarding claim 5, Hicks in view of Mande and further in view of Rawat and Mu teach the meter device according to claim 1, but the combination does not teach wherein the processor circuitry is to: Determine an area of the face; determine a first normalized distance based on a square root of the area of the face and the first distance; and determine a second normalized distance based on the square root of the area of the face and the second distance.
However, De Martino teaches wherein the processor is to determine an area of the face; determine a first normalized distance based on a square root of the area of the face and the first distance; and determine a second normalized distance based on the square root of the area of the face and the second distance (De Martino, “In one embodiment of the invention the face distances were normalized by the square root of the face area”, pg. 11, lines 5-6, Distances corresponding to the face are normalized by using a square root of the face area.).
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to have modified Hicks in view of Mande and further in view of Rawat and Mu to include normalization of the first and second distances relative to the face area, as taught by De Martino (De Martino, pg. 11, lines 5-6). The motivation for doing so would be to make the facial distances independent of the scale, as suggested by De Martino (De Martino, “seeking to make the proportional distances independent of scale”, pg. 11, line 6), thereby accounting for variations in facial size. Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine the teachings of Hicks in view of Mande and further in view of Rawat and Mu with De Martino to obtain the invention as specified in claim 5.
Claim 14 corresponds to claim 5, additionally reciting a non-transitory computer-readable medium. Hicks in view of Mande and further in view of Rawat, Mu, and De Martino teaches a non-transitory computer-readable medium (see analysis of claim 10) to perform the functions of claim 5. As indicated in the analysis of claim 8, Hicks in view of Mande and further in view of Rawat, Mu, and De Martino teaches all the limitations of claim 5. Therefore, claim 14 is rejected for the same reasons of obviousness as claim 5.
Regarding claim 28, Hicks in view of Mande and further in view of Rawat and Mu teach the meter device according to claim 1, but the combination does not teach wherein the processor is to: normalize the first and second distances with respect to a width and a height of a face in at least one image frame of the adjusted image data.
However, De Martino teaches wherein the processor is to: normalize the first and second distances with respect to a width and a height of a face in at least one image frame of the adjusted image data (De Martino, “In one embodiment of the invention the face distances were normalized by the square root of the face area”, pg. 11, lines 5-6, Distances corresponding to the face are normalized by using a square root of the face area. The area of the face is calculated with respect to).
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to have modified Hicks in view of Mande and further in view of Rawat and Mu to include normalization of the first and second distances relative to the face area, as taught by De Martino (De Martino, pg. 11, lines 5-6). The motivation for doing so would be to make the facial distances independent of the scale, as suggested by De Martino (De Martino, “seeking to make the proportional distances independent of scale”, pg. 11, line 6), thereby accounting for variations in facial size. Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine the teachings of Hicks in view of Mande and further in view of Rawat and Mu with De Martino to obtain the invention as specified in claim 28.
Claim 29 corresponds to claim 28, additionally reciting a non-transitory computer-readable medium. Hicks in view of Mande and further in view of Rawat, Mu, and De Martino teaches a non-transitory computer readable medium (see analysis of claim 10) to perform the functions of claim 28. As indicated in the analysis of claim 28, Hicks in view of Mande and further in view of Rawat, Mu, and De Martino teaches all the limitations of claim 28. Therefore, claim 29 is rejected for the same reasons of obviousness as claim 28.
Claim 30 corresponds to claim 28, additionally reciting a method. Hicks in view of Mande and further in view of Rawat Mu, and De Martino teaches a method to perform the functions of claim 28 (see analysis of claim 1). As indicated in the analysis of claim 28, Hicks in view of Mande and further in view of Rawat, Mu, and De Martino teaches all the limitations of claim 28. Therefore, claim 30 is rejected for the same reasons of obviousness as claim 28.
Claims 6-8 are rejected under 35 U.S.C. 103 as being unpatentable over Hicks (US 20160066034 A1), in view of Mande et al. (US 2020/0311392 A1), (hereinafter, Mande) and further in view of Rawat et al. (US 2021/0056667 A1), (hereinafter, Rawat), Mu (CN 110769148 A), and Rosebrock (“Simple object tracking with OpenCV ” PyImageSearchpyimagesearch.com/2018/07/23/
simple-object-tracking-with-opencv/), (hereinafter, Rosebrock).
Regarding claim 6, Hicks in view of Mande and further in view of Rawat and Mu teach the meter device according to claim 1, wherein the processor circuitry is to track the attentiveness metric of the face over a period of time (Hicks, “With the identification of the media and the amount of people in the room at a given date and time, the meter can indicate how many people were exposed to the specific media and/or associate the demographics of the people with the media to determine audience characteristics for the specific media.”, pg. 1, paragraph 0013, lines 18-23)
Hicks in view of Mande and further in view of Rawat and Mu does not teach using centroid tracking.
However, Rosebrock teaches using centroid tracking (Rosebrock, “In today's blog post, you will learn how to implement centroid tracking with OpenCV, an easy to understand, yet highly effective tracking algorithm.”, pg. 2, line 19).
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to have modified Hicks in view of Mande and further in view of Rawat and Mu by replacing the frame by frame face tracking with the centroid tracking algorithm of Rosebrock (Rosebrock, pg. 2, line 19). The motivation for doing so would have been to increase computational efficiency by object tracking using a center coordinate of a bounding box rather than a position of the detected face shape, as disclosed in Mande (Mande, “If a shape of a face is indeed detected in the new frame, the process proceeds to determine whether the position of the detected face shape in the new frame is within a predetermined divergence threshold from the position of the previously detected face.”, pg. 8, paragraph 0140, lines 18-23). Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine the teachings of Hicks in view of Mande and further in view of Rawat and Mu with Rosebrock to obtain the invention as specified in claim 6.
Regarding claim 7, Hicks in view of Mande and further in view of Rawat and Mu teaches the method according to claim 6, wherein the input image data includes a first image and a second image, the face is a first face detected in the first image, the attentiveness metric is a first attentiveness metric associated with the first face in the first image, the threshold is a first threshold (Mande, “an audience individual that was detected in other frames, e.g. in previous frames, may be tracked from one frame to another. In such cases, when a new frame is received, data relating to an estimated location of the tracked individual may also be received… The process then proceeds to determine a direction of attention of the individual associated with the detected face and updates the tracked sequence associated with the detected face which is the same tracked sequence of the tracked individual.”, pg. 8, paragraph 0140, Detected faces are tracked across frames of video data, and for each frame a direction of attention is determined.).
Hicks in view of Mande and further in view of Rawat and Mu does not teach the processor circuitry is to: obtain a first set of bounding box coordinates bounding the first face in the first image and a second set of bounding box coordinates bounding a second face detected in the second image; determine a first center coordinate of the first set of bounding box coordinates; assign an object identifier to the first center coordinate; determine a second center coordinate of the second set of bounding box coordinates; compare a third distance between the first center coordinate and the second center coordinate to a second threshold to determine whether the second center coordinate is associated with the first center coordinate; in response to the third distance satisfying the second threshold, assign the object identifier to the second center coordinate to identify the second face as corresponding to the first face; and determine a second attentiveness metric associated with the second face in the second image.
However, Rosebrock teaches the processor circuitry is to: obtain a first set of bounding box coordinates bounding the first face in the first image and a second set of bounding box coordinates bounding a second face detected in the second image (Rosebrock, “The centroid tracking algorithm assumes that we are passing in a set of bounding box (x, y) coordinates for each detected object in every single frame.”, pg. 5, lines 1-2, Each face of the frames will obtain bounding box coordinates.); determine a first center coordinate of the first set of bounding box coordinates; assign an object identifier to the first center coordinate; determine a second center coordinate of the second set of bounding box coordinates (Rosebrock, “Once we have the bounding box coordinates, we must compute the "centroid", or more simply, the center (x, y)-coordinates of the bounding box. Figure 1 above demonstrates accepting a set of bounding box coordinates and computing the centroid. Since these are the first initial set of bounding boxes presented to our algorithm, we will assign them unique IDs.”, pg. 5, lines 6-10, see Fig. 1, Center coordinates are determined for each bounding box, the initial frame obtains IDs.); compare a third distance between the first center coordinate and the second center coordinate to a second threshold to determine whether the second center coordinate is associated with the first center coordinate (Rosebrock, “For every subsequent frame in our video stream we apply Step #1 of computing object centroids; however, instead of assigning a new unique ID to each detected object (which would defeat the purpose of object tracking), we first need to determine if we can associate the new object centroids (yellow) with the old object centroids (purple). To accomplish this process, we compute the Euclidean distance (highlighted with green arrows) between each pair of existing object centroids and input object centroids.”, pg. 6, lines 1-6, “The primary assumption of the centroid tracking algorithm is that a given object will potentially move in between subsequent frames, but the distance between the centroids for frames F t and F t+i will be smaller than all other distances between objects. Therefore, if we choose to associate centroids with minimum distances between subsequent frames, we can build our object tracker.”, pg. 7, lines 1-5, see Fig. 2 and 3, Distances between faces of the frames are compared to determine associations.); in response to the third distance satisfying the second threshold, assign the object identifier to the second center coordinate to identify the second face as corresponding to the first face (Rosebrock, see Fig. 4, The associated centroids are assigned their respective IDs.); and determine a second attentiveness metric associated with the second face in the second image (Mande, “The process then proceeds to determine a direction of attention of the individual associated with the detected face and updates the tracked sequence associated with the detected face”, pg. 8, paragraph 0140, The direction of attention is determined for each new frame.).
Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to have modified Hicks in view of Mande and further in view of Rawat and Mu by replacing the frame by frame face tracking with the centroid tracking algorithm of Rosebrock (Rosebrock, pg. 2, line 19). The motivation for doing so would have been to increase computational efficiency by object tracking using a center coordinate of a bounding box rather than a position of the detected face shape, as disclosed in Mande (Mande, “If a shape of a face is indeed detected in the new frame, the process proceeds to determine whether the position of the detected face shape in the new frame is within a predetermined divergence threshold from the position of the previously detected face.”, pg. 8, paragraph 0140, lines 18-23). Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine the teachings of Hicks in view of Mande and further in view of Rawat and Mu with Rosebrock to obtain the invention as specified in claim 7.
Regarding claim 8, Hicks in view of Mande and further in view of Rawat, Mu, and Rosebrock teach the method according to claim 7, wherein the processor circuitry is to:
in response to the third distance not satisfying the second threshold, assign a unique object identifier to the second center coordinate, the unique object identifier identify the second face is different from the first face (Rosebrock, “In the event that there are more input detections than existing objects being tracked, we need to register the new object. "Registering" simply means that we are adding the new object to our list of tracked objects by: 1 Assigning it a new object ID”, pg. 8, lines 1-4, For detected faces that do not meet a distance requirement to a previous detected face, a new ID is assigned.); and
determine a third attentiveness metric associated with the second face in the second image (Mande, “In some examples, more than one face associated with more than one audience individual can be detected in a single frame.”, pg. 8, paragraph 0137, lines 4-7, The direction of attention is determined for each detected face in the frame.).
Claims 15, 16, and 17 correspond to claims 6, 7, and 8, respectively, additionally reciting a non-transitory computer-readable medium. Hicks in view of Mande and further in view of Rawat, Mu, and Rosebrock teaches a non-transitory computer-readable medium (see analysis of claim 10) to perform the functions of claims 6, 7, and 8, respectively. As indicated in the analysis of claims 6, 7, and 8, Hicks in view of Mande and further in view of Rawat, Mu, and Rosebrock teaches all the limitations of claims 6, 7, and 8, respectively. Therefore, claim 15, 16, and 17 are rejected for the same reasons of obviousness as claims 6, 7, and 8, respectively.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CONNOR LEVI HANSEN whose telephone number is (703)756-5533. The examiner can normally be reached Monday-Friday 9:00-5:00 (ET).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Sumati Lefkowitz can be reached at (571) 272-3638. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/CONNOR L HANSEN/Examiner, Art Unit 2672
/SUMATI LEFKOWITZ/Supervisory Patent Examiner, Art Unit 2672