DETAILED ACTION
Claims 1-3 and 5-23 are pending in this application and have been examined with the priority date of 06/02/2023 in accordance with the applicant’s claim to the parent application. Claims 1, 5, 6 and 17 have been amended, claim 4 has been canceled, and claim 23 is newly added.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Applicant’s claim for the benefit of a prior-filed application under 35 U.S.C. 119(e) or under 35 U.S.C. 120, 121, 365(c), or 386(c) is acknowledged.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 07/10/2024 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Response to Arguments
Claim interpretation
Applicant’s arguments (see Remarks filed 05/08/2026) regarding the claim interpretations under 35 U.S.C. 112(f) have been fully considered by the examiner and are not persuasive. The recited claim limitations of a “terminal device” and a “memory device” do not include the terms “means” or “step”, however they include the term “device” which is a generic placeholder as defined in MPEP 2181 section A. Prong 2 of the analysis requires the placeholder be modified by functional language, which both have. In Claim 11, the terminal device has video data transmitted to it, which constitutes function language. Further in claims 17, the memory device stores program instructions which also satisfies this prong. Prong 3 of the analysis states that the functional language not be modified by structure or materials for performing the acts of the claimed function. Claims 11 and 17 fails to provide sufficient material or structure for performing the acts claimed. Therefore, for at least the reasons listed above, the examiner maintains the claim interpretations under 35 U.S.C. 112(f).
35 U.S.C. 101
Applicant’s arguments (see Remarks filed 05/08/2026) have been fully considered by the examiner and are not persuasive. The examiner disagrees that the newly added limitations of amended claims 1, 5, 6 and 17 translate the claimed steps into practical application or amount to significantly more. Taking newly amended claim 1 as example, the claim recites;
A method, comprising: performing face detection on an image (Mental process in which a human could reasonably look at an image, inspect it visually and determine faces in the images),
determining a gaze direction within image content associated with a detected face (Mental process in which a human could reasonably look at an image, inspect it visually and determine which way a person is looking);
performing object detection in a region of the image corresponding to the direction of gaze, and when an object is detected within the region and aligned to the gaze direction of the detected face, (Mental process of visually assessing and image and determining that an object is being looked at by a human)
defining a cropping window for the image based on the detected face and the determined gaze direction of the detected face (Mental process of determining a face and the gaze direction, and then a step of mere data gathering in which a human could crop an image using a mouse and computer)
wherein the cropping window is defined to include the detected face and, when the object is detected to be aligned to the gaze direction of the detected face. the detected object; (step of mere data gathering of visually assessing and image and determining that an object is being looked at by a human, then placing a bounding box around both the face and the object when this has been determined)
and cropping the image according to the cropping window (step of mere data gathering in which a human could crop an image using a mouse and computer).
The limitations above are steps which could practically be performed as a mental process or step of mere data gathering performed by a human under step 2A prong 1 (MPEP 2106). Under step 2A prong 2, the claim recites no additional elements. Under step 2B, the claim does not include elements or limitations that meaningfully translate the claims into practical application or amount to significantly more than an abstract idea. (See MPEP section 2106)
Independent claims 6, 13 and 17 recite similar limitations which are drawn to abstract ideas, mental processes or steps of mere data gathering as noted by the examiner below. Further dependent claims 2-3, 5, 7-12, 14-16 and 18-23 do not add limitations which meaningfully translate the abstract ideas of claims 1, 6, 13 and 17 into practical application or amount to significantly more. For at least the reasons above, the examiner respectfully maintains the rejections under 35 U.S.C. 101.
35 U.S.C. 102
Applicant’s arguments (see Remarks filed 05/08/2026) with regard to the rejections made under 35 U.S.C. 102 in view of Ptucha to claims 1, 6 and 17 have been fully considered by the examiner and are persuasive in view the change of scope to claims 1, 6 and 17. In view of the newly amended limitations to claims 1, 6, and 17 new grounds of rejection are presented over Ptucha in view of Coughlan and Lubelsky as fully discussed below.
However, applicant’s arguments (see Remarks filed 05/08/2026) with regard to the rejections made under Ptucha to claim 13 have been fully considered by the examiner and are not persuasive. Applicant argues that Ptucha fails to teach the following limitations as recited by claim 13;
“determining whether a recognized object is present in a region aligned with the gaze direction,
if so, defining a cropping window to include the detected face and the recognized object.”
The examiner disagrees that Ptucha fails to meet the broadest reasonable interpretation of these limitations. The broadest reasonable interpretation of the limitation is determining if an object is located in an area of undisclosed size near where subject is looking. This limitation is broad and could be satisfied by any background or foreground object being included in a region of interest if the subject of interest is looking towards the region in which it is located. The examiner further notes that this limitation does not require the object itself to be directly matched with a gaze direction or line of sight, just that it is in a region of undisclosed size which is generally in the direction the subject is looking.
Ptucha teaches in paragraphs [0054] and [0072] that when a human subject in an image is looking to the left where an object in the background or foreground of an image is located, and the images is subsequently cropped to include this portion of the image (denoted as a low priority area), thereby including the background or foreground objects which the subject is looking at. Further, figures 6 and 7 shows this adjustment to include more background objects or other people in the image when a person is looking in the direction of that region. Additionally, Ptucha paragraph [0089] also details an example in which positional and head pose relationships between subjects are used to adjust the cropping. The example given in Ptucha [0089] is a situation in which a parent subjects face is positioned to be looking down at a child, the child would be included in the image based on the clustering of the two detected regions and the facial position of the parent holding the baby. This would be understood by one of ordinary skill in the art as being analogous to using a facial position or gaze direction of a person towards a general region of an object in an image to determine whether to include that object or other image subject in the image. For at least the reasons above, the examiner maintains the rejections made to claim 13 under 35 U.S.C. 102.
PNG
media_image1.png
154
280
media_image1.png
Greyscale
(Ptucha, [0054])
PNG
media_image2.png
400
288
media_image2.png
Greyscale
(Ptucha, [0072], emphasis added)
PNG
media_image3.png
322
280
media_image3.png
Greyscale
(Ptucha, [0089], emphasis added)
Claim Objections
Claim 17 is objected to because of the following informalities:
The claim recites “directions of the game” which the examiner believes is a typo and should read “directions of the gaze”. Appropriate correction is required,
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitations are:
“Terminal device” in claim 11
“Memory device” in claim 17
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-3 and 5-23 are rejected under 35 U.S.C. 101 because they are drawn to an abstract idea, mental process or step of mere data gathering.
Regarding claim 1, the claim recites the following limitations, which are drawn to a mental process or step of mere data gathering as noted below:
A method, comprising: performing face detection on an image (Mental process in which a human could reasonably look at an image, inspect it visually and determine faces in the images),
determining a gaze direction within image content associated with a detected face (Mental process in which a human could reasonably look at an image, inspect it visually and determine which way a person is looking);
performing object detection in a region of the image corresponding to the direction of gaze, and when an object is detected within the region and aligned to the gaze direction of the detected face, (Mental process of visually assessing and image and determining that an object is being looked at by a human)
defining a cropping window for the image based on the detected face and the determined gaze direction of the detected face (Mental process of determining a face and the gaze direction, and then a step of mere data gathering in which a human could crop an image using a mouse and computer)
wherein the cropping window is defined to include the detected face and, when the object is detected to be aligned to the gaze direction of the detected face. the detected object; (step of mere data gathering of visually assessing and image and determining that an object is being looked at by a human, then placing a bounding box around both the face and the object when this has been determined)
and cropping the image according to the cropping window (step of mere data gathering in which a human could crop an image using a mouse and computer).
The limitations above are steps which could practically be performed as a mental process or step of mere data gathering performed by a human under step 2A prong 1 (MPEP 2106). Under step 2A prong 2, the claim recites no additional elements. Under step 2B, the claim does not include elements that translate the claims into practical application or amount to significantly more than an abstract idea. See MPEP section 2106.
Dependent claims 2-5 do not add limitations that meaningfully translate the abstract idea into practical application or amount to significantly more.
Regarding claim 2, claim 2 recites the limitations; wherein the cropping window is defined to include the detected face in the cropping window when the gaze direction is estimated to be directed to a camera from which the image was derived. (Mental process of determining the gaze direction, and step of data gathering in which a cropping window is manually defined)
The recited limitations are drawn to steps which can practically be performed as a mental process or steps of mere data gathering.
Regarding claim 3, claim 3 recites the limitations; wherein the cropping window is defined to exclude the detected face from the cropping window when the gaze direction is estimated not to be directed to a camera from which the image was derived. (Mental process of determining the gaze direction, and step of data gathering in which a cropping window is manually defined)
The recited limitations are drawn to steps which can practically be performed as a mental process or steps of mere data gathering.
Regarding claim 5, claim 5 recites the limitations; further comprising: (Mental process of determining the gaze direction, and step of data gathering in which a cropping window is manually defined)
The recited limitations are drawn to steps which can practically be performed as a mental process or steps of mere data gathering.
Regarding claim 6, the claim recites the following limitations, which are drawn to a mental process or step of mere data gathering as noted below:
A method, comprising:
identifying face(s) from a stream of video; (Mental process in which a human could reasonably look at an image, inspect it visually and determine faces in the images)
for each identified face: determining a gaze of direction of the respective face; (Mental process in which a human could reasonably look at an image, inspect it visually and determine which way a person is looking)
determining whether the respective face is to be included in a cropping window based on its determined gaze of direction; (Mental process of determining a face and the gaze direction, and then a step of mere data gathering in which a human could crop an image using a mouse and computer)
for face(s) determined to be included, performing object detection in a region of the video corresponding to respective face's the direction of gaze, (Mental process of visually assessing and image and determining that an object is being looked at by a human)
and defining a cropping window for the video based on the face(s) determined to be included, the determined gaze directions of face(s) determined to be included, and object(s) detected to be aligned to the gaze directions of the face(s) determined to be included; (step of mere data gathering of visually assessing and image and determining that an object is being looked at by a human, then placing a bounding box around both the face and the object when this has been determined)
and cropping frames of the video based on (step of mere data gathering in which a human could crop an image using a mouse and computer)
The limitations above are steps which could practically be performed as a mental process or step of mere data gathering performed by a human under step 2A prong 1 (MPEP 2106). Under step 2A prong 2, the claim recites no additional elements. Under step 2B, the claim does not include elements that translate the claims into practical application or amount to significantly more than an abstract idea. See MPEP section 2106.
Dependent claims 7-12 do not add limitations that do not meaningfully translate the abstract idea into practical application or amount to significantly more.
Regarding claim 7, claim 7 recites the limitations; wherein the cropping window circumscribes all faces determined to be included in the cropping window. (Mental process of determining area belonging to the face, and step of data gathering in which a cropping window is manually defined)
The recited limitations are drawn to steps which can practically be performed as a mental process or steps of mere data gathering.
Regarding claim 8, claim 8 recites the limitations; further comprising, identifying face(s) located within a foreground location of the video, wherein the face(s) determined to be in the foreground location are determined to be included in the cropping window regardless of the respective face’s gaze of direction. (Mental process of determining area belonging to the face, and step of data gathering in which a cropping window is manually defined)
The recited limitations are drawn to steps which can practically be performed as a mental process or steps of mere data gathering.
Regarding claim 9, claim 9 recites the limitations; wherein one of the detected face are determined to be included in a cropping window only after its gaze direction looks at a camera that captured the video for a threshold amount of time. (Mental process of determining area belonging to the face, and step of data gathering in which a cropping window is manually defined)
The recited limitations are drawn to steps which can practically be performed as a mental process or steps of mere data gathering.
Regarding claim 10, claim 10 recites the limitations; wherein one of the detected face are determined not to be included in a cropping window after its gaze direction looks away from a camera that captured the video for a threshold amount of time. (Mental process of determining area belonging to the face, and step of data gathering in which a cropping window is manually defined)
The recited limitations are drawn to steps which can practically be performed as a mental process or steps of mere data gathering.
Regarding claim 11, claim 11 recites the limitations; further comprising, following the cropping, transmitting the cropped video to a distant terminal device. (step of data gathering/transmitting)
The recited limitations are drawn to steps which can practically be performed as a mental process or steps of mere data gathering.
Regarding claim 12, claim 12 recites the limitations; further comprising, following the cropping, transmitting the cropped video in a videoconference stream.(step of data gathering/transmitting)
The recited limitations are drawn to steps which can practically be performed as a mental process or steps of mere data gathering.
Regarding claim 13, the claim recites the following limitations, which are drawn to a mental process or step of mere data gathering as noted below:
An image processing method, comprising:
performing face recognition on an image, (Mental process in which a human could reasonably look at an image, inspect it visually and determine faces in the images)
performing object recognition on the image; (Mental process in which a human could reasonably look at an image, inspect it visually and determine objects in the images)
estimating a gaze direction of a detected face, (Mental process in which a human could reasonably look at an image, inspect it visually and determine which way a person is looking)
determining whether a recognized object is present in a region aligned with the gaze direction, if so, defining a cropping window to include the detected face and the recognized object. (Mental process of determining a face and the gaze direction, and then a step of mere data gathering in which a human could crop an image using a mouse and computer)
The limitations above are steps which could practically be performed as a mental process or step of mere data gathering performed by a human under step 2A prong 1 (MPEP 2106). Under step 2A prong 2, the claim recites no additional elements. Under step 2B, the claim does not include elements that translate the claims into practical application or amount to significantly more than an abstract idea. See MPEP section 2106.
Dependent claims 14-16 do not add limitations that do not meaningfully translate the abstract idea into practical application or amount to significantly more.
Regarding claim 14, claim 14 recites the limitations; further comprising, if a recognized object is not present in the region aligned with the gaze direction, defining a cropping window to include the detected face an at least a portion of the region aligned with the gaze direction, the detected face placed off-center within the cropping window and the region aligned with the gaze direction place in a center area of the cropping window. (Comprises a mental process in which a human could reasonably look at an image, determine gaze direction, then place a cropping box as a result)
The recited limitations are drawn to steps which can practically be performed as a mental process or steps of mere data gathering.
Regarding claim 15, claim 15 recites the limitations; wherein the detected face is placed in the cropping window according to a compositional rule. (Comprises a mental process in which a human could reasonably look at an image, then place a cropping box as a result)
The recited limitations are drawn to steps which can practically be performed as a mental process or steps of mere data gathering.
Regarding claim 16, claim 16 recites the limitations; wherein the recognized object is placed in the cropping window according to a compositional rule. (Comprises a mental process in which a human could reasonably look at an image, then place a cropping box as a result)
The recited limitations are drawn to steps which can practically be performed as a mental process or steps of mere data gathering.
Regarding claim 17, the claim recites the following limitations, which are drawn to a mental process or step of mere data gathering as noted below:
A system comprising:
a processing system, including a first processor and a second processor, (Additional elements of a processing system, and processors which are recited with a high level of generality and fail to translate the claims into practical application)
the first processor being a neural network processor trained to estimate gaze direction of face(s) detected from image data (Mental process in which a human could reasonably look at an image, inspect it visually and determine faces in the images and where they are looking),
the second processor to execute program instructions stored in a memory device; (Additional elements of processors and a memory which are recited with a high level of generality and fail to translate the claims into practical application)
the memory device, storing the program instructions that, when executed by the second processor, cause the second processor to: compare first detected face to a location of a second detected face; (Mental process of determining multiple faces and the gaze directions of those faces)
compare a direction of gaze identified by the first processor for the second detected face to a location of the first detected face; (Mental process of determining multiple faces and the gaze directions of those faces);
when the directions of game for the two detected faces are aligned with the locations of the counterpart detected faces, define a cropping window for an image to include the two faces, (Mental process of discerning if two individuals are looking at one another in a video, then a step of mere data gathering cropping the video accordingly)
and crop the image according to the cropping window. (Step of mere data gathering in which a human could crop an image)
The limitations above are steps which could practically be performed as a mental process or step of mere data gathering performed by a human under step 2A prong 1 (MPEP 2106). Under step 2A prong 2, the claim recites the additional elements of a processing system, a first and second processor, a neural network and a memory which are recited with a high level of generality and do not meaningfully translate the claims into practical application. Under step 2B, the claim does not include elements that translate the claims into practical application or amount to significantly more than an abstract idea. See MPEP section 2106.
Dependent claims 18-22 do not add limitations that do not meaningfully translate the abstract idea into practical application or amount to significantly more.
Regarding claim 18, claim 18 recites the limitations; further comprising an image signal processor having an output for identification of face(s) in the image data. (Mental process in which a face could be identified in an image visually and a person could send an output of this)
The recited limitations are drawn to steps which can practically be performed as a mental process or steps of mere data gathering.
Regarding claim 19, claim 19 recites the limitations; wherein the program instructions cause the cropping window to be defined to include the detected face in the cropping window when the gaze direction is estimated to be directed to a camera from which the image was derived. (Mental process in which a human could assess a gaze direction of a person and a define a cropping window as a result)
The recited limitations are drawn to steps which can practically be performed as a mental process or steps of mere data gathering.
Regarding claim 20, claim 20 recites the limitations; wherein the program instructions cause the cropping window to be defined to exclude the detected face from the cropping window when the gaze direction is estimated not to be directed to a camera from which the image was derived. (Mental process in which a human could assess a gaze direction of a person and a define a cropping window as a result)
The recited limitations are drawn to steps which can practically be performed as a mental process or steps of mere data gathering.
Regarding claim 21, claim 21 recites the limitations wherein the program instructions cause the second processor to: perform object detection in a region of the image corresponding to the direction of gaze, and when an object is detected within the region, define the cropping window to include the detected face and the detected object. (Mental process in which a human could assess a gaze direction of a person and a define a cropping window as a result)
The recited limitations are drawn to steps which can practically be performed as a mental process or steps of mere data gathering. The additional element of a second processor is recited with a high level of generality and does not meaningfully translate the claim into practical application or significantly more than an abstract idea.
Regarding claim 22, claim 22 recites the limitations; wherein the program instructions cause the second processor to: perform object detection in a region of the image corresponding to the direction of gaze, and when an object is not detected within the region, define the cropping window to place the detected face in an off center region of the cropping window with the region placed in a center of the cropping window. (Mental process in which a human could assess a gaze direction of a person and a define a cropping window as a result)
The recited limitations are drawn to steps which can practically be performed as a mental process or steps of mere data gathering. The additional element of a second processor is recited with a high level of generality and does not meaningfully translate the claim into practical application or significantly more than an abstract idea.
Regarding claim 23, claim 23 recites; The method of claim 1, comprising repeating the performing and determining for a plurality of images from video for at least one other face, and: (Mental process in which a human could assess an image and find a face of a person’s face and determine a gaze direction of a person)
cropping frames of the video based on positions of the face(s) determined to be included in the cropping window. (Step of mere data gathering a define a cropping window as a result of this above determination)
The recited limitations are drawn to steps which can practically be performed as a mental process or steps of mere data gathering without translation into practical application or significantly more.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 13-16 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Ptucha (US 20130108164 A1).
Regarding claim 13 Ptucha discloses; An image processing method, comprising:
performing face recognition on an image (Ptucha, [0050] the system performs face detection),
performing object recognition on the image (Ptucha, [0055] objects are detected in the foreground and background of the image based on previous knowledge by the algorithm, [0056] the objects are classified);
estimating a gaze direction of a detected face (Ptucha, [0064] the eye gaze direction is determined for the face detected in the image),
determining whether a recognized object is present in a region aligned with the gaze direction, if so, defining a cropping window to include the detected face and the recognized object (Ptucha, [0072] the eye gaze and facial pose vectors are used to perform the generation of a detecting box to include certain portions of the background or foreground as well as the face, where [0055]-[0056] the regions of background and foreground are evaluated for features/objects, figure 7 shows the placement of the face in a larger box to include more of the background/foreground objects, vs a situation where just the face is included in the center of the bounding box).
Regarding claim 14 Ptucha discloses; The method of claim 13, further comprising, if a recognized object is not present in the region aligned with the gaze direction, defining a cropping window to include the detected face an at least a portion of the region aligned with the gaze direction, the detected face placed off-center within the cropping window and the region aligned with the gaze direction place in a center area of the cropping window (Ptucha, [0072] the eye gaze and facial pose vectors are used to perform the generation of a detecting box to include certain portions of the background or foreground as well as the face, where [0055]-[0056] the regions of background and foreground are evaluated for features/objects, figure 7 shows the placement of the face in a larger box to include more of the background/foreground objects, vs a situation where just the face is included in the center of the bounding box).
Regarding claim 15 Ptucha discloses; The method of claim 13, wherein the detected face is placed in the cropping window according to a compositional rule (Ptucha, [0065] ‘rule of thirds’ compositional rule is used to place the subject/face of interest in the bounding box/cropping window).
Regarding claim 16 Ptucha discloses; The method of claim 13, wherein the recognized object is placed in the cropping window according to a compositional rule (Ptucha, [0072] the eye gaze and facial pose vectors are used to perform the generation of a detecting box to include certain portions of the background or foreground as well as the face, where to determine the box, an aspect ratio is determined based on the region inclusion priority (compositional rule)).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1, 5-8, and 11,are rejected under 35 U.S.C. 103 as being unpatentable over Ptucha (US 20130108164 A1) in view of Coughlan (US 7460150 B1).
Regarding claim 1 Ptucha discloses; A method, comprising:
performing face detection on an image (Ptucha, [0050] the system performs face detection),
determining a gaze direction within image content associated with a detected face (Ptucha, [0064] the eye gaze direction is determined for the face detected in the image);
[performing object detection in a region of the image corresponding to the direction of gaze, and when an object is detected within the region and aligned to the gaze direction of the detected face,]
defining a cropping window for the image based on the detected face and the determined gaze direction of the detected face (Ptucha, [0064] padded face boxes (bounding boxes) are determined based on the fitness score of the face detected and other factors including eye gaze direction, and head angles, [0050] cropping is performed based on the determined bounding boxes)
[ wherein the cropping window is defined to include the detected face and, when the object is detected to be aligned to the gaze direction of the detected face, the detected object;]
and cropping the image according to the cropping window (Ptucha, [0051] the image is cropped based on the determined crop box/bounding box).
Ptucha fails to teach; performing object detection in a region of the image corresponding to the direction of gaze, and when an object is detected within the region and aligned to the gaze direction of the detected face,
wherein the cropping window is defined to include the detected face and, when the object is detected to be aligned to the gaze direction of the detected face, the detected object;
However, in the same field of endeavor Coughlan teaches; performing object detection in a region of the image corresponding to the direction of gaze, and when an object is detected within the region and aligned to the gaze direction of the detected face (Coughlan, column 5 lines 20-35, the object of interest can be determined using the line of sight or gaze of the person when they are directed to a specific area, column 5 lines 5-20 when the gazes of detected faces in a video are all drawn to a specific location the system directs the camera to make sure this is included in the video/image),
PNG
media_image4.png
370
302
media_image4.png
Greyscale
(Coughlan, column 5 emphasis added)
wherein the cropping window is defined to include the detected face and, when the object is detected to be aligned to the gaze direction of the detected face, the detected object (Coughlan, column 3 lines 30-53, the person’s line of sight or gaze may be determined, and the area in the image corresponding to this gaze may be cropped to be adjusted to include this area in the image as well as the persons being imaged, column 5 line 65- column 6 line 20, if the gaze of persons in the video is directed to an object that object will also be included in the image/video frame);
PNG
media_image5.png
294
300
media_image5.png
Greyscale
(Coughlan, column 3, emphasis added)
PNG
media_image6.png
38
296
media_image6.png
Greyscale
PNG
media_image7.png
266
308
media_image7.png
Greyscale
(Coughlan, columns 5 and 6, emphasis added)
The combination of Ptucha and Coughlan would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The motivation for the addition of the object detection when gaze alignment is detected as taught by Coughlan lies in that this allows for objects of interest or events of interest to be captured during surveillance of a scene. (Coughlan, column 3 lines 30 – 53, and columns 5 and 6)
Regarding claim 5 the combination of Ptucha and Coughlan teaches; The method of claim 1, further comprising: (Ptucha, [0072] the eye gaze and facial pose vectors are used to perform the generation of a detecting box to include certain portions of the background or foreground as well as the face, where [0055]-[0056] the regions of background and foreground are evaluated for features/objects, figure 7 shows the placement of the face in a larger box to include more of the background/foreground objects, vs a situation where just the face is included in the center of the bounding box).
PNG
media_image8.png
442
538
media_image8.png
Greyscale
(Ptucha, Figure 7)
Regarding claim 6 the combination of Ptucha and Coughlan teaches; A method, comprising:
identifying face(s) from a stream of video (Ptucha, [0050] the system performs face detection);
for each identified face: determining a gaze of direction of the respective face (Ptucha, [0064] the eye gaze direction is determined for the face detected in the image);
determining whether the respective face is to be included in a cropping window based on its determined gaze of direction (Ptucha, [0064] padded face boxes (bounding boxes) are determined based on the fitness score of the face detected and other factors including eye gaze direction, and head angles, [0050] cropping is performed based on the determined bounding boxes);
for face(s) determined to be included, performing object detection in a region of the video corresponding to respective face's the direction of gaze (Coughlan, column 5 lines 20-35, the object of interest can be determined using the line of sight or gaze of the person when they are directed to a specific area, column 5 lines 5-20 when the gazes of detected faces in a video are all drawn to a specific location the system directs the camera to make sure this is included in the video/image),
and defining a cropping window for the video based on the face(s) determined to be included, the determined gaze directions of face(s) determined to be included, and object(s) detected to be aligned to the gaze directions of the face(s) determined to be included (Coughlan, column 3 lines 30-53, the person’s line of sight or gaze may be determined, and the area in the image corresponding to this gaze may be cropped to be adjusted to include this area in the image as well as the persons being imaged, column 5 line 65- column 6 line 20, if the gaze of persons in the video is directed to an object that object will also be included in the image/video frame);
and cropping frames of the video based on (Ptucha, [0051] the image is cropped based on the determined crop box/bounding box).
The combination of Ptucha and Coughlan would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The motivation for the addition of the object detection when gaze alignment is detected as taught by Coughlan lies in that this allows for objects of interest or events of interest to be captured during surveillance of a scene. (Coughlan, column 3 lines 30 – 53, and columns 5 and 6)
Regarding claim 7 the combination of Ptucha and Coughlan teaches; The method of claim 6, wherein the cropping window circumscribes all faces determined to be included in the cropping window (Ptucha, figure 7 and [0072] the bounding boxes are set around faces/limited to the face area).
Regarding claim 8 the combination of Ptucha and Coughlan teaches; The method of claim 6, further comprising, identifying face(s) located within a foreground location of the video, wherein the face(s) determined to be in the foreground location are determined to be included in the cropping window regardless of the respective face’s gaze of direction (Ptucha, [0058] the faces are selected and the faces width and determination of where the face is made, where the face size is used to determine whether to include or exclude the face, [0059] this is used to determine the face’s distance in the image(background or foreground determination)).
Regarding claim 11 the combination of Ptucha and Coughlan teaches; The method of claim 6, further comprising, following the cropping, transmitting the cropped video to a distant terminal device (Ptucha, [0036] the system used remote user devices where [0043] the image data can be conveyed to remote devices where [0050] remote devices may receive or generate the cropped images).
Regarding claim 23 the combination of Ptucha and Coughlan teaches; The method of claim 1, comprising repeating the performing and determining for a plurality of images from video for at least one other face (Ptucha, [0050] the system performs face detection, [0052] face detection is performed for all faces in the images, [0064] gaze detection and other analysis is performed for all faces in an image),
and: cropping frames of the video based on positions of the face(s) determined to be included in the cropping window (Ptucha, [0058] the figures are cropped to include all faces which need to be included in the image).
Claims 2-3, 9-10, and 12are rejected under 35 U.S.C. 103 as being unpatentable over Ptucha (US 20130108164 A1) in view of Coughlan (US 7460150 B1)and in further view of Li (US 202130342640 A1).
Regarding claim 2 the combination of Ptucha and Coughlan does not disclose; The method of claim 1, wherein the cropping window is defined to include the detected face in the cropping window when the gaze direction is estimated to be directed to a camera from which the image was derived.
However, in the same field of facial detection, Li teaches; wherein the cropping window is defined to include the detected face in the cropping window when the gaze direction is estimated to be directed to a camera from which the image was derived (Li, [0059] the ROI used in cropping the image is estimated, where the face orientation of the participant may be used to determine this, for example if the participant’s face is oriented such that it is not looking at the camera the system may exclude that face, and further if there is a substantial change in the face orientation of the participant (i.e. to be looking at the camera) the system may adjust the ROI to include the participant in the image).
The combination of Ptucha, Coughlan and Li would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. Ptucha teaches a method of face and gaze detection, and the use of bounding boxes to crop images around detected faces. Li teaches a similar method of face and gaze detection, where the gaze direction is used to determine whether to include the face in the image. The motivation for the combination lies in that using the gaze direction to determine whether to include a face in the image would allow for focus on people on a conference call who are engaged in the meeting or who are actively speaking on a video call. (Li, [0058]-[0060])
Regarding claim 3 the combination of Ptucha, Coughlan and Li teaches; The method of claim 1, wherein the cropping window is defined to exclude the detected face from the cropping window when the gaze direction is estimated not to be directed to a camera from which the image was derived (Li, [0059] the ROI used in cropping the image is estimated, where the face orientation of the participant may be used to determine this, for example if the participant’s face is oriented such that it is not looking at the camera the system may exclude that face, and further if there is a substantial change in the face orientation of the participant (i.e. to be looking at the camera) the system may adjust the ROI to include the participant in the image).
The combination of Ptucha, Coughlan and Li would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. Ptucha teaches a method of face and gaze detection, and the use of bounding boxes to crop images around detected faces. Li teaches a similar method of face and gaze detection, where the gaze direction is used to determine whether to include the face in the image. The motivation for the combination lies in that using the gaze direction to determine whether to include a face in the image would allow for focus on people on a conference call who are engaged in the meeting or who are actively speaking on a video call, and removing focus from those who are not active participants to improve video meeting focus. (Li, [0058]-[0060])
Regarding claim 9, the combination of Ptucha, Coughlan and Li teaches; The method of claim 6, wherein one of the detected face are determined to be included in a cropping window only after its gaze direction looks at a camera that captured the video for a threshold amount of time (Li, [0059] the ROI used in cropping the image is estimated, where the face orientation of the participant may be used to determine this, for example if the participant’s face is oriented such that it is not looking at the camera the system may exclude that face, and further if there is a substantial change in the face orientation of the participant (i.e. to be looking at the camera) the system may adjust the ROI to include the participant in the image).
The combination of Ptucha, Coughlan and Li would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. Ptucha teaches a method of face and gaze detection, and the use of bounding boxes to crop images around detected faces. Li teaches a similar method of face and gaze detection, where the gaze direction is used to determine whether to include the face in the image. The motivation for the combination lies in that using the gaze direction to determine whether to include a face in the image would allow for focus on people on a conference call who are engaged in the meeting or who are actively speaking on a video call. (Li, [0058]-[0060])
Regarding claim 10 the combination of Ptucha, Coughlan and Li teaches; The method of claim 6, wherein one of the detected face are determined not to be included in a cropping window after its gaze direction looks away from a camera that captured the video for a threshold amount of time (Li, [0059] the ROI used in cropping the image is estimated, where the face orientation of the participant may be used to determine this, for example if the participant’s face is oriented such that it is not looking at the camera the system may exclude that face, and further if there is a substantial change in the face orientation of the participant (i.e. to be looking at the camera) the system may adjust the ROI to include the participant in the image).
The combination of Ptucha, Coughlan and Li would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. Ptucha teaches a method of face and gaze detection, and the use of bounding boxes to crop images around detected faces. Li teaches a similar method of face and gaze detection, where the gaze direction is used to determine whether to include the face in the image. The motivation for the combination lies in that using the gaze direction to determine whether to include a face in the image would allow for focus on people on a conference call who are engaged in the meeting or who are actively speaking on a video call, and removing focus from those who are not active participants to improve video meeting focus. (Li, [0058]-[0060])
Regarding claim 12 the combination of Ptucha, Coughlan and Li teaches; The method of claim 6, further comprising, following the cropping, transmitting the cropped video in a videoconference stream (Li, [0050] the system transmits the processed data to the video conference software).
The combination of Ptucha, Coughlan and Li would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The motivation for the combination lies in that this allows the videoconference software to more effectively capture the participant’s interactions. (Li, [0045]-[0060])
Claims 17-18 and 21-22 are rejected under 35 U.S.C. 103 as being unpatentable over Ptucha (US 20130108164 A1) in view of Lubelsky (US 20180098027 A1).
Regarding claim 17 Ptucha discloses; A system comprising:
a processing system, including a first processor and a second processor (Ptucha, [0036] the system has a processor system, [0046] where the processor system may include multiple processors/microprocessors),
the first processor being a neural network processor trained to estimate gaze direction of face(s) detected from image data (Ptucha, [0053] the system uses neural networks for face detection and gaze direction determining, where [0036] processors are used to execute the algorithms used, i.e. the neural networks),
the second processor to execute program instructions stored in a memory device (Ptucha, [0036] processors are used to execute the algorithms used which are stored on a memory);
the memory device, storing the program instructions that, when executed by the second processor, cause the second processor to (Ptucha, [0036] processors are used to execute the algorithms used which are stored on a memory):
[compare first detected face to a location of a second detected face;
compare a direction of gaze identified by the first processor for the second detected face to a location of the first detected face;
when the directions of game for the two detected faces are aligned with the locations of the counterpart detected faces, define a cropping window for an image to include the two faces, ]
and crop the image according to the cropping window (Ptucha, [0051] the image is cropped based on the determined crop box/bounding box).
Ptucha does not teach;
compare first detected face to a location of a second detected face;
compare a direction of gaze identified by the first processor for the second detected face to a location of the first detected face;
when the directions of game for the two detected faces are aligned with the locations of the counterpart detected faces, define a cropping window for an image to include the two faces,
However in the same field of endeavor Lubelsky teaches;
compare first detected face to a location of a second detected face (Lubelsky, [0075] gaze directions of one or more participants in the room are used to determine a focus point using the gaze detection, facial detection and spatial locations of the participants, where in [0068] when two participants are detected as looking at eachother, indicating the gaze of one face is aligned and directed to the second face, the camera will pan to include both in the frame);
compare a direction of gaze identified by the first processor for the second detected face to a location of the first detected face (Lubelsky, [0075] gaze directions of one or more participants in the room are used to determine a focus point using the gaze detection, facial detection and spatial locations of the participants, indicating the location of the faces of all participants and their respective gazes are detected by the system, where in [0068] when two participants are detected as looking at each other, indicating the gaze of one face is aligned and directed to the second face, the camera will pan to include both in the frame);
when the directions of game for the two detected faces are aligned with the locations of the counterpart detected faces, define a cropping window for an image to include the two faces, (Lubelsky, [0068] the camera can adjust its angle to include two participants in the frame who are looking at each other for a discussion, which indicates that when two participants gaze directions/view angles are facing one another the camera adjusts the window to include them)
PNG
media_image9.png
360
324
media_image9.png
Greyscale
(Lubelsky, [0068])
The combination of Ptucha and Lubelsky would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The motivation to include the gaze detection method of Lubelsky is that is allows the camera to pick up when two participants in a meeting are discussing something, which would allow for more complete capturing of the meeting interactions and discussions. (Lubelsky, [0065]-[0075])
Regarding claim 18 the combination of Ptucha and Lubelsky teaches; The system of claim 17, further comprising an image signal processor having an output for identification of face(s) in the image data (Ptucha, [0050] the system outputs faces detected using bounding boxes, where [0036] processors perform the image processing).
Regarding claim 21 the combination of Ptucha and Lubelsky teaches; The system of claim 17, wherein the program instructions cause the second processor to: perform object detection in a region of the image corresponding to the direction of gaze, and when an object is detected within the region, define the cropping window to include the detected face and the detected object (Ptucha, [0072] the eye gaze and facial pose vectors are used to perform the generation of a detecting/bounding box to include certain portions of the background or foreground as well as the face, where [0055]-[0056] the regions of background and foreground are evaluated for features/objects).
Regarding claim 22 the combination of Ptucha and Lubelsky teaches; The system of claim 17, wherein the program instructions cause the second processor to: perform object detection in a region of the image corresponding to the direction of gaze, and when an object is not detected within the region, define the cropping window to place the detected face in an off center region of the cropping window with the region placed in a center of the cropping window (Ptucha, [0072] the eye gaze and facial pose vectors are used to perform the generation of the bounding box to include certain portions of the background or foreground as well as the face, where [0055]-[0056] the regions of the background and foreground as evaluated for features/objects, figure 7 shows the placements of a face in a larger box to include more of the background/foreground objects vs a situation where just the face is included in the center of the box).
6. Claims 19-20 rejected under 35 U.S.C. 103 as being unpatentable over Ptucha (US 20130108164 A1) in view of Lubelsky (US 20180098027 A1) and in further view of Li (US 202130342640 A1).
Regarding claim 19 the combination of Ptucha Lubelsky fails to teach; The system of claim 17, wherein the program instructions cause the cropping window to be defined to include the detected face in the cropping window when the gaze direction is estimated to be directed to a camera from which the image was derived
However, in the same field of endeavor of image processing Li teaches; The system of claim 17, wherein the program instructions cause the cropping window to be defined to include the detected face in the cropping window when the gaze direction is estimated to be directed to a camera from which the image was derived (Li, [0059] the ROI used in cropping the image is estimated, where the face orientation of the participant may be used to determine this, for example if the participant’s face is oriented such that it is not looking at the camera the system may exclude that face, and further if there is a substantial change in the face orientation of the participant (i.e. to be looking at the camera) the system may adjust the ROI to include the participant in the image).
The combination of Ptucha, Lubelsky and Li would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. Ptucha teaches a method of face and gaze detection, and the use of bounding boxes to crop images around detected faces. Li teaches a similar method of face and gaze detection, where the gaze direction is used to determine whether to include the face in the image. The motivation for the combination lies in that using the gaze direction to determine whether to include a face in the image would allow for focus on people on a conference call who are engaged in the meeting or who are actively speaking on a video call. (Li, [0058]-[0060])
Regarding claim 20 the combination of Ptucha, Lubelsky and Li teaches; The system of claim 17, wherein the program instructions cause the cropping window to be defined to exclude the detected face from the cropping window when the gaze direction is estimated not to be directed to a camera from which the image was derived time (Li, [0059] the ROI used in cropping the image is estimated, where the face orientation of the participant may be used to determine this, for example if the participant’s face is oriented such that it is not looking at the camera the system may exclude that face, and further if there is a substantial change in the face orientation of the participant (i.e. to be looking at the camera) the system may adjust the ROI to include the participant in the image).
The combination of Ptucha, Lubelsky and Li would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. Ptucha teaches a method of face and gaze detection, and the use of bounding boxes to crop images around detected faces. Li teaches a similar method of face and gaze detection, where the gaze direction is used to determine whether to include the face in the image. The motivation for the combination lies in that using the gaze direction to determine whether to include a face in the image would allow for focus on people on a conference call who are engaged in the meeting or who are actively speaking on a video call, and removing focus from those who are not active participants to improve video meeting focus. (Li, [0058]-[0060])
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. For a listing of analogous prior art as cited by the examiner see the attached PTO-892 Notice of References Cited.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JORDAN M ELLIOTT whose telephone number is (703)756-5463. The examiner can normally be reached M-F 8AM-5PM ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Emily Terrell can be reached at (571) 270-3717. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/J.M.E./Examiner, Art Unit 2666 /Molly Wilburn/Primary Examiner, Art Unit 2666