DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Objections to the Specification and Rejections under 35 USC 112(b) and 35 USC 101 are withdrawn.
Applicant's arguments regarding Prior Art rejections of claim 1 have been fully considered but they are not persuasive.
On Page 16-17 of Remarks, Applicant argues, “As discussed during the interview, applicant respectfully submits that these features are not disclosed in Liu. First, applicant submits that Liu only discloses the use of one loss and therefore does not disclose both the recited first loss and the related loss. Second, and similarly, applicant submits that Liu does not disclose the recited first loss. As discussed during the interview, the recited first loss is based on a single image (e.g., the first training image and the first numerical value) and is focused on the output/prediction space (e.g., comparing the model output to the ground truth data). Conversely, the loss in Liu is calculated based on relationship between generated feature amounts for multiple training images within the machine learning model and not the outputs.”
Liu does disclose the use of two losses, including both the first loss and the related loss. The “first loss” is mapped to the mean absolute error between the predicted age and the ground-truth, described by Liu in Section IV.D. and cited on Pages 8-9 of the Non-Final Rejection filed 3/10/2026. While Liu uses this first loss for validation rather than training, this does not change the fact that Liu calculates the ‘first loss’ as claimed. Also, as will be described in further detail below, Liu teaches the ‘related loss’ as claimed, wherein this ‘related loss’ is mapped to the distance-based losses exemplified in Figures 2-3 and described in Section III.A.: “The crucial part of our LSDML is to learn the network parameters f (·). To achieve this goal, we first pass a given mini-batch forward the deep network, and we select each quadruplet of (i, j,k,l), such that (xi,xj) ∈ P, (xi,xk) ∈ N and (xj,xl) ∈ N,whereP and N denote the positive and negative pair set, respectively. More details are illustrated in Fig. 3. Moreover, to achieve the discriminativeness of the feature similarity, our LSDML enforces each df (xi,xj) pair in positive set is close to each other, and at the same time df (xi,xj) and df (xk,xl) in negative set is pushed far away. As a result, the distance of inter-class pairs is minimized, and the distance of intra-class pairs is larger than a margin τ in the transformed subspace.”
Applicant's arguments regarding Prior Art rejections of claim 2 have been fully considered but they are not persuasive.
On Page 17 of Remarks, Applicant argues: “As discussed during the interview, applicant submits that Liu does not disclose the triplet- based learning or the intricacies thereto as recited in claim 2 because Liu is based on at least quadruplets (e.g., multiple face pairs which are each two or more pairs of images.).”
Quadruplet learning encompasses triplet-based learning, by definition, by including the same elements (plus a second negative sample).
Applicant's arguments regarding Prior Art rejections of claims 3-17, 19, and 20 have been fully considered but they are not persuasive.
On Page 17 of Remarks, Applicant argues, “As claims 3-17, 19, and 20 either depend on or recite features similar to those in claim 1, please see the discussion of claim 1 above.”
See above discussion of the claim 1 rejections.
Regarding Applicant’s arguments regarding new claim 21:
New claim 21 is rejected under 35 USC 112(a) and 112(b), as described below. In view of the below 35 USC 112(b) interpretation, no Prior Art teaches claim 21.
Claim Rejections - 35 USC § 112
The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112:
The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention.
Claim 21 is rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention. Claim 21 recites “the learning processing is executed based on metric learning using only the first training image and the second training image”. While Paragraph 79 does describe a scenario where a first and second image are used without using the third training image, this does not exclude scenarios where a fourth training image is used. In other words, not using a third image is not the same as only using a first and second. Moreover, even in the scenario where a first and second image are used without using the third training image or any other training images, the metric learning uses many other factors like loss values, numerical predictions, feature amounts, etc. Therefore, the claimed scenario of “using only the first training image and the second training image” is never described, as the first and second training images are used in conjunction with the above listed features.
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 10 and 21 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Regarding claim 10, this claim recites “a first feature amount” and “a second feature amount”, which have improper antecedence and are being interpreted as “the first feature amount” and “the second feature amount”.
Regarding claim 21, this claim recites “the learning processing is executed based on metric learning using only the first training image and the second training image”. This is functionally impossible. Machine model learning requires additional features such as one or more of nodes, calculations, errors, weights, encoding/decoding, outputs, numerical values, etc. While the claim appears to attempt specify the situation where no images other than the first and second training images are used, the current wording does not sufficiently capture this. In contrast, the current wording describes a scenario where two images are used for metric learning without any other learning features/elements (i.e. not just images, but all possible features). The claim is being interpreted in accordance with the current wording.
The following is a quotation of 35 U.S.C. 112(d):
(d) REFERENCE IN DEPENDENT FORMS.—Subject to subsection (e), a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers.
The following is a quotation of pre-AIA 35 U.S.C. 112, fourth paragraph:
Subject to the following paragraph [i.e., the fifth paragraph of pre-AIA 35 U.S.C. 112], a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers.
Claims 10 and 14 are rejected under 35 U.S.C. 112(d) or pre-AIA 35 U.S.C. 112, 4th paragraph, as being of improper dependent form for failing to further limit the subject matter of the claim upon which it depends, or for failing to include all the limitations of the claim upon which it depends. The limitations of Claim 10 are fully embodied in the limitations of amended Claim 1, without adding additional limitations. Similarly, the limitations of Claim 14 are fully embodied in the limitations of amended Claim 1, without adding additional limitations. Claim 14 introduces several instances of new terminology (such as first estimation result) which appears to be identical to elements of claim 1 (processing result), just with a different name. Applicant may cancel the claim(s), amend the claim(s) to place the claim(s) in proper dependent form, rewrite the claim(s) in independent form, or present a sufficient showing that the dependent claim(s) complies with the statutory requirements.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 10, 12, 14-17 and 19-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Liu (Label-Sensitive Deep Metric Learning for Facial Age Estimation) in view of Yao (US 20190102668 A1).
Regarding claim 1, Liu teaches “A learning system, comprising at least one processor” (Liu, Table VI and section E.5. Paragraph 2, “We also investigated the computational time with different deep networks based on the efficient Caffe [51] toolbox, and the whole architectures were built on a speed-up parallel computing GPU with NVIDIA GTX 1080. Table VI tabulates the comparisons of the computational time during the testing phase. From these results, we see that the deep architectures achieve the real-time age estimation with a GPU for the feature extraction procedure. In addition, we carefully implemented the OHRANK by following the details provided in [9]. The OHRANK takes 0.04 seconds by using an Intel i5-CPU@3.20GHz PC, which also satisfies the real-time requirements.”)
“configured to: acquire a first training image relating to a first object having a first numerical value; acquire a second training image relating to a second object having a second numerical value;” (Liu, Figure 2 and section III.A. Paragraph 1, “To learn robust and discriminative feature similarity for facial age estimation, the basic idea of our LSDML is to exploit the label correlation among face samples in the transformed subspace. Unlike recent deep metric learning methods [37], [38] which utilize hand-crafted features to be fed to the deep networks, our model jointly optimizes both tasks of learning similarity and embedding features for face representation in a unified deep architecture. Let X={(xi,yi)}Ni=1 denote the training set which consists of N samples, where xi∈RD denote the i th face image of D pixels and yi∈R1 is the groundtruth age value, respectively. Our model is to compare the distance of face pairs by computing the feature representation f(xi) for the i th face image xi via deep neural networks. In terms of network architecture, we employ the residual learning method to optimize the whole network parameters, which have achieved superior performance in a volume of visual recognition tasks [43]. To better measure the learned face descriptors, we apply L2 normalization on the obtained outcomes from the fully connected layers.” The faces in the images are mapped to the objects and the groundtruth ages are mapped to the numerical values. Figure 2 shows training data, each associated with an image representing an age (numerical value).)
“and execute learning processing of a learning model which estimates a numerical value to be estimated relating to an object to be estimated included in an estimation-target image, wherein the learning processing includes: calculating a first loss based on a processing result of the learning model for the first training image and the first numerical value, the first loss including at least one of a softmax loss, a mean loss, or a variance loss;” (Liu, Sections IV.D. and IV.E., “1) Mean Absolute Error: For the evaluation metrics, we utilized the mean absolute error (MAE) [1], [19], [25], [33] to measure the error between the predicted age and the ground-truth, which is computed as follows:
PNG
media_image1.png
109
348
media_image1.png
Greyscale
where y^ and y∗ denote predicted and ground-truth age value, respectively, and N denotes the number of the testing samples. 2) Cumulative Score Curve: We also applied the cumulative score (CS) [23], [24], [26], [33] curve to quantitatively evaluate the performance of age estimation methods. The cumulative prediction accuracy at the error ϵ is computed as:
PNG
media_image2.png
99
245
media_image2.png
Greyscale
where K is the total number of testing images, Kn is the number of testing images whose absolute error between the estimated age and the ground-truth age is not greater than n years.”; “In our settings, we performed five folds cross-validation of our proposed approach on the MORPH (Album2) dataset. Specifically, we divided the whole dataset into five equal-size folds. Then we used one fold (20% of total data) as the testing set and the other four folds (80% of total data) as the training set. We repeated this procedure ten times and finally averaged the results as the facial age estimation results.” Note that the MORPH dataset has faces (objects to be estimated) in 55000 face images (estimation-target images). The model of Liu estimates age (numerical value to be estimated) related to these faces. The MAE is mapped to the first loss.)
“calculating a related loss based on a relationship between a first feature amount obtained from the first training image and a second feature amount obtained from the second training image, the related loss including a cosine similarity loss or a distance-based loss;” (Liu, Figures 2-3 and section III.A. Paragraphs 1 and 3, “To learn robust and discriminative feature similarity for facial age estimation, the basic idea of our LSDML is to exploit the label correlation among face samples in the transformed subspace. Unlike recent deep metric learning methods [37], [38] which utilize hand-crafted features to be fed to the deep networks, our model jointly optimizes both tasks of learning similarity and embedding features for face representation in a unified deep architecture. Let X={(xi,yi)}Ni=1 denote the training set which consists of N samples, where xi∈RD denote the i th face image of D pixels and yi∈R1 is the groundtruth age value, respectively. Our model is to compare the distance of face pairs by computing the feature representation f(xi) for the i th face image xi via deep neural networks. In terms of network architecture, we employ the residual learning method to optimize the whole network parameters, which have achieved superior performance in a volume of visual recognition tasks [43]. To better measure the learned face descriptors, we apply L2 normalization on the obtained outcomes from the fully connected layers.”; “The crucial part of our LSDML is to learn the network parameters f(⋅) . To achieve this goal, we first pass a given mini-batch forward the deep network, and we select each quadruplet of (i,j,k,l) , such that (xi,xj)∈P , (xi,xk)∈N and (xj,xl)∈N , where P and N denote the positive and negative pair set, respectively. More details are illustrated in Fig. 3. Moreover, to achieve the discriminativeness of the feature similarity, our LSDML enforces each df(xi,xj) pair in positive set is close to each other, and at the same time df(xi,xj) and df(xk,xl) in negative set is pushed far away. As a result, the distance of inter-class pairs is minimized, and the distance of intra-class pairs is larger than a margin τ in the transformed subspace.” Note that Liu calculates feature amounts for each training image based on the model, and calculates the distances between these feature amounts (related losses).)
While Liu discloses updating parameters of the learning model based on the related loss (see excerpt directly above), Liu does not expressly disclose updating parameters of the learning model based on the first loss. Rather, Liu describes using the first loss for evaluation.
Yao discloses updating parameters of a learning model based on a mean-based error relating to a prediction (Yao, Paragraph 6, “In accordance with a further aspect of the present disclosure, there is provided a method of learning an action model for an object, such as a vehicle, in an environment using a neural network. A subsequent state of the object in the environment, s′, is predicted from a current training state, s, from sample data set D {(s, a, s′)}, for at least two corresponding training actions, a. A reward is calculated for the subsequent state in accordance with a reward function. A predicted subsequent state, s′*, that produces a maximized reward is selected. A training error is calculated as the difference between the selected predicted subsequent state, s′*, and a corresponding subsequent state of the object in the environment, s′ from the sample data set D. Parameters of the neural network to minimize a mean square error (MSE) of the training error are updated.”
It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify Liu to include learning model parameter updating based on the first loss, as taught by the mean-based learning model parameter updating of Yao.
The motivation for doing so would have been to improve the learning model using an error value already calculated. Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Liu with the above teaching of Yao to fully disclose, “and updating parameters of the learning model based on the first loss and the related loss,”
Liu in view of Yao further disclose, “wherein the learning processing is executed based on metric learning using the first training image and the second training image.” (Liu, Figures 2-3 and Section III.A. Paragraph 1, “To learn robust and discriminative feature similarity for facial age estimation, the basic idea of our LSDML is to exploit the label correlation among face samples in the transformed subspace. Unlike recent deep metric learning methods [37], [38] which utilize hand-crafted features to be fed to the deep networks, our model jointly optimizes both tasks of learning similarity and embedding features for face representation in a unified deep architecture. Let X={(xi,yi)}Ni=1 denote the training set which consists of N samples, where xi∈RD denote the i th face image of D pixels and yi∈R1 is the groundtruth age value, respectively. Our model is to compare the distance of face pairs by computing the feature representation f(xi) for the i th face image xi via deep neural networks. In terms of network architecture, we employ the residual learning method to optimize the whole network parameters, which have achieved superior performance in a volume of visual recognition tasks [43]. To better measure the learned face descriptors, we apply L2 normalization on the obtained outcomes from the fully connected layers.” Accordingly, the metric learning of Liu uses the training images.)
Regarding claim 10, this claim is rejected under 35 USC 112(d) as failing to further limit claim 1. Accordingly, the rejection of claim 1 is applied here.
Regarding claim 12, Liu in view of Yao teaches “The learning system according to claim 1,”
“wherein the at least one processor is configured to: acquire a first estimation result obtained by the learning model based on the first training image; acquire a second estimation result obtained by the learning model based on the second training image; and execute the learning processing based on the first estimation result and the second estimation result.” (Liu, Figure 3 and section III.A. Paragraph 1, “To learn robust and discriminative feature similarity for facial age estimation, the basic idea of our LSDML is to exploit the label correlation among face samples in the transformed subspace. Unlike recent deep metric learning methods [37], [38] which utilize hand-crafted features to be fed to the deep networks, our model jointly optimizes both tasks of learning similarity and embedding features for face representation in a unified deep architecture. Let X={(xi,yi)}Ni=1 denote the training set which consists of N samples, where xi∈RD denote the i th face image of D pixels and yi∈R1 is the groundtruth age value, respectively. Our model is to compare the distance of face pairs by computing the feature representation f(xi) for the i th face image xi via deep neural networks. In terms of network architecture, we employ the residual learning method to optimize the whole network parameters, which have achieved superior performance in a volume of visual recognition tasks [43]. To better measure the learned face descriptors, we apply L2 normalization on the obtained outcomes from the fully connected layers.” Note that Liu calculates feature representations (estimation results) for each training image using the model, and uses these feature representations for the learning processing as described above and in Figure 3.)
Regarding claim 14, this claim is rejected under 35 USC 112(d) as failing to further limit claim 1. Accordingly, the rejection of claim 1 is applied here.
Regarding claim 15, Liu in view of Yao teaches “The learning system according to claim 14,”
“wherein the at least one processor is configured to execute the learning processing based on the first loss, a weighting coefficient relating to the related loss, and the related loss.” (Liu, Figure 2, Algorithm 1, Section III.A. Paragraph 3, “The crucial part of our LSDML is to learn the network parameters f(⋅) . To achieve this goal, we first pass a given mini-batch forward the deep network, and we select each quadruplet of (i,j,k,l) , such that (xi,xj)∈P , (xi,xk)∈N and (xj,xl)∈N , where P and N denote the positive and negative pair set, respectively. More details are illustrated in Fig. 3. Moreover, to achieve the discriminativeness of the feature similarity, our LSDML enforces each df(xi,xj) pair in positive set is close to each other, and at the same time df(xi,xj) and df(xk,xl) in negative set is pushed far away. As a result, the distance of inter-class pairs is minimized, and the distance of intra-class pairs is larger than a margin τ in the transformed subspace.” Note that the training to pull certain sets together and push other sets apart amounts to applying weight coefficients based on related losses. The use of the first loss and related loss, as claimed, are addressed in the rejection of claim 1.)
Regarding claim 16, Liu in view of Yao teaches “The learning system according to claim 14,”
“wherein the at least one processor is configured to: calculate a plurality of first losses each of which is based on the first estimation result and the first numerical value; and execute the learning processing based on the plurality of first losses and the related loss.” (This claim amounts to the performance of claim 1, but with the distinction that there are a plurality of first loss values used for the learning processing, along with the related loss. This is fully embodied in Figures 2-3 of Liu which show that there are a plurality of first loss values used for the learning processing.)
Regarding claim 17, Liu in view of Yao teaches “The learning system according to claim 14,”
“wherein the at least one processor is configured to: calculate a plurality of related losses each of which is based on the first processing result and the second processing result; and execute the learning processing based on the first loss and the plurality of related losses.” (This claim amounts to the performance of claim 1, but with the distinction that there are a plurality of related losses used for the learning processing, along with the first loss. This is fully embodied in Figures 2-3 of Liu which show that there are a plurality of related values used for the learning processing.)
Regarding claim 19, Claim 19 recites a method with steps corresponding to the elements of the system recited in Claim 1. Therefore, the recited steps of this claim are mapped to the analogous elements in the corresponding system claim. The rationale and motivation to combine the references apply here.
Regarding claim 20, Claim 20 recites a non-transitory computer-readable information storage medium storing a program with instructions corresponding to the steps recited in Claim 1. Therefore, the recited programming instructions of this claim are mapped to the analogous steps in the corresponding method claim. The rationale and motivation to combine the references apply here. Additionally, as described in the rejection of Claim 1, Liu teaches neural network operations performed by a processor of a computer. These operations are therefore program instructions that are inherently stored for execution.
Claim(s) 2-9 is/are rejected under 35 U.S.C. 103 as being unpatentable over Liu in view of Yao further in view of Rodriguez (WO 2020104542 A1).
Regarding claim 2, Liu in view of Yao teaches “The learning system according to claim 1,”
While Liu in view of Yao discloses that the learning model outputs, as an estimation result of an estimation-target age relating to an estimation-target person included in an estimation-target image, a predicted age (Liu, Figure 1 and Section IV.D.), that a third training image is obtained (see third and fourth training images exemplified in Figure 2), that this predicted age is calculated for a plurality of training images (i.e. the first, second, third) (Liu, Figure 1), and that learning processing is executed such that differences between predicted ages approach a margin corresponding to difference between ages (Liu, Figures 2-3 and described in Section III.A.: “The crucial part of our LSDML is to learn the network parameters f (·). To achieve this goal, we first pass a given mini-batch forward the deep network, and we select each quadruplet of (i, j,k,l), such that (xi,xj) ∈ P, (xi,xk) ∈ N and (xj,xl) ∈ N,whereP and N denote the positive and negative pair set, respectively. More details are illustrated in Fig. 3. Moreover, to achieve the discriminativeness of the feature similarity, our LSDML enforces each df (xi,xj) pair in positive set is close to each other, and at the same time df (xi,xj) and df (xk,xl) in negative set is pushed far away. As a result, the distance of inter-class pairs is minimized, and the distance of intra-class pairs is larger than a margin τ in the transformed subspace.”), they do not expressly disclose that a probability distribution indicating age probability within a range is calculated, or that the predicted age is calculated as an average based on these distributions.
Rodriguez discloses a probability distribution indicating age probability within a range is calculated, and that the predicted age is calculated as an average based on these distributions (Rodriguez, Page 8 Lines 6-9, “The output 210 of the age estimation system 200 is a probability distribution over all possible ages. An estimated age is computed from this as the probability-weighted average of these ages and the confidence score 212 is computed as the standard deviation of the distribution (the lower the standard deviation, the more confident the system is about the age estimation).”)
It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to replace the estimated age of Liu in view of Yao with average age calculated based on the calculated probability distribution of predicted ages within a predetermined range as taught by Rodriguez.
The motivation for doing so would have been to enable calculating a confidence represented by the standard deviation, as is performed by Rodriguez. This provides additional key information to be used for model evaluation in addition to training. Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Liu in view of Yao with the above teaching of Rodriguez to fully disclose, “wherein the learning model outputs, as an estimation result of an estimation-target age relating to an estimation-target person included in an estimation-target image, a probability distribution indicating, for each age within a predetermined range, a probability that the estimation- target person is the age, and wherein the at least one processor configured to: acquire a third training image relating to a third object having a third numerical value, calculate a first average age based on the age and the probability indicated in a first distribution, the first distribution being the probability distribution output from the learning model when the first training image is input to the learning model, calculate a second average age based on the age and the probability indicated in a second distribution, the second distribution being the probability distribution output from the learning model when the second training image is input to the learning model, calculate a third average age based on the age and the probability indicated in a third distribution, the third distribution being the probability distribution output from the learning model when the third training image is input to the learning model, and execute the learning processing such that a difference between a difference between any two of the first average age, the second average age, and the third average age and another difference between any other two thereof approaches a margin corresponding to a difference between a difference between any two of the first age, the second age, and the third age and another difference between any other two thereof.”
Regarding claim 3, Liu in view of Yao further in view of Rodriguez teaches “The learning system according to claim 2,”
“wherein the first numerical value is the same as the second numerical value, wherein the first numerical value is different from the third numerical value,” (Liu, Figure 2 shows that in the mini-batch samples, two samples (i, j) represent the same age (numerical value) and the third and fourth samples (k, l) represent a different age (numerical value).)
“and wherein the at least one processor is configured to: acquire a first processing result obtained by the learning model based on the first training image; acquire a second processing result obtained by the learning model based on the second training image; acquire a third processing result obtained by the learning model based on the third training image; and execute the learning processing so that a difference between the first processing result and the second processing result becomes smaller and a difference between the first processing result and the third processing result becomes larger.” (Liu, Figures 2-3 and section III.A. Paragraphs 1 and 3, “To learn robust and discriminative feature similarity for facial age estimation, the basic idea of our LSDML is to exploit the label correlation among face samples in the transformed subspace. Unlike recent deep metric learning methods [37], [38] which utilize hand-crafted features to be fed to the deep networks, our model jointly optimizes both tasks of learning similarity and embedding features for face representation in a unified deep architecture. Let X={(xi,yi)}Ni=1 denote the training set which consists of N samples, where xi∈RD denote the i th face image of D pixels and yi∈R1 is the groundtruth age value, respectively. Our model is to compare the distance of face pairs by computing the feature representation f(xi) for the i th face image xi via deep neural networks. In terms of network architecture, we employ the residual learning method to optimize the whole network parameters, which have achieved superior performance in a volume of visual recognition tasks [43]. To better measure the learned face descriptors, we apply L2 normalization on the obtained outcomes from the fully connected layers.”; “The crucial part of our LSDML is to learn the network parameters f(⋅) . To achieve this goal, we first pass a given mini-batch forward the deep network, and we select each quadruplet of (i,j,k,l) , such that (xi,xj)∈P , (xi,xk)∈N and (xj,xl)∈N , where P and N denote the positive and negative pair set, respectively. More details are illustrated in Fig. 3. Moreover, to achieve the discriminativeness of the feature similarity, our LSDML enforces each df(xi,xj) pair in positive set is close to each other, and at the same time df(xi,xj) and df(xk,xl) in negative set is pushed far away. As a result, the distance of inter-class pairs is minimized, and the distance of intra-class pairs is larger than a margin τ in the transformed subspace.” Note that Liu calculates feature representations (processing results) for each training image based on the model, and uses these for the learning processing as described above and in Figure 3. As described above and in Figure 2, right, the difference between the first and second is minimized, while the distance between the first and third is increased.)
Regarding claim 4, Liu in view of Yao further in view of Rodriguez teaches “The learning system according to claim 2,”
“wherein the first numerical value is different from both the second numerical value and the third numerical value, wherein a difference between the first numerical value and the third numerical value is larger than a difference between the first numerical value and the second numerical value,” (Liu, Figure 2: the first numerical value is mapped to i, the second is mapped to k, and the third is mapped to j. Numerical value i is different from k and l, and the difference between i and l is larger than the difference between i and k.)
“and wherein the at least one processor is configured to: acquire a first processing result obtained by the learning model based on the first training image; acquire a second processing result obtained by the learning model based on the second training image; acquire a third processing result obtained by the learning model based on the third training image; and execute the learning processing so that a difference between the first processing result and the third processing result becomes larger than a difference between the first processing result and the second processing result.” (Liu, Figures 2-3 and section III.A. Paragraph 1, “To learn robust and discriminative feature similarity for facial age estimation, the basic idea of our LSDML is to exploit the label correlation among face samples in the transformed subspace. Unlike recent deep metric learning methods [37], [38] which utilize hand-crafted features to be fed to the deep networks, our model jointly optimizes both tasks of learning similarity and embedding features for face representation in a unified deep architecture. Let X={(xi,yi)}Ni=1 denote the training set which consists of N samples, where xi∈RD denote the i th face image of D pixels and yi∈R1 is the groundtruth age value, respectively. Our model is to compare the distance of face pairs by computing the feature representation f(xi) for the i th face image xi via deep neural networks. In terms of network architecture, we employ the residual learning method to optimize the whole network parameters, which have achieved superior performance in a volume of visual recognition tasks [43]. To better measure the learned face descriptors, we apply L2 normalization on the obtained outcomes from the fully connected layers.”; Note that Liu calculates feature representations (processing results) for each training image based on the model, and uses these for the learning processing as described above and in Figure 3. As shown in Figure 2, right, the difference between the first and third is greater than the distance between the first and second.)
Regarding claim 5, Liu in view of Yao further in view of Rodriguez teaches “The learning system according to claim 3,”
“wherein the at least one processor is configured to execute the learning processing so that the difference between the first processing result and the third processing result becomes a difference corresponding to a difference between the first numerical value and the third numerical value.” (Liu, Figure 2, Algorithm 1, Section III.A. Paragraph 3, “The crucial part of our LSDML is to learn the network parameters f(⋅) . To achieve this goal, we first pass a given mini-batch forward the deep network, and we select each quadruplet of (i,j,k,l) , such that (xi,xj)∈P , (xi,xk)∈N and (xj,xl)∈N , where P and N denote the positive and negative pair set, respectively. More details are illustrated in Fig. 3. Moreover, to achieve the discriminativeness of the feature similarity, our LSDML enforces each df(xi,xj) pair in positive set is close to each other, and at the same time df(xi,xj) and df(xk,xl) in negative set is pushed far away. As a result, the distance of inter-class pairs is minimized, and the distance of intra-class pairs is larger than a margin τ in the transformed subspace.” Note that parameter optimization for minimizing loss, as shown in algorithm 1, amounts to making the processing result differences correspond to the real ground truth differences. This is also embodied in Figure 2, which shows differences in the feature space correspond to real age differences.)
Regarding claim 6, Liu in view of Yao further in view of Rodriguez teaches “The learning system according to claim 3,”
“wherein the first processing result is a first estimation result obtained by the learning model, wherein the second processing result is a second estimation result obtained by the learning model, wherein the third processing result is a third estimation result obtained by the learning model, and wherein the at least one processor is configured to execute the learning processing based on the first estimation result, the second estimation result, and the third estimation result.” (As described in the rejection of claim 3, the processing results for the three images are mapped to feature representations. These feature representations further correspond to the estimation results of claim 6. These feature representations are used for the learning processing, as recited in the rejection of claim 3 and in figure 3.)
Regarding claim 7, Liu in view of Yao further in view of Rodriguez teaches “The learning system according to claim 6,”
“wherein the first estimation result is a first distribution including each of a plurality of numerical values and a first probability that the first object has the each of the plurality of numerical values, wherein the second estimation result is a second distribution including each of the plurality of numerical values and a second probability that the second object has the each of the plurality of numerical values, wherein the third estimation result is a third distribution including each of the plurality of numerical values and a third probability that the third object has the each of the plurality of numerical values, and wherein the at least one processor is configured to execute the learning processing based on the first distribution, the second distribution, and the third distribution.” (As the reference are combined in the rejection of claim 2, Liu in view of Yao further in view of Rodriguez fully disclose the limitations of this claim. In particular, the combination describes that the estimations results are distributions including numerical values and probabilities, as claimed, and that the learning is based on all of the distributions, as further claimed.)
Regarding claim 8, Liu in view of Yao further in view of Rodriguez teaches “The learning system according to claim 3,”
“wherein the at least one processor is configured to execute the learning processing so that a difference between the second processing result and the third processing result becomes larger.” (Liu, Algorithm 1, Section III.A. Paragraph 3, “The crucial part of our LSDML is to learn the network parameters f(⋅) . To achieve this goal, we first pass a given mini-batch forward the deep network, and we select each quadruplet of (i,j,k,l) , such that (xi,xj)∈P , (xi,xk)∈N and (xj,xl)∈N , where P and N denote the positive and negative pair set, respectively. More details are illustrated in Fig. 3. Moreover, to achieve the discriminativeness of the feature similarity, our LSDML enforces each df(xi,xj) pair in positive set is close to each other, and at the same time df(xi,xj) and df(xk,xl) in negative set is pushed far away. As a result, the distance of inter-class pairs is minimized, and the distance of intra-class pairs is larger than a margin τ in the transformed subspace.” Note that parameter optimization for minimizing loss, as shown in algorithm 1, amounts to making the processing results correspond to the ground truth. In this case, making the difference between the second and third processing results larger, as the difference between the ground truth ages is larger. This is also embodied in Figure 2, which shows how in the feature space, the feature representations of farther ages move farther apart, and the arrow direction confirms this in the right panel.)
Regarding claim 9, Liu in view of Yao further in view of Rodriguez teaches “The learning system according to claim 8,”
“wherein the learning module at least one processor is configured to execute the learning processing so that the difference between the second processing result and the third processing result becomes a difference corresponding to a difference between the second numerical value and the third numerical value.” (Liu, Figure 2, Algorithm 1, Section III.A. Paragraph 3, “The crucial part of our LSDML is to learn the network parameters f(⋅) . To achieve this goal, we first pass a given mini-batch forward the deep network, and we select each quadruplet of (i,j,k,l) , such that (xi,xj)∈P , (xi,xk)∈N and (xj,xl)∈N , where P and N denote the positive and negative pair set, respectively. More details are illustrated in Fig. 3. Moreover, to achieve the discriminativeness of the feature similarity, our LSDML enforces each df(xi,xj) pair in positive set is close to each other, and at the same time df(xi,xj) and df(xk,xl) in negative set is pushed far away. As a result, the distance of inter-class pairs is minimized, and the distance of intra-class pairs is larger than a margin τ in the transformed subspace.” Note that parameter optimization for minimizing loss, as shown in algorithm 1, amounts to making the processing result differences correspond to the real ground truth differences. Accordingly, feature representations are optimized to have differences correlated to ground truth differences, including the differences for the second and third processing result. This is also embodied in Figure 2, which shows differences in the feature space correspond to real age differences.)
Claim(s) 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Liu in view of Yao further in view of Ustinova (US 20200218971 A1).
Regarding claim 11, Liu in view of Yao teaches “The learning system according to claim 10,”
While Liu in view of Yao calculates similarity between the first and second feature amounts and executes learning processing based on the similarity (Liu, Figure 3 shows learning based on similarity and Section III.A. Paragraph 2:
PNG
media_image3.png
423
731
media_image3.png
Greyscale
Here, Liu describes similarity calculation as Euclidian distance between the first and second feature amounts.), Liu in view of Yao does not expressly disclose the use of cosine similarities.
Ustinova discloses the use of cosine similarities with respect to image representations (Ustinova, Paragraph 80, “These maps are combined into a single vector of deep representation of the original image with the length of 500 elements using a fully connected layer, in which each element of the output vector is linked to each element of the map of attributes of each part. Subsequently, L2-normalization is performed. The obtained deep representations of input images along with the marks of classes are used to determine the measures of similarity for all possible pairs through the calculation of a cosine measure of similarity between deep representations, then the probability distributions of similarity measures of positive and negative pairs are formed on the basis of marks of classes using histograms. The obtained distributions are used to calculate the loss function proposed, then the back propagation of the derived loss is performed to adjust the neural network parameters. The process of deep neural network learning at all bases differed only by the loss function selected. The learning with a binomial loss function was performed with two values of losses for negative pairs: c=10 and c=25.”)
It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to replace the Euclidian distance-based similarity calculation of Liu in view of Yao with the cosine similarity calculation of Ustinova.
The motivation for doing so would have been to use a well-known and conventional alternative (cosine similarity) to compare the feature representations. Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Liu in view of Yao with the above teaching of Ustinova to fully disclose, “wherein the at least one processor is configured to: calculate a cosine similarity based on the first feature amount and the second feature amount; and execute the learning processing based on the cosine similarity.”
Claim(s) 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Liu in view of Yao further in view of Rodriguez further in view of Wang (CN 111242091 A).
Regarding claim 13, Liu in view of Yao teach “The learning system according to claim 12,”
While Liu in view of Yao teach numerical age estimations related to the first and second object (Liu, Figures 1-3 and Section IV.D.), then calculating a loss and performing training based on the loss (see claim 1 rejection), Liu in view of Yao do not expressly disclose “wherein the first estimation result is a first distribution including each of a plurality of numerical values and a first probability that the first object has the each of the plurality of numerical values, wherein the second estimation result is a second distribution including each of the plurality of numerical values and a second probability that the second object has the each of the plurality of numerical values, and wherein the at least one processor is configured to: calculate a Kullback-Leibler divergence based on the first distribution and the second distribution; and execute the learning processing based on the Kullback-Leibler divergence.”
Rodriguez teaches age estimation results that are distributions including a plurality of numerical values and associated probabilities related to objects (Rodriguez, Page 8 Lines 6-9, “The output 210 of the age estimation system 200 is a probability distribution over all possible ages. An estimated age is computed from this as the probability-weighted average of these ages and the confidence score 212 is computed as the standard deviation of the distribution (the lower the standard deviation, the more confident the system is about the age estimation).”)
Wang teaches age recognition model training based on KL divergence (Wang, First full Paragraph of Page 5, “The embodiment of the invention claims a age recognition model training plan, by obtaining a plurality of sample face image of the preset age section, each sample face image corresponding to a real age value, for several sample face each sample face image in the image, obtaining the area sample face images of a plurality of specific areas corresponding to the sample face image, according to the initial age recognition model, the predicted age value obtaining a sample face image and face image respectively corresponding to each area sample, according to each prediction age value and true age value. calculating the initial age recognition model of the KL divergence loss value and loss value between the average absolute error value; the lower and of the sum value is in the predetermined range, the initial age recognition model as the final age recognition model. The embodiment of the invention uses KL based on loss and mae loss of small network for age estimation, small network model to realize the fast and accurate age estimation, and reinforcing step area selection operation on the sample image can be a certain meaning, so that training time acquiring enough feature detail information, auxiliary network training, the prediction result has more robustness.”)
It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to replace the estimated age of Liu in view of Yao with calculated probability distribution of predicted ages as taught by Rodriguez, and replace the MAE loss-based training of Liu in view of Yao with the KL divergence-based training of Wang.
The motivation for doing so would have been to enable calculating a confidence represented by the standard deviation, as is performed by Rodriguez, then acquire a more comprehensive loss value (KL divergence) compared to MAE. Further, one skilled in the art could have combined the elements as described above by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Liu in view of Yao with the above teachings of Rodriguez and Wang to fully disclose, “wherein the first estimation result is a first distribution including each of a plurality of numerical values and a first probability that the first object has the each of the plurality of numerical values, wherein the second estimation result is a second distribution including each of the plurality of numerical values and a second probability that the second object has the each of the plurality of numerical values, and wherein the at least one processor is configured to: calculate a Kullback-Leibler divergence based on the first distribution and the second distribution; and execute the learning processing based on the Kullback-Leibler divergence.”
Regarding claim 21, claim 21 is not rejected under Prior Art.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to AARON JOSEPH SORRIN whose telephone number is (703)756-1565. The examiner can normally be reached Monday - Friday 9am - 5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Sumati Lefkowitz can be reached at (571) 272-3638. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/AARON JOSEPH SORRIN/Examiner, Art Unit 2672
/SUMATI LEFKOWITZ/Supervisory Patent Examiner, Art Unit 2672