Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Notice to Applicants
This communication is in response to the amendment filed on 06/29/2026.
Claims 1-20 are pending.
Response to Arguments
Applicant’s arguments and amendments with respect to claim(s) 1-20 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Applicant’s amendments for the 112b rejections have been fully considered and they are persuasive and sufficient. The 112b rejections have been withdrawn.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 1-3, 7, 9-12, 16 and 18- 20 are rejected under 35 U.S.C. 103 as being unpatentable over Pan et al. (U.S. Publication No. 2023/0196710) (hereafter, "Pan") in view of YANG et al. (U.S. Publication No. 2024/0249515) (hereafter, "YANG").
Regarding claim 1, Pan teaches an image processing method comprising ([0018] The framework, in an embodiment is an input-dependent dynamic inference framework for the vision transformer ... The framework can be applicable to any types of models and any types of tasks such as image processing (e.g., classifying objects in an image), video processing (e.g., classifying actions in the video) and natural language processing): determining, by a patch sampler model ([0020] FIG. 1 is a diagram illustrating an interpretability-aware redundancy reduction framework; [0022] the framework can include vision transformers … a transformer architecture can include multi-headed self-attention (MSA) 102 and feed forward network (FFN) 104; [0023] the framework which can provide interpretability-aware redundancy reduction can include a plurality of multi-headed interpreters 106 to learn and determine patches of an image), selection probabilities of a first plurality of patches included in an image ([0033] the multi-head interpreter 106 can use a policy token 116 to estimate the importance of the input token. For instance, in vision transformers, image patches can be referred to as tokens, e.g., patch tokens ... The multi-head interpreter 106 receives patch tokens 108 … policy token can be multiplied with the input tokens (e.g., after linear projection) for obtaining the importance scores in Eq. (1); [0034] the patch tokens 108 can be evaluated by the multi-head interpreter 106 for the informative score Iij; FIG. 1; [0027] Multi-head interpreter 106 receives input tokens, e.g., patches of images, 108); selecting, by the patch sampler model ([0020] an interpretability-aware redundancy reduction framework; [0033] the multi-head interpreter 106), a second plurality of patches from among the first plurality of patches of the image based on the selection probabilities; and ([0032] Before input to the MSA and FFN, the patch tokens are evaluated by the multi-head interpreter to drop some uninformative patches; [0033] policy token can be multiplied with the input tokens (e.g., after linear projection) for obtaining the importance scores in Eq. (1). Policy queries 120 refer to input tokens after linear transformation, e.g., linear transformation of input patch tokens produces policy queries. Policy queries 120 output from the linear projection 118 are combined with the policy token 116 (e.g., by dot product computation) and input to activation functions, e.g., a sigmoid function. Based on the activations, the multi-head interpreter 106 outputs a vector, e.g., a binary vector indicating patch tokens to drop and keep) determining a classification of the image by processing the second plurality of patches through an encoder ([0023] the framework which can provide interpretability-aware redundancy reduction can include a plurality of multi-headed interpreters 106 to learn and determine patches of an image, which are considered informative for classifying the image ... The framework learns which patches are uninformative for a particular task and removes those patches from classification computation or transformer encoding; [0035] all of the MSA-FFN blocks in the original vision transformer can be evenly assigned into D groups in the framework, where each group contains L MSA-FFN blocks and one multi-head interpreter ... the interpreter 106 at the early stage may learn to select the patches containing all of the necessary contextual information for the correct final prediction; [0042] the framework can interpret where the informative region for the correct prediction is and can localize the salient object on the input images ... Highlighted informative regions can be used as interpretation or explanation for the prediction, e.g., object classification); wherein the patch sampler model ([0020] an interpretability-aware redundancy reduction framework; [0033] the multi-head interpreter 106) is trained based on a sampling loss ([0036] the framework may optimize the multi-head interpreters 106 by using a reinforce method where the reward 112 considers both the efficiency and accuracy, and fine-tune the MSA-FFN blocks 102-104 with gradients computed based on cross-entropy loss; [0053]) which indicates a difference ([0037] Iij is defined in the Equation 1 and Xi denotes the i-th token in the token sequence X. These actions are associated with the reward function: R(u) = 1-(|u|0/N)2 if correct, -τ otherwise (2) where (|u|0/N)2 measures the percentage of the patches kept, and τ is the value of penalty for the error prediction which controls the trade-off between the efficiency and the accuracy of the network. This reward function encourages the multi-head interpreter to predict the correct results with as few patch tokens as possible) between the selection probabilities and attention scores of the image obtained via the encoder ([0033] the importance scores in Eq. (1) … the multi-head interpreter 106 outputs a vector; [0027] Multi-head interpreter 106 receives input tokens, e.g., patches of images, 108 and outputs a vector 110, e.g., binary vector of zeroes and ones, indicating which patches to keep and which patches to remove, for the transformer encoding or processing for prediction 114, e.g., by MSA 102 and FFN 104 … the multi-head interpreter 106 implements reinforcement learning, e.g., reward-based 112, to learn the binary vector 110; [0028] Vision transformer can include multi-head self-attention layer (MSA) 102, which learns relationships between every two different patches among all the input tokens … the query Qi computes the dot products with all the keys K and these dot products are scaled and normalized by the softmax layer to get the attention weights); and wherein the patch sampler model is trained to minimize the sampling loss ([0036] the framework may optimize the multi-head interpreters 106 by using a reinforce method where the reward 112 considers both the efficiency and accuracy, and fine-tune the MSA-FFN blocks 102-104 with gradients computed based on cross-entropy loss; [0037] These actions are associated with the reward function: R(u) = 1-(|u|0/N)2 if correct, -τ otherwise (2) where (|u|0/N)2 measures the percentage of the patches kept, and τ is the value of penalty for the error prediction which controls the trade-off between the efficiency and the accuracy of the network; [0047] A pseudo-code that shows details of a training process of the framework is shown in Algorithm 1. By way of example, D, the number of groups is set to 3. Wp denotes the parameters of the multi-head interpreters, Wb denotes the parameters of the MSA-FFN blocks.
Algorithm 1 Optimize multi-head interpreters and MSA-FFN blocks
Input: A token sequence X right after the positional embedding and its label Y.
for i ← 1 to D do
for j ← 1 to 10 do
for j ← 11 to 30 do
for each iteration do
for each iteration do
R ← Reward(X, Y | Wp1:i, Wb)
L ← CrossEntropyLoss(X, Y | Wp1:i, Wb)
Compute_Policy_Gradient(R)
Compute_Gradient(L)
Wpi ← Update_Parameters(Wpi)
Wb1:D ←Updated_Parameters(Wb1:D)
end for
end for
end for
end for
In machine learning, maximizing a reward (REINFORCE) or minimizing cross-entropy loss are standard optimization processes. Minimizing a loss function is mathematically equivalent to maximizing a reward/objective function), such that the sampling loss between a distribution of the selection probabilities of the second plurality of patches ([0037] given a sequence of patch tokens X ∈ RN×d input to the jth multi-head interpreter, the multi-head interpreter generates policies for each input token of dropping or keeping it as Bernoulli distribution by: πw(ui|Xi)=Iijui*(1−Iij)1−ui, where ui=1 means to keep the token and ui=0 means to discard the token, Iij is defined in the Equation 1 and Xi denotes the i-th token in the token sequence X; [0033] the multi-head interpreter 106 can use a policy token 116 to estimate the importance of the input token. For instance, in vision transformers, image patches can be referred to as tokens, e.g., patch tokens ... The multi-head interpreter 106 receives patch tokens 108 … policy token can be multiplied with the input tokens (e.g., after linear projection) for obtaining the importance scores in Eq. (1); [0034] the patch tokens 108 can be evaluated by the multi-head interpreter 106 for the informative score Iij) and … attention scores of the second plurality of patches ([0028] Vision transformer can include multi-head self-attention layer (MSA) 102, which learns relationships between every two different patches among all the input tokens … the query Qi computes the dot products with all the keys K and these dot products are scaled and normalized by the softmax layer to get the attention weights; [0035] The framework ... focus on the parameters inside each group, e.g., optimizing parameters like policy tokens, parameters in multi-headed self-attention and feedforward network inside each group).
Pan does not expressly teach … a distribution of … is reduced to a predetermined threshold or converges.
However, YANG teaches the sampling loss between ([0045] The relative entropy loss is determined based on a difference between an expected distribution 122 corresponding to a bag label labeled by the patch bag 120 and the foregoing predicted attention distribution 121; [0133] Step 3035: Determine, based on a difference between the attention distribution and the expected distribution, the relative entropy loss corresponding to the sample patch bag; [0134] after the attention distribution and the expected distribution are determined, the difference between the attention distribution and the expected distribution is determined, so that the relative entropy loss corresponding to the sample patch bag is obtained) the selection probabilities of the … plurality of patches and ([0045] The relative entropy loss is determined based on a difference between an expected distribution 122 corresponding to a bag label labeled by the patch bag 120 and the foregoing predicted attention distribution 121; [0116] Step 3034: Determine, based on the bag label, an expected distribution corresponding to the sample patch bag; [0119] 1. The expected distribution of the sample patch bag is determined, in response to the bag label indicating that the image content does not exist in the sample image patches in the sample patch bag, as a uniform distribution; [0122] 2. A patch label corresponding to the sample image patches is obtained in the bag label in response to the bag label indicating that the image content exists in the sample image patches in the sample patch bag, and the expected distribution of the sample patch bag is determined based on the patch labels) a distribution of attention scores of the … plurality of patches ([0045] an attention distribution 121 of the patch bag 120 is classified and predicted by an attention layer 123 according to the bag feature and the patch features; [0111] Step 3033: Perform attention analysis on the first fully connected feature by an attention layer to obtain an attention distribution corresponding to the sample patch bag; [0113] the foregoing attention layer applies the attention mechanism to selectively focus on a feature belonging to the positive classification or the negative classification in the first fully connected feature, so that the attention distribution corresponding to the sample patch bag is obtained) is reduced to a predetermined threshold or converges ([0045] The relative entropy loss is determined based on a difference between an expected distribution 122 corresponding to a bag label labeled by the patch bag 120 and the foregoing predicted attention distribution 121; [0133] Step 3035: Determine, based on a difference between the attention distribution and the expected distribution, the relative entropy loss corresponding to the sample patch bag; [0134] after the attention distribution and the expected distribution are determined, the difference between the attention distribution and the expected distribution is determined, so that the relative entropy loss corresponding to the sample patch bag is obtained; In machine learning, minimizing a relative entropy loss inherently means driving the system toward convergence).
It would have been obvious before the effective filing date of the claimed invention to one having ordinary skill in the art to modify the method and device of Pan to incorporate the step/system of calculating a distribution of attention scores, establishing an expected target distribution (selection probabilities) and calculating a relative entropy loss (sampling loss) based on their difference taught by YANG.
The suggestion/motivation for doing so would have been to improve the accuracy of recognizing image content ([0006] provide a method and apparatus for training an image recognition model, a device, and a medium, capable of improving accuracy of image recognition; [0016] During a training process of an image recognition model, for an image that requires patch recognition, the image recognition model is trained using sample image patches and a patch bag separately. While overall accuracy of recognizing the sample image is improved, accuracy of recognizing image content in the sample image patches is also improved. This not only avoids a problem of an erroneous recognition result of the entire image due to incorrect recognition of a single image patch, but also increases a screening negative rate when the image recognition model is used to recognize a lesion in a pathological image, to improve efficiency and accuracy of lesion recognition). Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predicted results. Therefore, it would have been obvious to combine Pan and YANG to obtain the invention as specified in claim 1.
Regarding claim 2, the combination of Pan and YANG teaches all the limitations of claim 1 above. Pan teaches wherein the attention scores are computed based on attention weights of the second plurality of patches, and ([0028] Vision transformer can include multi-head self-attention layer (MSA) 102, which learns relationships between every two different patches among all the input tokens. There can be h self-attention heads inside the MSA 102. In each self-attention head, the input token X, is first projected to a query Qi, a key Ki, and a value V, by three different linear transformations. Then, the query Qi computes the dot products with all the keys K and these dot products are scaled and normalized by the softmax layer to get the attention weights; [0032] Before input to the MSA and FFN, the patch tokens are evaluated by the multi-head interpreter to drop some uninformative patches; FIG. 1) wherein each of the attention weights indicates an importance of a corresponding patch of the second plurality of patches ([0028] In each self-attention head … the query Qi computes the dot products with all the keys K and these dot products are scaled and normalized by the softmax layer to get the attention weights … it outputs the token Yi by weighted sum of all the values V with the obtained attention weights; [0052] the reduced sequence of patch tokens).
Regarding claim 3, the combination of Pan and YANG teaches all the limitations of claim 1 above. Pan teaches wherein each of the attention scores is computed by: for a specific patch among the second plurality of patches ([0028] Vision transformer can include multi-head self-attention layer (MSA) 102, which learns relationships between every two different patches … In each self-attention head, the input token X, is first projected to a query Qi, a key Ki, and a value V, by three different linear transformations; [0032] Before input to the MSA and FFN, the patch tokens are evaluated by the multi-head interpreter to drop some uninformative patches; FIG. 1), computing a plurality of individual attention scores ([0028] the query Qi computes the dot products with all the keys K and these dot products are scaled and normalized by the softmax layer to get the attention weights) from a plurality of transformer layers of the encoder ([0028] Vision transformer can include multi-head self-attention layer (MSA) 102, which learns relationships between every two different patches among all the input tokens. There can be h self-attention heads inside the MSA 102); and aggregating the plurality of individual attention scores to obtain the attention score for the specific patch ([0028] After that, it outputs the token Yi by weighted sum of all the values V with the obtained attention weights).
Regarding claim 7, the combination of Pan and YANG teaches all the limitations of claim 1 above. Pan teaches wherein the patch sampler model is trained using a reinforce algorithm ([0048] the multi-head interpreters can be trained using REINFORCE, which does not require gradients for the backbone network and saves much of computation. Briefly, REINFORCE is a policy gradient algorithm in reinforcement learning) based on the selection probabilities as an action ([0034] the patch tokens 108 can be evaluated by the multi-head interpreter 106 for the informative score Iij, where i and j represent the position of the input token and the group respectively; [0037] during the training phase, given a sequence of patch tokens X ∈ RN×d input to the jth multi-head interpreter, the multi-head interpreter generates policies for each input token of dropping or keeping it as Bernoulli distribution by: πw(ui|Xi) = Iijui*(1−Iij)1−ui, where ui=1 means to keep the token and ui=0 means to discard the token, Iij is defined in the Equation 1 and Xi denotes the i-th token in the token sequence X) and the attention scores as a reward ([0006] The processor can also be configured to fine-tune the attention-based deep learning neural network to recognize the object in the image using the reduced sequence of patch tokens; [0028] the query Qi computes the dot products with all the keys K and these dot products are scaled and normalized by the softmax layer to get the attention weights. After that, it outputs the token Yi by weighted sum of all the values V with the obtained attention weights).
Regarding claim 9, the combination of Pan and YANG teaches all the limitations of claim 1 above. Pan teaches wherein the encoder is trained based on a cross entropy loss ([0036] the framework may optimize the multi-head interpreters 106 by using a reinforce method where the reward 112 considers both the efficiency and accuracy, and fine-tune the MSA-FFN blocks 102-104 with gradients computed based on cross-entropy loss).
With respect to claim 10, arguments analogous to those presented for claim 1, are applicable.
With respect to claim 11, arguments analogous to those presented for claim 2, are applicable.
With respect to claim 12, arguments analogous to those presented for claim 3, are applicable.
With respect to claim 16, arguments analogous to those presented for claim 7, are applicable.
With respect to claim 18, arguments analogous to those presented for claim 9, are applicable.
With respect to claim 19, arguments analogous to those presented for claim 1, are applicable.
With respect to claim 20, arguments analogous to those presented for claim 2, are applicable.
Claim 8 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Pan et al. (U.S. Publication No. 2023/0196710) (hereafter, "Pan") in view of YANG et al. (U.S. Publication No. 2024/0249515) (hereafter, "YANG") and further in view of KIMURA (U.S. Publication No. 2024/0071130).
Regarding claim 8, the combination of Pan and YANG teaches all the limitations of claim 1 above. Pan teaches wherein the patch sampler model ([0020] an interpretability-aware redundancy reduction framework; [0033] the multi-head interpreter 106).
Pan does not expressly teach … is trained using a Mean Squared Error (MSE) loss or a Kullback-Leibler (KL) Divergence loss.
However, KIMURA teaches … is trained using a Mean Squared Error (MSE) loss or a Kullback-Leibler (KL) Divergence loss ([0049] The DNN and the class token are learnable parameters and are trained through a learning process to be described later in FIG. 6; [0051] The DNN has a configuration in which transformer encoder blocks ... step S104 functions as an encoding step of obtaining the encoded representation sequence 304 by calculating an attention map indicating the degree of association between tokens to update the token sequence; [0102] one inter-token association degree map is calculated for all the attention maps corresponding to one face image Ii from a set of region information and inter-region association degree map; [0106] the attention loss calculation unit 604 calculates an attention loss value LAttn from the attention map Ai and the association degree map Ri of each face image, for example, using the following Formula 6; [0107] [Formula 6] LAttn=ΣiΣlΣhΣtΣu|Ailh(t,u)−Ri(t,u)|, if Ri(t,u)≠NaN; [0108] although the absolute error is used as the loss value in Formula 6, the squared error may be used).
It would have been obvious before the effective filing date of the claimed invention to one having ordinary skill in the art to modify the device and method of Pan to incorporate the step/system of training the DNN (transformer encoder) using the attention loss which is calculated by squared errors by KIMURA.
The suggestion/motivation for doing so would have been to improve the efficiency and accuracy of image recognition by ensuring each image token contains a comprehensive representation ([0118] the representative vector is improved to function more effectively as a value representing the feature of each person, and the feature vectors output by the feature vector calculation unit 100 are improved to resemble each other if they are the feature vectors of the same person. Therefore, each element of the attention map output by each self-attention of the feature vector calculation unit 100 is improved so as to approach the target value set in the inter-region association degree map; [0121] in self-attention, it is difficult to pay attention to a region unnecessary for face authentication, and more attention is paid to an important region for face authentication. Therefore, the performance of face authentication is improved by important information for face authentic; [0084] The representative vector method is a method of learning face authentication in which a feature amount vector representing each person is set and used together to improve learning efficiency). Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predicted results. Therefore, it would have been obvious to combine Pan and KIMURA to obtain the invention as specified in claim 8.
With respect to claim 17, arguments analogous to those presented for claim 8, are applicable.
Allowable Subject Matter
Claim 4-6 and 13-15 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DANIEL C. CHANG whose telephone number is (571)270-1277. The examiner can normally be reached Monday-Thursday and Alternate Fridays 8:00-5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chan S. Park can be reached at (571) 272-7409. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DANIEL C CHANG/Examiner, Art Unit 2669 /CHAN S PARK/Supervisory Patent Examiner, Art Unit 2669