DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of the Claims
Claims 1-10, as originally filed, are currently pending and have been considered below.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: “input unit configured to receive,” “detection unit configured to detect,” “jigsaw generation unit configured to transform” and “training unit configured to specify” in claims 6 and 9.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-10 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
The scope of the claims cannot be properly determined as they do not conform with current U.S. practice regarding the use of abbreviations or reference characters. When establishing abbreviations or reference characters in the claims the abbreviation should first be stated in full immediately followed by the abbreviation intended to be subsequently used. The abbreviation must be established for each independent claim.
Claim 1 recites the limitation, “A method of detecting an OOD object.” This limitation is indefinite and should recite, i.e., “A method of detecting an out-of-distribution (OOD) object.”
Claim 6 recites the limitation, “A system for detecting an OOD object.” This limitation is indefinite and should recite, i.e., “A system for detecting an out-of-distribution (OOD) object.”
Claim 7 recites the limitation, “recognizing an OOD object” in line 5 of the claim. This limitation is indefinite and should recite, i.e., “recognizing an out-of-distribution (OOD) object.”
Claim 8 recites the limitation, “specifying the jigsaw image as a proxy OOD.” This limitation is indefinite and should recite, i.e., “specifying the jigsaw image as a proxy out-of-distribution (OOD) image.”
Claim 9 recites the limitation, “specify the jigsaw image as a proxy OOD.” This limitation is indefinite and should recite, i.e., “specify the jigsaw image as a proxy out-of-distribution (OOD) image.”
Claim 10 recites the limitation, “specify the jigsaw image as a proxy OOD.” This limitation is indefinite and should recite, i.e., “specify the jigsaw image as a proxy out-of-distribution (OOD) image.”
Claims 2-5 are rejected for being dependent on a rejected base claim.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 7 and 10 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claim(s) does/do not fall within at least one of the four categories of patent eligible subject matter because claims 7 and 10 are directed to a program stored on a computer-readable recording medium. The United States Patent and Trademark Office (USPTO) is obliged to give claims their broadest reasonable interpretation consistent with the specification during proceedings before the USPTO. See In re Zletz, 893F.2d 319 (Fed. Cir. 1989) (during patent examination the pending claims must be interpreted as broadly as their terms reasonably allow). The broadest reasonable interpretation of claims 7 and 10 drawn to a program stored on a computer readable recording medium covers forms of non-transitory tangible media and transitory propagating signals per se in view of the ordinary and customary meaning of computer readable storage media as well as software per se. See MPEP 2111.01. When the broadest reasonable interpretation of a claim covers a signal per se or software per se, the claim must be rejected under 35 U.S.C. § 101 as covering non-statutory subject matter. See In re Nuijten, 500 F.3d 1346, 1356-57 (Fed. Cir. 2007) (transitory embodiments are not directed to statutory subject matter). See also Kappos memo on Subject Matter Eligibility of Computer Readable Media, February 23, 2010, 1351 OG 212. Applicants are advised to amend claims 7 and 10 to recite "A non-transitory computer readable storage medium storing instructions... " or “A non-transitory computer readable storage medium storing program instructions...” in order to overcome the rejection.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1-10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yu Y, Shin S, Lee S, Jun C, Lee K. Block Selection Method for Using Feature Norm in Out-of-distribution Detection. arXiv preprint arXiv:2212.02295v3. 2023 Mar 2, hereinafter, “Yu”, and further in view of Chi Zhang, Yue Song, Nicu Sebe, Yao Zhao, Wei Wang, "Why SAM finetuning can benefit Out-of-Distribution Detection?" ICLR 2024. 25 Mar 2024, hereinafter, “Zhang”.
As per claim 1, Yu discloses a method of detecting an OOD object using an OOD object detection system (Yu, Abstract, Detecting out-of-distribution (OOD) inputs during the inference stage; Yu, page 1, 1. Introduction, separate the in-distribution (ID) and out-of-distribution (OOD) data ... the norm of the feature map for the OOD and ID is quite separable ... This motivates a simple and effective OOD detection framework), comprising:
receiving an input image (Yu, page 2, Figure 2. Illustration of our proposed out-of-distribution detection framework ... During inference time, for a given input image, the OOD score is calculated); and
recognizing an OOD object from the input image using a pre-trained deep learning model (Yu, page 2, Figure 2. Illustration of our proposed out-of-distribution detection framework ... We use NormRatio of ID and pseudo OOD (i.e., Jigsaw puzzle images) to find which block is suitable for OOD detection (a). During inference time, for a given input image, the OOD score is calculated by FeatureNorm on the selected block (b). If FeatureNorm for a given input is smaller than the threshold, the given input is classified as OOD; Yu, page 3, 3.1. Overview of OOD detection framework, After the training is done, we select the block by NormRatio for OOD detection (Figure 2; left). Then, we use the norm of the feature map FeatureNorm obtained from the selected feature map for OOD detection during the inference stage (Figure 2; right)),
wherein the pre-trained deep learning model is trained according to a method of training a deep learning model (Yu, page 4, 4. Experiments, Setup We use commonly utilized CNN architectures: ResNet18, VGG11 and WideResNet with depth 28 and width 10 (WRN28) for the CIFAR10 benchmark. The ResNet18 and VGG11 are trained with batch size 128 for 100 epochs ... The WRN28 is trained with batch size 128 for 200 epochs; Yu, page 4, Table 1. Summary of the selected blocks for each architecture), the method of training a deep learning model including:
receiving an original image (Yu, page 2, Figure 2. Illustration of our proposed out-of-distribution detection framework ... During inference time, for a given input image, the OOD score is calculated);
transforming unique features represented from the original image to generate a jigsaw image (Yu, page 3, 3.1. Overview of OOD detection framework, we first generate the Jigsaw puzzle image as pseudo OOD from ID training samples and calculate NormRatio of training samples and corresponding pseudo OOD; Yu, page 4, 3.3. NormRatio: measure of block’s suitability, we generate the 3x Jigsaw puzzle image, which is semantically shifted ... using training samples);
specifying the jigsaw image as a proxy OOD (Yu, Abstract, we create Jigsaw puzzle images as pseudo OOD from ID training samples); and
training the deep learning model using the original image to recognize the OOD object from the input image (Yu, page 3, 3.1. Overview of OOD detection framework, we first generate the Jigsaw puzzle image as pseudo OOD from ID training samples … After the training is done, we select the block by NormRatio for OOD detection (Figure 2; left). Then, we use the norm of the feature map FeatureNorm obtained from the selected feature map for OOD detection during the inference stage (Figure 2; right). Specifically, we first generate the Jigsaw puzzle image as pseudo OOD from ID training samples and calculate NormRatio of training samples and corresponding pseudo OOD).
Yu does not explicitly disclose the following limitations as further recited however Zhang discloses
training the deep learning model using the original image and the jigsaw image to recognize the OOD object from the input image (Zhang, page 2, 1 Introduction, We propose our Sharpness-aware Fine-Tuning (SFT) method, in which the pretrained model is fine-tuned by SAM using pseudo-OOD data within 1 epoch (50 steps more precisely). Fig. 1 depicts the overall procedure of the fine-tuning process. Our SFT generates pseudo-OOD data using both Jigsaw Puzzle Patch Shuffling and RGB Channel Shuffling. As can be seen in Fig. 2, ID data and OOD data are better distinguished; Zhang, page 4, 3 Methodology, Preliminary: OOD detection task. The OOD detection problem usually is defined as a binary classification problem between ID and OOD data ... given the data distributions Din and Dout, we fine-tune the pre-trained model ... To craft Dout, we use both Jigsaw Puzzle Patch Shuffling and RGB Channel Shuffling methods to generate pseudo-OOD data. More specifically, Jigsaw Puzzle Patch Shuffling divides the original ID image into small patches and then reassembles these patches randomly to generate new images ... At each fine-tuning iteration, we use both Jigsaw Puzzle Patch Shuffling and RGB Channel Shuffling to generate pseudo-OOD data).
It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to combine the teachings of Zhang with Yu because they are in the same field of endeavor. One skilled in the art would have been motivated to include using the OOD data to train the model as taught by Zhang in the system of Yu in order to improve the ability to distinguish between ID data and OOD data (Zhang, page 2, 1 Introduction).
As per claim 2, Yu and Zhang disclose the method of claim 1, wherein the generating of the jigsaw image includes:
dividing the original image to generate a plurality of fragment images (Yu, page 2, Figure 2. Illustration of our proposed out-of-distribution detection framework ... We use NormRatio of ID and pseudo OOD (i.e., Jigsaw puzzle images); Yu, page 4, 3.3. NormRatio: measure of block’s suitability, we generate the 3x Jigsaw puzzle image, which is semantically shifted; Zhang, page 2, 1 Introduction, Fig. 1 depicts the overall procedure of the fine-tuning process. Our SFT generates pseudo-OOD data using … Jigsaw Puzzle Patch Shuffling); and
changing positions of the plurality of fragment images to generate the jigsaw image (Yu, page 2, Figure 2. Illustration of our proposed out-of-distribution detection framework ... We use NormRatio of ID and pseudo OOD (i.e., Jigsaw puzzle images); Yu, page 4, 3.3. NormRatio: measure of block’s suitability, we generate the 3x Jigsaw puzzle image, which is semantically shifted; Zhang, page 2, 1 Introduction, Fig. 1 depicts the overall procedure of the fine-tuning process. Our SFT generates pseudo-OOD data using … Jigsaw Puzzle Patch Shuffling).
As per claim 3, Yu and Zhang disclose the method of claim 1, wherein the training of the deep learning model includes:
specifying ground-truth ID object data, which is label data for the original image, as an ID object (Yu, page 2, 2. Preliminaries, for the given training dataset Din ... where xi ... is the input RGB image and yi ... is the corresponding label with K class categories ... Out-of-distribution detection ... the network classify known images correctly and detect the OOD image as “unknown”. For an OOD detection problem with image classification networks, the given test image x is considered as an OOD image when x semantically ... or non-semantically ... differs from the images of the Din. The decision of the OOD detection is a binary classification with a scoring function that produce ID-ness for the given image);
specifying the jigsaw image as a proxy OOD (Yu, page 3, 3.1. Overview of OOD detection framework, we first generate the Jigsaw puzzle image as pseudo OOD from ID training samples and calculate NormRatio of training samples and corresponding pseudo OOD); and
training the deep learning model using the original image, the ground-truth ID object data, and the jigsaw image (Zhang, page 2, 1 Introduction, We propose our Sharpness-aware Fine-Tuning (SFT) method, in which the pretrained model is fine-tuned by SAM using pseudo-OOD data within 1 epoch (50 steps more precisely). Fig. 1 depicts the overall procedure of the fine-tuning process. Our SFT generates pseudo-OOD data using both Jigsaw Puzzle Patch Shuffling and RGB Channel Shuffling. As can be seen in Fig. 2, ID data and OOD data are better distinguished; Zhang, page 4, 3 Methodology, Preliminary: OOD detection task. The OOD detection problem usually is defined as a binary classification problem between ID and OOD data ... given the data distributions Din and Dout, we fine-tune the pre-trained model ... To craft Dout, we use both Jigsaw Puzzle Patch Shuffling and RGB Channel Shuffling methods to generate pseudo-OOD data. More specifically, Jigsaw Puzzle Patch Shuffling divides the original ID image into small patches and then reassembles these patches randomly to generate new images ... At each fine-tuning iteration, we use both Jigsaw Puzzle Patch Shuffling and RGB Channel Shuffling to generate pseudo-OOD data).
As per claim 4, Yu and Zhang disclose the method of claim 1, wherein the deep learning model is trained to:
identify whether the OOD object is recognized from the input image (Yu, page 2, 2. Preliminaries, for the given training dataset Din ... where xi ... is the input RGB image and yi ... is the corresponding label with K class categories ... Out-of-distribution detection ... the network classify known images correctly and detect the OOD image as “unknown”); and
recognize an ID object from the input image based on the identification result (Yu, page 2, 2. Preliminaries, for the given training dataset Din ... where xi ... is the input RGB image and yi ... is the corresponding label with K class categories ... Out-of-distribution detection ... the network classify known images correctly and detect the OOD image as “unknown”. For an OOD detection problem with image classification networks, the given test image x is considered as an OOD image when x semantically ... or non-semantically ... differs from the images of the Din. The decision of the OOD detection is a binary classification with a scoring function that produce ID-ness for the given image).
As per claim 5, Yu and Zhang disclose the method of claim 4, wherein the deep learning model is trained to:
output an recognition result for the OOD object as an output corresponding to the input image when it is identified that the OOD object is recognized from the input image (Yu, page 2, 2. Preliminaries, for the given training dataset Din ... where xi ... is the input RGB image and yi ... is the corresponding label with K class categories ... Out-of-distribution detection ... the network classify known images correctly and detect the OOD image as “unknown”); and
recognize an ID object from the input image and output the recognized ID object when it is identified that the recognition of the OOD object from the input image has failed (Yu, page 2, 2. Preliminaries, for the given training dataset Din ... where xi ... is the input RGB image and yi ... is the corresponding label with K class categories ... Out-of-distribution detection ... the network classify known images correctly and detect the OOD image as “unknown”. For an OOD detection problem with image classification networks, the given test image x is considered as an OOD image when x semantically ... or non-semantically ... differs from the images of the Din. The decision of the OOD detection is a binary classification with a scoring function that produce ID-ness for the given image; Zhang, page 4, 3 Methodology, Preliminary: OOD detection task. The OOD detection problem usually is defined as a binary classification problem between ID and OOD data. For a given model f which is trained on ID data Did, OOD detection aim to design the score function G: [Equation 1] where x denotes the sample to be tested, S(.) is the seeking score function. γ is the chosen threshold to distinguish whether the given sample is ID or OOD ... Sharpness-aware Fine-Tuning (SFT) ... we fine-tune the pre-trained model using SAM to push the ID and OOD score distributions away; Zhang, page 2, Figure 1, SAM incorporates two additional terms which can reflect the upper bound of the change of Energy score: the gradient norm ∥∇f(wt)∥2 and the maximum output logit max(gi) ... the gradient norm distributions of ID and OOD data ∥∇f(wt)∥2 remain similar before and after fine-tuning, while the maximum output logit distribution of OOD data max(gi) decreases more than that of ID data, making the OOD score distribution flat and move more to the left. As demonstrated by the example at the bottom, the upper bound change finally leads to better score distributions for OOD and eventually helps to better distinguish the ID and OOD samples).
As per claim 6, Yu discloses a system for detecting an OOD object, comprising:
an input unit configured to receive an input image (Yu, page 2, Figure 2. Illustration of our proposed out-of-distribution detection framework ... During inference time, for a given input image, the OOD score is calculated); and
a detection unit configured to detect an OOD object from the input image using a pre-trained deep learning model (Yu, page 2, Figure 2. Illustration of our proposed out-of-distribution detection framework ... We use NormRatio of ID and pseudo OOD (i.e., Jigsaw puzzle images) to find which block is suitable for OOD detection (a). During inference time, for a given input image, the OOD score is calculated by FeatureNorm on the selected block (b). If FeatureNorm for a given input is smaller than the threshold, the given input is classified as OOD; Yu, page 3, 3.1. Overview of OOD detection framework, After the training is done, we select the block by NormRatio for OOD detection (Figure 2; left). Then, we use the norm of the feature map FeatureNorm obtained from the selected feature map for OOD detection during the inference stage (Figure 2; right)),
wherein the pre-trained deep learning model is trained according to a method of training a deep learning model (Yu, page 4, 4. Experiments, Setup We use commonly utilized CNN architectures: ResNet18, VGG11 and WideResNet with depth 28 and width 10 (WRN28) for the CIFAR10 benchmark. The ResNet18 and VGG11 are trained with batch size 128 for 100 epochs ... The WRN28 is trained with batch size 128 for 200 epochs; Yu, page 4, Table 1. Summary of the selected blocks for each architecture), the method of training a deep learning model including:
receiving an original image (Yu, page 2, Figure 2. Illustration of our proposed out-of-distribution detection framework ... During inference time, for a given input image, the OOD score is calculated);
transforming unique features represented from the original image to generate a jigsaw image (Yu, page 3, 3.1. Overview of OOD detection framework, we first generate the Jigsaw puzzle image as pseudo OOD from ID training samples and calculate NormRatio of training samples and corresponding pseudo OOD; Yu, page 4, 3.3. NormRatio: measure of block’s suitability, we generate the 3x Jigsaw puzzle image, which is semantically shifted ... using training samples);
specifying the jigsaw image as a proxy OOD (Yu, Abstract, we create Jigsaw puzzle images as pseudo OOD from ID training samples); and
training the deep learning model using the original image to recognize the OOD object from the input image (Yu, page 3, 3.1. Overview of OOD detection framework, we first generate the Jigsaw puzzle image as pseudo OOD from ID training samples … After the training is done, we select the block by NormRatio for OOD detection (Figure 2; left). Then, we use the norm of the feature map FeatureNorm obtained from the selected feature map for OOD detection during the inference stage (Figure 2; right). Specifically, we first generate the Jigsaw puzzle image as pseudo OOD from ID training samples and calculate NormRatio of training samples and corresponding pseudo OOD).
Yu does not explicitly disclose the following limitations as further recited however Zhang discloses
training the deep learning model using the original image and the jigsaw image to recognize the OOD object from the input image (Zhang, page 2, 1 Introduction, We propose our Sharpness-aware Fine-Tuning (SFT) method, in which the pretrained model is fine-tuned by SAM using pseudo-OOD data within 1 epoch (50 steps more precisely). Fig. 1 depicts the overall procedure of the fine-tuning process. Our SFT generates pseudo-OOD data using both Jigsaw Puzzle Patch Shuffling and RGB Channel Shuffling. As can be seen in Fig. 2, ID data and OOD data are better distinguished; Zhang, page 4, 3 Methodology, Preliminary: OOD detection task. The OOD detection problem usually is defined as a binary classification problem between ID and OOD data ... given the data distributions Din and Dout, we fine-tune the pre-trained model ... To craft Dout, we use both Jigsaw Puzzle Patch Shuffling and RGB Channel Shuffling methods to generate pseudo-OOD data. More specifically, Jigsaw Puzzle Patch Shuffling divides the original ID image into small patches and then reassembles these patches randomly to generate new images ... At each fine-tuning iteration, we use both Jigsaw Puzzle Patch Shuffling and RGB Channel Shuffling to generate pseudo-OOD data).
It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to combine the teachings of Zhang with Yu because they are in the same field of endeavor. One skilled in the art would have been motivated to include using the OOD data to train the model as taught by Zhang in the system of Yu in order to improve the ability to distinguish between ID data and OOD data (Zhang, page 2, 1 Introduction).
As per claim 7, Yu discloses a program stored on a computer-readable recording medium, and executed by one or more processes in an electronic device, the program comprising instructions to allow the program to perform:
receiving an input image (Yu, page 2, Figure 2. Illustration of our proposed out-of-distribution detection framework ... During inference time, for a given input image, the OOD score is calculated); and
recognizing an OOD object from the input image using a pre-trained deep learning model (Yu, page 2, Figure 2. Illustration of our proposed out-of-distribution detection framework ... We use NormRatio of ID and pseudo OOD (i.e., Jigsaw puzzle images) to find which block is suitable for OOD detection (a). During inference time, for a given input image, the OOD score is calculated by FeatureNorm on the selected block (b). If FeatureNorm for a given input is smaller than the threshold, the given input is classified as OOD; Yu, page 3, 3.1. Overview of OOD detection framework, After the training is done, we select the block by NormRatio for OOD detection (Figure 2; left). Then, we use the norm of the feature map FeatureNorm obtained from the selected feature map for OOD detection during the inference stage (Figure 2; right)),
wherein the pre-trained deep learning model is trained according to a method of training a deep learning model (Yu, page 4, 4. Experiments, Setup We use commonly utilized CNN architectures: ResNet18, VGG11 and WideResNet with depth 28 and width 10 (WRN28) for the CIFAR10 benchmark. The ResNet18 and VGG11 are trained with batch size 128 for 100 epochs ... The WRN28 is trained with batch size 128 for 200 epochs; Yu, page 4, Table 1. Summary of the selected blocks for each architecture), the method of training a deep learning model including:
receiving an original image (Yu, page 2, Figure 2. Illustration of our proposed out-of-distribution detection framework ... During inference time, for a given input image, the OOD score is calculated);
transforming unique features represented from the original image to generate a jigsaw image (Yu, page 3, 3.1. Overview of OOD detection framework, we first generate the Jigsaw puzzle image as pseudo OOD from ID training samples and calculate NormRatio of training samples and corresponding pseudo OOD; Yu, page 4, 3.3. NormRatio: measure of block’s suitability, we generate the 3x Jigsaw puzzle image, which is semantically shifted ... using training samples);
specifying the jigsaw image as a proxy OOD (Yu, Abstract, we create Jigsaw puzzle images as pseudo OOD from ID training samples); and
training the deep learning model using the original image to recognize the OOD object from the input image (Yu, page 3, 3.1. Overview of OOD detection framework, we first generate the Jigsaw puzzle image as pseudo OOD from ID training samples … After the training is done, we select the block by NormRatio for OOD detection (Figure 2; left). Then, we use the norm of the feature map FeatureNorm obtained from the selected feature map for OOD detection during the inference stage (Figure 2; right). Specifically, we first generate the Jigsaw puzzle image as pseudo OOD from ID training samples and calculate NormRatio of training samples and corresponding pseudo OOD).
Yu does not explicitly disclose the following limitations as further recited however Zhang discloses
training the deep learning model using the original image and the jigsaw image to recognize the OOD object from the input image (Zhang, page 2, 1 Introduction, We propose our Sharpness-aware Fine-Tuning (SFT) method, in which the pretrained model is fine-tuned by SAM using pseudo-OOD data within 1 epoch (50 steps more precisely). Fig. 1 depicts the overall procedure of the fine-tuning process. Our SFT generates pseudo-OOD data using both Jigsaw Puzzle Patch Shuffling and RGB Channel Shuffling. As can be seen in Fig. 2, ID data and OOD data are better distinguished; Zhang, page 4, 3 Methodology, Preliminary: OOD detection task. The OOD detection problem usually is defined as a binary classification problem between ID and OOD data ... given the data distributions Din and Dout, we fine-tune the pre-trained model ... To craft Dout, we use both Jigsaw Puzzle Patch Shuffling and RGB Channel Shuffling methods to generate pseudo-OOD data. More specifically, Jigsaw Puzzle Patch Shuffling divides the original ID image into small patches and then reassembles these patches randomly to generate new images ... At each fine-tuning iteration, we use both Jigsaw Puzzle Patch Shuffling and RGB Channel Shuffling to generate pseudo-OOD data).
It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to combine the teachings of Zhang with Yu because they are in the same field of endeavor. One skilled in the art would have been motivated to include using the OOD data to train the model as taught by Zhang in the system of Yu in order to improve the ability to distinguish between ID data and OOD data (Zhang, page 2, 1 Introduction).
As per claim 8, Yu discloses a method of training a deep learning model using a system for training a deep learning model, comprising:
receiving an original image (Yu, page 2, Figure 2. Illustration of our proposed out-of-distribution detection framework ... During inference time, for a given input image, the OOD score is calculated);
transforming unique features represented from the original image to generate a jigsaw image (Yu, page 3, 3.1. Overview of OOD detection framework, we first generate the Jigsaw puzzle image as pseudo OOD from ID training samples and calculate NormRatio of training samples and corresponding pseudo OOD; Yu, page 4, 3.3. NormRatio: measure of block’s suitability, we generate the 3x Jigsaw puzzle image, which is semantically shifted ... using training samples);
specifying the jigsaw image as a proxy OOD (Yu, Abstract, we create Jigsaw puzzle images as pseudo OOD from ID training samples); and
training the deep learning model using the original image to recognize the OOD object from the input image (Yu, page 3, 3.1. Overview of OOD detection framework, we first generate the Jigsaw puzzle image as pseudo OOD from ID training samples … After the training is done, we select the block by NormRatio for OOD detection (Figure 2; left). Then, we use the norm of the feature map FeatureNorm obtained from the selected feature map for OOD detection during the inference stage (Figure 2; right). Specifically, we first generate the Jigsaw puzzle image as pseudo OOD from ID training samples and calculate NormRatio of training samples and corresponding pseudo OOD).
Yu does not explicitly disclose the following limitations as further recited however Zhang discloses
training the deep learning model using the original image and the jigsaw image to recognize the OOD object from the input image (Zhang, page 2, 1 Introduction, We propose our Sharpness-aware Fine-Tuning (SFT) method, in which the pretrained model is fine-tuned by SAM using pseudo-OOD data within 1 epoch (50 steps more precisely). Fig. 1 depicts the overall procedure of the fine-tuning process. Our SFT generates pseudo-OOD data using both Jigsaw Puzzle Patch Shuffling and RGB Channel Shuffling. As can be seen in Fig. 2, ID data and OOD data are better distinguished; Zhang, page 4, 3 Methodology, Preliminary: OOD detection task. The OOD detection problem usually is defined as a binary classification problem between ID and OOD data ... given the data distributions Din and Dout, we fine-tune the pre-trained model ... To craft Dout, we use both Jigsaw Puzzle Patch Shuffling and RGB Channel Shuffling methods to generate pseudo-OOD data. More specifically, Jigsaw Puzzle Patch Shuffling divides the original ID image into small patches and then reassembles these patches randomly to generate new images ... At each fine-tuning iteration, we use both Jigsaw Puzzle Patch Shuffling and RGB Channel Shuffling to generate pseudo-OOD data).
It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to combine the teachings of Zhang with Yu because they are in the same field of endeavor. One skilled in the art would have been motivated to include using the OOD data to train the model as taught by Zhang in the system of Yu in order to improve the ability to distinguish between ID data and OOD data (Zhang, page 2, 1 Introduction).
As per claim 9, Yu discloses a system for training a deep learning model, comprising:
an input unit configured to receive an original image (Yu, page 2, Figure 2. Illustration of our proposed out-of-distribution detection framework ... During inference time, for a given input image, the OOD score is calculated);
a jigsaw generation unit configured to transform unique features represented from the original image to generate a jigsaw image (Yu, page 3, 3.1. Overview of OOD detection framework, we first generate the Jigsaw puzzle image as pseudo OOD from ID training samples and calculate NormRatio of training samples and corresponding pseudo OOD; Yu, page 4, 3.3. NormRatio: measure of block’s suitability, we generate the 3x Jigsaw puzzle image, which is semantically shifted ... using training samples); and
a training unit configured to specify the jigsaw image as a proxy OOD (Yu, Abstract, we create Jigsaw puzzle images as pseudo OOD from ID training samples) and
train the deep learning model using the original image to recognize the OOD object from the input image (Yu, page 3, 3.1. Overview of OOD detection framework, we first generate the Jigsaw puzzle image as pseudo OOD from ID training samples … After the training is done, we select the block by NormRatio for OOD detection (Figure 2; left). Then, we use the norm of the feature map FeatureNorm obtained from the selected feature map for OOD detection during the inference stage (Figure 2; right). Specifically, we first generate the Jigsaw puzzle image as pseudo OOD from ID training samples and calculate NormRatio of training samples and corresponding pseudo OOD).
Yu does not explicitly disclose the following limitations as further recited however Zhang discloses
train the deep learning model using the original image and the jigsaw image to recognize the OOD object from the input image (Zhang, page 2, 1 Introduction, We propose our Sharpness-aware Fine-Tuning (SFT) method, in which the pretrained model is fine-tuned by SAM using pseudo-OOD data within 1 epoch (50 steps more precisely). Fig. 1 depicts the overall procedure of the fine-tuning process. Our SFT generates pseudo-OOD data using both Jigsaw Puzzle Patch Shuffling and RGB Channel Shuffling. As can be seen in Fig. 2, ID data and OOD data are better distinguished; Zhang, page 4, 3 Methodology, Preliminary: OOD detection task. The OOD detection problem usually is defined as a binary classification problem between ID and OOD data ... given the data distributions Din and Dout, we fine-tune the pre-trained model ... To craft Dout, we use both Jigsaw Puzzle Patch Shuffling and RGB Channel Shuffling methods to generate pseudo-OOD data. More specifically, Jigsaw Puzzle Patch Shuffling divides the original ID image into small patches and then reassembles these patches randomly to generate new images ... At each fine-tuning iteration, we use both Jigsaw Puzzle Patch Shuffling and RGB Channel Shuffling to generate pseudo-OOD data).
It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to combine the teachings of Zhang with Yu because they are in the same field of endeavor. One skilled in the art would have been motivated to include using the OOD data to train the model as taught by Zhang in the system of Yu in order to improve the ability to distinguish between ID data and OOD data (Zhang, page 2, 1 Introduction).
As per claim 10, Yu discloses a program stored on a computer-readable recording medium, and executed by one or more processes in an electronic device, the program comprising instructions to allow the program to perform:
receiving an original image (Yu, page 2, Figure 2. Illustration of our proposed out-of-distribution detection framework ... During inference time, for a given input image, the OOD score is calculated);
transforming unique features represented from the original image to generate a jigsaw image (Yu, page 3, 3.1. Overview of OOD detection framework, we first generate the Jigsaw puzzle image as pseudo OOD from ID training samples and calculate NormRatio of training samples and corresponding pseudo OOD; Yu, page 4, 3.3. NormRatio: measure of block’s suitability, we generate the 3x Jigsaw puzzle image, which is semantically shifted ... using training samples);
specifying the jigsaw image as a proxy OOD (Yu, Abstract, we create Jigsaw puzzle images as pseudo OOD from ID training samples); and
training the deep learning model using the original image to recognize the OOD object from the input image (Yu, page 3, 3.1. Overview of OOD detection framework, we first generate the Jigsaw puzzle image as pseudo OOD from ID training samples … After the training is done, we select the block by NormRatio for OOD detection (Figure 2; left). Then, we use the norm of the feature map FeatureNorm obtained from the selected feature map for OOD detection during the inference stage (Figure 2; right). Specifically, we first generate the Jigsaw puzzle image as pseudo OOD from ID training samples and calculate NormRatio of training samples and corresponding pseudo OOD).
Yu does not explicitly disclose the following limitations as further recited however Zhang discloses
training the deep learning model using the original image and the jigsaw image to recognize the OOD object from the input image (Zhang, page 2, 1 Introduction, We propose our Sharpness-aware Fine-Tuning (SFT) method, in which the pretrained model is fine-tuned by SAM using pseudo-OOD data within 1 epoch (50 steps more precisely). Fig. 1 depicts the overall procedure of the fine-tuning process. Our SFT generates pseudo-OOD data using both Jigsaw Puzzle Patch Shuffling and RGB Channel Shuffling. As can be seen in Fig. 2, ID data and OOD data are better distinguished; Zhang, page 4, 3 Methodology, Preliminary: OOD detection task. The OOD detection problem usually is defined as a binary classification problem between ID and OOD data ... given the data distributions Din and Dout, we fine-tune the pre-trained model ... To craft Dout, we use both Jigsaw Puzzle Patch Shuffling and RGB Channel Shuffling methods to generate pseudo-OOD data. More specifically, Jigsaw Puzzle Patch Shuffling divides the original ID image into small patches and then reassembles these patches randomly to generate new images ... At each fine-tuning iteration, we use both Jigsaw Puzzle Patch Shuffling and RGB Channel Shuffling to generate pseudo-OOD data).
It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to combine the teachings of Zhang with Yu because they are in the same field of endeavor. One skilled in the art would have been motivated to include using the OOD data to train the model as taught by Zhang in the system of Yu in order to improve the ability to distinguish between ID data and OOD data (Zhang, page 2, 1 Introduction).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Chen, Jiankang, et al. "Tagfog: Textual anchor guidance and fake outlier generation for visual out-of-distribution detection." Proceedings of the AAAI Conference on Artificial Intelligence. 2024 Mar 24 (Vol. 38, No. 2, pp. 1100-1109), discloses a system and method that uses an original image and a jigsaw image derived from the original image as a pseudo out-of-distribution image to train a deep learning model.
Han et al., Korean Patent No. KR 102170620B1 discloses a system and method that cuts and mixes patches from original images to create mixed images to train a deep neural network to improve out-of-distribution recognition.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to TRACY MANGIALASCHI whose telephone number is (571)270-5189. The examiner can normally be reached M-F, 9:30AM TO 6:00PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vu Le can be reached at (571) 272-7332. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/TRACY MANGIALASCHI/Primary Examiner, Art Unit 2668