DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application is being examined under the pre-AIA first to invent provisions.
Claim Rejections - 35 USC § 112b
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 10 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 10 recites the limitation "the action" in line 1. There is insufficient antecedent basis for this limitation in the claim.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
In reference to claim 1:
Step 1 - Is the claim to a process, machine, manufacture or composition of matter?
Yes, the claim is directed to a process
Step 2A Prong 1 - Does the claim recite an abstract idea, law of nature, or natural phenomenon?
“A computer-implemented method for fine-tuning a model, comprising: generating a label space for a target domain;” which is an abstract idea because it is directed to a mental process, an observation, evaluation, judgement, or opinion. The limitation as drafted, and under a broadest reasonable interpretation, can be performed in the human mind, or by a human using a pen and paper (MPEP 2106.04(a)(2)(Ill)(c)). For example, a person could generate a label space for a target domain.
“generating text pseudo-labels for images in an unlabeled dataset from the target domain based on the label space [using a pre-trained vision language model; and]” which is an abstract idea because it is directed to a mental process, an observation, evaluation, judgement, or opinion. The limitation as drafted, and under a broadest reasonable interpretation, can be performed in the human mind, or by a human using a pen and paper (MPEP 2106.04(a)(2)(Ill)(c)). For example, a person could generate text pseudo-labels for images in an unlabeled dataset from the target domain based on the label space.
Step 2A Prong 2 - Does the claim recite additional elements that integrate the judicial exception into a practical application?
“using a pre-trained vision language model; and” is merely reciting the words "apply it" (or an equivalent) with the judicial exception, or merely including instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
“fine-tuning the pre-trained vision language model for the target domain using the images with the text pseudo-labels.” is merely reciting the words "apply it" (or an equivalent) with the judicial exception, or merely including instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
The claim does not include additional elements that are integrated into a practical application.
Step 2B - Does the claim recite additional elements that amount to significantly more than the judicial exception?
“using a pre-trained vision language model; and” is merely reciting the words "apply it" (or an equivalent) with the judicial exception, or merely including instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
“fine-tuning the pre-trained vision language model for the target domain using the images with the text pseudo-labels.” is merely reciting the words "apply it" (or an equivalent) with the judicial exception, or merely including instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
In reference to claim 2:
Step 1 - Is the claim to a process, machine, manufacture or composition of matter?
Yes, the claim is directed to a process
Step 2A Prong 1 - Does the claim recite an abstract idea, law of nature, or natural phenomenon?
“The method of claim 1, wherein generating the label space includes pruning a list of tokens to remove tokens that represent punctuation, special symbols, and plural forms.” which is an abstract idea because it is directed to a mental process, an observation, evaluation, judgement, or opinion. The limitation as drafted, and under a broadest reasonable interpretation, can be performed in the human mind, or by a human using a pen and paper (MPEP 2106.04(a)(2)(Ill)(c)). For example, a person could prune/not consider tokens that represent punctuation, special symbols, and plural forms.
Step 2A Prong 2 - Does the claim recite additional elements that integrate the judicial exception into a practical application?
No
The claim does not include additional elements that are integrated into a practical application.
Step 2B - Does the claim recite additional elements that amount to significantly more than the judicial exception?
No
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
In reference to claim 3:
Step 1 - Is the claim to a process, machine, manufacture or composition of matter?
Yes, the claim is directed to a process
Step 2A Prong 1 - Does the claim recite an abstract idea, law of nature, or natural phenomenon?
“The method of claim 1, wherein generating the label space includes determining similarity values between tokens in a list of tokens” which is an abstract idea because it is directed to a mental process, an observation, evaluation, judgement, or opinion. The limitation as drafted, and under a broadest reasonable interpretation, can be performed in the human mind, or by a human using a pen and paper (MPEP 2106.04(a)(2)(Ill)(c)). For example, a person could determine a similarity value between tokens in a list of tokens.
“and generating a bag of words representation based on tokens that have an above-threshold similarity value.” which is an abstract idea because it is directed to a mental process, an observation, evaluation, judgement, or opinion. The limitation as drafted, and under a broadest reasonable interpretation, can be performed in the human mind, or by a human using a pen and paper (MPEP 2106.04(a)(2)(Ill)(c)). For example, a person could generate a bag of words representation based on tokens that have an above-threshold similarity value.
Step 2A Prong 2 - Does the claim recite additional elements that integrate the judicial exception into a practical application?
No
The claim does not include additional elements that are integrated into a practical application.
Step 2B - Does the claim recite additional elements that amount to significantly more than the judicial exception?
No
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
In reference to claim 4:
Step 1 - Is the claim to a process, machine, manufacture or composition of matter?
Yes, the claim is directed to a process
Step 2A Prong 1 - Does the claim recite an abstract idea, law of nature, or natural phenomenon?
“The method of claim 3, wherein generating the label space further includes combining bag of words representations for all images in the unlabeled dataset to generate a bag of words representation for the entire unlabeled dataset.” which is an abstract idea because it is directed to a mental process, an observation, evaluation, judgement, or opinion. The limitation as drafted, and under a broadest reasonable interpretation, can be performed in the human mind, or by a human using a pen and paper (MPEP 2106.04(a)(2)(Ill)(c)). For example, a person could combine bag of words representations to generate a bag of words representation for the entire unlabeled dataset.
Step 2A Prong 2 - Does the claim recite additional elements that integrate the judicial exception into a practical application?
No
The claim does not include additional elements that are integrated into a practical application.
Step 2B - Does the claim recite additional elements that amount to significantly more than the judicial exception?
No
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
In reference to claim 5:
Step 1 - Is the claim to a process, machine, manufacture or composition of matter?
Yes, the claim is directed to a process
Step 2A Prong 1 - Does the claim recite an abstract idea, law of nature, or natural phenomenon?
“The method of claim 4, wherein generating the label space further includes pruning words from the bag of words representation for the entire unlabeled dataset based on a frequency of occurrence.” which is an abstract idea because it is directed to a mental process, an observation, evaluation, judgement, or opinion. The limitation as drafted, and under a broadest reasonable interpretation, can be performed in the human mind, or by a human using a pen and paper (MPEP 2106.04(a)(2)(Ill)(c)). For example, a person could prune/not consider tokens based on a frequency of occurrence.
Step 2A Prong 2 - Does the claim recite additional elements that integrate the judicial exception into a practical application?
No
The claim does not include additional elements that are integrated into a practical application.
Step 2B - Does the claim recite additional elements that amount to significantly more than the judicial exception?
No
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
In reference to claim 6:
Step 1 - Is the claim to a process, machine, manufacture or composition of matter?
Yes, the claim is directed to a process
Step 2A Prong 1 - Does the claim recite an abstract idea, law of nature, or natural phenomenon?
“The method of claim 4, wherein generating the label space further includes pruning words from the bag of words representation for the entire unlabeled dataset based on an above-threshold degree of similarity to other words from the bag of words representation.” which is an abstract idea because it is directed to a mental process, an observation, evaluation, judgement, or opinion. The limitation as drafted, and under a broadest reasonable interpretation, can be performed in the human mind, or by a human using a pen and paper (MPEP 2106.04(a)(2)(Ill)(c)). For example, a person could prune/not consider words from bag of words based on an above threshold degree of similarity to other words from the bag of words representation.
Step 2A Prong 2 - Does the claim recite additional elements that integrate the judicial exception into a practical application?
No
The claim does not include additional elements that are integrated into a practical application.
Step 2B - Does the claim recite additional elements that amount to significantly more than the judicial exception?
No
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
In reference to claim 7:
Claim 7 is directed to a judicial exception from claim(s) depended on and does not recite additional elements that integrate the judicial exception into a practical application and amount to significantly more than the judicial exception.
In reference to claim 8:
Claim 8 is directed to a judicial exception from claim(s) depended on and does not recite additional elements that integrate the judicial exception into a practical application and amount to significantly more than the judicial exception.
In reference to claim 9:
Step 1 - Is the claim to a process, machine, manufacture or composition of matter?
Yes, the claim is directed to a process
Step 2A Prong 1 - Does the claim recite an abstract idea, law of nature, or natural phenomenon?
“The method of claim 1, further comprising processing a new image in the target domain [using the fine-tuned vision language model] to generate a label for the new image and performing an action responsive to the label.” which is an abstract idea because it is directed to a mental process, an observation, evaluation, judgement, or opinion. The limitation as drafted, and under a broadest reasonable interpretation, can be performed in the human mind, or by a human using a pen and paper (MPEP 2106.04(a)(2)(Ill)(c)). For example, a person could process a new image in the target domain to generate a label for the new image and perform an action responsive to the label.
Step 2A Prong 2 - Does the claim recite additional elements that integrate the judicial exception into a practical application?
“using the fine-tuned vision language model” is merely reciting the words "apply it" (or an equivalent) with the judicial exception, or merely including instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
The claim does not include additional elements that are integrated into a practical application.
Step 2B - Does the claim recite additional elements that amount to significantly more than the judicial exception?
“using the fine-tuned vision language model” is merely reciting the words "apply it" (or an equivalent) with the judicial exception, or merely including instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
In reference to claim 10:
Claim 10 is directed to a judicial exception from claim(s) depended on and does not recite additional elements that integrate the judicial exception into a practical application and amount to significantly more than the judicial exception.
In reference to claim 11:
Step 1 - Is the claim to a process, machine, manufacture or composition of matter?
Yes, the claim is directed to a process
Step 2A Prong 1 - Does the claim recite an abstract idea, law of nature, or natural phenomenon?
“A computer-implemented method for fine-tuning a model, comprising: generating a label space for a target domain, including:” which is an abstract idea because it is directed to a mental process, an observation, evaluation, judgement, or opinion. The limitation as drafted, and under a broadest reasonable interpretation, can be performed in the human mind, or by a human using a pen and paper (MPEP 2106.04(a)(2)(Ill)(c)). For example, a person could generate a label space for a target domain.
“determining similarity values between tokens in a list of tokens and generating a bag of words representation based on tokens that have an above-threshold similarity value;” which is an abstract idea because it is directed to a mental process, an observation, evaluation, judgement, or opinion. The limitation as drafted, and under a broadest reasonable interpretation, can be performed in the human mind, or by a human using a pen and paper (MPEP 2106.04(a)(2)(Ill)(c)). For example, a person could determine similarity values between tokens in a list of tokens and generate a bag of words representations based on the tokens that have an above threshold similarity value.
“combining bag of words representations for all images in an unlabeled dataset to generate a bag of words representation for the entire unlabeled dataset;” which is an abstract idea because it is directed to a mental process, an observation, evaluation, judgement, or opinion. The limitation as drafted, and under a broadest reasonable interpretation, can be performed in the human mind, or by a human using a pen and paper (MPEP 2106.04(a)(2)(Ill)(c)). For example, a person could combine bag of word representations to generate a bag of words representation for the entire unlabeled dataset.
“pruning words from the bag of words representation for the entire unlabeled dataset based on a frequency of occurrence; and” which is an abstract idea because it is directed to a mental process, an observation, evaluation, judgement, or opinion. The limitation as drafted, and under a broadest reasonable interpretation, can be performed in the human mind, or by a human using a pen and paper (MPEP 2106.04(a)(2)(Ill)(c)). For example, a person could prune/not consider words from the bag of words based on a frequency of occurrence.
“pruning words from the bag of words representation for the entire unlabeled dataset based on an above-threshold degree of similarity to other words from the bag of words representation;” which is an abstract idea because it is directed to a mental process, an observation, evaluation, judgement, or opinion. The limitation as drafted, and under a broadest reasonable interpretation, can be performed in the human mind, or by a human using a pen and paper (MPEP 2106.04(a)(2)(Ill)(c)). For example, a person could prune/not consider words from the bag of words based on an above threshold degree of similarity to other words from the bag of words.
“generating text pseudo-labels for images in an unlabeled dataset from the target domain based on the label space [using a pre-trained vision language model; and]” which is an abstract idea because it is directed to a mental process, an observation, evaluation, judgement, or opinion. The limitation as drafted, and under a broadest reasonable interpretation, can be performed in the human mind, or by a human using a pen and paper (MPEP 2106.04(a)(2)(Ill)(c)). For example, a person could generate text pseudo-labels for images in an unlabeled dataset from the target domain based on the label space.
Step 2A Prong 2 - Does the claim recite additional elements that integrate the judicial exception into a practical application?
“using a pre-trained vision language model; and” is merely reciting the words "apply it" (or an equivalent) with the judicial exception, or merely including instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
“fine-tuning the pre-trained vision language model for the target domain using the images with the text pseudo-labels.” is merely reciting the words "apply it" (or an equivalent) with the judicial exception, or merely including instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
The claim does not include additional elements that are integrated into a practical application.
Step 2B - Does the claim recite additional elements that amount to significantly more than the judicial exception?
“using a pre-trained vision language model; and” is merely reciting the words "apply it" (or an equivalent) with the judicial exception, or merely including instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
“fine-tuning the pre-trained vision language model for the target domain using the images with the text pseudo-labels.” is merely reciting the words "apply it" (or an equivalent) with the judicial exception, or merely including instructions to implement an abstract idea on a computer, or merely using The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
In reference to claim 12:
Step 1 - Is the claim to a process, machine, manufacture or composition of matter?
Yes, the claim is directed to a manufacture
Step 2A Prong 1 - Does the claim recite an abstract idea, law of nature, or natural phenomenon?
“generate a label space for a target domain;” which is an abstract idea because it is directed to a mental process, an observation, evaluation, judgement, or opinion. The limitation as drafted, and under a broadest reasonable interpretation, can be performed in the human mind, or by a human using a pen and paper (MPEP 2106.04(a)(2)(Ill)(c)). For example, a person could generate a label space for a target domain.
“generate text pseudo-labels for images in an unlabeled dataset from the target domain based on the label space [using a pre-trained vision language model; and]” which is an abstract idea because it is directed to a mental process, an observation, evaluation, judgement, or opinion. The limitation as drafted, and under a broadest reasonable interpretation, can be performed in the human mind, or by a human using a pen and paper (MPEP 2106.04(a)(2)(Ill)(c)). For example, a person could generate text pseudo-labels for images in an unlabeled dataset from the target domain based on the label space.
Step 2A Prong 2 - Does the claim recite additional elements that integrate the judicial exception into a practical application?
“A system for fine-tuning a model, comprising: a hardware processor; and a memory that stores a computer program which, when executed by the hardware processor, causes the hardware processor to:” is merely reciting the words "apply it" (or an equivalent) with the judicial exception, or merely including instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
“using a pre-trained vision language model; and” is merely reciting the words "apply it" (or an equivalent) with the judicial exception, or merely including instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
“fine-tune the pre-trained vision language model for the target domain using the images with the text pseudo-labels.” is merely reciting the words "apply it" (or an equivalent) with the judicial exception, or merely including instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
The claim does not include additional elements that are integrated into a practical application.
Step 2B - Does the claim recite additional elements that amount to significantly more than the judicial exception?
“A system for fine-tuning a model, comprising: a hardware processor; and a memory that stores a computer program which, when executed by the hardware processor, causes the hardware processor to:” is merely reciting the words "apply it" (or an equivalent) with the judicial exception, or merely including instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
“using a pre-trained vision language model; and” is merely reciting the words "apply it" (or an equivalent) with the judicial exception, or merely including instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
“fine-tune the pre-trained vision language model for the target domain using the images with the text pseudo-labels.” is merely reciting the words "apply it" (or an equivalent) with the judicial exception, or merely including instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
In reference to claim 13:
Step 1 - Is the claim to a process, machine, manufacture or composition of matter?
Yes, the claim is directed to a manufacture
Step 2A Prong 1 - Does the claim recite an abstract idea, law of nature, or natural phenomenon?
“The system of claim 12, wherein the computer program further causes the hardware processor to prune a list of tokens to remove tokens that represent punctuation, special symbols, and plural forms.” which is an abstract idea because it is directed to a mental process, an observation, evaluation, judgement, or opinion. The limitation as drafted, and under a broadest reasonable interpretation, can be performed in the human mind, or by a human using a pen and paper (MPEP 2106.04(a)(2)(Ill)(c)). For example, a person could prune/not consider tokens that represent punctuation, special symbols, and plural forms.
Step 2A Prong 2 - Does the claim recite additional elements that integrate the judicial exception into a practical application?
No
The claim does not include additional elements that are integrated into a practical application.
Step 2B - Does the claim recite additional elements that amount to significantly more than the judicial exception?
No
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
In reference to claim 14:
Step 1 - Is the claim to a process, machine, manufacture or composition of matter?
Yes, the claim is directed to a process
Step 2A Prong 1 - Does the claim recite an abstract idea, law of nature, or natural phenomenon?
“The system of claim 12, wherein the computer program further causes the hardware processor to determine similarity values between tokens in a list of tokens” which is an abstract idea because it is directed to a mental process, an observation, evaluation, judgement, or opinion. The limitation as drafted, and under a broadest reasonable interpretation, can be performed in the human mind, or by a human using a pen and paper (MPEP 2106.04(a)(2)(Ill)(c)). For example, a person could determine a similarity value between tokens in a list of tokens.
“and generating a bag of words representation based on tokens that have an above-threshold similarity value.” which is an abstract idea because it is directed to a mental process, an observation, evaluation, judgement, or opinion. The limitation as drafted, and under a broadest reasonable interpretation, can be performed in the human mind, or by a human using a pen and paper (MPEP 2106.04(a)(2)(Ill)(c)). For example, a person could generate a bag of words representation based on tokens that have an above-threshold similarity value.
Step 2A Prong 2 - Does the claim recite additional elements that integrate the judicial exception into a practical application?
No
The claim does not include additional elements that are integrated into a practical application.
Step 2B - Does the claim recite additional elements that amount to significantly more than the judicial exception?
No
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
In reference to claim 15:
Step 1 - Is the claim to a process, machine, manufacture or composition of matter?
Yes, the claim is directed to a manufacture
Step 2A Prong 1 - Does the claim recite an abstract idea, law of nature, or natural phenomenon?
“The system of claim 14, wherein the computer program further causes the hardware processor to combine bag of words representations for all images in the unlabeled dataset to generate a bag of words representation for the entire unlabeled dataset.” which is an abstract idea because it is directed to a mental process, an observation, evaluation, judgement, or opinion. The limitation as drafted, and under a broadest reasonable interpretation, can be performed in the human mind, or by a human using a pen and paper (MPEP 2106.04(a)(2)(Ill)(c)). For example, a person could combine bag of words representations to generate a bag of words representation for the entire unlabeled dataset.
Step 2A Prong 2 - Does the claim recite additional elements that integrate the judicial exception into a practical application?
No
The claim does not include additional elements that are integrated into a practical application.
Step 2B - Does the claim recite additional elements that amount to significantly more than the judicial exception?
No
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
In reference to claim 16:
Step 1 - Is the claim to a process, machine, manufacture or composition of matter?
Yes, the claim is directed to a manufacture
Step 2A Prong 1 - Does the claim recite an abstract idea, law of nature, or natural phenomenon?
“The system of claim 15, wherein the computer program further causes the hardware processor to prune words from the bag of words representation for the entire unlabeled dataset based on a frequency of occurrence.” which is an abstract idea because it is directed to a mental process, an observation, evaluation, judgement, or opinion. The limitation as drafted, and under a broadest reasonable interpretation, can be performed in the human mind, or by a human using a pen and paper (MPEP 2106.04(a)(2)(Ill)(c)). For example, a person could prune/not consider tokens based on a frequency of occurrence.
Step 2A Prong 2 - Does the claim recite additional elements that integrate the judicial exception into a practical application?
No
The claim does not include additional elements that are integrated into a practical application.
Step 2B - Does the claim recite additional elements that amount to significantly more than the judicial exception?
No
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
In reference to claim 17:
Step 1 - Is the claim to a process, machine, manufacture or composition of matter?
Yes, the claim is directed to a manufacture
Step 2A Prong 1 - Does the claim recite an abstract idea, law of nature, or natural phenomenon?
“The system of claim 15, wherein the computer program further causes the hardware processor to prune words from the bag of words representation for the entire unlabeled dataset based on an above-threshold degree of similarity to other words from the bag of words representation.” which is an abstract idea because it is directed to a mental process, an observation, evaluation, judgement, or opinion. The limitation as drafted, and under a broadest reasonable interpretation, can be performed in the human mind, or by a human using a pen and paper (MPEP 2106.04(a)(2)(Ill)(c)). For example, a person could prune/not consider words from bag of words based on an above threshold degree of similarity to other words from the bag of words representation.
Step 2A Prong 2 - Does the claim recite additional elements that integrate the judicial exception into a practical application?
No
The claim does not include additional elements that are integrated into a practical application.
Step 2B - Does the claim recite additional elements that amount to significantly more than the judicial exception?
No
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
In reference to claim 18:
Claim 18 is directed to a judicial exception from claim(s) depended on and does not recite additional elements that integrate the judicial exception into a practical application and amount to significantly more than the judicial exception.
In reference to claim 19:
Claim 19 is directed to a judicial exception from claim(s) depended on and does not recite additional elements that integrate the judicial exception into a practical application and amount to significantly more than the judicial exception.
In reference to claim 20:
Step 1 - Is the claim to a process, machine, manufacture or composition of matter?
Yes, the claim is directed to a manufacture
Step 2A Prong 1 - Does the claim recite an abstract idea, law of nature, or natural phenomenon?
“The system of claim 12, wherein the computer program further causes the hardware processor to process a new image in the target domain [using the fine-tuned vision language model] to generate a label for the new image” which is an abstract idea because it is directed to a mental process, an observation, evaluation, judgement, or opinion. The limitation as drafted, and under a broadest reasonable interpretation, can be performed in the human mind, or by a human using a pen and paper (MPEP 2106.04(a)(2)(Ill)(c)). For example, a person could process a new image in the target domain to generate a label for the new image and perform an action responsive to the label.
Step 2A Prong 2 - Does the claim recite additional elements that integrate the judicial exception into a practical application?
“using the fine-tuned vision language model” is merely reciting the words "apply it" (or an equivalent) with the judicial exception, or merely including instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
“performing an action responsive to the label, the action being selected from the group consisting of controlling access to a secure area, performing an automated security action, and sending relief to people suffering from a natural disaster.” (insignificant extra-solution activity mere data gathering MPEP 2106.05(g))
The claim does not include additional elements that are integrated into a practical application.
Step 2B - Does the claim recite additional elements that amount to significantly more than the judicial exception?
“using the fine-tuned vision language model” is merely reciting the words "apply it" (or an equivalent) with the judicial exception, or merely including instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
“performing an action responsive to the label, the action being selected from the group consisting of controlling access to a secure area, performing an automated security action, and sending relief to people suffering from a natural disaster.” (well-understood, routine, conventional MPEP 2106.05(d))
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Tony Huang et al; “Unsupervised Prompt Learning for Vision-Language Models” published Aug 22, 2022 (hereinafter “Huang”) in view of Akshay Soni et al; US 20180053097 A1 filed on Aug 16, 2016 (hereinafter “Soni”)
Regarding claim 1, Huang teaches A computer-implemented method for fine-tuning a model, comprising: […] generating text pseudo-labels for images in an unlabeled dataset from the target domain based on the label space using a pre-trained vision language model; and (Huang Page 3 Paragraph 2; “we utilize a pre-trained vision-language model (e.g. CLIP) to generate pseudo labels for unlabeled images from the target dataset.” Examiner notes that text pseudo labels (pseudo labels) is generated for images in an unlabeled dataset (unlabeled images) from the target domain based on the label space (target dataset) using a pre-trained vision language model)
fine-tuning the pre-trained vision language model for the target domain using the images with the text pseudo-labels. (Huang Page 3 Paragraph 2; “the correlation between confidence scores and pseudo label accuracy is relatively low; 2) vision-language models have biased per-class accuracy, we thus select top-K confident samples for each class, instead of keeping all samples with confidence scores higher than a pre-defined threshold, for the subsequent prompt representation optimization.” Examiner notes that the pretrained vision language model (Vision language model) is fine-tuned (select top-K confident samples for each class, … for the subsequent prompt representation optimization) for the target domain using the images with the text pseudo labels (correlation between confidence scores and pseudo label accuracy))
Huang does not teach generating a label space for a target domain;
However, Soni does teach generating a label space for a target domain; (Soni Paragraph 0008; “In some embodiments, generating a label space further comprises obtaining a plurality of data samples from at least a knowledge base… and generating the label space based on the second label matrix and the trained one or more parameters.” Examiner notes that a label space is generated for a target domain (knowledge base))
It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Huang and Soni. Huang teaches an unsupervised prompt learning for vision language models. Soni teaches a method for multi-label prediction. One of ordinary skill would have motivation to combine Huang and Soni to accurately and efficiently recognize and label newly available data “there is a need to provide a solution to accurately and efficiently recognize and label newly available data points to tackle the above-mentioned challenges.” (Soni Paragraph 0005).
Regarding claim 12, Huang does not teach A system for fine-tuning a model, comprising: comprising a hardware processor; and a memory that stores a computer program which, when executed by the hardware processor, causes the hardware processor to:
However, Soni does teach A system for fine-tuning a model, comprising: comprising a hardware processor; and a memory that stores a computer program which, when executed by the hardware processor, causes the hardware processor to: (Soni Paragraph 0062; “FIG. 9 depicts a general mobile device architecture on which the present teaching can be implemented.”)
PNG
media_image1.png
461
266
media_image1.png
Greyscale
Claim 12 is a system claim of method claim 1 and is accordingly rejected using substantially similar rationale as to that which is set for with respect to claim 1
It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Huang and Soni. Huang teaches an unsupervised prompt learning for vision language models. Soni teaches a method for multi-label prediction. One of ordinary skill would have motivation to combine Huang and Soni to accurately and efficiently recognize and label newly available data “there is a need to provide a solution to accurately and efficiently recognize and label newly available data points to tackle the above-mentioned challenges.” (Soni Paragraph 0005).
Claim(s) 2 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Tony Huang et al; “Unsupervised Prompt Learning for Vision-Language Models” published Aug 22, 2022 (hereinafter “Huang”) in view of Akshay Soni et al; US 20180053097 A1 filed on Aug 16, 2016 (hereinafter “Soni”) in further view of Ashutosh Tripathi; “A Quick Guide to Tokenization, Lemmatization, Stop Words, and Phrase Matching using spaCy | NLP | Part 2” available online Aug 11, 2020 (hereinafter “Tripathi”)
Regarding claim 2, Huang does not teach The method of claim 1, wherein generating the label space includes pruning a list of tokens to remove tokens that represent punctuation, special symbols, and plural forms.
However, Tripathi does teach The method of claim 1, wherein generating the label space includes pruning a list of tokens to remove tokens that represent punctuation, special symbols, and plural forms. (Tripathi Section “Tokenization” Paragraph 2; “Prefix, Suffix and Infix check: Punctuation like commas, periods, hyphens or quotes to be treated as tokens and separated out… Prefix: Look for Character(s) at the beginning ▸ $ ( “Suffix: Look for Character(s) at the end ▸ mm ) , . ! ” mm is an example of unit Infix: Look for Character(s) in between ▸ - -- / ...” Tripathi Section Lemmatization Paragraph 1; “Lemmatization looks beyond word reduction, and considers a language’s full vocabulary to apply a morphological analysis to words. The lemma of ‘was’ is ‘be’, lemma of “rats” is “rat” and the lemma of ‘mice’ is ‘mouse’. Further, the lemma of ‘meeting’ might be ‘meet’ or ‘meeting’ depending on its use in a sentence.” Examiner notes that tokens that represent punctuation and special symbols; Lemmatization is used to remove plural forms)
It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Huang, Soni, and Tripathi. Huang teaches an unsupervised prompt learning for vision language models. Soni teaches a method for multi-label prediction. Tripathi teaches Tokenization, Lemmatization, Stop Words, and Phrase Matching using spaCy. One of ordinary skill would have motivation to combine Huang, Soni, and Tripathi to efficiently build a natural language understanding system ““spaCy” is designed specifically for production use. It helps you build applications that process and “understand” large volumes of text. It can be used to build information extraction or natural language understanding systems, or to pre-process text for deep learning.” (Tripathi Paragraph 1).
Regarding claim 13, claim 13 has similar limitations as of claim 2, except it is a system claim, therefore it is rejected under the same rationale as claim 2.
Claim(s) 3 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Tony Huang et al; “Unsupervised Prompt Learning for Vision-Language Models” published Aug 22, 2022 (hereinafter “Huang”) in view of Akshay Soni et al; US 20180053097 A1 filed on Aug 16, 2016 (hereinafter “Soni”) in further view of Haytham Yaghi et al; US 20240202443 A1 filed on Dec 15, 2022 (hereinafter “Yaghi”)
Regarding claim 3, Huang does not teach The method of claim 1, wherein generating the label space includes determining similarity values between tokens in a list of tokens and generating a bag of words representation based on tokens that have an above-threshold similarity value.
However, Yaghi does teach The method of claim 1, wherein generating the label space includes determining similarity values between tokens in a list of tokens and generating a bag of words representation based on tokens that have an above-threshold similarity value. (Yaghi Paragraph 0052; “the system may calculate bag-of-words vector representations of the textual data and each datum within the dataset and compare these (e.g., by calculating a normalized inner product between the vector representations). Generating similarity metrics between textual data without use of a natural language processor may enable faster, more efficient calculations, as less computational power may be required for vector manipulation and analysis when compared to natural language processing or other machine learning techniques. In some embodiments, the system may determine similarity metrics between two different labels, such as those determined at different times for the same training datum. For example, the system may determine a label similarity metric by comparing the text within labels, and use this similarity metric to determine whether to generate a warning based on a label similarity threshold.” Examiner notes similarity values (similarity metrics) is determined/generated between tokens in a list of tokens (textual data) and generating a bag or words representation based on tokens (calculate bag of words vector representations of the textual data) that have an above threshold similarity value (use this similarity metric to determine whether to generate a warning based on a label similarity threshold))
It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Huang, Soni, and Yaghi. Huang teaches an unsupervised prompt learning for vision language models. Soni teaches a method for multi-label prediction. Yaghi teaches a method for generating recommendations for unlabeled data for machine learning model training using natural language processing. One of ordinary skill would have motivation to combine Huang, Soni, and Yaghi to enable faster and more efficient calculations for comparing textual data “Generating similarity metrics between textual data without use of a natural language processor may enable faster, more efficient calculations, as less computational power may be required for vector manipulation and analysis when compared to natural language processing or other machine learning techniques.” (Yaghi Paragraph 0052).
Regarding claim 14, claim 14 has similar limitations as of claim 3, except it is a system claim, therefore it is rejected under the same rationale as claim 3.
Claim(s) 4 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Tony Huang et al; “Unsupervised Prompt Learning for Vision-Language Models” published Aug 22, 2022 (hereinafter “Huang”) in view of Akshay Soni et al; US 20180053097 A1 filed on Aug 16, 2016 (hereinafter “Soni”) in further view of Haytham Yaghi et al; US 20240202443 A1 filed on Dec 15, 2022 (hereinafter “Yaghi”) in further view of Kye-Hyeon Kim et al; US 10579907 B1 filed on Jan 31, 2019 (hereinafter “Kim”)
Regarding claim 4, Huang does not teach The method of claim 3, wherein generating the label space further includes combining bag of words representations for all images in the unlabeled dataset to generate a bag of words representation for the entire unlabeled dataset.
However, Kim does teach The method of claim 3, wherein generating the label space further includes combining bag of words representations for all images in the unlabeled dataset to generate a bag of words representation for the entire unlabeled dataset. (Kim Fig 3 and Column 10 Line 7; “And the similar-image selection network 200 may generate at least one bag-of-words histogram of the original images by applying at least one operation, of creating at least one bag of words to all of the original images, by referring to the top-k pieces of the class information corresponding to each of the manipulated images outputted from the image classification CNN 210.” Examiner notes that bag of word representations for all images (generate at least one bag-of-words histogram of the original images) is combined to generate a bag of words representation for the entire unlabeled dataset (Fig 3 shows combined bag of words of top k classes for unlabeled images))
PNG
media_image2.png
443
757
media_image2.png
Greyscale
It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Huang, Soni, Yaghi, and Kim. Huang teaches an unsupervised prompt learning for vision language models. Soni teaches a method for multi-label prediction. Yaghi teaches a method for generating recommendations for unlabeled data for machine learning model training using natural language processing. Kim teaches a method for evaluating a reliability of labeling training images to be used for learning a deep learning network. One of ordinary skill would have motivation to combine Huang, Soni, Yaghi, and Kim to optimize sampling processes and reduce annotation costs “This method can be used to recognize surroundings by applying a bag-of-words model, to optimize sampling processes for selecting a valid image among similar images, and to reduce annotation costs.” (Kim Abstract).
Regarding claim 15, claim 15 has similar limitations as of claim 4, except it is a system claim, therefore it is rejected under the same rationale as claim 4.
Claim(s) 5 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Tony Huang et al; “Unsupervised Prompt Learning for Vision-Language Models” published Aug 22, 2022 (hereinafter “Huang”) in view of Akshay Soni et al; US 20180053097 A1 filed on Aug 16, 2016 (hereinafter “Soni”) in further view of Haytham Yaghi et al; US 20240202443 A1 filed on Dec 15, 2022 (hereinafter “Yaghi”) in further view of Kye-Hyeon Kim et al; US 10579907 B1 filed on Jan 31, 2019 (hereinafter “Kim”) in further view of Tzvia Bader et al; US 20210183526 A1 filed on Dec 16, 2020 (hereinafter “Bader”)
Regarding claim 5, Huang does not teach The method of claim 4, wherein generating the label space further includes pruning words from the bag of words representation for the entire unlabeled dataset based on a frequency of occurrence.
However, Bader does teach The method of claim 4, wherein generating the label space further includes pruning words from the bag of words representation for the entire unlabeled dataset based on a frequency of occurrence. (Bader Paragraph 0062; “high-frequency words often provide little information. In various embodiments, words with a frequency above a certain threshold may be subsampled to increase training speed. In various embodiments, high-frequency words may be removed.” Examiner notes that words from the bag or words representation for the entire unlabeled dataset (words) is pruned/removed based on frequency of occurrence (words with a frequency above a certain threshold))
It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Huang, Soni, Yaghi, Kim, and Bader. Huang teaches an unsupervised prompt learning for vision language models. Soni teaches a method for multi-label prediction. Yaghi teaches a method for generating recommendations for unlabeled data for machine learning model training using natural language processing. Kim teaches a method for evaluating a reliability of labeling training images to be used for learning a deep learning network. Bader teaches removing high frequency words. One of ordinary skill would have motivation to combine Huang, Soni, Yaghi, Kim, and Bader to increase training speed “words with a frequency above a certain threshold may be subsampled to increase training speed.” (Bader Paragraph 0062).
Regarding claim 16, claim 16 has similar limitations as of claim 5, except it is a system claim, therefore it is rejected under the same rationale as claim 5.
Claim(s) 6 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Tony Huang et al; “Unsupervised Prompt Learning for Vision-Language Models” published Aug 22, 2022 (hereinafter “Huang”) in view of Akshay Soni et al; US 20180053097 A1 filed on Aug 16, 2016 (hereinafter “Soni”) in further view of Haytham Yaghi et al; US 20240202443 A1 filed on Dec 15, 2022 (hereinafter “Yaghi”) in further view of Kye-Hyeon Kim et al; US 10579907 B1 filed on Jan 31, 2019 (hereinafter “Kim”) in further view of Yao Sun; US 20140195348 A1 filed on Jan 8, 2014 (hereinafter “Sun”)
Regarding claim 6, Huang does not teach The method of claim 4, wherein generating the label space further includes pruning words from the bag of words representation for the entire unlabeled dataset based on an above-threshold degree of similarity to other words from the bag of words representation.
However, Sun does teach The method of claim 4, wherein generating the label space further includes pruning words from the bag of words representation for the entire unlabeled dataset based on an above-threshold degree of similarity to other words from the bag of words representation. (Sun Paragraph 0068; “Synonym removal: Assuming the threshold similarity for a synonym is 95%, upon calculating the similarity of all pairs of tokenized words among the above tokenized word collection” Examiner notes that the words from the bag of words representation for the entire unlabeled dataset (tokenized words among the above tokenized word collection) is pruned/removed based on an above threshold degree of similarity to other words from the bag of words representation (Assuming the threshold similarity for a synonym is 95%, upon calculating the similarity of all pairs of tokenized words among the above tokenized word collection))
It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Huang, Soni, Yaghi, Kim, and Sun. Huang teaches an unsupervised prompt learning for vision language models. Soni teaches a method for multi-label prediction. Yaghi teaches a method for generating recommendations for unlabeled data for machine learning model training using natural language processing. Kim teaches a method for evaluating a reliability of labeling training images to be used for learning a deep learning network. Sun teaches removing synonyms in a word collection. One of ordinary skill would have motivation to combine Huang, Soni, Yaghi, Kim, and Sun to optimize computational costs by removing redundant information “the computer removes redundant information from search behavioral data. For example, the computer may remove duplicate words or synonyms, and/or merge synonyms or near-synonyms.” (Sun Paragraph 0051).
Regarding claim 17, claim 17 has similar limitations as of claim 6, except it is a system claim, therefore it is rejected under the same rationale as claim 6.
Claim(s) 7-8 and 18-19 are rejected under 35 U.S.C. 103 as being unpatentable over Tony Huang et al; “Unsupervised Prompt Learning for Vision-Language Models” published Aug 22, 2022 (hereinafter “Huang”) in view of Akshay Soni et al; US 20180053097 A1 filed on Aug 16, 2016 (hereinafter “Soni”) in further view of Yurii Huts et al; US 20230067026 A1 filed on Feb 17, 2021 (hereinafter “Huts”)
Regarding claim 7, Huang does not teach The method of claim 1, wherein the pre-trained vision language model is pre-trained on images from an original domain that is different from the target domain.
However, Huts does teach The method of claim 1, wherein the pre-trained vision language model is pre-trained on images from an original domain that is different from the target domain. (Huts Paragraph 0222; “the inventors have observed that neural networks trained to extract image features in for a task T1 in a computer vision domain D1 can be used to extract image features for a different task T2 in the domain D1, in a different computer vision domain D2, or in other fields of data analytics that can derive useful information from image data.” Examiner notes that the pre-trained vision language model is pretrain on images form an original domain (neural networks trained to extract image features in for a task T1 in a computer vision domain D1) that is different from the target domain (can be used to extract image features for a different task T2… in a different computer vision domain D2))
It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Huang, Soni, and Huts. Huang teaches an unsupervised prompt learning for vision language models. Soni teaches a method for multi-label prediction. Huts teaches repurposing neural networks for different tasks. One of ordinary skill would have motivation to combine Huang, Soni, and Huts to repurpose neural networks in a way to eliminate the inefficiencies of training a new network “This successful repurposing of neural networks both eliminates the inefficiencies of training a new neural network for the task T2, and enables users to rely on their particular, preferred machine-learning models for modeling tasks or data analytic tasks, even when the tasks involve analysis of image data.” (Huts Paragraph 0222).
Regarding claim 18, claim 18 has similar limitations as of claim 7, except it is a system claim, therefore it is rejected under the same rationale as claim 7.
Regarding claim 8, Huang does not teach The method of claim 7, wherein the target domain differs from the original domain in at least one respect selected from the group consisting of color range, visual angle, content, environment, and camera settings.
However, Huts does teach The method of claim 7, wherein the target domain differs from the original domain in at least one respect selected from the group consisting of color range, visual angle, content, environment, and camera settings. (Huts Paragraph 0123; “The image feature extraction model may be “pre-trained” in the sense that it has been trained to extract features suitable for performing a particular computer vision task (e.g., detecting cats in images), whereas the model development system 100 may be developing a model 130 that performs a different computer vision task (e.g., detecting fractures in medical images) or data analytics task (e.g., estimating the value of a house based in part on images thereof).” Examiner notes that the target domain (detecting fractures in medical images) differs from the original domain (detecting cats in images) in respect to content)
It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Huang, Soni, and Huts. Huang teaches an unsupervised prompt learning for vision language models. Soni teaches a method for multi-label prediction. Huts teaches repurposing neural networks for different tasks. One of ordinary skill would have motivation to combine Huang, Soni, and Huts to repurpose neural networks in a way to eliminate the inefficiencies of training a new network “This successful repurposing of neural networks both eliminates the inefficiencies of training a new neural network for the task T2, and enables users to rely on their particular, preferred machine-learning models for modeling tasks or data analytic tasks, even when the tasks involve analysis of image data.” (Huts Paragraph 0222).
Regarding claim 19, claim 19 has similar limitations as of claim 8, except it is a system claim, therefore it is rejected under the same rationale as claim 8.
Claim(s) 9, 10, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Tony Huang et al; “Unsupervised Prompt Learning for Vision-Language Models” published Aug 22, 2022 (hereinafter “Huang”) in view of Akshay Soni et al; US 20180053097 A1 filed on Aug 16, 2016 (hereinafter “Soni”) in further view of Angel Alejandro Rodriguez; US 20240220851 A1 filed on Dec 29, 2022 (hereinafter “Rodriguez”)
Regarding claim 9, Huang teaches The method of claim 1, further comprising processing a new image in the target domain using the fine-tuned vision language model to generate a label for the new image (Huang Page 5 Paragraph 2; “Once the optimization of prompt representation V is finished, given the target dataset, we feed {Vc} Cc=1 into the CLIP’s text encoder to generate class embeddings for all categories.” Examiner notes that the fine tune vision language model (optimization of prompt representation that is used by CLIP) processes a new image in the target domain (given the target dataset) to generate a label for the new image (class embeddings for all categories))
Huang does not teach and performing an action responsive to the label.
However, Rodriguez does teach and performing an action responsive to the label. (Rodriguez Paragraph 0042; “Alarm unit 38 may receive one or more premises security surveillance video feeds, images, audio recordings, or sensor data from one or more premises device 16, may utilize the ML model to identify objects in the received premises device 16 data (e.g., objects in one or more image frames of a video feed) and to label or classifies the identified objects, and may associate the labels with the objects and/or received premises device 16 data. Based on the rule configuration file, the alarm unit 38 maps the labels to corresponding alarm events and/or premises security actions, and performs one or more premises security actions (e.g., trigger an alarm, suppress an alarm, send a notification to a user device 14, send a notification to remote monitoring center 22, etc.) based on the mapping.” Examiner notes that responsive to the label (labels to corresponding alarm events and/or premises security actions) performs an action (performs one or more premises security actions (e.g., trigger an alarm, suppress an alarm, send a notification to a user device 14, send a notification to remote monitoring center 22, etc))
It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Huang, Soni, and Rodriguez. Huang teaches an unsupervised prompt learning for vision language models. Soni teaches a method for multi-label prediction. Rodriguez teaches A user device for a premises security system comprising a plurality of premises devices. One of ordinary skill would have motivation to combine Huang, Soni, and Rodriguez to leverage machine learning for detecting specific objects and improve security “This model may facilitate security devices detecting specific objects, providing users with improved and/or heightened security for the events and objects with which they are concerned.” (Rodriguez Paragraph 0008).
Regarding claim 10, Huang does not teach The method of claim 1, wherein the action is selected from the group consisting of controlling access to a secure area, performing an automated security action, and sending relief to people suffering from a natural disaster.
However, Rodriguez does teach The method of claim 1, wherein the action is selected from the group consisting of controlling access to a secure area, performing an automated security action, and sending relief to people suffering from a natural disaster. (Rodriguez Paragraph 0042; “Alarm unit 38 may receive one or more premises security surveillance video feeds, images, audio recordings, or sensor data from one or more premises device 16, may utilize the ML model to identify objects in the received premises device 16 data (e.g., objects in one or more image frames of a video feed) and to label or classifies the identified objects, and may associate the labels with the objects and/or received premises device 16 data. Based on the rule configuration file, the alarm unit 38 maps the labels to corresponding alarm events and/or premises security actions, and performs one or more premises security actions (e.g., trigger an alarm, suppress an alarm, send a notification to a user device 14, send a notification to remote monitoring center 22, etc.) based on the mapping.” Examiner notes that responsive to the label (labels to corresponding alarm events and/or premises security actions) performs an action (performs one or more premises security actions (e.g., trigger an alarm, suppress an alarm, send a notification to a user device 14, send a notification to remote monitoring center 22, etc))
It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Huang, Soni, and Rodriguez. Huang teaches an unsupervised prompt learning for vision language models. Soni teaches a method for multi-label prediction. Rodriguez teaches A user device for a premises security system comprising a plurality of premises devices. One of ordinary skill would have motivation to combine Huang, Soni, and Rodriguez to leverage machine learning for detecting specific objects and improve security “This model may facilitate security devices detecting specific objects, providing users with improved and/or heightened security for the events and objects with which they are concerned.” (Rodriguez Paragraph 0008).
Regarding claim 20, Huang teaches The system of claim 12, wherein the computer program further causes the hardware processor to process a new image in the target domain using the fine-tuned vision language model to generate a label for the new image and (Huang Page 5 Paragraph 2; “Once the optimization of prompt representation V is finished, given the target dataset, we feed {Vc} Cc=1 into the CLIP’s text encoder to generate class embeddings for all categories.” Examiner notes that the fine tune vision language model (optimization of prompt representation that is used by CLIP) processes a new image in the target domain (given the target dataset) to generate a label for the new image (class embeddings for all categories))
Huang does not teach performing an action responsive to the label, the action being selected from the group consisting of controlling access to a secure area, performing an automated security action, and sending relief to people suffering from a natural disaster.
However, Rodriguez does teach performing an action responsive to the label, the action being selected from the group consisting of controlling access to a secure area, performing an automated security action, and sending relief to people suffering from a natural disaster. (Rodriguez Paragraph 0042; “Alarm unit 38 may receive one or more premises security surveillance video feeds, images, audio recordings, or sensor data from one or more premises device 16, may utilize the ML model to identify objects in the received premises device 16 data (e.g., objects in one or more image frames of a video feed) and to label or classifies the identified objects, and may associate the labels with the objects and/or received premises device 16 data. Based on the rule configuration file, the alarm unit 38 maps the labels to corresponding alarm events and/or premises security actions, and performs one or more premises security actions (e.g., trigger an alarm, suppress an alarm, send a notification to a user device 14, send a notification to remote monitoring center 22, etc.) based on the mapping.” Examiner notes that responsive to the label (labels to corresponding alarm events and/or premises security actions) performs an action (performs one or more premises security actions (e.g., trigger an alarm, suppress an alarm, send a notification to a user device 14, send a notification to remote monitoring center 22, etc))
It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Huang, Soni, and Rodriguez. Huang teaches an unsupervised prompt learning for vision language models. Soni teaches a method for multi-label prediction. Rodriguez teaches A user device for a premises security system comprising a plurality of premises devices. One of ordinary skill would have motivation to combine Huang, Soni, and Rodriguez to leverage machine learning for detecting specific objects and improve security “This model may facilitate security devices detecting specific objects, providing users with improved and/or heightened security for the events and objects with which they are concerned.” (Rodriguez Paragraph 0008).
Claim(s) 11 is rejected under 35 U.S.C. 103 as being unpatentable over Tony Huang et al; “Unsupervised Prompt Learning for Vision-Language Models” published Aug 22, 2022 (hereinafter “Huang”) in view of Akshay Soni et al; US 20180053097 A1 filed on Aug 16, 2016 (hereinafter “Soni”) in further view of Haytham Yaghi et al; US 20240202443 A1 filed on Dec 15, 2022 (hereinafter “Yaghi”) in further view of Kye-Hyeon Kim et al; US 10579907 B1 filed on Jan 31, 2019 (hereinafter “Kim”) in further view of Tzvia Bader et al; US 20210183526 A1 filed on Dec 16, 2020 (hereinafter “Bader”) in further view of Yao Sun; US 20140195348 A1 filed on Jan 8, 2014 (hereinafter “Sun”)
Regarding claim 11, Huang teaches generating text pseudo-labels for images in an unlabeled dataset from the target domain based on the label space using a pre-trained vision language model; (Huang Page 3 Paragraph 2; “we utilize a pre-trained vision-language model (e.g. CLIP) to generate pseudo labels for unlabeled images from the target dataset.” Examiner notes that text pseudo labels (pseudo labels) is generated for images in an unlabeled dataset (unlabeled images) from the target domain based on the label space (target dataset) using a pre-trained vision language model)
and fine-tuning the pre-trained vision language model for the target domain using the images with the text pseudo-labels. (Huang Page 3 Paragraph 2; “the correlation between confidence scores and pseudo label accuracy is relatively low; 2) vision-language models have biased per-class accuracy, we thus select top-K confident samples for each class, instead of keeping all samples with confidence scores higher than a pre-defined threshold, for the subsequent prompt representation optimization.” Examiner notes that the pretrained vision language model (Vision language model) is fine-tuned (select top-K confident samples for each class, … for the subsequent prompt representation optimization) for the target domain using the images with the text pseudo labels (correlation between confidence scores and pseudo label accuracy))
Huang does not teach A computer-implemented method for fine-tuning a model, comprising: generating a label space for a target domain, including:
However, Soni does teach A computer-implemented method for fine-tuning a model, comprising: generating a label space for a target domain, including: (Soni Paragraph 0008; “In some embodiments, generating a label space further comprises obtaining a plurality of data samples from at least a knowledge base… and generating the label space based on the second label matrix and the trained one or more parameters.” Examiner notes that a label space is generated for a target domain (knowledge base))
It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Huang and Soni. Huang teaches an unsupervised prompt learning for vision language models. Soni teaches a method for multi-label prediction. One of ordinary skill would have motivation to combine Huang and Soni to accurately and efficiently recognize and label newly available data “there is a need to provide a solution to accurately and efficiently recognize and label newly available data points to tackle the above-mentioned challenges.” (Soni Paragraph 0005).
Huang in view of Soni does not teach determining similarity values between tokens in a list of tokens and generating a bag of words representation based on tokens that have an above-threshold similarity value;
However, Yaghi does teach determining similarity values between tokens in a list of tokens and generating a bag of words representation based on tokens that have an above-threshold similarity value; (Yaghi Paragraph 0052; “the system may calculate bag-of-words vector representations of the textual data and each datum within the dataset and compare these (e.g., by calculating a normalized inner product between the vector representations). Generating similarity metrics between textual data without use of a natural language processor may enable faster, more efficient calculations, as less computational power may be required for vector manipulation and analysis when compared to natural language processing or other machine learning techniques. In some embodiments, the system may determine similarity metrics between two different labels, such as those determined at different times for the same training datum. For example, the system may determine a label similarity metric by comparing the text within labels, and use this similarity metric to determine whether to generate a warning based on a label similarity threshold.” Examiner notes similarity values (similarity metrics) is determined/generated between tokens in a list of tokens (textual data) and generating a bag or words representation based on tokens (calculate bag of words vector representations of the textual data) that have an above threshold similarity value (use this similarity metric to determine whether to generate a warning based on a label similarity threshold))
It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Huang, Soni, and Yaghi. Huang teaches an unsupervised prompt learning for vision language models. Soni teaches a method for multi-label prediction. Yaghi teaches a method for generating recommendations for unlabeled data for machine learning model training using natural language processing. One of ordinary skill would have motivation to combine Huang, Soni, and Yaghi to enable faster and more efficient calculations for comparing textual data “Generating similarity metrics between textual data without use of a natural language processor may enable faster, more efficient calculations, as less computational power may be required for vector manipulation and analysis when compared to natural language processing or other machine learning techniques.” (Yaghi Paragraph 0052).
Huang in view of Soni in further view of Yaghi does not teach combining bag of words representations for all images in an unlabeled dataset to generate a bag of words representation for the entire unlabeled dataset;
However, Kim does teach combining bag of words representations for all images in an unlabeled dataset to generate a bag of words representation for the entire unlabeled dataset; (Kim Fig 3 and Column 10 Line 7; “And the similar-image selection network 200 may generate at least one bag-of-words histogram of the original images by applying at least one operation, of creating at least one bag of words to all of the original images, by referring to the top-k pieces of the class information corresponding to each of the manipulated images outputted from the image classification CNN 210.” Examiner notes that bag of word representations for all images (generate at least one bag-of-words histogram of the original images) is combined to generate a bag of words representation for the entire unlabeled dataset (Fig 3 shows combined bag of words of top k classes for unlabeled images))
PNG
media_image2.png
443
757
media_image2.png
Greyscale
It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Huang, Soni, Yaghi, and Kim. Huang teaches an unsupervised prompt learning for vision language models. Soni teaches a method for multi-label prediction. Yaghi teaches a method for generating recommendations for unlabeled data for machine learning model training using natural language processing. Kim teaches a method for evaluating a reliability of labeling training images to be used for learning a deep learning network. One of ordinary skill would have motivation to combine Huang, Soni, Yaghi, and Kim to optimize sampling processes and reduce annotation costs “This method can be used to recognize surroundings by applying a bag-of-words model, to optimize sampling processes for selecting a valid image among similar images, and to reduce annotation costs.” (Kim Abstract).
Huang in view of Soni in further view of Yaghi in further view of Kim does not teach pruning words from the bag of words representation for the entire unlabeled dataset based on a frequency of occurrence; and
However, Bader does teach pruning words from the bag of words representation for the entire unlabeled dataset based on a frequency of occurrence; and (Bader Paragraph 0062; “high-frequency words often provide little information. In various embodiments, words with a frequency above a certain threshold may be subsampled to increase training speed. In various embodiments, high-frequency words may be removed.” Examiner notes that words from the bag or words representation for the entire unlabeled dataset (words) is pruned/removed based on frequency of occurrence (words with a frequency above a certain threshold))
It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Huang, Soni, Yaghi, Kim, and Bader. Huang teaches an unsupervised prompt learning for vision language models. Soni teaches a method for multi-label prediction. Yaghi teaches a method for generating recommendations for unlabeled data for machine learning model training using natural language processing. Kim teaches a method for evaluating a reliability of labeling training images to be used for learning a deep learning network. Bader teaches removing high frequency words. One of ordinary skill would have motivation to combine Huang, Soni, Yaghi, Kim, and Bader to increase training speed “words with a frequency above a certain threshold may be subsampled to increase training speed.” (Bader Paragraph 0062).
Huang in view of Soni in further view of Yaghi in further view of Kim in further view of Bader does not teach The method of claim 4, wherein generating the label space further includes pruning words from the bag of words representation for the entire unlabeled dataset based on an above-threshold degree of similarity to other words from the bag of words representation.
However, Sun does teach The method of claim 4, wherein generating the label space further includes pruning words from the bag of words representation for the entire unlabeled dataset based on an above-threshold degree of similarity to other words from the bag of words representation. (Sun Paragraph 0068; “Synonym removal: Assuming the threshold similarity for a synonym is 95%, upon calculating the similarity of all pairs of tokenized words among the above tokenized word collection” Examiner notes that the words from the bag of words representation for the entire unlabeled dataset (tokenized words among the above tokenized word collection) is pruned/removed based on an above threshold degree of similarity to other words from the bag of words representation (Assuming the threshold similarity for a synonym is 95%, upon calculating the similarity of all pairs of tokenized words among the above tokenized word collection))
It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Huang, Soni, Yaghi, Kim, Bader, and Sun. Huang teaches an unsupervised prompt learning for vision language models. Soni teaches a method for multi-label prediction. Yaghi teaches a method for generating recommendations for unlabeled data for machine learning model training using natural language processing. Kim teaches a method for evaluating a reliability of labeling training images to be used for learning a deep learning network. Bader teaches removing high frequency words. Sun teaches removing synonyms in a word collection. One of ordinary skill would have motivation to combine Huang, Soni, Yaghi, Kim, Bader, and Sun to optimize computational costs by removing redundant information “the computer removes redundant information from search behavioral data. For example, the computer may remove duplicate words or synonyms, and/or merge synonyms or near-synonyms.” (Sun Paragraph 0051).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DANIEL DUC TRAN whose telephone number is (571)272-6870. The examiner can normally be reached Mon-Fri 8:00-5:00 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached at (571) 270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/D.D.T./Examiner, Art Unit 2147
/VIKER A LAMARDO/Supervisory Patent Examiner, Art Unit 2147