DETAILED ACTION
This communication is in response to the Application No. 18/423,153 filed January 25, 2024
in which Claims 1 - 20 are presented for examination.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim 1-20 are rejected under 35 U.S.C. 101 because these claimed inventions are directed to an
abstract idea without significantly more.
Regarding Claim 1:
Step 1: Claim 1 is a method type claim. Therefore, Claims 1-7 fall within one of the four statutory
categories (i.e., process, machine, manufacture, or composition of matter).
2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance
of the limitation in the mind but for the recitation of generic computer components, then it falls within
the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable
interpretation, covers performance of the limitation by mathematical calculation but for the recitation
of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract
ideas.
generating a textual prompt, wherein the textual prompt includes a human-centric evaluation criterion (mental process – generating a textual prompt that includes a human-centric evaluation criterion may be performed mentally or using pen and paper by a user observing/analyzing the human-centric evaluation criterion and accordingly using judgment/evaluation to create or write the textual prompt including the human-centric evaluation criterion)
generating […] an alignment score based on the textual prompt and a two-dimensional (2D) input image, wherein the 2D input image is a visual representation of an architectural space, and the alignment score is a quantitative measure of how accurately the textual prompt describes the 2D input image based on the human-centric evaluation criterion (mental process – generating an alignment score based on the textual prompt and a two-dimensional (2D) input image may be performed mentally or using pen and paper by a user observing/analyzing the textual prompt and the 2D input image and accordingly using judgment/evaluation to determine how accurately the textual prompt describes the 2D input image based on the human-centric evaluation criterion)
assigning the human-centric evaluation criterion and the alignment score to the 2D input image (mental process – assigning the human-centric evaluation criterion and the alignment score to the 2D input image may be performed mentally or using pen and paper by a user identifying the applicable criterion and score and associating them with the corresponding 2D input image)
Step 2A Prong 2: This judicial exception is not integrated into a practical application.
[…], via a machine learning model, […] (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using a machine learning model without significantly more)
displaying […] one or more of the 2D input image, the human-centric evaluation criterion, and the alignment score (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g))
[…], via a graphical user interface, […] (recited at a high-level of generality (i.e., a graphical user interface, a generic processor, computer-readable storage medium, a communication interface, a user interface and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components)
Step 2B: The claim does not include additional elements considered individually and in combination that
are sufficient to amount to significantly more than the judicial exception.
[…], via a machine learning model, […] (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using a machine learning model without significantly more)
displaying […] one or more of the 2D input image, the human-centric evaluation criterion, and the alignment score (MPEP 2106.05(g) indicates that merely outputting the results of an abstract analysis is insignificant extra-solution activity and does not amount to an inventive concept. Here, the displaying merely outputs the results of the preceding analysis to a user. Thereby, this additional limitation does not amount to significantly more than the judicial exception.)
[…], via a graphical user interface, […] (recited at a high-level of generality (i.e., a graphical user interface, a generic processor, computer-readable storage medium, a communication interface, a user interface and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components)
For the reasons above, Claim 1 is rejected as being directed to an abstract idea without
significantly more. This rejection applies equally to dependent claims 1 - 7. The additional limitations of the dependent claims are addressed below.
Regarding Claim 2:
Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 2 depends on.
Step 2A Prong 2 & Step 2B:
wherein the human-centric evaluation criterion includes one of the terms “social,” “isolating,” “tranquil,” “distracting,” “inspirational,” or “boring.” (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the human-centric evaluation criterion includes one of the terms “social,” “isolating,” “tranquil,” “distracting,” “inspirational,” or “boring” does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h))
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the
abstract idea into practical application because it does not impose any meaningful limits on practicing
the abstract idea, as discussed above in the rejection of claim 1. The claim does not include additional
elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 3:
Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 3 depends on.
Step 2A Prong 2 & Step 2B:
wherein the 2D input image is one of a photograph of an architectural space, a 2D rendering of an architectural space, or a still image representing a single frame included in a video recording of an architectural space (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the 2D input image is one of a photograph of an architectural space, a 2D rendering of an architectural space, or a still image representing a single frame included in a video recording of an architectural space does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h))
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the
abstract idea into practical application because it does not impose any meaningful limits on practicing
the abstract idea, as discussed above in the rejection of claim 1. The claim does not include additional
elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 4:
Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 4 depends on.
initializing, […], a plurality of learnable parameters […] (mental process - initializing the plurality of learnable parameters may be performed mentally or using pen and paper by a user observing/analyzing relevant information and accordingly using judgment/evaluation to determine and assign initial parameter values based on said analysis)
iteratively adjusting one or more of the plurality of learnable parameters based on a first plurality of alignment scores generated for a first plurality of image-text pairs included in a training data set (mental process - iteratively adjusting one or more of the plurality of learnable parameters based on the first plurality of alignment scores may be performed mentally or using pen and paper by a user observing/analyzing the alignment scores and accordingly using judgment/evaluation to repeatedly modify the learnable parameters based on said analysis)
calculating an accuracy […] based on a second plurality of alignment scores generated for a second plurality of image-text pairs included in a testing data set (mathematical concept - calculating the accuracy based on the second plurality of alignment scores involves performing mathematical calculations on numerical alignment score values to determine a quantitative measure of accuracy)
continuing or terminating the iterative adjustment of the one or more of the plurality of learnable parameters based on the calculated accuracy (mental process - continuing or terminating the iterative adjustment based on the calculated accuracy may be performed mentally by a user observing/analyzing the calculated accuracy and accordingly using judgment/evaluation to determine whether to continue or terminate the adjustment)
Step 2A Prong 2 & Step 2B:
[…] based on a pre-trained model […] included in the machine learning model (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using a pre-trained model without significantly more)
[…] associated with the machine learning model […] (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of a machine learning model without significantly more)
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the
abstract idea into practical application because it does not impose any meaningful limits on practicing
the abstract idea, as discussed above in the rejection of claim 1. The claim does not include additional
elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 5:
Step 2A Prong 1: See the rejection of Claim 4 above, which Claim 5 depends on.
Step 2A Prong 2 & Step 2B:
wherein each image-text pair included in the first and second pluralities of image-text pairs includes an image representing an architectural space and a ground truth human-centric evaluation criterion associated with the image (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that each image-text pair included in the first and second pluralities of image-text pairs includes an image representing an architectural space and a ground truth human-centric evaluation criterion associated with the image does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h))
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the
abstract idea into practical application because it does not impose any meaningful limits on practicing
the abstract idea, as discussed above in the rejection of claim 4. The claim does not include additional
elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 6:
Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 6 depends on.
Step 2A Prong 2 & Step 2B:
displaying […] simultaneous visual representations of a plurality of architectural spaces and a plurality of alignment scores associated with the plurality of architectural spaces (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g) - MPEP 2106.05(g) indicates that merely outputting the results of an abstract analysis is insignificant extra-solution activity and does not amount to an inventive concept. Here, the displaying merely outputs the results of the preceding analysis to a user. Thereby, this additional limitation does not amount to significantly more than the judicial exception.)
[…], via the graphical user interface, […] (recited at a high-level of generality (i.e., a graphical user interface, a generic processor, computer-readable storage medium, a communication interface, a user interface and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components)
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the
abstract idea into practical application because it does not impose any meaningful limits on practicing
the abstract idea, as discussed above in the rejection of claim 1. The claim does not include additional
elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 7:
Step 2A Prong 1: See the rejection of Claim 6 above, which Claim 7 depends on.
Step 2A Prong 2 & Step 2B:
wherein the simultaneous visual representations of the plurality of architectural spaces are arranged […] based on the plurality of alignment scores associated with the plurality of architectural spaces (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the simultaneous visual representations of the plurality of architectural spaces are arranged based on the plurality of alignment scores associated with the plurality of architectural spaces does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h))
[…]on the graphical user interface […] (recited at a high-level of generality (i.e., a graphical user interface, a generic processor, computer-readable storage medium, a communication interface, a user interface and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components)
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the
abstract idea into practical application because it does not impose any meaningful limits on practicing
the abstract idea, as discussed above in the rejection of claim 6. The claim does not include additional
elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 8:
Step 1: Claim 8 is a non-transitory type claim. Therefore, Claims 8-14 fall within one of the four statutory
categories (i.e., process, machine, manufacture, or composition of matter).
2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance
of the limitation in the mind but for the recitation of generic computer components, then it falls within
the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable
interpretation, covers performance of the limitation by mathematical calculation but for the recitation
of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract
ideas.
generating a textual prompt, wherein the textual prompt includes a human-centric evaluation criterion (mental process – generating a textual prompt that includes a human-centric evaluation criterion may be performed mentally or using pen and paper by a user observing/analyzing the human-centric evaluation criterion and accordingly using judgment/evaluation to create or write the textual prompt including the human-centric evaluation criterion)
generating […] an alignment score based on the textual prompt and a two-dimensional (2D) input image, wherein the 2D input image is a visual representation of an architectural space, and the alignment score is a quantitative measure of how accurately the textual prompt describes the 2D input image based on the human-centric evaluation criterion (mental process – generating an alignment score based on the textual prompt and a two-dimensional (2D) input image may be performed mentally or using pen and paper by a user observing/analyzing the textual prompt and the 2D input image and accordingly using judgment/evaluation to determine how accurately the textual prompt describes the 2D input image based on the human-centric evaluation criterion)
assigning the human-centric evaluation criterion and the alignment score to the 2D input image (mental process – assigning the human-centric evaluation criterion and the alignment score to the 2D input image may be performed mentally or using pen and paper by a user identifying the applicable criterion and score and associating them with the corresponding 2D input image)
Step 2A Prong 2: This judicial exception is not integrated into a practical application.
[…] one or more processors […] (recited at a high-level of generality (i.e., a graphical user interface, a generic processor, computer-readable storage medium, a communication interface, a user interface and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components)
[…], via a machine learning model, […] (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using a machine learning model without significantly more)
displaying […] one or more of the 2D input image, the human-centric evaluation criterion, and the alignment score (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g))
[…], via a graphical user interface, […] (recited at a high-level of generality (i.e., a graphical user interface, a generic processor, computer-readable storage medium, a communication interface, a user interface and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components)
Step 2B: The claim does not include additional elements considered individually and in combination that
are sufficient to amount to significantly more than the judicial exception.
[…], via a machine learning model, […] (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using a machine learning model without significantly more)
displaying […] one or more of the 2D input image, the human-centric evaluation criterion, and the alignment score (MPEP 2106.05(g) indicates that merely outputting the results of an abstract analysis is insignificant extra-solution activity and does not amount to an inventive concept. Here, the displaying merely outputs the results of the preceding analysis to a user. Thereby, this additional limitation does not amount to significantly more than the judicial exception.)
[…], via a graphical user interface, […] (recited at a high-level of generality (i.e., a graphical user interface, a generic processor, computer-readable storage medium, a communication interface, a user interface and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components)
For the reasons above, Claim 8 is rejected as being directed to an abstract idea without
significantly more. This rejection applies equally to dependent claims 8 - 14. The additional limitations of the dependent claims are addressed below.
Regarding Claim 9:
Step 2A Prong 1: See the rejection of Claim 8 above, which Claim 9 depends on.
Step 2A Prong 2 & Step 2B:
wherein the human-centric evaluation criterion includes one of the terms “social,” “isolating,” “tranquil,” “distracting,” “inspirational,” or “boring.” (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the human-centric evaluation criterion includes one of the terms “social,” “isolating,” “tranquil,” “distracting,” “inspirational,” or “boring” does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h))
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the
abstract idea into practical application because it does not impose any meaningful limits on practicing
the abstract idea, as discussed above in the rejection of claim 8. The claim does not include additional
elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 10:
Step 2A Prong 1: See the rejection of Claim 8 above, which Claim 10 depends on.
Step 2A Prong 2 & Step 2B:
wherein the 2D input image is one of a photograph of an architectural space, a 2D rendering of an architectural space, or a still image representing a single frame included in a video recording of an architectural space (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the 2D input image is one of a photograph of an architectural space, a 2D rendering of an architectural space, or a still image representing a single frame included in a video recording of an architectural space does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h))
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the
abstract idea into practical application because it does not impose any meaningful limits on practicing
the abstract idea, as discussed above in the rejection of claim 8. The claim does not include additional
elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 11:
Step 2A Prong 1: See the rejection of Claim 8 above, which Claim 11 depends on.
initializing, […], a plurality of learnable parameters […] (mental process - initializing the plurality of learnable parameters may be performed mentally or using pen and paper by a user observing/analyzing relevant information and accordingly using judgment/evaluation to determine and assign initial parameter values based on said analysis)
iteratively adjusting one or more of the plurality of learnable parameters based on a first plurality of alignment scores generated for a first plurality of image-text pairs included in a training data set (mental process - iteratively adjusting one or more of the plurality of learnable parameters based on the first plurality of alignment scores may be performed mentally or using pen and paper by a user observing/analyzing the alignment scores and accordingly using judgment/evaluation to repeatedly modify the learnable parameters based on said analysis)
calculating an accuracy […] based on a second plurality of alignment scores generated for a second plurality of image-text pairs included in a testing data set (mathematical concept - calculating the accuracy based on the second plurality of alignment scores involves performing mathematical calculations on numerical alignment score values to determine a quantitative measure of accuracy)
continuing or terminating the iterative adjustment of the one or more of the plurality of learnable parameters based on the calculated accuracy (mental process - continuing or terminating the iterative adjustment based on the calculated accuracy may be performed mentally by a user observing/analyzing the calculated accuracy and accordingly using judgment/evaluation to determine whether to continue or terminate the adjustment)
Step 2A Prong 2 & Step 2B:
[…] based on a pre-trained model […] included in the machine learning model (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using a pre-trained model without significantly more)
[…] associated with the machine learning model […] (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of a machine learning model without significantly more)
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the
abstract idea into practical application because it does not impose any meaningful limits on practicing
the abstract idea, as discussed above in the rejection of claim 8. The claim does not include additional
elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 12:
Step 2A Prong 1: See the rejection of Claim 11 above, which Claim 12 depends on.
Step 2A Prong 2 & Step 2B:
wherein each image-text pair included in the first and second pluralities of image-text pairs includes an image representing an architectural space and a ground truth human-centric evaluation criterion associated with the image (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that each image-text pair included in the first and second pluralities of image-text pairs includes an image representing an architectural space and a ground truth human-centric evaluation criterion associated with the image does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h))
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the
abstract idea into practical application because it does not impose any meaningful limits on practicing
the abstract idea, as discussed above in the rejection of claim 11. The claim does not include additional
elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 13:
Step 2A Prong 1: See the rejection of Claim 8 above, which Claim 13 depends on.
Step 2A Prong 2 & Step 2B:
displaying […] simultaneous visual representations of a plurality of architectural spaces and a plurality of alignment scores associated with the plurality of architectural spaces (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g) - MPEP 2106.05(g) indicates that merely outputting the results of an abstract analysis is insignificant extra-solution activity and does not amount to an inventive concept. Here, the displaying merely outputs the results of the preceding analysis to a user. Thereby, this additional limitation does not amount to significantly more than the judicial exception.)
[…], via the graphical user interface, […] (recited at a high-level of generality (i.e., a graphical user interface, a generic processor, computer-readable storage medium, a communication interface, a user interface and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components)
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the
abstract idea into practical application because it does not impose any meaningful limits on practicing
the abstract idea, as discussed above in the rejection of claim 8. The claim does not include additional
elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 14:
Step 2A Prong 1: See the rejection of Claim 13 above, which Claim 14 depends on.
Step 2A Prong 2 & Step 2B:
wherein the simultaneous visual representations of the plurality of architectural spaces are arranged […] based on the plurality of alignment scores associated with the plurality of architectural spaces (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the simultaneous visual representations of the plurality of architectural spaces are arranged based on the plurality of alignment scores associated with the plurality of architectural spaces does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h))
[…]on the graphical user interface […] (recited at a high-level of generality (i.e., a graphical user interface, a generic processor, computer-readable storage medium, a communication interface, a user interface and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components)
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the
abstract idea into practical application because it does not impose any meaningful limits on practicing
the abstract idea, as discussed above in the rejection of claim 13. The claim does not include additional
elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 15:
Step 1: Claim 15 is a system type claim. Therefore, Claims 15-20 fall within one of the four statutory
categories (i.e., process, machine, manufacture, or composition of matter).
2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance
of the limitation in the mind but for the recitation of generic computer components, then it falls within
the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable
interpretation, covers performance of the limitation by mathematical calculation but for the recitation
of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract
ideas.
generate a textual prompt, wherein the textual prompt includes a human-centric evaluation criterion (mental process – generating a textual prompt that includes a human-centric evaluation criterion may be performed mentally or using pen and paper by a user observing/analyzing the human-centric evaluation criterion and accordingly using judgment/evaluation to create or write the textual prompt including the human-centric evaluation criterion)
generate […] an alignment score based on the textual prompt and a two-dimensional (2D) input image, wherein the 2D input image is a visual representation of an architectural space, and the alignment score is a quantitative measure of how accurately the textual prompt describes the 2D input image based on the human-centric evaluation criterion (mental process – generating an alignment score based on the textual prompt and a two-dimensional (2D) input image may be performed mentally or using pen and paper by a user observing/analyzing the textual prompt and the 2D input image and accordingly using judgment/evaluation to determine how accurately the textual prompt describes the 2D input image based on the human-centric evaluation criterion)
assign the human-centric evaluation criterion and the alignment score to the 2D input image (mental process – assigning the human-centric evaluation criterion and the alignment score to the 2D input image may be performed mentally or using pen and paper by a user identifying the applicable criterion and score and associating them with the corresponding 2D input image)
Step 2A Prong 2: This judicial exception is not integrated into a practical application.
one or more memories […]; one or more processors [...](recited at a high-level of generality (i.e., a graphical user interface, a generic processor, computer-readable storage medium, a communication interface, a user interface and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components)
[…], via a machine learning model, […] (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using a machine learning model without significantly more)
display […] one or more of the 2D input image, the human-centric evaluation criterion, and the alignment score (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g))
[…], via a graphical user interface, […] (recited at a high-level of generality (i.e., a graphical user interface, a generic processor, computer-readable storage medium, a communication interface, a user interface and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components)
Step 2B: The claim does not include additional elements considered individually and in combination that
are sufficient to amount to significantly more than the judicial exception.
[…], via a machine learning model, […] (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using a machine learning model without significantly more)
display […] one or more of the 2D input image, the human-centric evaluation criterion, and the alignment score (MPEP 2106.05(g) indicates that merely outputting the results of an abstract analysis is insignificant extra-solution activity and does not amount to an inventive concept. Here, the displaying merely outputs the results of the preceding analysis to a user. Thereby, this additional limitation does not amount to significantly more than the judicial exception.)
[…], via a graphical user interface, […] (recited at a high-level of generality (i.e., a graphical user interface, a generic processor, computer-readable storage medium, a communication interface, a user interface and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components)
For the reasons above, Claim 15 is rejected as being directed to an abstract idea without
significantly more. This rejection applies equally to dependent claims 15 - 20. The additional limitations of the dependent claims are addressed below.
Regarding Claim 16:
Step 2A Prong 1: See the rejection of Claim 15 above, which Claim 16 depends on.
Step 2A Prong 2 & Step 2B:
wherein the human-centric evaluation criterion includes one of the terms “social,” “isolating,” “tranquil,” “distracting,” “inspirational,” or “boring.” (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the human-centric evaluation criterion includes one of the terms “social,” “isolating,” “tranquil,” “distracting,” “inspirational,” or “boring” does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h))
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the
abstract idea into practical application because it does not impose any meaningful limits on practicing
the abstract idea, as discussed above in the rejection of claim 15. The claim does not include additional
elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 17:
Step 2A Prong 1: See the rejection of Claim 15 above, which Claim 17 depends on.
Step 2A Prong 2 & Step 2B:
wherein the 2D input image is one of a photograph of an architectural space, a 2D rendering of an architectural space, or a still image representing a single frame included in a video recording of an architectural space (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the 2D input image is one of a photograph of an architectural space, a 2D rendering of an architectural space, or a still image representing a single frame included in a video recording of an architectural space does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h))
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the
abstract idea into practical application because it does not impose any meaningful limits on practicing
the abstract idea, as discussed above in the rejection of claim 15. The claim does not include additional
elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 18:
Step 2A Prong 1: See the rejection of Claim 15 above, which Claim 18 depends on.
initialize, […], a plurality of learnable parameters […] (mental process - initializing the plurality of learnable parameters may be performed mentally or using pen and paper by a user observing/analyzing relevant information and accordingly using judgment/evaluation to determine and assign initial parameter values based on said analysis)
iteratively adjust one or more of the plurality of learnable parameters based on a first plurality of alignment scores generated for a first plurality of image-text pairs included in a training data set (mental process - iteratively adjusting one or more of the plurality of learnable parameters based on the first plurality of alignment scores may be performed mentally or using pen and paper by a user observing/analyzing the alignment scores and accordingly using judgment/evaluation to repeatedly modify the learnable parameters based on said analysis)
calculate an accuracy […] based on a second plurality of alignment scores generated for a second plurality of image-text pairs included in a testing data set (mathematical concept - calculating the accuracy based on the second plurality of alignment scores involves performing mathematical calculations on numerical alignment score values to determine a quantitative measure of accuracy)
continue or terminate the iterative adjustment of the one or more of the plurality of learnable parameters based on the calculated accuracy (mental process - continuing or terminating the iterative adjustment based on the calculated accuracy may be performed mentally by a user observing/analyzing the calculated accuracy and accordingly using judgment/evaluation to determine whether to continue or terminate the adjustment)
Step 2A Prong 2 & Step 2B:
[…] based on a pre-trained model […] included in the machine learning model (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using a pre-trained model without significantly more)
[…] associated with the machine learning model […] (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of a machine learning model without significantly more)
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the
abstract idea into practical application because it does not impose any meaningful limits on practicing
the abstract idea, as discussed above in the rejection of claim 15. The claim does not include additional
elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 19:
Step 2A Prong 1: See the rejection of Claim 18 above, which Claim 19 depends on.
Step 2A Prong 2 & Step 2B:
wherein each image-text pair included in the first and second pluralities of image-text pairs includes an image representing an architectural space and a ground truth human-centric evaluation criterion associated with the image (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that each image-text pair included in the first and second pluralities of image-text pairs includes an image representing an architectural space and a ground truth human-centric evaluation criterion associated with the image does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h))
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the
abstract idea into practical application because it does not impose any meaningful limits on practicing
the abstract idea, as discussed above in the rejection of claim 18. The claim does not include additional
elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 20:
Step 2A Prong 1: See the rejection of Claim 15 above, which Claim 20 depends on.
Step 2A Prong 2 & Step 2B:
display […] simultaneous visual representations of a plurality of architectural spaces and a plurality of alignment scores associated with the plurality of architectural spaces (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g) - MPEP 2106.05(g) indicates that merely outputting the results of an abstract analysis is insignificant extra-solution activity and does not amount to an inventive concept. Here, the displaying merely outputs the results of the preceding analysis to a user. Thereby, this additional limitation does not amount to significantly more than the judicial exception.)
[…], via the graphical user interface, […] (recited at a high-level of generality (i.e., a graphical user interface, a generic processor, computer-readable storage medium, a communication interface, a user interface and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components)
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the
abstract idea into practical application because it does not impose any meaningful limits on practicing
the abstract idea, as discussed above in the rejection of claim 15. The claim does not include additional
elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-3, 6-10, 13-17, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Radford et al. (hereinafter Radford, a non-patent literature reference titled “Learning transferable visual models from natural language supervision.”), in view of Yao et al. (hereinafter Yao, a non-patent literature reference titled “A human-machine adversarial scoring framework for urban perception assessment using street-view images.”), and further in view of Hoser et al. (hereinafter Hoser, a non-patent literature reference titled “Analysis of text-to-image artificial intelligence Systems in terms of contribution to interior coloring.”).
Regarding Claim 1, Radford teaches:
generating a textual prompt, wherein the textual prompt includes […] (Radford, Page 4 – Section 2.5, “For each dataset, we use the names of all the classes in the dataset as the set of potential text pairings and predict the most probable (image, text) pair according to CLIP. We additionally experiment with providing CLIP with text prompts to help specify the task as well as ensembling multiple of these templates in order to boost performance”, thus generating a textual prompt, wherein the textual prompt includes […] is disclosed, because Radford teaches providing CLIP with text prompts to specify a task and using the names of classes as potential text pairings for comparison with an image. Radford’s text prompts correspond to the textual prompt because they are textual inputs provided to CLIP to define the task to be performed. Radford’s class names correspond to information included in the textual prompt because the class names identify the visual concepts or descriptions against which the input image is evaluated)
generating, via a machine learning model, an alignment score based on the textual prompt and a two-dimensional (2D) input image, wherein the 2D input image […], and the alignment score is a quantitative measure of how accurately the textual prompt describes the 2D input image […] (Radford, Page 3 – Section 2.2, “Given a batch of N (image, text) pairs, CLIP is trained to predict which of the N × N possible (image, text) pairings across a batch actually occurred. To do this, CLIP learns a multi-modal embedding space by jointly training an image encoder and text encoder to maximize the cosine similarity of the image and text embeddings of the N real pairs in the batch while minimizing the cosine similarity of the embeddings of the N2 −N incorrect pairings. We optimize a symmetric cross entropy loss over these similarity scores”, & Page 4 – Section 2.5, “For each dataset, we use the names of all the classes in the dataset as the set of potential text pairings and predict the most probable (image, text) pair according to CLIP”, thus generating, via a machine learning model, an alignment score based on the textual prompt and a two-dimensional (2D) input image, wherein the 2D input image […], and the alignment score is a quantitative measure of how accurately the textual prompt describes the 2D input image […] is disclosed, because Radford teaches that CLIP uses an image encoder and a text encoder to evaluate image-text pairings and to generate similarity scores for those pairings. Radford’s image encoder and text encoder together correspond to the machine learning model because they process the image and text inputs and produce a comparison result. Radford’s image and text pairings correspond to the textual prompt and the 2D input image because CLIP operates on paired text and image inputs. Radford’s cosine similarity and similarity scores correspond to the alignment score because they provide a numerical indication of the degree of correspondence between the text embedding and the image embedding. Radford’s prediction of the most probable image-text pair further shows that the score reflects how well the text describes the image, with a higher similarity indicating a closer match between the textual prompt and the 2D input image)
[…] and the alignment score to the 2D input image (Radford, Page 3 – Section 2.2, “Given a batch of N (image, text) pairs, CLIP is trained to predict which of the N × N possible (image, text) pairings across a batch actually occurred. To do this, CLIP learns a multi-modal embedding space by jointly training an image encoder and text encoder to maximize the cosine similarity of the image and text embeddings of the N real pairs in the batch while minimizing the cosine similarity of the embeddings of the N2 −N incorrect pairings. We optimize a symmetric cross entropy loss over these similarity scores”, thus […] and the alignment score to the 2D input image is disclosed, because Radford calculates a similarity score for an image-text pairing. The similarity score corresponds to the alignment score associated with the input image because it quantifies the correspondence between that image and the paired text)
Radford does not explicitly teach […] includes a human-centric evaluation criterion, […] based on the human-centric evaluation criterion, assigning the human-centric evaluation criterion […], displaying, via a graphical user interface, one or more of [… 2D input image…], the human-centric evaluation criterion, and […alignment score…], and […] is a visual representation of an architectural space […].
However, Yao teaches:
[…] includes a human-centric evaluation criterion (Yao, Page 1 – Section 1, “Urban perceptions, which are the psychological feelings held by residents about an urban locale (Tuan 2013, Ordonez and Berg 2014), provide an important basis for understanding the ways in which urban environments interact with public mental health”, & Page 5 – Section 2.2.2, “This study focuses on six categories of urban perceptions: wealthy, safety, lively, beautiful, boring and depressing, as in Place Pulse 2.0 (Dubey et al. 2016). These six perception datasets from previous studies are assembled through training on mostly western urban street scenery with CNNs and annotators’ votes on pairwise image comparisons (Dubey et al. 2016)”, thus […] includes a human-centric evaluation criterion is disclosed, because Yao teaches evaluating urban environments according to human perceptions that represent psychological feelings held by residents about an urban locale. Yao further identifies specific perception categories, including wealthy, safety, lively, beautiful, boring, and depressing. Yao’s urban perceptions correspond to the human-centric evaluation criterion because they represent human psychological responses to the visual environment. Yao’s categories of wealthy, safety, lively, beautiful, boring, and depressing correspond to specific human-centric evaluation criteria because each category provides a human-perception attribute by which an image of an environment may be evaluated)
[…] based on the human-centric evaluation criterion (Yao, Page 5 – Section 2.2.2, “This study focuses on six categories of urban perceptions: wealthy, safety, lively, beautiful, boring and depressing, as in Place Pulse 2.0 (Dubey et al. 2016). These six perception datasets from previous studies are assembled through training on mostly western urban street scenery with CNNs and annotators’ votes on pairwise image comparisons (Dubey et al. 2016). Our proposed framework, however, collects real score annotations on each image from volunteers”, & Page 5 – Section 2.2.2, “Our human-machine adversarial scoring module processed the FCN sematic segmentation result and the city-scale human perceptions from local volunteers. Furthermore, the module mined the relationship between the visual scenery and perception directly to expedite image classification according to human perception”, thus […] based on the human-centric evaluation criterion is disclosed, because Yao teaches scoring images according to human-perception categories, including wealthy, safety, lively, beautiful, boring, and depressing. Yao’s perception categories correspond to the human-centric evaluation criterion, and its scoring module relates visual scenery to those perceptions to classify the image according to human perception)
assigning the human-centric evaluation criterion […] (Yao, Page 5 – Section 2.2.2, “Our human-machine adversarial scoring module processed the FCN sematic segmentation result and the city-scale human perceptions from local volunteers. Furthermore, the module mined the relationship between the visual scenery and perception directly to expedite image classification according to human perception. We use RF algorithm to determine the final image classification, and therefore the results are subject to RF fitting and limited to one perception label per image”, thus assigning the human-centric evaluation criterion […] is disclosed, because Yao teaches classifying each image according to human perception and assigning one perception label to each image. Yao’s perception label corresponds to the human-centric evaluation criterion because it identifies the human-perception category associated with the image)
displaying, via a graphical user interface, one or more of [… 2D input image…], the human-centric evaluation criterion, and […alignment score…] (Yao, Page 6 – Section 2.2.2, “Figure 3 illustrates the process of human-machine adversarial scoring. Volunteers score a displayed street-view image in terms of the six types of perception in a range of 1–100 with 0 being the lowest and 100 being the highest level of a perception”, & Page 6 – Section 2.2.2.1, “Once a user has scored the first 50 photos, the scoring software establishes a random forest set to fit the scoring process. Then, as users have rated subsequent photos, the software offers a recommendation score based on the rules learned from the previous user rating actions”, thus Yao teaches scoring software that presents a displayed street-view image to a user for evaluation according to six types of human perception and provides a recommendation score during that process. Yao’s scoring software corresponds to the graphical user interface, the displayed street-view image corresponds to the 2D input image, the six perception types correspond to the human-centric evaluation criterion, and the recommendation score corresponds to score information provided through the interface)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine Radford with Yao by incorporating Yao’s human-perception evaluation criteria into Radford’s CLIP image-text matching framework. Radford teaches using natural language prompts and class names to specify a task and generating similarity scores that quantify the correspondence between image and text inputs. Yao teaches evaluating visual environments according to human-perception categories and recognizes that traditional evaluation of human perceptions is difficult, costly, and time consuming, stating that a framework that can boost work efficiency is needed to optimize the urban perception assessment process. Therefore, a POSITA would have been motivated to use Yao’s human-perception categories as the class information or textual concepts supplied to Radford’s CLIP so that Radford’s image-text similarity technique could efficiently evaluate images according to human perception. Such a combination would have provided an automated and more efficient technique for quantitatively determining how well an image corresponds to a human-perception criterion, thereby addressing Yao’s identified need to improve the efficiency of urban perception assessment (Yao, Page 1 – Section 1, “Traditionally, the evaluation of human perceptions towards their visual surroundings remains difficult due to the lack of high-throughput methods, inadequate sample problems and being restricted to inter views and questionnaires (Hannay 1983, Halpern 1995, Kabisch et al. 2015, Dadvand et al. 2016). Given the costly and time-consuming nature of these investigation methods, a framework that can boost work efficiency is needed to optimize the urban perception assessment process”)
Radford combined with Yao does not explicitly teach […] is a visual representation of an architectural space […].
However, Hoser teaches:
[…] is a visual representation of an architectural space […] (Hoser, Page 2 – Section 1, “Following this, RGB color codes were extracted from the cumulative 16 images obtained, and these 16 distinct images were employed to colorize the interior of a 3D model representing a preschool space. Consequently, the same set of 16 images, pertaining to the identical spatial context, was presented to a group of 62 experts consisting of specialized architects/interior architects, and preschool educators”, & Page 5 – Section 3, “Afterward, the colors obtained from the visuals were transferred to the 3D model as a material color (Figure 7). Thus, a total of 16 views of the same space were obtained from 4 different systems. These views are listed below (Figure 8)”, thus […] is a visual representation of an architectural space […] is disclosed, because Hoser teaches generating multiple image views of the same modeled preschool interior. The preschool interior corresponds to the architectural space, while the resulting images and views correspond to visual representations of that space)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to further combine Radford and Yao with Hoser by applying the combined image-text and human perception evaluation technique to visual representations of architectural spaces. Radford and Yao together teach evaluating images based on textual concepts and human perception criteria, while Hoser teaches applying artificial intelligence techniques to interior architectural spaces and evaluating images of those spaces according to user-oriented criteria. Therefore, a POSITA would have been motivated to use Hoser’s architectural space images as the visual inputs evaluated by the Radford-Yao technique so that the depicted architectural spaces could be quantitatively assessed according to human centric criteria, thereby providing a more systematic and efficient technique for evaluating how users perceive visual representations of architectural and interior spaces and facilitating the application of artificial-intelligence-based evaluation to architecture and interior design (Hoser, Page 2 – Section 1, “artificial intelligence-supported data processing systems, transitioning from text to image, can contribute to interior spatial coloring in a scientifically accurate manner, thereby enriching the fields of architecture and interior design”, & “the same set of 16 images, pertaining to the identical spatial context, was presented to a group of 62 experts consisting of specialized architects/interior architects, and preschool educators. Through a survey employing Likert-type questions, these experts were prompted to assess the volumes in terms of “academic instruction” and “entertainment””)
Regarding Claim 2, Radford and Yao combined with Hoser teaches all the limitations of claim 1 as cited above and Yao further teaches:
wherein the human-centric evaluation criterion includes one of the terms “social,” “isolating,” “tranquil,” “distracting,” “inspirational,” or “boring” (Yao, Page 5 – Section 2.2.2., “This study focuses on six categories of urban perceptions: wealthy, safety, lively, beautiful, boring and depressing, as in Place Pulse 2.0 (Dubey et al. 2016). These six perception datasets from previous studies are assembled through training on mostly western urban street scenery with CNNs and annotators’ votes on pairwise image comparisons (Dubey et al. 2016)”, thus thus wherein the human-centric evaluation criterion includes one of the terms “social,” “isolating,” “tranquil,” “distracting,” “inspirational,” or “boring” is disclosed, because Yao identifies “boring” as one of its categories of human perception. Yao’s “boring” perception category corresponds to the recited human-centric evaluation criterion because it is a human-perception attribute used to evaluate the depicted environment)
Regarding Claim 3, Radford and Yao combined with Hoser teaches all the limitations of claim 1 as cited above and Hoser further teaches:
wherein […the 2D input image…] is one of a photograph of an architectural space, a 2D rendering of an architectural space, or a still image representing a single frame included in a video recording of an architectural space (Hoser, Hoser, Page 5 – Section 3.1, “The survey study includes a mixed group of graduates and academics in the fields of pre-school teaching and architecture/interior architecture. The appropriateness of the color schemes in the images in the contexts of ‘education’ and ‘entertainment’ for a pre-school educational space was measured using the ‘relational screening model’”, & Page 5 – Section 3, “Afterward, the colors obtained from the visuals were transferred to the 3D model as a material color (Figure 7). Thus, a total of 16 views of the same space were obtained from 4 different systems. These views are listed below (Figure 8)”, thus wherein […the 2D input image…] is one of a photograph of an architectural space, a 2D rendering of an architectural space, or a still image representing a single frame included in a video recording of an architectural space is disclosed, because Hoser teaches images representing a pre-school educational architectural space and multiple views obtained from a 3D model of that space. Hoser’s pre-school educational space corresponds to the architectural space, and the resulting image views of the 3D-modeled space correspond to 2D renderings of the architectural space)
Regarding Claim 6, Radford and Yao combined with Hoser teaches all the limitations of claim 1 as cited above and Yao further teaches:
displaying, via the graphical user interface, […] (Yao, Page 6 – Section 2.2.2, “Volunteers score a displayed street-view image in terms of the six types of perception in a range of 1–100 with 0 being the lowest and 100 being the highest level of a perception”, & Page 6 – Section 2.2.2.1, “Then, as users have rated subsequent photos, the software offers a recommendation score based on the rules learned from the previous user rating actions”, thus displaying, via the graphical user interface, […] is disclosed because Yao teaches user-facing scoring software that displays images for evaluation and provides corresponding score information. Yao’s scoring software corresponds to the graphical user interface)
Hoser further teaches:
[…] simultaneous visual representations of a plurality of architectural spaces […] associated with the plurality of architectural spaces (Hoser, Page 5 – Section 3, “Afterward, the colors obtained from the visuals were transferred to the 3D model as a material color (Figure 7). Thus, a total of 16 views of the same space were obtained from 4 different systems. These views are listed below (Figure 8)”, & Page 5 – Section 3.1, “The appropriateness of the color schemes in the images in the contexts of “education” and “entertainment” for a pre-school educational space was measured using the ‘relational screening model’”, thus […] simultaneous visual representations […] associated with the plurality of architectural spaces is disclosed because Hoser presents multiple visual representations of an architectural interior together and evaluates the respective images in relation to the depicted architectural space. Hoser’s displayed image views correspond to the simultaneous visual representations, and the depicted pre-school interior corresponds to the architectural space associated with those representations)
Radford further teaches:
[…] and a plurality of alignment scores […] (Page 3 – Section 2.2, “CLIP learns a multi-modal embedding space by jointly training an image encoder and text encoder to maximize the cosine similarity of the image and text embeddings of the N real pairs while minimizing the cosine similarity of the embeddings of the N² − N incorrect pairings. We optimize a symmetric cross entropy loss over these similarity scores”, thus […] and a plurality of alignment scores […] is disclosed because Radford generates similarity scores for multiple image-text pairings. Radford’s similarity scores correspond to the plurality of alignment scores because each score quantifies the correspondence between an image representation and its textual representation)
Regarding Claim 7, Radford and Yao combined with Hoser teaches all the limitations of claim 6 as cited above and Hoser further teaches:
wherein the simultaneous visual representations of the plurality of architectural spaces […] associated with the plurality of architectural spaces (Hoser, Page 5 – Section 3, “Afterward, the colors obtained from the visuals were transferred to the 3D model as a material color (Figure 7). Thus, a total of 16 views of the same space were obtained from 4 different systems. These views are listed below (Figure 8)”, thus wherein the simultaneous visual representations […] associated with the plurality of architectural spaces is disclosed because Hoser teaches multiple visual representations of an architectural interior. Hoser’s image views correspond to the visual representations associated with the depicted architectural space)
Yao further teaches:
[…] are arranged on the graphical user interface […] (Yao, Page 18 – Discussion, “Future studies may revise the method, such as sorting the images according to existing residential segregation phenomenon. A revised method may sort images into various neighborhood contexts in advance and make batches of display images highly consistent. The revised method can keep records of image displaying orders”, thus […] are arranged on the graphical user interface […] is disclosed because Yao teaches sorting images and controlling their displaying order. Yao’s sorting and displaying order corresponds to arranging the visual representations on the graphical user interface)
Radford further teaches:
[…] based on the plurality of alignment scores […] (Radford, Page 3 – Section 2.2, “CLIP learns a multi-modal embedding space by jointly training an image encoder and text encoder to maximize the cosine similarity of the image and text embeddings of the N real pairs in the batch while minimizing the cosine similarity of the embeddings of the N² − N incorrect pairings. We optimize a symmetric cross entropy loss over these similarity scores”, & Page 4 – Section 2.5, “For each dataset, we use the names of all the classes in the dataset as the set of potential text pairings and predict the most probable (image, text) pair according to CLIP”, thus […] based on the plurality of alignment scores […] is disclosed because Radford teaches generating similarity scores for multiple image-text pairings and using the scores to determine the relative correspondence of the pairings. Radford’s similarity scores correspond to the plurality of alignment scores)
Regarding Claim 8, Radford teaches:
generating a textual prompt, wherein the textual prompt includes […] (Radford, Page 4 – Section 2.5, “For each dataset, we use the names of all the classes in the dataset as the set of potential text pairings and predict the most probable (image, text) pair according to CLIP. We addi tionally experiment with providing CLIP with text prompts to help specify the task as well as ensembling multiple of these templates in order to boost performance”, thus generating a textual prompt, wherein the textual prompt includes […] is disclosed, because Radford teaches providing CLIP with text prompts to specify a task and using the names of classes as potential text pairings for comparison with an image. Radford’s text prompts correspond to the textual prompt because they are textual inputs provided to CLIP to define the task to be performed. Radford’s class names correspond to information included in the textual prompt because the class names identify the visual concepts or descriptions against which the input image is evaluated)
generating, via a machine learning model, an alignment score based on the textual prompt and a two-dimensional (2D) input image, wherein the 2D input image […], and the alignment score is a quantitative measure of how accurately the textual prompt describes the 2D input image […] (Radford, Page 3 – Section 2.2, “Given a batch of N (image, text) pairs, CLIP is trained to predict which of the N × N possible (image, text) pairings across a batch actually occurred. To do this, CLIP learns a multi-modal embedding space by jointly training an image encoder and text encoder to maximize the cosine similarity of the image and text embeddings of the N real pairs in the batch while minimizing the cosine similarity of the embeddings of the N2 −N incorrect pairings. We optimize a symmetric cross entropy loss over these similarity scores”, & Page 4 – Section 2.5, “For each dataset, we use the names of all the classes in the dataset as the set of potential text pairings and predict the most probable (image, text) pair according to CLIP”, thus generating, via a machine learning model, an alignment score based on the textual prompt and a two-dimensional (2D) input image, wherein the 2D input image […], and the alignment score is a quantitative measure of how accurately the textual prompt describes the 2D input image […] is disclosed, because Radford teaches that CLIP uses an image encoder and a text encoder to evaluate image-text pairings and to generate similarity scores for those pairings. Radford’s image encoder and text encoder together correspond to the machine learning model because they process the image and text inputs and produce a comparison result. Radford’s image and text pairings correspond to the textual prompt and the 2D input image because CLIP operates on paired text and image inputs. Radford’s cosine similarity and similarity scores correspond to the alignment score because they provide a numerical indication of the degree of correspondence between the text embedding and the image embedding. Radford’s prediction of the most probable image-text pair further shows that the score reflects how well the text describes the image, with a higher similarity indicating a closer match between the textual prompt and the 2D input image)
[…] and the alignment score to the 2D input image (Radford, Page 3 – Section 2.2, “Given a batch of N (image, text) pairs, CLIP is trained to predict which of the N × N possible (image, text) pairings across a batch actually occurred. To do this, CLIP learns a multi-modal embedding space by jointly training an image encoder and text encoder to maximize the cosine similarity of the image and text embeddings of the N real pairs in the batch while minimizing the cosine similarity of the embeddings of the N2 −N incorrect pairings. We optimize a symmetric cross entropy loss over these similarity scores”, thus […] and the alignment score to the 2D input image is disclosed, because Radford calculates a similarity score for an image-text pairing. The similarity score corresponds to the alignment score associated with the input image because it quantifies the correspondence between that image and the paired text)
Radford does not explicitly teach […] includes a human-centric evaluation criterion, […] based on the human-centric evaluation criterion, assigning the human-centric evaluation criterion […], displaying, via a graphical user interface, one or more of [… 2D input image…], the human-centric evaluation criterion, and […alignment score…], and […] is a visual representation of an architectural space […].
However, Yao teaches:
[…] includes a human-centric evaluation criterion (Yao, Page 1 – Section 1, “Urban perceptions, which are the psychological feelings held by residents about an urban locale (Tuan 2013, Ordonez and Berg 2014), provide an important basis for understanding the ways in which urban environments interact with public mental health”, & Page 5 – Section 2.2.2, “This study focuses on six categories of urban perceptions: wealthy, safety, lively, beautiful, boring and depressing, as in Place Pulse 2.0 (Dubey et al. 2016). These six perception datasets from previous studies are assembled through training on mostly western urban street scenery with CNNs and annotators’ votes on pairwise image comparisons (Dubey et al. 2016)”, thus […] includes a human-centric evaluation criterion is disclosed, because Yao teaches evaluating urban environments according to human perceptions that represent psychological feelings held by residents about an urban locale. Yao further identifies specific perception categories, including wealthy, safety, lively, beautiful, boring, and depressing. Yao’s urban perceptions correspond to the human-centric evaluation criterion because they represent human psychological responses to the visual environment. Yao’s categories of wealthy, safety, lively, beautiful, boring, and depressing correspond to specific human-centric evaluation criteria because each category provides a human-perception attribute by which an image of an environment may be evaluated)
[…] based on the human-centric evaluation criterion (Yao, Page 5 – Section 2.2.2, “This study focuses on six categories of urban perceptions: wealthy, safety, lively, beautiful, boring and depressing, as in Place Pulse 2.0 (Dubey et al. 2016). These six perception datasets from previous studies are assembled through training on mostly western urban street scenery with CNNs and annotators’ votes on pairwise image comparisons (Dubey et al. 2016). Our proposed framework, however, collects real score annotations on each image from volunteers”, & Page 5 – Section 2.2.2, “Our human-machine adversarial scoring module processed the FCN sematic segmentation result and the city-scale human perceptions from local volunteers. Furthermore, the module mined the relationship between the visual scenery and perception directly to expedite image classification according to human perception”, thus […] based on the human-centric evaluation criterion is disclosed, because Yao teaches scoring images according to human-perception categories, including wealthy, safety, lively, beautiful, boring, and depressing. Yao’s perception categories correspond to the human-centric evaluation criterion, and its scoring module relates visual scenery to those perceptions to classify the image according to human perception)
assigning the human-centric evaluation criterion […] (Yao, Page 5 – Section 2.2.2, “Our human-machine adversarial scoring module processed the FCN sematic segmentation result and the city-scale human perceptions from local volunteers. Furthermore, the module mined the relationship between the visual scenery and perception directly to expedite image classification according to human perception. We use RF algorithm to determine the final image classification, and therefore the results are subject to RF fitting and limited to one perception label per image”, thus assigning the human-centric evaluation criterion […] is disclosed, because Yao teaches classifying each image according to human perception and assigning one perception label to each image. Yao’s perception label corresponds to the human-centric evaluation criterion because it identifies the human-perception category associated with the image)
displaying, via a graphical user interface, one or more of [… 2D input image…], the human-centric evaluation criterion, and […alignment score…] (Yao, Page 6 – Section 2.2.2, “Figure 3 illustrates the process of human-machine adversarial scoring. Volunteers score a displayed street-view image in terms of the six types of perception in a range of 1–100 with 0 being the lowest and 100 being the highest level of a perception”, & Page 6 – Section 2.2.2.1, “Once a user has scored the first 50 photos, the scoring software establishes a random forest set to fit the scoring process. Then, as users have rated subsequent photos, the software offers a recommendation score based on the rules learned from the previous user rating actions”, thus Yao teaches scoring software that presents a displayed street-view image to a user for evaluation according to six types of human perception and provides a recommendation score during that process. Yao’s scoring software corresponds to the graphical user interface, the displayed street-view image corresponds to the 2D input image, the six perception types correspond to the human-centric evaluation criterion, and the recommendation score corresponds to score information provided through the interface)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine Radford with Yao by incorporating Yao’s human-perception evaluation criteria into Radford’s CLIP image-text matching framework. Radford teaches using natural language prompts and class names to specify a task and generating similarity scores that quantify the correspondence between image and text inputs. Yao teaches evaluating visual environments according to human-perception categories and recognizes that traditional evaluation of human perceptions is difficult, costly, and time consuming, stating that a framework that can boost work efficiency is needed to optimize the urban perception assessment process. Therefore, a POSITA would have been motivated to use Yao’s human-perception categories as the class information or textual concepts supplied to Radford’s CLIP so that Radford’s image-text similarity technique could efficiently evaluate images according to human perception. Such a combination would have provided an automated and more efficient technique for quantitatively determining how well an image corresponds to a human-perception criterion, thereby addressing Yao’s identified need to improve the efficiency of urban perception assessment (Yao, Page 1 – Section 1, “Traditionally, the evaluation of human perceptions towards their visual surroundings remains difficult due to the lack of high-throughput methods, inadequate sample problems and being restricted to inter views and questionnaires (Hannay 1983, Halpern 1995, Kabisch et al. 2015, Dadvand et al. 2016). Given the costly and time-consuming nature of these investigation methods, a framework that can boost work efficiency is needed to optimize the urban perception assessment process”)
Radford combined with Yao does not explicitly teach […] is a visual representation of an architectural space […].
However, Hoser teaches:
[…] is a visual representation of an architectural space […] (Hoser, Page 2 – Section 1, “Following this, RGB color codes were extracted from the cumulative 16 images obtained, and these 16 distinct images were employed to colorize the interior of a 3D model representing a preschool space. Consequently, the same set of 16 images, pertaining to the identical spatial context, was presented to a group of 62 experts consisting of specialized architects/interior architects, and preschool educators”, & Page 5 – Section 3, “Afterward, the colors obtained from the visuals were transferred to the 3D model as a material color (Figure 7). Thus, a total of 16 views of the same space were obtained from 4 different systems. These views are listed below (Figure 8)”, thus […] is a visual representation of an architectural space […] is disclosed, because Hoser teaches generating multiple image views of the same modeled preschool interior. The preschool interior corresponds to the architectural space, while the resulting images and views correspond to visual representations of that space)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to further combine Radford and Yao with Hoser by applying the combined image-text and human perception evaluation technique to visual representations of architectural spaces. Radford and Yao together teach evaluating images based on textual concepts and human perception criteria, while Hoser teaches applying artificial intelligence techniques to interior architectural spaces and evaluating images of those spaces according to user-oriented criteria. Therefore, a POSITA would have been motivated to use Hoser’s architectural space images as the visual inputs evaluated by the Radford-Yao technique so that the depicted architectural spaces could be quantitatively assessed according to human centric criteria, thereby providing a more systematic and efficient technique for evaluating how users perceive visual representations of architectural and interior spaces and facilitating the application of artificial-intelligence-based evaluation to architecture and interior design (Hoser, Page 2 – Section 1, “artificial intelligence-supported data processing systems, transitioning from text to image, can contribute to interior spatial coloring in a scientifically accurate manner, thereby enriching the fields of architecture and interior design”, & “the same set of 16 images, pertaining to the identical spatial context, was presented to a group of 62 experts consisting of specialized architects/interior architects, and preschool educators. Through a survey employing Likert-type questions, these experts were prompted to assess the volumes in terms of “academic instruction” and “entertainment””)
Regarding Claim 9, Radford and Yao combined with Hoser teaches all the limitations of claim 8 as cited above and Yao further teaches:
wherein the human-centric evaluation criterion includes one of the terms “social,” “isolating,” “tranquil,” “distracting,” “inspirational,” or “boring” (Yao, Page 5 – Section 2.2.2., “This study focuses on six categories of urban perceptions: wealthy, safety, lively, beautiful, boring and depressing, as in Place Pulse 2.0 (Dubey et al. 2016). These six perception datasets from previous studies are assembled through training on mostly western urban street scenery with CNNs and annotators’ votes on pairwise image comparisons (Dubey et al. 2016)”, thus thus wherein the human-centric evaluation criterion includes one of the terms “social,” “isolating,” “tranquil,” “distracting,” “inspirational,” or “boring” is disclosed, because Yao identifies “boring” as one of its categories of human perception. Yao’s “boring” perception category corresponds to the recited human-centric evaluation criterion because it is a human-perception attribute used to evaluate the depicted environment)
Regarding Claim 10, Radford and Yao combined with Hoser teaches all the limitations of claim 8 as cited above and Hoser further teaches:
wherein […the 2D input image…] is one of a photograph of an architectural space, a 2D rendering of an architectural space, or a still image representing a single frame included in a video recording of an architectural space (Hoser, Hoser, Page 5 – Section 3.1, “The survey study includes a mixed group of graduates and academics in the fields of pre-school teaching and architecture/interior architecture. The appropriateness of the color schemes in the images in the contexts of ‘education’ and ‘entertainment’ for a pre-school educational space was measured using the ‘relational screening model’”, & Page 5 – Section 3, “Afterward, the colors obtained from the visuals were transferred to the 3D model as a material color (Figure 7). Thus, a total of 16 views of the same space were obtained from 4 different systems. These views are listed below (Figure 8)”, thus wherein […the 2D input image…] is one of a photograph of an architectural space, a 2D rendering of an architectural space, or a still image representing a single frame included in a video recording of an architectural space is disclosed, because Hoser teaches images representing a pre-school educational architectural space and multiple views obtained from a 3D model of that space. Hoser’s pre-school educational space corresponds to the architectural space, and the resulting image views of the 3D-modeled space correspond to 2D renderings of the architectural space)
Regarding Claim 13, Radford and Yao combined with Hoser teaches all the limitations of claim 8 as cited above and Yao further teaches:
displaying, via the graphical user interface, […] (Yao, Page 6 – Section 2.2.2, “Volunteers score a displayed street-view image in terms of the six types of perception in a range of 1–100 with 0 being the lowest and 100 being the highest level of a perception”, & Page 6 – Section 2.2.2.1, “Then, as users have rated subsequent photos, the software offers a recommendation score based on the rules learned from the previous user rating actions”, thus displaying, via the graphical user interface, […] is disclosed because Yao teaches user-facing scoring software that displays images for evaluation and provides corresponding score information. Yao’s scoring software corresponds to the graphical user interface)
Hoser further teaches:
[…] simultaneous visual representations of a plurality of architectural spaces […] associated with the plurality of architectural spaces (Hoser, Page 5 – Section 3, “Afterward, the colors obtained from the visuals were transferred to the 3D model as a material color (Figure 7). Thus, a total of 16 views of the same space were obtained from 4 different systems. These views are listed below (Figure 8)”, & Page 5 – Section 3.1, “The appropriateness of the color schemes in the images in the contexts of “education” and “entertainment” for a pre-school educational space was measured using the ‘relational screening model’”, thus […] simultaneous visual representations […] associated with the plurality of architectural spaces is disclosed because Hoser presents multiple visual representations of an architectural interior together and evaluates the respective images in relation to the depicted architectural space. Hoser’s displayed image views correspond to the simultaneous visual representations, and the depicted pre-school interior corresponds to the architectural space associated with those representations)
Radford further teaches:
[…] and a plurality of alignment scores […] (Page 3 – Section 2.2, “CLIP learns a multi-modal embedding space by jointly training an image encoder and text encoder to maximize the cosine similarity of the image and text embeddings of the N real pairs while minimizing the cosine similarity of the embeddings of the N² − N incorrect pairings. We optimize a symmetric cross entropy loss over these similarity scores”, thus […] and a plurality of alignment scores […] is disclosed because Radford generates similarity scores for multiple image-text pairings. Radford’s similarity scores correspond to the plurality of alignment scores because each score quantifies the correspondence between an image representation and its textual representation)
Regarding Claim 14, Radford and Yao combined with Hoser teaches all the limitations of claim 13 as cited above and Hoser further teaches:
wherein the simultaneous visual representations of the plurality of architectural spaces […] associated with the plurality of architectural spaces (Hoser, Page 5 – Section 3, “Afterward, the colors obtained from the visuals were transferred to the 3D model as a material color (Figure 7). Thus, a total of 16 views of the same space were obtained from 4 different systems. These views are listed below (Figure 8)”, thus wherein the simultaneous visual representations […] associated with the plurality of architectural spaces is disclosed because Hoser teaches multiple visual representations of an architectural interior. Hoser’s image views correspond to the visual representations associated with the depicted architectural space)
Yao further teaches:
[…] are arranged on the graphical user interface […] (Yao, Page 18 – Discussion, “Future studies may revise the method, such as sorting the images according to existing residential segregation phenomenon. A revised method may sort images into various neighborhood contexts in advance and make batches of display images highly consistent. The revised method can keep records of image displaying orders”, thus […] are arranged on the graphical user interface […] is disclosed because Yao teaches sorting images and controlling their displaying order. Yao’s sorting and displaying order corresponds to arranging the visual representations on the graphical user interface)
Radford further teaches:
[…] based on the plurality of alignment scores […] (Radford, Page 3 – Section 2.2, “CLIP learns a multi-modal embedding space by jointly training an image encoder and text encoder to maximize the cosine similarity of the image and text embeddings of the N real pairs in the batch while minimizing the cosine similarity of the embeddings of the N² − N incorrect pairings. We optimize a symmetric cross entropy loss over these similarity scores”, & Page 4 – Section 2.5, “For each dataset, we use the names of all the classes in the dataset as the set of potential text pairings and predict the most probable (image, text) pair according to CLIP”, thus […] based on the plurality of alignment scores […] is disclosed because Radford teaches generating similarity scores for multiple image-text pairings and using the scores to determine the relative correspondence of the pairings. Radford’s similarity scores correspond to the plurality of alignment scores)
Regarding Claim 15, Radford teaches:
generate a textual prompt, wherein the textual prompt includes […] (Radford, Page 4 – Section 2.5, “For each dataset, we use the names of all the classes in the dataset as the set of potential text pairings and predict the most probable (image, text) pair according to CLIP. We addi tionally experiment with providing CLIP with text prompts to help specify the task as well as ensembling multiple of these templates in order to boost performance”, thus generating a textual prompt, wherein the textual prompt includes […] is disclosed, because Radford teaches providing CLIP with text prompts to specify a task and using the names of classes as potential text pairings for comparison with an image. Radford’s text prompts correspond to the textual prompt because they are textual inputs provided to CLIP to define the task to be performed. Radford’s class names correspond to information included in the textual prompt because the class names identify the visual concepts or descriptions against which the input image is evaluated)
generate, via a machine learning model, an alignment score based on the textual prompt and a two-dimensional (2D) input image, wherein the 2D input image […], and the alignment score is a quantitative measure of how accurately the textual prompt describes the 2D input image […] (Radford, Page 3 – Section 2.2, “Given a batch of N (image, text) pairs, CLIP is trained to predict which of the N × N possible (image, text) pairings across a batch actually occurred. To do this, CLIP learns a multi-modal embedding space by jointly training an image encoder and text encoder to maximize the cosine similarity of the image and text embeddings of the N real pairs in the batch while minimizing the cosine similarity of the embeddings of the N2 −N incorrect pairings. We optimize a symmetric cross entropy loss over these similarity scores”, & Page 4 – Section 2.5, “For each dataset, we use the names of all the classes in the dataset as the set of potential text pairings and predict the most probable (image, text) pair according to CLIP”, thus generating, via a machine learning model, an alignment score based on the textual prompt and a two-dimensional (2D) input image, wherein the 2D input image […], and the alignment score is a quantitative measure of how accurately the textual prompt describes the 2D input image […] is disclosed, because Radford teaches that CLIP uses an image encoder and a text encoder to evaluate image-text pairings and to generate similarity scores for those pairings. Radford’s image encoder and text encoder together correspond to the machine learning model because they process the image and text inputs and produce a comparison result. Radford’s image and text pairings correspond to the textual prompt and the 2D input image because CLIP operates on paired text and image inputs. Radford’s cosine similarity and similarity scores correspond to the alignment score because they provide a numerical indication of the degree of correspondence between the text embedding and the image embedding. Radford’s prediction of the most probable image-text pair further shows that the score reflects how well the text describes the image, with a higher similarity indicating a closer match between the textual prompt and the 2D input image)
[…] and the alignment score to the 2D input image (Radford, Page 3 – Section 2.2, “Given a batch of N (image, text) pairs, CLIP is trained to predict which of the N × N possible (image, text) pairings across a batch actually occurred. To do this, CLIP learns a multi-modal embedding space by jointly training an image encoder and text encoder to maximize the cosine similarity of the image and text embeddings of the N real pairs in the batch while minimizing the cosine similarity of the embeddings of the N2 −N incorrect pairings. We optimize a symmetric cross entropy loss over these similarity scores”, thus […] and the alignment score to the 2D input image is disclosed, because Radford calculates a similarity score for an image-text pairing. The similarity score corresponds to the alignment score associated with the input image because it quantifies the correspondence between that image and the paired text)
Radford does not explicitly teach […] includes a human-centric evaluation criterion, […] based on the human-centric evaluation criterion, assign the human-centric evaluation criterion […], display, via a graphical user interface, one or more of [… 2D input image…], the human-centric evaluation criterion, and […alignment score…], and […] is a visual representation of an architectural space […].
However, Yao teaches:
[…] includes a human-centric evaluation criterion (Yao, Page 1 – Section 1, “Urban perceptions, which are the psychological feelings held by residents about an urban locale (Tuan 2013, Ordonez and Berg 2014), provide an important basis for understanding the ways in which urban environments interact with public mental health”, & Page 5 – Section 2.2.2, “This study focuses on six categories of urban perceptions: wealthy, safety, lively, beautiful, boring and depressing, as in Place Pulse 2.0 (Dubey et al. 2016). These six perception datasets from previous studies are assembled through training on mostly western urban street scenery with CNNs and annotators’ votes on pairwise image comparisons (Dubey et al. 2016)”, thus […] includes a human-centric evaluation criterion is disclosed, because Yao teaches evaluating urban environments according to human perceptions that represent psychological feelings held by residents about an urban locale. Yao further identifies specific perception categories, including wealthy, safety, lively, beautiful, boring, and depressing. Yao’s urban perceptions correspond to the human-centric evaluation criterion because they represent human psychological responses to the visual environment. Yao’s categories of wealthy, safety, lively, beautiful, boring, and depressing correspond to specific human-centric evaluation criteria because each category provides a human-perception attribute by which an image of an environment may be evaluated)
[…] based on the human-centric evaluation criterion (Yao, Page 5 – Section 2.2.2, “This study focuses on six categories of urban perceptions: wealthy, safety, lively, beautiful, boring and depressing, as in Place Pulse 2.0 (Dubey et al. 2016). These six perception datasets from previous studies are assembled through training on mostly western urban street scenery with CNNs and annotators’ votes on pairwise image comparisons (Dubey et al. 2016). Our proposed framework, however, collects real score annotations on each image from volunteers”, & Page 5 – Section 2.2.2, “Our human-machine adversarial scoring module processed the FCN sematic segmentation result and the city-scale human perceptions from local volunteers. Furthermore, the module mined the relationship between the visual scenery and perception directly to expedite image classification according to human perception”, thus […] based on the human-centric evaluation criterion is disclosed, because Yao teaches scoring images according to human-perception categories, including wealthy, safety, lively, beautiful, boring, and depressing. Yao’s perception categories correspond to the human-centric evaluation criterion, and its scoring module relates visual scenery to those perceptions to classify the image according to human perception)
assign the human-centric evaluation criterion […] (Yao, Page 5 – Section 2.2.2, “Our human-machine adversarial scoring module processed the FCN sematic segmentation result and the city-scale human perceptions from local volunteers. Furthermore, the module mined the relationship between the visual scenery and perception directly to expedite image classification according to human perception. We use RF algorithm to determine the final image classification, and therefore the results are subject to RF fitting and limited to one perception label per image”, thus assigning the human-centric evaluation criterion […] is disclosed, because Yao teaches classifying each image according to human perception and assigning one perception label to each image. Yao’s perception label corresponds to the human-centric evaluation criterion because it identifies the human-perception category associated with the image)
display, via a graphical user interface, one or more of [… 2D input image…], the human-centric evaluation criterion, and […alignment score…] (Yao, Page 6 – Section 2.2.2, “Figure 3 illustrates the process of human-machine adversarial scoring. Volunteers score a displayed street-view image in terms of the six types of perception in a range of 1–100 with 0 being the lowest and 100 being the highest level of a perception”, & Page 6 – Section 2.2.2.1, “Once a user has scored the first 50 photos, the scoring software establishes a random forest set to fit the scoring process. Then, as users have rated subsequent photos, the software offers a recommendation score based on the rules learned from the previous user rating actions”, thus Yao teaches scoring software that presents a displayed street-view image to a user for evaluation according to six types of human perception and provides a recommendation score during that process. Yao’s scoring software corresponds to the graphical user interface, the displayed street-view image corresponds to the 2D input image, the six perception types correspond to the human-centric evaluation criterion, and the recommendation score corresponds to score information provided through the interface)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine Radford with Yao by incorporating Yao’s human-perception evaluation criteria into Radford’s CLIP image-text matching framework. Radford teaches using natural language prompts and class names to specify a task and generating similarity scores that quantify the correspondence between image and text inputs. Yao teaches evaluating visual environments according to human-perception categories and recognizes that traditional evaluation of human perceptions is difficult, costly, and time consuming, stating that a framework that can boost work efficiency is needed to optimize the urban perception assessment process. Therefore, a POSITA would have been motivated to use Yao’s human-perception categories as the class information or textual concepts supplied to Radford’s CLIP so that Radford’s image-text similarity technique could efficiently evaluate images according to human perception. Such a combination would have provided an automated and more efficient technique for quantitatively determining how well an image corresponds to a human-perception criterion, thereby addressing Yao’s identified need to improve the efficiency of urban perception assessment (Yao, Page 1 – Section 1, “Traditionally, the evaluation of human perceptions towards their visual surroundings remains difficult due to the lack of high-throughput methods, inadequate sample problems and being restricted to inter views and questionnaires (Hannay 1983, Halpern 1995, Kabisch et al. 2015, Dadvand et al. 2016). Given the costly and time-consuming nature of these investigation methods, a framework that can boost work efficiency is needed to optimize the urban perception assessment process”)
Radford combined with Yao does not explicitly teach […] is a visual representation of an architectural space […].
However, Hoser teaches:
[…] is a visual representation of an architectural space […] (Hoser, Page 2 – Section 1, “Following this, RGB color codes were extracted from the cumulative 16 images obtained, and these 16 distinct images were employed to colorize the interior of a 3D model representing a preschool space. Consequently, the same set of 16 images, pertaining to the identical spatial context, was presented to a group of 62 experts consisting of specialized architects/interior architects, and preschool educators”, & Page 5 – Section 3, “Afterward, the colors obtained from the visuals were transferred to the 3D model as a material color (Figure 7). Thus, a total of 16 views of the same space were obtained from 4 different systems. These views are listed below (Figure 8)”, thus […] is a visual representation of an architectural space […] is disclosed, because Hoser teaches generating multiple image views of the same modeled preschool interior. The preschool interior corresponds to the architectural space, while the resulting images and views correspond to visual representations of that space)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to further combine Radford and Yao with Hoser by applying the combined image-text and human perception evaluation technique to visual representations of architectural spaces. Radford and Yao together teach evaluating images based on textual concepts and human perception criteria, while Hoser teaches applying artificial intelligence techniques to interior architectural spaces and evaluating images of those spaces according to user-oriented criteria. Therefore, a POSITA would have been motivated to use Hoser’s architectural space images as the visual inputs evaluated by the Radford-Yao technique so that the depicted architectural spaces could be quantitatively assessed according to human centric criteria, thereby providing a more systematic and efficient technique for evaluating how users perceive visual representations of architectural and interior spaces and facilitating the application of artificial-intelligence-based evaluation to architecture and interior design (Hoser, Page 2 – Section 1, “artificial intelligence-supported data processing systems, transitioning from text to image, can contribute to interior spatial coloring in a scientifically accurate manner, thereby enriching the fields of architecture and interior design”, & “the same set of 16 images, pertaining to the identical spatial context, was presented to a group of 62 experts consisting of specialized architects/interior architects, and preschool educators. Through a survey employing Likert-type questions, these experts were prompted to assess the volumes in terms of “academic instruction” and “entertainment””)
Regarding Claim 16, Radford and Yao combined with Hoser teaches all the limitations of claim 15 as cited above and Yao further teaches:
wherein the human-centric evaluation criterion includes one of the terms “social,” “isolating,” “tranquil,” “distracting,” “inspirational,” or “boring” (Yao, Page 5 – Section 2.2.2., “This study focuses on six categories of urban perceptions: wealthy, safety, lively, beautiful, boring and depressing, as in Place Pulse 2.0 (Dubey et al. 2016). These six perception datasets from previous studies are assembled through training on mostly western urban street scenery with CNNs and annotators’ votes on pairwise image comparisons (Dubey et al. 2016)”, thus thus wherein the human-centric evaluation criterion includes one of the terms “social,” “isolating,” “tranquil,” “distracting,” “inspirational,” or “boring” is disclosed, because Yao identifies “boring” as one of its categories of human perception. Yao’s “boring” perception category corresponds to the recited human-centric evaluation criterion because it is a human-perception attribute used to evaluate the depicted environment)
Regarding Claim 17, Radford and Yao combined with Hoser teaches all the limitations of claim 15 as cited above and Hoser further teaches:
wherein […the 2D input image…] is one of a photograph of an architectural space, a 2D rendering of an architectural space, or a still image representing a single frame included in a video recording of an architectural space (Hoser, Hoser, Page 5 – Section 3.1, “The survey study includes a mixed group of graduates and academics in the fields of pre-school teaching and architecture/interior architecture. The appropriateness of the color schemes in the images in the contexts of ‘education’ and ‘entertainment’ for a pre-school educational space was measured using the ‘relational screening model’”, & Page 5 – Section 3, “Afterward, the colors obtained from the visuals were transferred to the 3D model as a material color (Figure 7). Thus, a total of 16 views of the same space were obtained from 4 different systems. These views are listed below (Figure 8)”, thus wherein […the 2D input image…] is one of a photograph of an architectural space, a 2D rendering of an architectural space, or a still image representing a single frame included in a video recording of an architectural space is disclosed, because Hoser teaches images representing a pre-school educational architectural space and multiple views obtained from a 3D model of that space. Hoser’s pre-school educational space corresponds to the architectural space, and the resulting image views of the 3D-modeled space correspond to 2D renderings of the architectural space)
Regarding Claim 20, Radford and Yao combined with Hoser teaches all the limitations of claim 15 as cited above and Yao further teaches:
display, via the graphical user interface, […] (Yao, Page 6 – Section 2.2.2, “Volunteers score a displayed street-view image in terms of the six types of perception in a range of 1–100 with 0 being the lowest and 100 being the highest level of a perception”, & Page 6 – Section 2.2.2.1, “Then, as users have rated subsequent photos, the software offers a recommendation score based on the rules learned from the previous user rating actions”, thus displaying, via the graphical user interface, […] is disclosed because Yao teaches user-facing scoring software that displays images for evaluation and provides corresponding score information. Yao’s scoring software corresponds to the graphical user interface)
Hoser further teaches:
[…] simultaneous visual representations of a plurality of architectural spaces […] associated with the plurality of architectural spaces (Hoser, Page 5 – Section 3, “Afterward, the colors obtained from the visuals were transferred to the 3D model as a material color (Figure 7). Thus, a total of 16 views of the same space were obtained from 4 different systems. These views are listed below (Figure 8)”, & Page 5 – Section 3.1, “The appropriateness of the color schemes in the images in the contexts of “education” and “entertainment” for a pre-school educational space was measured using the ‘relational screening model’”, thus […] simultaneous visual representations […] associated with the plurality of architectural spaces is disclosed because Hoser presents multiple visual representations of an architectural interior together and evaluates the respective images in relation to the depicted architectural space. Hoser’s displayed image views correspond to the simultaneous visual representations, and the depicted pre-school interior corresponds to the architectural space associated with those representations)
Radford further teaches:
[…] and a plurality of alignment scores […] (Page 3 – Section 2.2, “CLIP learns a multi-modal embedding space by jointly training an image encoder and text encoder to maximize the cosine similarity of the image and text embeddings of the N real pairs while minimizing the cosine similarity of the embeddings of the N² − N incorrect pairings. We optimize a symmetric cross entropy loss over these similarity scores”, thus […] and a plurality of alignment scores […] is disclosed because Radford generates similarity scores for multiple image-text pairings. Radford’s similarity scores correspond to the plurality of alignment scores because each score quantifies the correspondence between an image representation and its textual representation)
Claims 4-5, 11-12, and 18-19 are rejected under 35 U.S.C. 103 as being unpatentable over Radford et al. (hereinafter Radford, a non-patent literature reference titled “Learning transferable visual models from natural language supervision.”), in view of Yao et al. (hereinafter Yao, a non-patent literature reference titled “A human-machine adversarial scoring framework for urban perception assessment using street-view images.”), in view of Hoser et al. (hereinafter Hoser, a non-patent literature reference titled “Analysis of text-to-image artificial intelligence Systems in terms of contribution to interior coloring.”), and further in view of Zhou et al. (hereinafter Zhou, a non-patent literature reference titled “Learning to prompt for vision-language models.”).
Regarding Claim 4, Radford and Yao combined with Hoser teaches all the limitations of claim 1 as cited above and Radford further teaches:
[…] based on a first plurality of alignment scores generated for a first plurality of image-text pairs included in a training data set (Radford, Page 3 – Section 2.2, “Given a batch of N (image, text) pairs, CLIP is trained to predict which of the N × N possible (image, text) pairings across a batch actually occurred. To do this, CLIP learns a multi-modal embedding space by jointly training an image encoder and text encoder to maximize the cosine similarity of the image and text embeddings of the N real pairs in the batch while minimizing the cosine similarity of the embeddings of the N² − N incorrect pairings. We optimize a symmetric cross entropy loss over these similarity scores”, thus […] based on a first plurality of alignment scores generated for a first plurality of image-text pairs included in a training data set is disclosed, because Radford teaches training CLIP using a batch containing multiple image-text pairs and calculating cosine similarity scores for the possible pairings. Radford’s batch of N image-text pairs corresponds to the first plurality of image-text pairs included in the training data set, and the cosine similarity scores correspond to the first plurality of alignment scores because they quantify the correspondence between the respective image and text embeddings)
[…] based on a second plurality of alignment scores generated for a second plurality of image-text pairs included in a testing data set (Radford, Page 3 – Section 2.2, “Given a batch of N (image, text) pairs, CLIP is trained to predict which of the N × N possible (image, text) pairings across a batch actually occurred. To do this, CLIP learns a multi-modal embedding space by jointly training an image encoder and text encoder to maximize the cosine similarity of the image and text embeddings of the N real pairs in the batch while minimizing the cosine similarity of the embeddings of the N² − N incorrect pairings. We optimize a symmetric cross entropy loss over these similarity scores”, & Page 4 – Section 2.5, “For each dataset, we use the names of all the classes in the dataset as the set of potential text pairings and predict the most probable (image, text) pair according to CLIP”, thus […] based on a second plurality of alignment scores generated for a second plurality of image-text pairs included in a testing data set is disclosed, because Radford teaches evaluating images from downstream datasets against multiple textual class descriptions and determining the most probable image-text pairing based on cosine similarity scores. Radford’s evaluated images and potential textual class pairings correspond to the second plurality of image-text pairs, and the cosine similarity scores used to determine the most probable pair correspond to the second plurality of alignment scores)
Yao further teaches:
continuing or terminating […accuracy guided parameter adjustment…] (Yao, Page 7 – Section 2.2.2.2, “If the recommendation scores of more than five pictures seriously deviate from a user’s score by more than 10 points, the embedded random forest module will be retrained and self-correct the fitting model. Otherwise, if the OOB validation error of the fitting model is less than 10 points, the user scoring procedure stops and outputs a human-machine adversarial scoring dataset”, thus continuing or terminating […accuracy guided parameter adjustment…] is disclosed, because Yao teaches retraining and self correcting the model when performance is insufficient and stopping the procedure when the validation error satisfies the specified accuracy threshold)
Radford and Yao combined with Hoser does not explicitly teach initializing, based on a pre-trained model, a plurality of learnable parameters included in […a machine learning model…], iteratively adjusting one or more of the plurality of learnable parameters […], calculating an accuracy associated with […a machine learning model…], […] the iterative adjustment of the one or more of the plurality of learnable parameters based on the calculated accuracy.
However, Zhou teaches:
initializing, based on a pre-trained model, a plurality of learnable parameters included in […a machine learning model…] (Zhou, Page 2 – Section 1, “Concretely, CoOp models a prompt’s context words with learnable vectors, which could be initialized with either random values or pre-trained word embeddings”, & Page 5 – Section 3.2, “We propose Context Optimization (CoOp), which avoids manual prompt tuning by modeling context words with continuous vectors that are end-to-end learned from data while the massive pre-trained parameters are frozen”, thus initializing, based on a pre-trained model, a plurality of learnable parameters included in […a machine learning model…] is disclosed, because Zhou teaches learnable context vectors that may be initialized using pre-trained word embeddings. Zhou’s context vectors correspond to the plurality of learnable parameters, and the pre-trained word embeddings used to initialize those vectors correspond to initialization based on a pre-trained model)
iteratively adjusting one or more of the plurality of learnable parameters […] (Zhou, Page 5 – Section 3.2, “Training is performed to minimize the standard classification loss based on the cross-entropy, and the gradients can be back-propagated all the way through the text encoder g(⋅), making use of the rich knowledge encoded in the parameters to optimize the context”, & Page 7 – Section 4.1, “Training is done with SGD and an initial learning rate of 0.002, which is decayed by the cosine annealing rule. The maximum epoch is set to 200 for 16/8 shots, 100 for 4/2 shots, and 50 for 1 shot”, thus iteratively adjusting one or more of the plurality of learnable parameters […] is disclosed, because Zhou teaches repeatedly optimizing the learnable context vectors through gradient back propagation and SGD over multiple training epochs. Zhou’s learnable context vectors correspond to the plurality of learnable parameters, and their repeated optimization during training corresponds to iterative adjustment)
calculating an accuracy associated with […a machine learning model…] (Zhou, Page 2 – Section 1, Figure 1, which reports an “Accuracy” value for the evaluated models, & Page 6 – Section 4.1, “We follow the few-shot evaluation protocol adopted in CLIP (Radford et al., 2021), using 1, 2, 4, 8 and 16 shots for training respectively and deploying models in the full test sets. The average results over three runs are reported for comparison”, thus calculating an accuracy associated with […a machine learning model…] is disclosed, because Zhou teaches evaluating the trained model on test data and calculating and reporting its resulting accuracy. Zhou’s reported test performance corresponds to the accuracy associated with the machine learning model)
[…] the iterative adjustment of the one or more of the plurality of learnable parameters based on the calculated accuracy (Zhou, Page 6 – Section 4.1, “We follow the few-shot evaluation protocol adopted in CLIP (Radford et al., 2021), using 1, 2, 4, 8 and 16 shots for training respectively and deploying models in the full test sets. The average results over three runs are reported for comparison”, & Page 7 – Section 4.1, “Training is done with SGD and an initial learning rate of 0.002, which is decayed by the cosine annealing rule. The maximum epoch is set to 200 for 16/8 shots, 100 for 4/2 shots, and 50 for 1 shot”, thus Zhou teaches iteratively adjusting the learnable context vectors during training and evaluating the resulting model performance on test data. Zhou’s repeated SGD optimization corresponds to the iterative adjustment of the learnable parameters, while the reported test-set performance corresponds to the calculated accuracy)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to further combine Radford, Yao, and Hoser with Zhou by incorporating Zhou’s prompt learning and context optimization technique into Radford’s pretrained CLIP image-text evaluation framework. Radford teaches a pretrained vision language model that compares image and textual representations, while Yao and Hoser provide the human perception criteria and architectural space imagery to which that image-text evaluation technique may be applied. Zhou specifically teaches adapting pretrained vision language models, including CLIP models, for downstream tasks using learnable context vectors that are optimized from training data while the pretrained model parameters remain fixed. Therefore, a POSITA would have been motivated to use Zhou’s learnable context vectors and training technique with the Radford, Yao, and Hoser combination to adapt the pretrained CLIP model to the particular task of evaluating architectural space images according to human centric criteria, thereby automating prompt optimization, reducing the effort associated with manual prompt engineering, and improving the downstream performance and efficiency of the image-text evaluation model (Zhou, Page 10 – Section 5, “these models, also called vision foundation models given their “critically central yet incomplete” nature (Bommasani et al., 2021), need to be adapted using automated techniques for better downstream performance and efficiency. Our research provides timely insights on how CLIP like models can be turned into a data-efficient learner by using prompt learning, and reveals that despite being a learning-based approach, CoOp performs much better in domain generalization than manual prompts”)
Regarding Claim 5, Radford, Yao and Hoser combined with Zhou teaches all the limitations of claim 4 as cited above and Hoser further teaches:
wherein each image-text pair included in […the first and second pluralities of image-text pairs…] includes an image representing an architectural space […] (Hoser, Page 5 – Section 3, “Afterward, the colors obtained from the visuals were transferred to the 3D model as a material color (Figure 7). Thus, a total of 16 views of the same space were obtained from 4 different systems. These views are listed below (Figure 8)”, thus wherein each image-text pair included in […the first and second pluralities of image-text pairs…] includes an image representing an architectural space […] is disclosed, because Hoser teaches multiple image views representing the same modeled interior space. Hoser’s image views correspond to images representing an architectural space)
Yao further teaches:
[…] and a ground truth human-centric evaluation criterion associated with the image (Yao, Page 5 – Section 2.2.2, “This study focuses on six categories of urban perceptions: wealthy, safety, lively, beautiful, boring and depressing” & “Our proposed framework, however, collects real score annotations on each image from volunteers” & “We use RF algorithm to determine the final image classification, and therefore the results are subject to RF fitting and limited to one perception label per image”, & Page 7 – Section 2.2.3, “this study uses the Pearson correlation coefficient (Pearson R), standard R2, root mean squared error (RMSE) and mean absolute error (MAE) to quantify the accuracy between the predictions and the ground-truth values”, thus […] and a ground truth human centric evaluation criterion associated with the image is disclosed, because Yao teaches associating each image with a human-perception label and volunteer-provided annotation that serves as ground truth information for evaluating model predictions. Yao’s perception categories correspond to the human centric evaluation criterion, and the human annotations and corresponding ground-truth values establish the ground truth criterion associated with the image)
Regarding Claim 11, Radford and Yao combined with Hoser teaches all the limitations of claim 8 as cited above and Radford further teaches:
[…] based on a first plurality of alignment scores generated for a first plurality of image-text pairs included in a training data set (Radford, Page 3 – Section 2.2, “Given a batch of N (image, text) pairs, CLIP is trained to predict which of the N × N possible (image, text) pairings across a batch actually occurred. To do this, CLIP learns a multi-modal embedding space by jointly training an image encoder and text encoder to maximize the cosine similarity of the image and text embeddings of the N real pairs in the batch while minimizing the cosine similarity of the embeddings of the N² − N incorrect pairings. We optimize a symmetric cross entropy loss over these similarity scores”, thus […] based on a first plurality of alignment scores generated for a first plurality of image-text pairs included in a training data set is disclosed, because Radford teaches training CLIP using a batch containing multiple image-text pairs and calculating cosine similarity scores for the possible pairings. Radford’s batch of N image-text pairs corresponds to the first plurality of image-text pairs included in the training data set, and the cosine similarity scores correspond to the first plurality of alignment scores because they quantify the correspondence between the respective image and text embeddings)
[…] based on a second plurality of alignment scores generated for a second plurality of image-text pairs included in a testing data set (Radford, Page 3 – Section 2.2, “Given a batch of N (image, text) pairs, CLIP is trained to predict which of the N × N possible (image, text) pairings across a batch actually occurred. To do this, CLIP learns a multi-modal embedding space by jointly training an image encoder and text encoder to maximize the cosine similarity of the image and text embeddings of the N real pairs in the batch while minimizing the cosine similarity of the embeddings of the N² − N incorrect pairings. We optimize a symmetric cross entropy loss over these similarity scores”, & Page 4 – Section 2.5, “For each dataset, we use the names of all the classes in the dataset as the set of potential text pairings and predict the most probable (image, text) pair according to CLIP”, thus […] based on a second plurality of alignment scores generated for a second plurality of image-text pairs included in a testing data set is disclosed, because Radford teaches evaluating images from downstream datasets against multiple textual class descriptions and determining the most probable image-text pairing based on cosine similarity scores. Radford’s evaluated images and potential textual class pairings correspond to the second plurality of image-text pairs, and the cosine similarity scores used to determine the most probable pair correspond to the second plurality of alignment scores)
Yao further teaches:
continuing or terminating […accuracy guided parameter adjustment…] (Yao, Page 7 – Section 2.2.2.2, “If the recommendation scores of more than five pictures seriously deviate from a user’s score by more than 10 points, the embedded random forest module will be retrained and self-correct the fitting model. Otherwise, if the OOB validation error of the fitting model is less than 10 points, the user scoring procedure stops and outputs a human-machine adversarial scoring dataset”, thus continuing or terminating […accuracy guided parameter adjustment…] is disclosed, because Yao teaches retraining and self correcting the model when performance is insufficient and stopping the procedure when the validation error satisfies the specified accuracy threshold)
Radford and Yao combined with Hoser does not explicitly teach initializing, based on a pre-trained model, a plurality of learnable parameters included in […a machine learning model…], iteratively adjusting one or more of the plurality of learnable parameters […], calculating an accuracy associated with […a machine learning model…], […] the iterative adjustment of the one or more of the plurality of learnable parameters based on the calculated accuracy.
However, Zhou teaches:
initializing, based on a pre-trained model, a plurality of learnable parameters included in […a machine learning model…] (Zhou, Page 2 – Section 1, “Concretely, CoOp models a prompt’s context words with learnable vectors, which could be initialized with either random values or pre-trained word embeddings”, & Page 5 – Section 3.2, “We propose Context Optimization (CoOp), which avoids manual prompt tuning by modeling context words with continuous vectors that are end-to-end learned from data while the massive pre-trained parameters are frozen”, thus initializing, based on a pre-trained model, a plurality of learnable parameters included in […a machine learning model…] is disclosed, because Zhou teaches learnable context vectors that may be initialized using pre-trained word embeddings. Zhou’s context vectors correspond to the plurality of learnable parameters, and the pre-trained word embeddings used to initialize those vectors correspond to initialization based on a pre-trained model)
iteratively adjusting one or more of the plurality of learnable parameters […] (Zhou, Page 5 – Section 3.2, “Training is performed to minimize the standard classification loss based on the cross-entropy, and the gradients can be back-propagated all the way through the text encoder g(⋅), making use of the rich knowledge encoded in the parameters to optimize the context”, & Page 7 – Section 4.1, “Training is done with SGD and an initial learning rate of 0.002, which is decayed by the cosine annealing rule. The maximum epoch is set to 200 for 16/8 shots, 100 for 4/2 shots, and 50 for 1 shot”, thus iteratively adjusting one or more of the plurality of learnable parameters […] is disclosed, because Zhou teaches repeatedly optimizing the learnable context vectors through gradient back propagation and SGD over multiple training epochs. Zhou’s learnable context vectors correspond to the plurality of learnable parameters, and their repeated optimization during training corresponds to iterative adjustment)
calculating an accuracy associated with […a machine learning model…] (Zhou, Page 2 – Section 1, Figure 1, which reports an “Accuracy” value for the evaluated models, & Page 6 – Section 4.1, “We follow the few-shot evaluation protocol adopted in CLIP (Radford et al., 2021), using 1, 2, 4, 8 and 16 shots for training respectively and deploying models in the full test sets. The average results over three runs are reported for comparison”, thus calculating an accuracy associated with […a machine learning model…] is disclosed, because Zhou teaches evaluating the trained model on test data and calculating and reporting its resulting accuracy. Zhou’s reported test performance corresponds to the accuracy associated with the machine learning model)
[…] the iterative adjustment of the one or more of the plurality of learnable parameters based on the calculated accuracy (Zhou, Page 6 – Section 4.1, “We follow the few-shot evaluation protocol adopted in CLIP (Radford et al., 2021), using 1, 2, 4, 8 and 16 shots for training respectively and deploying models in the full test sets. The average results over three runs are reported for comparison”, & Page 7 – Section 4.1, “Training is done with SGD and an initial learning rate of 0.002, which is decayed by the cosine annealing rule. The maximum epoch is set to 200 for 16/8 shots, 100 for 4/2 shots, and 50 for 1 shot”, thus Zhou teaches iteratively adjusting the learnable context vectors during training and evaluating the resulting model performance on test data. Zhou’s repeated SGD optimization corresponds to the iterative adjustment of the learnable parameters, while the reported test-set performance corresponds to the calculated accuracy)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to further combine Radford, Yao, and Hoser with Zhou by incorporating Zhou’s prompt learning and context optimization technique into Radford’s pretrained CLIP image-text evaluation framework. Radford teaches a pretrained vision language model that compares image and textual representations, while Yao and Hoser provide the human perception criteria and architectural space imagery to which that image-text evaluation technique may be applied. Zhou specifically teaches adapting pretrained vision language models, including CLIP models, for downstream tasks using learnable context vectors that are optimized from training data while the pretrained model parameters remain fixed. Therefore, a POSITA would have been motivated to use Zhou’s learnable context vectors and training technique with the Radford, Yao, and Hoser combination to adapt the pretrained CLIP model to the particular task of evaluating architectural space images according to human centric criteria, thereby automating prompt optimization, reducing the effort associated with manual prompt engineering, and improving the downstream performance and efficiency of the image-text evaluation model (Zhou, Page 10 – Section 5, “these models, also called vision foundation models given their “critically central yet incomplete” nature (Bommasani et al., 2021), need to be adapted using automated techniques for better downstream performance and efficiency. Our research provides timely insights on how CLIP like models can be turned into a data-efficient learner by using prompt learning, and reveals that despite being a learning-based approach, CoOp performs much better in domain generalization than manual prompts”)
Regarding Claim 12, Radford, Yao and Hoser combined with Zhou teaches all the limitations of claim 11 as cited above and Hoser further teaches:
wherein each image-text pair included in […the first and second pluralities of image-text pairs…] includes an image representing an architectural space […] (Hoser, Page 5 – Section 3, “Afterward, the colors obtained from the visuals were transferred to the 3D model as a material color (Figure 7). Thus, a total of 16 views of the same space were obtained from 4 different systems. These views are listed below (Figure 8)”, thus wherein each image-text pair included in […the first and second pluralities of image-text pairs…] includes an image representing an architectural space […] is disclosed, because Hoser teaches multiple image views representing the same modeled interior space. Hoser’s image views correspond to images representing an architectural space)
Yao further teaches:
[…] and a ground truth human-centric evaluation criterion associated with the image (Yao, Page 5 – Section 2.2.2, “This study focuses on six categories of urban perceptions: wealthy, safety, lively, beautiful, boring and depressing” & “Our proposed framework, however, collects real score annotations on each image from volunteers” & “We use RF algorithm to determine the final image classification, and therefore the results are subject to RF fitting and limited to one perception label per image”, & Page 7 – Section 2.2.3, “this study uses the Pearson correlation coefficient (Pearson R), standard R2, root mean squared error (RMSE) and mean absolute error (MAE) to quantify the accuracy between the predictions and the ground-truth values”, thus […] and a ground truth human centric evaluation criterion associated with the image is disclosed, because Yao teaches associating each image with a human-perception label and volunteer-provided annotation that serves as ground truth information for evaluating model predictions. Yao’s perception categories correspond to the human centric evaluation criterion, and the human annotations and corresponding ground-truth values establish the ground truth criterion associated with the image)
Regarding Claim 18, Radford and Yao combined with Hoser teaches all the limitations of claim 15 as cited above and Radford further teaches:
[…] based on a first plurality of alignment scores generated for a first plurality of image-text pairs included in a training data set (Radford, Page 3 – Section 2.2, “Given a batch of N (image, text) pairs, CLIP is trained to predict which of the N × N possible (image, text) pairings across a batch actually occurred. To do this, CLIP learns a multi-modal embedding space by jointly training an image encoder and text encoder to maximize the cosine similarity of the image and text embeddings of the N real pairs in the batch while minimizing the cosine similarity of the embeddings of the N² − N incorrect pairings. We optimize a symmetric cross entropy loss over these similarity scores”, thus […] based on a first plurality of alignment scores generated for a first plurality of image-text pairs included in a training data set is disclosed, because Radford teaches training CLIP using a batch containing multiple image-text pairs and calculating cosine similarity scores for the possible pairings. Radford’s batch of N image-text pairs corresponds to the first plurality of image-text pairs included in the training data set, and the cosine similarity scores correspond to the first plurality of alignment scores because they quantify the correspondence between the respective image and text embeddings)
[…] based on a second plurality of alignment scores generated for a second plurality of image-text pairs included in a testing data set (Radford, Page 3 – Section 2.2, “Given a batch of N (image, text) pairs, CLIP is trained to predict which of the N × N possible (image, text) pairings across a batch actually occurred. To do this, CLIP learns a multi-modal embedding space by jointly training an image encoder and text encoder to maximize the cosine similarity of the image and text embeddings of the N real pairs in the batch while minimizing the cosine similarity of the embeddings of the N² − N incorrect pairings. We optimize a symmetric cross entropy loss over these similarity scores”, & Page 4 – Section 2.5, “For each dataset, we use the names of all the classes in the dataset as the set of potential text pairings and predict the most probable (image, text) pair according to CLIP”, thus […] based on a second plurality of alignment scores generated for a second plurality of image-text pairs included in a testing data set is disclosed, because Radford teaches evaluating images from downstream datasets against multiple textual class descriptions and determining the most probable image-text pairing based on cosine similarity scores. Radford’s evaluated images and potential textual class pairings correspond to the second plurality of image-text pairs, and the cosine similarity scores used to determine the most probable pair correspond to the second plurality of alignment scores)
Yao further teaches:
continue or terminate […accuracy guided parameter adjustment…] (Yao, Page 7 – Section 2.2.2.2, “If the recommendation scores of more than five pictures seriously deviate from a user’s score by more than 10 points, the embedded random forest module will be retrained and self-correct the fitting model. Otherwise, if the OOB validation error of the fitting model is less than 10 points, the user scoring procedure stops and outputs a human-machine adversarial scoring dataset”, thus continuing or terminating […accuracy guided parameter adjustment…] is disclosed, because Yao teaches retraining and self correcting the model when performance is insufficient and stopping the procedure when the validation error satisfies the specified accuracy threshold)
Radford and Yao combined with Hoser does not explicitly teach initialize, based on a pre-trained model, a plurality of learnable parameters included in […a machine learning model…], iteratively adjust one or more of the plurality of learnable parameters […], calculate an accuracy associated with […a machine learning model…], […] the iterative adjustment of the one or more of the plurality of learnable parameters based on the calculated accuracy.
However, Zhou teaches:
initialize, based on a pre-trained model, a plurality of learnable parameters included in […a machine learning model…] (Zhou, Page 2 – Section 1, “Concretely, CoOp models a prompt’s context words with learnable vectors, which could be initialized with either random values or pre-trained word embeddings”, & Page 5 – Section 3.2, “We propose Context Optimization (CoOp), which avoids manual prompt tuning by modeling context words with continuous vectors that are end-to-end learned from data while the massive pre-trained parameters are frozen”, thus initializing, based on a pre-trained model, a plurality of learnable parameters included in […a machine learning model…] is disclosed, because Zhou teaches learnable context vectors that may be initialized using pre-trained word embeddings. Zhou’s context vectors correspond to the plurality of learnable parameters, and the pre-trained word embeddings used to initialize those vectors correspond to initialization based on a pre-trained model)
iteratively adjust one or more of the plurality of learnable parameters […] (Zhou, Page 5 – Section 3.2, “Training is performed to minimize the standard classification loss based on the cross-entropy, and the gradients can be back-propagated all the way through the text encoder g(⋅), making use of the rich knowledge encoded in the parameters to optimize the context”, & Page 7 – Section 4.1, “Training is done with SGD and an initial learning rate of 0.002, which is decayed by the cosine annealing rule. The maximum epoch is set to 200 for 16/8 shots, 100 for 4/2 shots, and 50 for 1 shot”, thus iteratively adjusting one or more of the plurality of learnable parameters […] is disclosed, because Zhou teaches repeatedly optimizing the learnable context vectors through gradient back propagation and SGD over multiple training epochs. Zhou’s learnable context vectors correspond to the plurality of learnable parameters, and their repeated optimization during training corresponds to iterative adjustment)
calculate an accuracy associated with […a machine learning model…] (Zhou, Page 2 – Section 1, Figure 1, which reports an “Accuracy” value for the evaluated models, & Page 6 – Section 4.1, “We follow the few-shot evaluation protocol adopted in CLIP (Radford et al., 2021), using 1, 2, 4, 8 and 16 shots for training respectively and deploying models in the full test sets. The average results over three runs are reported for comparison”, thus calculating an accuracy associated with […a machine learning model…] is disclosed, because Zhou teaches evaluating the trained model on test data and calculating and reporting its resulting accuracy. Zhou’s reported test performance corresponds to the accuracy associated with the machine learning model)
[…] the iterative adjustment of the one or more of the plurality of learnable parameters based on the calculated accuracy (Zhou, Page 6 – Section 4.1, “We follow the few-shot evaluation protocol adopted in CLIP (Radford et al., 2021), using 1, 2, 4, 8 and 16 shots for training respectively and deploying models in the full test sets. The average results over three runs are reported for comparison”, & Page 7 – Section 4.1, “Training is done with SGD and an initial learning rate of 0.002, which is decayed by the cosine annealing rule. The maximum epoch is set to 200 for 16/8 shots, 100 for 4/2 shots, and 50 for 1 shot”, thus Zhou teaches iteratively adjusting the learnable context vectors during training and evaluating the resulting model performance on test data. Zhou’s repeated SGD optimization corresponds to the iterative adjustment of the learnable parameters, while the reported test-set performance corresponds to the calculated accuracy)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to further combine Radford, Yao, and Hoser with Zhou by incorporating Zhou’s prompt learning and context optimization technique into Radford’s pretrained CLIP image-text evaluation framework. Radford teaches a pretrained vision language model that compares image and textual representations, while Yao and Hoser provide the human perception criteria and architectural space imagery to which that image-text evaluation technique may be applied. Zhou specifically teaches adapting pretrained vision language models, including CLIP models, for downstream tasks using learnable context vectors that are optimized from training data while the pretrained model parameters remain fixed. Therefore, a POSITA would have been motivated to use Zhou’s learnable context vectors and training technique with the Radford, Yao, and Hoser combination to adapt the pretrained CLIP model to the particular task of evaluating architectural space images according to human centric criteria, thereby automating prompt optimization, reducing the effort associated with manual prompt engineering, and improving the downstream performance and efficiency of the image-text evaluation model (Zhou, Page 10 – Section 5, “these models, also called vision foundation models given their “critically central yet incomplete” nature (Bommasani et al., 2021), need to be adapted using automated techniques for better downstream performance and efficiency. Our research provides timely insights on how CLIP like models can be turned into a data-efficient learner by using prompt learning, and reveals that despite being a learning-based approach, CoOp performs much better in domain generalization than manual prompts”)
Regarding Claim 19, Radford, Yao and Hoser combined with Zhou teaches all the limitations of claim 18 as cited above and Hoser further teaches:
wherein each image-text pair included in […the first and second pluralities of image-text pairs…] includes an image representing an architectural space […] (Hoser, Page 5 – Section 3, “Afterward, the colors obtained from the visuals were transferred to the 3D model as a material color (Figure 7). Thus, a total of 16 views of the same space were obtained from 4 different systems. These views are listed below (Figure 8)”, thus wherein each image-text pair included in […the first and second pluralities of image-text pairs…] includes an image representing an architectural space […] is disclosed, because Hoser teaches multiple image views representing the same modeled interior space. Hoser’s image views correspond to images representing an architectural space)
Yao further teaches:
[…] and a ground truth human-centric evaluation criterion associated with the image (Yao, Page 5 – Section 2.2.2, “This study focuses on six categories of urban perceptions: wealthy, safety, lively, beautiful, boring and depressing” & “Our proposed framework, however, collects real score annotations on each image from volunteers” & “We use RF algorithm to determine the final image classification, and therefore the results are subject to RF fitting and limited to one perception label per image”, & Page 7 – Section 2.2.3, “this study uses the Pearson correlation coefficient (Pearson R), standard R2, root mean squared error (RMSE) and mean absolute error (MAE) to quantify the accuracy between the predictions and the ground-truth values”, thus […] and a ground truth human centric evaluation criterion associated with the image is disclosed, because Yao teaches associating each image with a human-perception label and volunteer-provided annotation that serves as ground truth information for evaluating model predictions. Yao’s perception categories correspond to the human centric evaluation criterion, and the human annotations and corresponding ground-truth values establish the ground truth criterion associated with the image)
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. US 20230351753 is pertinent because it teaches determining relevance scores between text and visual content using learned embeddings, training model parameters using labeled text and visual pairs, and initializing model parameters using pretrained CLIP weights. Because applicant’s disclosure similarly concerns image-text alignment scores and training a machine learning model using image-text pairs, the reference is relevant to the invention but is not relied upon in the rejection.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MAHLIET ADMASU whose telephone number is (571)272-0034. The examiner can normally be reached Mon-Fri, 8am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571)270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/M.T.A./Examiner, Art Unit 2123
/ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123