DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This action is in response to the original application filed on Mar. 27th, 2024.
Drawings
The drawings are objected to as failing to comply with 37 CFR 1.84(p)(4) because reference character “150” has been used to designate both Model Weights and Storage. Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitations uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitations are:
“a parser configured to analyze a technical document”
“a data preparator configured to convert the multiple text fragments into a dataset”
“a machine learning model block configured to receive the dataset and perform a training mode”
“wherein the machine learning model block is configured to”
“the documents processor is configured to receive the unlabeled document”
“the parser is configured to extract the image and convert the image into meaningful text data”
“wherein the data preparator is configured to connect text data with a label”
“wherein the machine learning model block is configured to execute each training epoch”
“wherein the machine learning model block is configured to generate an acknowledgement signal”
in claims 1 and 4-7.
Because these claim limitations are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, they are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have these limitations interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitations to avoid them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitations recites sufficient structure to perform the claimed function so as to avoid them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1 and 4-7 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-Al A 35 U.S.C. 112, the applicant), regards as the invention.
Claim Limitations:
Claim 1
“a parser configured to analyze a technical document”
“a data preparator configured to convert the multiple text fragments into a dataset”
“a machine learning model block configured to receive the dataset and perform a training mode”
“wherein the machine learning model block is configured to”
“the documents processor is configured to receive the unlabeled document”
Claim 4
“the parser is configured to extract the image and convert the image into meaningful text data”
Claim 5
“wherein the data preparator is configured to connect text data with a label”
Claim 6
“wherein the machine learning model block is configured to execute each training epoch”
Claim 7
“wherein the machine learning model block is configured to generate an acknowledgement signal”
invoke 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. However, the written description fails to disclose the corresponding structure, material, or acts for performing the entire claimed function and to clearly link the structure, material, or acts to the function. The disclosure of this claim is devoid of any structure that performs the function in the claim. This claim discloses an apparatus which does not further teach the structure which the functions are performed on. Therefore, the claim is indefinite and is rejected under 35 U.S.C. 112(b) or pre- AIA 35 U.S.C. 112, second paragraph.
Applicant may:
Amend the claim so that the claim limitation will no longer be interpreted as a limitation under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph;
Amend the written description of the specification such that it expressly recites what structure, material, or acts perform the entire claimed function, without introducing any new matter (35 U.S.C. 132(a)); or
Amend the written description of the specification such that it clearly links the structure, material, or acts disclosed therein to the function recited in the claim, without introducing any new matter (35 U.S.C. 132(a)).
If applicant is of the opinion that the written description of the specification already implicitly or inherently discloses the corresponding structure, material, or acts and clearly links them to the function so that one of ordinary skill in the art would recognize what structure, material, or acts perform the claimed function, applicant should clarify the record by either:
Amending the written description of the specification such that it expressly recites the corresponding structure, material, or acts for performing the claimed function and clearly links or associates the structure, material, or acts to the claimed function, without introducing any new matter (35 U.S.C. 132(a)); or
Stating on the record what the corresponding structure, material, or acts, which are implicitly or inherently set forth in the written description of the specification, perform the claimed function. For more information, see 37 CFR 1.75(d) and M PEP§§ 608.0l(o) and 2181.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C 101 because the claimed invention is directed to an abstract idea without significantly more. The analysis of the claims will follow the 2019 Revised Patent Subject Matter Eligibility Guidance, 84 Fed. Reg. 50 (“2019 PEG”).
Claim 1
Step 1 – Is the claim to a process, machine, manufacture or composition of matter?
Claim 1 recites, "A system for analyzing documents to be used for verifying a system-on-a-chip (SoC), the system comprising:" therefore it is directed to the statutory category of a machine.
Step 2A Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
The claim recites, inter alia:
“a parser configured to analyze a technical document associated with the Soc to generate multiple text fragments;” Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mental process of evaluating and observing data, which is an evaluation or observation that is practically capable of being performed in the human mind with the assistance of pen and paper. A human is able to use a generic computer to evaluate and process text data using conventional computing methods. The limitation is merely applying an abstract idea on generic computer system. See MPEP 2106.04(a)(2)(III)(c).
“a data preparator configured to convert the multiple text fragments into a dataset suitable for a machine learning model;” Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mental process of evaluating and observing data, which is an evaluation or observation that is practically capable of being performed in the human mind with the assistance of pen and paper. A human is able to use a generic computer to evaluate and process text data using conventional computing methods. The limitation is merely applying an abstract idea on generic computer system. The limitation is merely applying an abstract idea on generic computer system. See MPEP 2106.04(a)(2)(III)(c).
“convert the source unlabeled document based on the dataset with the prediction values into a source labeled document used for verifying the soc.” Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mental process of evaluating and observing data, which is an evaluation or observation that is practically capable of being performed in the human mind with the assistance of pen and paper. A human is able to use a generic computer to evaluate and process text data using conventional computing methods. The limitation is merely applying an abstract idea on generic computer system. The limitation is merely applying an abstract idea on generic computer system. See MPEP 2106.04(a)(2)(III)(c).
Step 2A Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
The claim recites the additional elements, “a machine learning model block configured to receive the dataset and perform a training mode including multiple training epochs or an inference mode; and” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
“a documents processor, wherein the technical document includes at least one of a labeled document for the training mode and an unlabeled document for the inference mode,” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
“wherein the machine learning model block is configured to: in each training epoch of the training mode, train the machine learning model based on the dataset such that a label is appended to each text fragment, and” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
“in the inference mode, make a prediction for each text fragment of the dataset based on the training result of the machine learning model to generate the dataset with prediction values, and” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
“wherein the documents processor is configured to receive the unlabeled document and the dataset with the prediction values, and” is an insignificant extra-solution activity required for any uses of the mental processes (see MPEP § 2106.05(g)) As such, the claim is ineligible.
Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception?
Finally, the claim taken as a whole does not contain an inventive concept which provides significantly more than the abstract idea. The additional elements, “a machine learning model block configured to receive the dataset and perform a training mode including multiple training epochs or an inference mode; and” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
“a documents processor, wherein the technical document includes at least one of a labeled document for the training mode and an unlabeled document for the inference mode,” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
“wherein the machine learning model block is configured to: in each training epoch of the training mode, train the machine learning model based on the dataset such that a label is appended to each text fragment, and” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
“in the inference mode, make a prediction for each text fragment of the dataset based on the training result of the machine learning model to generate the dataset with prediction values, and” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
“wherein the documents processor is configured to receive the unlabeled document and the dataset with the prediction values, and” is an insignificant extra-solution activity required for any uses of abstract ideas (see MPEP § 2106.05(g)), and is a well-understood, routine, conventional activity (see MPEP § 2106.05(d)(II)(i); “Receiving or transmitting data over a network, e.g., using the Internet to gather data”.
Taken alone or in combination, the additional elements of the claim do not provide an inventive concept and thus the claim is subject-matter ineligible.
Claim 2
Step 1 – Is the claim to a process, machine, manufacture or composition of matter?
A machine, as above.
Step 2A Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
The claim recites the abstract ideas of the preceding claims from which it depends.
Step 2A Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
The claim recites the additional elements, “wherein the technical document includes at least one of a specification, a manual, a user guide and a standard, which are each associated with the Soc.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception?
Finally, the claim taken as a whole does not contain an inventive concept which provides significantly more than the abstract idea. The additional elements, “wherein the technical document includes at least one of a specification, a manual, a user guide and a standard, which are each associated with the Soc.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Taken alone or in combination, the additional elements of the claim do not provide an inventive concept and thus the claim is subject-matter ineligible.
Claim 3
Step 1 – Is the claim to a process, machine, manufacture or composition of matter?
A machine, as above.
Step 2A Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
The claim recites the abstract ideas of the preceding claims from which it depends.
Step 2A Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
The claim recites the additional elements, “wherein each text fragment includes text data for at least one of a sentence, a paragraph, and a page.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception?
Finally, the claim taken as a whole does not contain an inventive concept which provides significantly more than the abstract idea. The additional elements, “wherein each text fragment includes text data for at least one of a sentence, a paragraph, and a page.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Taken alone or in combination, the additional elements of the claim do not provide an inventive concept and thus the claim is subject-matter ineligible.
Claim 4
Step 1 – Is the claim to a process, machine, manufacture or composition of matter?
A machine, as above.
Step 2A Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
The claim recites the abstract ideas of the preceding claims from which it depends.
Step 2A Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
The claim recites the additional elements, “wherein, when the technical document includes image, the parser is configured to extract the image and convert the image into meaningful text data.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception?
Finally, the claim taken as a whole does not contain an inventive concept which provides significantly more than the abstract idea. The additional elements, “wherein, when the technical document includes image, the parser is configured to extract the image and convert the image into meaningful text data.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Taken alone or in combination, the additional elements of the claim do not provide an inventive concept and thus the claim is subject-matter ineligible.
Claim 5
Step 1 – Is the claim to a process, machine, manufacture or composition of matter?
A machine, as above.
Step 2A Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
The claim recites, inter alia:
“wherein the data preparator is configured to connect text data with a label among the multiple text fragments to generate the dataset.” Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mental process of evaluating and observing data, which is an evaluation or observation that is practically capable of being performed in the human mind with the assistance of pen and paper. A human is able to evaluate text data and apply labels to that data. The limitation is merely applying an abstract idea on generic computer system. See MPEP 2106.04(a)(2)(III)(c).
Step 2A Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
This claim does not recite any additional limitations which integrate the abstract idea into a practical application.
Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception?
Finally, the claim taken as a whole does not contain an inventive concept which provides significantly more than the abstract idea and thus the claim is subject-matter ineligible.
Claim 6
Step 1 – Is the claim to a process, machine, manufacture or composition of matter?
A machine, as above.
Step 2A Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
The claim recites the abstract ideas of the preceding claims from which it depends.
Step 2A Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
The claim recites the additional elements, “wherein the machine learning model block is configured to execute each training epoch based on the dataset and update one or more model weights associated with particular performance metrics of the machine learning model.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception?
Finally, the claim taken as a whole does not contain an inventive concept which provides significantly more than the abstract idea. The additional elements, “wherein the machine learning model block is configured to execute each training epoch based on the dataset and update one or more model weights associated with particular performance metrics of the machine learning model.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Taken alone or in combination, the additional elements of the claim do not provide an inventive concept and thus the claim is subject-matter ineligible.
Claim 7
Step 1 – Is the claim to a process, machine, manufacture or composition of matter?
A machine, as above.
Step 2A Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
The claim recites the abstract ideas of the preceding claims from which it depends.
Step 2A Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
The claim recites the additional elements, “wherein the machine learning model block is configured to generate an acknowledgement signal indicating that the model weights are updated.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception?
Finally, the claim taken as a whole does not contain an inventive concept which provides significantly more than the abstract idea. The additional elements, “wherein the machine learning model block is configured to generate an acknowledgement signal indicating that the model weights are updated.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Taken alone or in combination, the additional elements of the claim do not provide an inventive concept and thus the claim is subject-matter ineligible.
Claim 8
Step 1 – Is the claim to a process, machine, manufacture or composition of matter?
A machine, as above.
Step 2A Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
The claim recites the abstract ideas of the preceding claims from which it depends.
Step 2A Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
The claim recites the additional elements, “wherein the label includes at least one of a binary value and a probability value.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception?
Finally, the claim taken as a whole does not contain an inventive concept which provides significantly more than the abstract idea. The additional elements, “wherein the label includes at least one of a binary value and a probability value.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Taken alone or in combination, the additional elements of the claim do not provide an inventive concept and thus the claim is subject-matter ineligible.
Claim 9
Step 1 – Is the claim to a process, machine, manufacture or composition of matter?
A machine, as above.
Step 2A Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
The claim recites the abstract ideas of the preceding claims from which it depends.
Step 2A Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
The claim recites the additional elements, “wherein the binary value includes one of a first binary value indicating that each text fragment does not have the required characteristic, and a second binary value indicating that each text fragment has the required characteristic.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception?
Finally, the claim taken as a whole does not contain an inventive concept which provides significantly more than the abstract idea. The additional elements, “wherein the binary value includes one of a first binary value indicating that each text fragment does not have the required characteristic, and a second binary value indicating that each text fragment has the required characteristic.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Taken alone or in combination, the additional elements of the claim do not provide an inventive concept and thus the claim is subject-matter ineligible.
Claim 10
Step 1 – Is the claim to a process, machine, manufacture or composition of matter?
A machine, as above.
Step 2A Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
The claim recites the abstract ideas of the preceding claims from which it depends.
Step 2A Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
The claim recites the additional elements, “wherein the probability value indicates the extent of which each text fragment has the required characteristic.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception?
Finally, the claim taken as a whole does not contain an inventive concept which provides significantly more than the abstract idea. The additional elements, “wherein the probability value indicates the extent of which each text fragment has the required characteristic.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Taken alone or in combination, the additional elements of the claim do not provide an inventive concept and thus the claim is subject-matter ineligible.
Claim 11
Step 1 – Is the claim to a process, machine, manufacture or composition of matter?
Claim 11 recites, "A method for analyzing documents to be used for verifying a system-on-a-chip (SoC), the method comprising:" therefore it is directed to the statutory category of a process.
Step 2A Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
The claim recites, inter alia:
“analyzing, by a parser, a technical document associated with the Soc to generate multiple text fragments;” Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mental process of evaluating and observing data, which is an evaluation or observation that is practically capable of being performed in the human mind with the assistance of pen and paper. A human is able to use a generic computer to evaluate and process text data using conventional computing methods. The limitation is merely applying an abstract idea on generic computer system. See MPEP 2106.04(a)(2)(III)(c).
“converting, by a data preparator, the multiple text fragments into a dataset suitable for a machine learning model;” Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mental process of evaluating and observing data, which is an evaluation or observation that is practically capable of being performed in the human mind with the assistance of pen and paper. A human is able to use a generic computer to evaluate and process text data using conventional computing methods. The limitation is merely applying an abstract idea on generic computer system. See MPEP 2106.04(a)(2)(III)(c).
“converting the source unlabeled document based on the dataset with the prediction values into a source labeled document used for verifying the Soc.” Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mental process of evaluating and observing data, which is an evaluation or observation that is practically capable of being performed in the human mind with the assistance of pen and paper. A human is able to use a generic computer to evaluate and process text data using conventional computing methods. The limitation is merely applying an abstract idea on generic computer system. See MPEP 2106.04(a)(2)(III)(c).
Step 2A Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
The claim recites the additional elements, “A method for analyzing documents to be used for verifying a system-on-a-chip (SoC), the method comprising:” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
“receiving, by a machine learning model block, the dataset and performing a training mode including multiple training epochs or an inference mode,” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
“wherein the technical document includes at least one of a labeled document for the training mode and an unlabeled document for the inference mode,” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
“wherein each training epoch of the training mode includes training the machine learning model based on the dataset such that a label is appended to each text fragment, and” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
“wherein the inference mode includes making a prediction for each text fragment of the dataset based on the training result of the machine learning model to generate the dataset with prediction values; and” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
“receiving, by a documents processor, the unlabeled document and the dataset with the prediction values, and” is an insignificant extra-solution activity required for any uses of the mental processes (see MPEP § 2106.05(g)) As such, the claim is ineligible.
Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception?
Finally, the claim taken as a whole does not contain an inventive concept which provides significantly more than the abstract idea. The additional elements, “A method for analyzing documents to be used for verifying a system-on-a-chip (SoC), the method comprising:” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
“receiving, by a machine learning model block, the dataset and performing a training mode including multiple training epochs or an inference mode,” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
“wherein the technical document includes at least one of a labeled document for the training mode and an unlabeled document for the inference mode,” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
“wherein each training epoch of the training mode includes training the machine learning model based on the dataset such that a label is appended to each text fragment, and” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
“wherein the inference mode includes making a prediction for each text fragment of the dataset based on the training result of the machine learning model to generate the dataset with prediction values; and” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
“receiving, by a documents processor, the unlabeled document and the dataset with the prediction values, and” is an insignificant extra-solution activity required for any uses of abstract ideas (see MPEP § 2106.05(g)), and is a well-understood, routine, conventional activity (see MPEP § 2106.05(d)(II)(i); “Receiving or transmitting data over a network, e.g., using the Internet to gather data”.
Taken alone or in combination, the additional elements of the claim do not provide an inventive concept and thus the claim is subject-matter ineligible.
Claim 12
Step 1 – Is the claim to a process, machine, manufacture or composition of matter?
A process, as above.
Step 2A Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
The claim recites the abstract ideas of the preceding claims from which it depends.
Step 2A Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
The claim recites the additional elements, “wherein the technical document includes at least one of a specification, a manual, a user guide and a standard, which are each associated with the Soc.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception?
Finally, the claim taken as a whole does not contain an inventive concept which provides significantly more than the abstract idea. The additional elements, “wherein the technical document includes at least one of a specification, a manual, a user guide and a standard, which are each associated with the Soc.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Taken alone or in combination, the additional elements of the claim do not provide an inventive concept and thus the claim is subject-matter ineligible.
Claim 13
Step 1 – Is the claim to a process, machine, manufacture or composition of matter?
A process, as above.
Step 2A Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
The claim recites the abstract ideas of the preceding claims from which it depends.
Step 2A Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
The claim recites the additional elements, “wherein each text fragment includes text data for at least one of a sentence, a paragraph, and a page.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception?
Finally, the claim taken as a whole does not contain an inventive concept which provides significantly more than the abstract idea. The additional elements, “wherein each text fragment includes text data for at least one of a sentence, a paragraph, and a page.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Taken alone or in combination, the additional elements of the claim do not provide an inventive concept and thus the claim is subject-matter ineligible.
Claim 14
Step 1 – Is the claim to a process, machine, manufacture or composition of matter?
A process, as above.
Step 2A Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
The claim recites the abstract ideas of the preceding claims from which it depends.
Step 2A Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
The claim recites the additional elements, “wherein the analyzing of the technical document includes extracting, when the technical document includes image, the image and convert the image into meaningful text data.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception?
Finally, the claim taken as a whole does not contain an inventive concept which provides significantly more than the abstract idea. The additional elements, “wherein the analyzing of the technical document includes extracting, when the technical document includes image, the image and convert the image into meaningful text data.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Taken alone or in combination, the additional elements of the claim do not provide an inventive concept and thus the claim is subject-matter ineligible.
Claim 15
Step 1 – Is the claim to a process, machine, manufacture or composition of matter?
A process, as above.
Step 2A Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
The claim recites, inter alia:
“wherein the converting of the multiple text fragments includes connecting text data with a label among the multiple text fragments to generate the dataset.” Under its broadest reasonable interpretation in light of the specification, this limitation encompasses the mental process of evaluating and observing data, which is an evaluation or observation that is practically capable of being performed in the human mind with the assistance of pen and paper. A human is able to evaluate text data and apply labels to that data. The limitation is merely applying an abstract idea on generic computer system. See MPEP 2106.04(a)(2)(III)(c).
Step 2A Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
This claim does not recite any additional limitations which integrate the abstract idea into a practical application.
Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception?
Finally, the claim taken as a whole does not contain an inventive concept which provides significantly more than the abstract idea and thus the claim is subject-matter ineligible.
Claim 16
Step 1 – Is the claim to a process, machine, manufacture or composition of matter?
A process, as above.
Step 2A Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
The claim recites the abstract ideas of the preceding claims from which it depends.
Step 2A Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
The claim recites the additional elements, “wherein each training epoch is executed based on the dataset and updates one or more model weights associated with particular performance metrics of the machine learning model.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception?
Finally, the claim taken as a whole does not contain an inventive concept which provides significantly more than the abstract idea. The additional elements, “wherein each training epoch is executed based on the dataset and updates one or more model weights associated with particular performance metrics of the machine learning model.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Taken alone or in combination, the additional elements of the claim do not provide an inventive concept and thus the claim is subject-matter ineligible.
Claim 17
Step 1 – Is the claim to a process, machine, manufacture or composition of matter?
A process, as above.
Step 2A Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
The claim recites the abstract ideas of the preceding claims from which it depends.
Step 2A Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
The claim recites the additional elements, “generating, by the machine learning model block, an acknowledgement signal indicating that the model weights are updated.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception?
Finally, the claim taken as a whole does not contain an inventive concept which provides significantly more than the abstract idea. The additional elements, “generating, by the machine learning model block, an acknowledgement signal indicating that the model weights are updated.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Taken alone or in combination, the additional elements of the claim do not provide an inventive concept and thus the claim is subject-matter ineligible.
Claim 18
Step 1 – Is the claim to a process, machine, manufacture or composition of matter?
A process, as above.
Step 2A Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
The claim recites the abstract ideas of the preceding claims from which it depends.
Step 2A Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
The claim recites the additional elements, “wherein the label includes at least one of a binary value and a probability value.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception?
Finally, the claim taken as a whole does not contain an inventive concept which provides significantly more than the abstract idea. The additional elements, “wherein the label includes at least one of a binary value and a probability value.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Taken alone or in combination, the additional elements of the claim do not provide an inventive concept and thus the claim is subject-matter ineligible.
Claim 19
Step 1 – Is the claim to a process, machine, manufacture or composition of matter?
A process, as above.
Step 2A Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
The claim recites the abstract ideas of the preceding claims from which it depends.
Step 2A Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
The claim recites the additional elements, “wherein the binary value includes one of a first binary value indicating that each text fragment does not have the required characteristic, and a second binary value indicating that each text fragment has the required characteristic.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception?
Finally, the claim taken as a whole does not contain an inventive concept which provides significantly more than the abstract idea. The additional elements, “wherein the binary value includes one of a first binary value indicating that each text fragment does not have the required characteristic, and a second binary value indicating that each text fragment has the required characteristic.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Taken alone or in combination, the additional elements of the claim do not provide an inventive concept and thus the claim is subject-matter ineligible.
Claim 20
Step 1 – Is the claim to a process, machine, manufacture or composition of matter?
A process, as above.
Step 2A Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
The claim recites the abstract ideas of the preceding claims from which it depends.
Step 2A Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
The claim recites the additional elements, “wherein the probability value indicates the extent of which each text fragment has the required characteristic.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception?
Finally, the claim taken as a whole does not contain an inventive concept which provides significantly more than the abstract idea. The additional elements, “wherein the probability value indicates the extent of which each text fragment has the required characteristic.” amounts to generic computer components used as a tool to perform an existing process. Thus, the additional element amounts to no more than a recitation of the words "apply it" (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer (see MPEP § 2106.05(f)).
Taken alone or in combination, the additional elements of the claim do not provide an inventive concept and thus the claim is subject-matter ineligible.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 3-11, and 13-20 are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al, (Wang et al, “LiLT: A Simple yet Effective Language-Independent Layout Transformer for Structured Document Understanding”, 2022, hereinafter “Wang”) in view of Liu et al, (Liu et al, “RoBERTa: A Robustly Optimized BERT Pretraining Approach” , 2019, hereinafter “Liu”).
Regarding claim 1, Wang discloses, “A system for analyzing documents to be used for verifying a system-on-a-chip (SoC), the system comprising:” (Introduction, pp. 7747; “Based on this inspiration, in this paper, we propose a simple yet effective Language-independent Layout Transformer (LiLT) for structured document understanding. In our framework, the text and layout information are first decoupled and joint optimized during pre-training, and then re-coupled for fine-tuning. To ensure that the two modalities have sufficient language-independent interaction, we further propose a novel bi-directional attention complementation mechanism (BiACM) to enhance the cross-modality cooperation.” Wang proposes a system that is able to use machine learning methods to analyze and evaluate structured documents and images.)
“a parser configured to analyze a technical document associated with the Soc to generate multiple text fragments;” (Text Embedding, pp. 7748; “Following the common practice (Devlin et al., 2019; Xu et al., 2020), in the text flow, all text strings in the OCR results are first tokenized and concatenated as a sequence St by sorting the corresponding text bounding boxes from the top-left to bottom-right. Intuitively, the special tokens [CLS] and [SEP] are also added at the beginning and end of the sequence respectively. After this, St will be truncated or padded with extra [PAD] tokens until its length equals the maximum sequence length N.” This system will evaluate document images and text. This system will OCR documents and tokenize the data. This teaches a method that can evaluate Text data and parse, or tokenize, that data into tokens or text fragments.)
“a data preparator configured to convert the multiple text fragments into a dataset suitable for a machine learning model;” (Layout Embedding, pp. 7748-7749; “As for the layout flow, we construct a 2D position sequence Sl with the same length as the token sequence St using the corresponding text bounding boxes. To be specific, we normalize and discretize all box coordinates to integers in the range [0, 1000], and use four embedding layers to generate x-axis, y-axis, height, and width features separately. Given the normalized bounding boxes
B
=
(
x
m
i
n
,
x
m
a
x
,
Y
m
i
n
,
x
m
a
x
,
w
i
d
t
h
,
h
e
i
g
h
t
), the 2D positional embedding
P
2
D
∈
R
N
×
d
l
(where dL is the number of layout feature dimension) is constructed as follows: [See Equation (2)].” Wang proposes a system that will also evaluate the layout of a document. This system will tokenize the text and embed the layout information. Wang teaches a system that tis able to take document data and embed, or convert it, into a useable form for further machine learning inference.)
“a machine learning model block configured to receive the dataset and perform a training mode including multiple training epochs or an inference mode; and” (Figure, 2, pp. 7749; This image shows the overall framework of the proposed work. This shows the transformer layers, which is interpreted to be the machine learning block. This system requires training of the individual models and of the overall model using different training methods including self-supervised learning. This system is also designed to execute inferences where it will take documents images and process them.)
PNG
media_image1.png
553
892
media_image1.png
Greyscale
“a documents processor, wherein the technical document includes at least one of a labeled document for the training mode and an unlabeled document for the inference mode,” (Pre-training Tasks, pp. 7750; “We conduct three self-supervised pre-training tasks to guide the model to autonomously learn joint representations with cross-modal cooperation.” This article discloses a training mode which requires multiple different pre training tasks.) and (Appendix, pp. 7757; “RVL-CDIP RVL-CDIP (Harley et al., 2015) consists of 400,000 gray-scale images of English documents, with 8:1:1 for the training set, validation set, and test set. A multi-class single-label classification task is defined on RVL-CDIP. The images are categorized into 16 classes, with 25,000 images per class. The evaluation metric is the overall classification accuracy as shown in Table 5. Text and layout information are extracted by TextIn API.” This article uses different datasets for training and testing their model. This is an example of one of the databases uses which contains labeled image data contained in training datasets.) and (LiLT, pp. 7748; “Figure 2 shows the overall illustration of our method. Given an input document image, we first use off-the-shelf OCR engines to get text bounding boxes and contents. Then, the text and layout information are separately embedded and fed into the corresponding Transformer-based architecture to obtain enhanced features. Bi-directional attention complementation mechanism (BiACM) is introduced to accomplish the cross-modality interaction of text and layout clues. Finally, the encoded text and layout features are concatenated and additional heads are added upon them, for the self-supervised pre-training or the downstream fine-tuning.” The proposed system LiLT will function in inference mode. This model, after it is trained, will take in a document image, which is an unlabeled document, and process it.)
“in the inference mode, make a prediction for each text fragment of the dataset based on the training result of the machine learning model to generate the dataset with prediction values, and” (BiACM, pp. 7749; “The text embedding ET and layout embedding EL are fed into their respective sub-models to generate high-level enhanced features. However, it will considerably ignore the cross-modal interaction process if we simply combine the text and layout features at the encoder output only. The network also needs to comprehensively analyze them at earlier stages. In view of this, we propose a new bi-directional attention complementation mechanism (BiACM) to strengthen the cross-modality interaction across the entire encoding pipeline.” Wang teaches a system that is able to take in structured documents and evaluate them. This system is trained and uses pretrained models during the inference mode to produce output data.)
PNG
media_image1.png
553
892
media_image1.png
Greyscale
“wherein the documents processor is configured to receive the unlabeled document and the dataset with the prediction values, and” (Fig. 2, pp. 7749; This system, which is interpreted to the the document processor, is able to take a document image on the left and process it. This system also uses pretrained systems such as RoBERTa models which it will feed the OCR datasets into the transformation layers and then into the LLM’s for further evaluation.)
“convert the source unlabeled document based on the dataset with the prediction values into a source labeled document used for verifying the soc.” (Text Embedding, pp. 7748; “Following the common practice (Devlin et al., 2019; Xu et al., 2020), in the text flow, all text strings in the OCR results are first tokenized and concatenated as a sequence St by sorting the corresponding text bounding boxes from the top-left to bottom-right. Intuitively, the special tokens [CLS] and [SEP] are also added at the beginning and end of the sequence respectively. After this, St will be truncated or padded with extra [PAD] tokens until its length equals the maximum sequence length N. Finally, we sum the token embedding
E
t
o
k
e
n
of St and the 1D positional embedding
P
1
D
to obtain the text embedding
E
T
∈
R
N
×
d
T
as: [See Equation (1)] where dT is the number of text feature dimension and LN is the layer normalization (Ba et al., 2016).” During the inference the system will process the document image and OCR the text. The system will then embed the tokens to be used in different transformer layers and other LLMs for further processing. After this the labeled tokens are further evaluated for key points and other relational data.)
Wang fails to explicitly disclose:
“wherein the machine learning model block is configured to: in each training epoch of the training mode, train the machine learning model based on the dataset such that a label is appended to each text fragment, and”
However, Liu discloses, “wherein the machine learning model block is configured to: in each training epoch of the training mode, train the machine learning model based on the dataset such that a label is appended to each text fragment, and” (RoBERTa, pp. 6; “In the previous section we propose modifications to the BERT pretraining procedure that improve end-task performance. We now aggregate these improvements and evaluate their combined impact. We call this configuration RoBERTa for Robustly optimized BERT approach. Specifically, RoBERTa is trained with dynamic masking (Section 4.1), FULL-SENTENCES without NSP loss (Section 4.2), large mini-batches (Section 4.3) and a larger byte-level BPE (Section 4.4).” Wang uses a RoBERTa to evaluate the text after it is extracted from the document images and processed by the transformer layers. RoBERT is pretrained using many training iterations. RoBERTa will also label the text data as it evaluates the unlabeled training set. See Section 4.3 & 4.4, pp. 5-6.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Wang and Liu. Wang teaches a system that uses multiple different LLM’s to evaluate structured documents. Liu teaches a language model that uses different training methods and modifies the BERT structure for processing text. One of ordinary skill would have motivation to combine both of these articles because the authors of Wang uses the model disclosed in Liu as part of their LiLT model. If one of ordinary skill in the art was researching articles on language models and structured documents evaluation one would reasonably discover one or both of these articles, “In this paper, we initialize the text flow from the existing pre-trained English RoBERTaBASE (Liu et al., 2019b) for our document pre-training, and combine LiLTBASE with the pre-trained InfoXLMBASE (Chi et al., 2021)/a new pre-trained RoBERTaBASE for multilingual/monolingual finetuning.” (Wang, Pre-training Setting, pp. 7750).
Regarding claim 3, Wang discloses “wherein each text fragment includes text data for at least one of a sentence, a paragraph, and a page.” (Text Embedding, pp. 7748; “Following the common practice (Devlin et al., 2019; Xu et al., 2020), in the text flow, all text strings in the OCR results are first tokenized and concatenated as a sequence St by sorting the corresponding text bounding boxes from the top-left to bottom-right. Intuitively, the special tokens [CLS] and [SEP] are also added at the beginning and end of the sequence respectively. After this, St will be truncated or padded with extra [PAD] tokens until its length equals the maximum sequence length N.” This article teaches a system that will tokenize the text from the input document image. This process converts text fragments into data, or tokens, which the machine learning blocks can use.)
Regarding claim 4, Wang discloses, “wherein, when the technical document includes image, the parser is configured to extract the image and convert the image into meaningful text data.” (LiLT, pp. 7748; “Figure 2 shows the overall illustration of our method. Given an input document image, we first use off-the-shelf OCR engines to get text bounding boxes and contents. Then, the text and layout information are separately embedded and fed into the corresponding Transformer-based architecture to obtain enhanced features. Bi-directional attention complementation mechanism (BiACM) is introduced to accomplish the cross-modality interaction of text and layout clues. Finally, the encoded text and layout features are concatenated and additional heads are added upon them, for the self-supervised pre-training or the downstream fine-tuning.” As stated above the process in Wang will intake document images and evaluate them.)
Regarding claim 5,
PNG
media_image2.png
452
827
media_image2.png
Greyscale
Wang discloses, “wherein the data preparator is configured to connect text data with a label among the multiple text fragments to generate the dataset.” (Figure. 2, pp. 7749; This system will use the OCR engines to process the document initially and then the system will organize and tokenize the text. The system will also embed the text and layout information for the machine learning models to use.)
Regarding claim 6, Liu discloses, “wherein the machine learning model block is configured to execute each training epoch based on the dataset and update one or more model weights associated with particular performance metrics of the machine learning model.” (Optimization, pp. 2; “BERT is optimized with Adam (Kingma and Ba, 2015) using the following parameters: β1 = 0.9, β2 = 0.999, e = 1e-6 and L2 weight decay of 0.01. The learning rate is warmed up over the first 10,000 steps to a peak value of 1e-4, and then linearly decayed. BERT trains with a dropout of 0.1 on all layers and attention weights, and a GELU activation function (Hendrycks and Gimpel, 2016). Models are pretrained for S = 1,000,000 updates, with minibatches containing B = 256 sequences of maximum length T = 512 tokens.” The model in this article is further trained using an Adam Optimizer. This is an iterative process which would update the weights and parameters based on this input training dataset. The Adam Optimizer is a process that uses performance metrices to train the model. The model RoBERTa is also trained using this process.)
Regarding claim 7, Liu discloses, “wherein the machine learning model block is configured to generate an acknowledgement signal indicating that the model weights are updated.” (RoBERTa, pp. 6-7; “To help disentangle the importance of these factors from other modeling choices (e.g., the pretraining objective), we begin by training RoBERTa following the BERTLARGE architecture (L = 24, H = 1024, A = 16, 355M parameters). We pretrain for 100K steps over a comparable BOOKCORPUS plus WIKIPEDIA dataset as was used in Devlin et al. (2019). We pretrain our model using 1024 V100 GPUs for approximately one day.” The process in Liu uses many training iterations to train their base model. This process requires time and when complete can produce an output signifying the completion of the training.)
Regarding claim 8, Wang discloses, “wherein the label includes at least one of a binary value and a probability value.” (Cross-modal Alignment Identification, pp. 7750; “We collect those encoded features of token-box pairs that are masked and further replaced (misaligned) or kept unchanged (aligned) by MVLM and KPL, and build an additional head upon them to identify whether each pair is aligned. To achieve this, the model is required to learn the cross-modal perception capacity. CAI is a binary classification task, and a cross-entropy loss is applied for it.” During the fine-tuning process discloses in Wang, the model will evaluate the mutated data and determine key points and relational data of the document. This process will evaluate the token and use a binary classification model to determine key points and/or relational data.) and (XFUND, pp. 7757; “We focus on the semantic entity recognition (SER) and relation extraction (RE) tasks defined in the original paper (Xu et al., 2021b). Relation extraction aims to predict the relation between any two given semantic entities, and we mainly focus on the key-value relation extraction.” This is a description of the relational data evaluated after the document has been scanned and embedded. They Key-value relation information is interpreted to be the probability value.)
Regarding claim 9, Wang discloses, “wherein the binary value includes one of a first binary value indicating that each text fragment does not have the required characteristic, and a second binary value indicating that each text fragment has the required characteristic.” (Key Point Location, pp. 7750; “We propose this task to make the model better understand layout information in the structured documents. KPL equally divides the entire layout into several regions (we set 7×7=49 regions by default) and randomly masks some of the input bounding boxes. The model is required to predict which regions the key points (top-left corner, bottom-right corner, and center point) of each box belong to using separate heads. To deal with it, the model is required to fully understand the text content and know where to put a specific word/sentence when the surrounding ones are given. We mask 15% boxes, among which 80% are replaced by (0,0,0,0,0,0), 10% are replaced by random boxes sampled from the same batch, and 10% remain the same. Cross entropy loss is adopted.” This system will evaluate embedded datasets of document data for relational data. This will evaluate the tokens and notate, using a 0 or a 1, see fig 2, to signify if a key entry is found in the document.)
Regarding claim 10, Wang discloses, “wherein the probability value indicates the extent of which each text fragment has the required characteristic.” (Fig. 2, pp. 7748; As seen on the right side of the image in the fine tuning tasks the embedded set of datapoints are evaluated for relational data and key-value data. The system will denote a 0 or 1 depending on the token that is evaluated.)
PNG
media_image2.png
452
827
media_image2.png
Greyscale
Regarding claim 11, Wang discloses, “A method for analyzing documents to be used for verifying a system-on-a-chip (SoC), the method comprising:” (Introduction, pp. 7747; “Based on this inspiration, in this paper, we propose a simple yet effective Language-independent Layout Transformer (LiLT) for structured document understanding. In our framework, the text and layout information are first decoupled and joint optimized during pre-training, and then re-coupled for fine-tuning.” This article discloses a process for evaluating structured documents.)
“analyzing, by a parser, a technical document associated with the Soc to generate multiple text fragments;” (Text Embedding, pp. 7748; “Following the common practice (Devlin et al., 2019; Xu et al., 2020), in the text flow, all text strings in the OCR results are first tokenized and concatenated as a sequence St by sorting the corresponding text bounding boxes from the top-left to bottom-right. Intuitively, the special tokens [CLS] and [SEP] are also added at the beginning and end of the sequence respectively. After this, St will be truncated or padded with extra [PAD] tokens until its length equals the maximum sequence length N.” This system will evaluate document images and text. This system will OCR documents and tokenize the data. This teaches a method that can evaluate Text data and parse, or tokenize, that data into tokens or text fragments.)
“converting, by a data preparator, the multiple text fragments into a dataset suitable for a machine learning model;” (Layout Embedding, pp. 7748-7749; “As for the layout flow, we construct a 2D position sequence Sl with the same length as the token sequence St using the corresponding text bounding boxes. To be specific, we normalize and discretize all box coordinates to integers in the range [0, 1000], and use four embedding layers to generate x-axis, y-axis, height, and width features separately. Given the normalized bounding boxes
B
=
(
x
m
i
n
,
x
m
a
x
,
Y
m
i
n
,
x
m
a
x
,
w
i
d
t
h
,
h
e
i
g
h
t
), the 2D positional embedding
P
2
D
∈
R
N
×
d
l
(where dL is the number of layout feature dimension) is constructed as follows: [See Equation (2)].” Wang proposes a system that will also evaluate the layout of a document. This system will tokenize the text and embed the layout information. Wang teaches a system that tis able to take document data and embed, or convert it, into a useable form for further machine learning inference.)
“receiving, by a machine learning model block, the dataset and performing a training mode including multiple training epochs or an inference mode,” (Figure, 2, pp. 7749; This image shows the overall framework of the proposed work. This shows the transformer layers, which is interpreted to be the machine learning block. This system requires training of the individual models and of the overall model using different training methods including self-supervised learning. This system is also designed to execute inferences where it will take
PNG
media_image1.png
553
892
media_image1.png
Greyscale
documents images and process them.)
“wherein the technical document includes at least one of a labeled document for the training mode and an unlabeled document for the inference mode,” (Pre-training Tasks, pp. 7750; “We conduct three self-supervised pre-training tasks to guide the model to autonomously learn joint representations with cross-modal cooperation.” This article discloses a training mode which requires multiple different pre training tasks.) and (Appendix, pp. 7757; “RVL-CDIP RVL-CDIP (Harley et al., 2015) consists of 400,000 gray-scale images of English documents, with 8:1:1 for the training set, validation set, and test set. A multi-class single-label classification task is defined on RVL-CDIP. The images are categorized into 16 classes, with 25,000 images per class. The evaluation metric is the overall classification accuracy as shown in Table 5. Text and layout information are extracted by TextIn API.” This article uses different datasets for training and testing their model. This is an example of one of the databases uses which contains labeled image data contained in training datasets.) and (LiLT, pp. 7748; “Figure 2 shows the overall illustration of our method. Given an input document image, we first use off-the-shelf OCR engines to get text bounding boxes and contents. Then, the text and layout information are separately embedded and fed into the corresponding Transformer-based architecture to obtain enhanced features. Bi-directional attention complementation mechanism (BiACM) is introduced to accomplish the cross-modality interaction of text and layout clues. Finally, the encoded text and layout features are concatenated and additional heads are added upon them, for the self-supervised pre-training or the downstream fine-tuning.” The proposed system LiLT will function in inference mode. This model, after it is trained, will take in a document image, which is an unlabeled document, and process it.)
“wherein the inference mode includes making a prediction for each text fragment of the dataset based on the training result of the machine learning model to generate the dataset with prediction values; and” (BiACM, pp. 7749; “The text embedding ET and layout embedding EL are fed into their respective sub-models to generate high-level enhanced features. However, it will considerably ignore the cross-modal interaction process if we simply combine the text and layout features at the encoder output only. The network also needs to comprehensively analyze them at earlier stages. In view of this, we propose a new bi-directional attention complementation mechanism (BiACM) to strengthen the cross-modality interaction across the entire encoding pipeline.” Wang teaches a system that is able to take in structured documents and evaluate them. This system is trained and uses pretrained models during the inference mode to produce output data.)
PNG
media_image1.png
553
892
media_image1.png
Greyscale
“receiving, by a documents processor, the unlabeled document and the dataset with the prediction values, and” (Fig. 2, pp. 7749; This system, which is interpreted to the the document processor, is able to take a document image on the left and process it. This system also uses pretrained systems such as RoBERTa models which it will feed the OCR datasets into the transformation layers and then into the LLM’s for further evaluation.)
“converting the source unlabeled document based on the dataset with the prediction values into a source labeled document used for verifying the Soc.” (Text Embedding, pp. 7748; “Following the common practice (Devlin et al., 2019; Xu et al., 2020), in the text flow, all text strings in the OCR results are first tokenized and concatenated as a sequence St by sorting the corresponding text bounding boxes from the top-left to bottom-right. Intuitively, the special tokens [CLS] and [SEP] are also added at the beginning and end of the sequence respectively. After this, St will be truncated or padded with extra [PAD] tokens until its length equals the maximum sequence length N. Finally, we sum the token embedding
E
t
o
k
e
n
of St and the 1D positional embedding
P
1
D
to obtain the text embedding
E
T
∈
R
N
×
d
T
as: [See Equation (1)] where dT is the number of text feature dimension and LN is the layer normalization (Ba et al., 2016).” During the inference the system will process the document image and OCR the text. The system will then embed the tokens to be used in different transformer layers and other LLMs for further processing. After this the labeled tokens are further evaluated for key points and other relational data.)
Wang fails to explicitly disclose:
“wherein each training epoch of the training mode includes training the machine learning model based on the dataset such that a label is appended to each text fragment, and”
However, Liu discloses, “wherein each training epoch of the training mode includes training the machine learning model based on the dataset such that a label is appended to each text fragment, and” (RoBERTa, pp. 6; “In the previous section we propose modifications to the BERT pretraining procedure that improve end-task performance. We now aggregate these improvements and evaluate their combined impact. We call this configuration RoBERTa for Robustly optimized BERT approach. Specifically, RoBERTa is trained with dynamic masking (Section 4.1), FULL-SENTENCES without NSP loss (Section 4.2), large mini-batches (Section 4.3) and a larger byte-level BPE (Section 4.4).” Wang uses a RoBERTa to evaluate the text after it is extracted from the document images and processed by the transformer layers. RoBERT is pretrained using many training iterations. RoBERTa will also label the text data as it evaluates the unlabeled training set. See Section 4.3 & 4.4, pp. 5-6.)
Regarding claim 13, Wang discloses, “wherein each text fragment includes text data for at least one of a sentence, a paragraph, and a page.” (Text Embedding, pp. 7748; “Following the common practice (Devlin et al., 2019; Xu et al., 2020), in the text flow, all text strings in the OCR results are first tokenized and concatenated as a sequence St by sorting the corresponding text bounding boxes from the top-left to bottom-right. Intuitively, the special tokens [CLS] and [SEP] are also added at the beginning and end of the sequence respectively. After this, St will be truncated or padded with extra [PAD] tokens until its length equals the maximum sequence length N.” This article teaches a system that will tokenize the text from the input document image. This process converts text fragments into data, or tokens, which the machine learning blocks can use.)
Regarding claim 14, Wang discloses, “wherein the analyzing of the technical document includes extracting, when the technical document includes image, the image and convert the image into meaningful text data.” (LiLT, pp. 7748; “Figure 2 shows the overall illustration of our method. Given an input document image, we first use off-the-shelf OCR engines to get text bounding boxes and contents. Then, the text and layout information are separately embedded and fed into the corresponding Transformer-based architecture to obtain enhanced features. Bi-directional attention complementation mechanism (BiACM) is introduced to accomplish the cross-modality interaction of text and layout clues. Finally, the encoded text and layout features are concatenated and additional heads are added upon them, for the self-supervised pre-training or the downstream fine-tuning.” As stated above the process in Wang will intake document images and evaluate them.)
Regarding claim 15, Wang discloses, “wherein the converting of the multiple text fragments includes connecting text data with a label among the multiple text fragments to generate the dataset.” (Figure. 2, pp. 7749; This system will use the OCR engines to process the document initially and then the system will organize and tokenize the text. The system will also
PNG
media_image2.png
452
827
media_image2.png
Greyscale
embed the text and layout information for the machine learning models to use.)
Regarding claim 16, Liu discloses, “wherein each training epoch is executed based on the dataset and updates one or more model weights associated with particular performance metrics of the machine learning model.” (Optimization, pp. 2; “BERT is optimized with Adam (Kingma and Ba, 2015) using the following parameters: β1 = 0.9, β2 = 0.999, e = 1e-6 and L2 weight decay of 0.01. The learning rate is warmed up over the first 10,000 steps to a peak value of 1e-4, and then linearly decayed. BERT trains with a dropout of 0.1 on all layers and attention weights, and a GELU activation function (Hendrycks and Gimpel, 2016). Models are pretrained for S = 1,000,000 updates, with minibatches containing B = 256 sequences of maximum length T = 512 tokens.” The model in this article is further trained using an Adam Optimizer. This is an iterative process which would update the weights and parameters based on this input training dataset. The Adam Optimizer is a process that uses performance metrices to train the model. The model RoBERTa is also trained using this process.)
Regarding claim 17, Liu discloses, “generating, by the machine learning model block, an acknowledgement signal indicating that the model weights are updated.” (RoBERTa, pp. 6-7; “To help disentangle the importance of these factors from other modeling choices (e.g., the pretraining objective), we begin by training RoBERTa following the BERTLARGE architecture (L = 24, H = 1024, A = 16, 355M parameters). We pretrain for 100K steps over a comparable BOOKCORPUS plus WIKIPEDIA dataset as was used in Devlin et al. (2019). We pretrain our model using 1024 V100 GPUs for approximately one day.” The process in Liu uses many training iterations to train their base model. This process requires time and when complete can produce an output signifying the completion of the training.)
Regarding claim 18, Wang discloses, “wherein the label includes at least one of a binary value and a probability value.” (Cross-modal Alignment Identification, pp. 7750; “We collect those encoded features of token-box pairs that are masked and further replaced (misaligned) or kept unchanged (aligned) by MVLM and KPL, and build an additional head upon them to identify whether each pair is aligned. To achieve this, the model is required to learn the cross-modal perception capacity. CAI is a binary classification task, and a cross-entropy loss is applied for it.” During the fine-tuning process discloses in Wang, the model will evaluate the mutated data and determine key points and relational data of the document. This process will evaluate the token and use a binary classification model to determine key points and/or relational data.) and (XFUND, pp. 7757; “We focus on the semantic entity recognition (SER) and relation extraction (RE) tasks defined in the original paper (Xu et al., 2021b). Relation extraction aims to predict the relation between any two given semantic entities, and we mainly focus on the key-value relation extraction.” This is a description of the relational data evaluated after the document has been scanned and embedded. They Key-value relation information is interpreted to be the probability value.)
Regarding claim 19, Wang discloses, “wherein the binary value includes one of a first binary value indicating that each text fragment does not have the required characteristic, and a second binary value indicating that each text fragment has the required characteristic.” (Key Point Location, pp. 7750; “We propose this task to make the model better understand layout information in the structured documents. KPL equally divides the entire layout into several regions (we set 7×7=49 regions by default) and randomly masks some of the input bounding boxes. The model is required to predict which regions the key points (top-left corner, bottom-right corner, and center point) of each box belong to using separate heads. To deal with it, the model is required to fully understand the text content and know where to put a specific word/sentence when the surrounding ones are given. We mask 15% boxes, among which 80% are replaced by (0,0,0,0,0,0), 10% are replaced by random boxes sampled from the same batch, and 10% remain the same. Cross entropy loss is adopted.” This system will evaluate embedded datasets of document data for relational data. This will evaluate the tokens and notate, using a 0 or a 1, see fig 2, to signify if a key entry is found in the document.)
PNG
media_image2.png
452
827
media_image2.png
Greyscale
Regarding claim 20, Wang discloses, “wherein the probability value indicates the extent of which each text fragment has the required characteristic.” (Fig. 2, pp. 7748; As seen on the right side of the image in the fine tuning tasks the embedded set of datapoints are evaluated for relational data and key-value data. The system will denote a 0 or 1 depending on the token that is evaluated.)
Claims 2 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Wanf and Liu in view of Tsigkanos et al, (Tsigkanos et al, “Variable Discovery with Large Language Models for Metamorphic Testing of Scientific Software”, 2023, hereinafter “Tsigkanos”).
Regarding claim 2, Tsigkanos discloses, “wherein the technical document includes at least one of a specification, a manual, a user guide and a standard, which are each associated with the Soc.” (State of the Art, pp. 323; “With these approaches, metamorphic relations are learned directly from the behavior of the program under test, which means that all found relations can only be used for regression testing of future program versions. Our approach is to learn metamorphic relations from auxiliary documents such as user manuals, and is thus more general.” Tsigkanos teaches a system that uses LLM’s to evaluate technical documents and extract relevant information. The model in this article uses BERT based models, similar to RoBERTa proposed in Liu. This system will evaluate software technical documents to ensure the software is compatible with given systems)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Wang, Liu and Tsigkanos. Wang teaches a system that uses multiple different LLM’s to evaluate structured documents. Liu teaches a language model that uses different training methods and modifies the BERT structure for processing text. Tsigkanos teaches a system that is able to evaluate technical documents for software products using language models. One of ordinary skill would have motivation to combine both of these articles because the authors of Wang uses the model disclosed in Liu as part of their LiLT model. If one of ordinary skill in the art was researching articles on language models and structured documents evaluation, one would reasonably discover one or both of these articles. Next, one would have motivation to modify or combine another machine learning model that is able to use LLM’s and similar training processes to evaluate documents to locate key points in technical documents related to software, “Experiment results are summarized in Table 1. We obtain a best accuracy of 0.88 (SWMM case), 0.91 (SWAT case) and 0.90 (MODFLOW case). Accuracy, up to the second digit is the same for both partial and exact matches. Table 1 further shows in detail the number of true positives for exact and partial matches and accuracy achieved over the unique lemmatized words of the source document, for best and worst temperature values.” (Tsigkanos, Results, pp. 329-330)
Regarding claim 12, Tsigkanos discloses, “wherein the technical document includes at least one of a specification, a manual, a user guide and a standard, which are each associated with the Soc.” (State of the Art, pp. 323; “With these approaches, metamorphic relations are learned directly from the behavior of the program under test, which means that all found relations can only be used for regression testing of future program versions. Our approach is to learn metamorphic relations from auxiliary documents such as user manuals, and is thus more general.” Tsigkanos teaches a system that uses LLM’s to evaluate technical documents and extract relevant information. The model in this article uses BERT based models, similar to RoBERTa proposed in Liu. This system will evaluate software technical documents to ensure the software is compatible with given systems)
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PAUL MICHAEL GALVIN-SIEBENALER whose telephone number is (571)272-1257. The examiner can normally be reached Monday - Friday 8AM to 5PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached at (571) 270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PAUL M GALVIN-SIEBENALER/Examiner, Art Unit 2147
/ERIC NILSSON/Primary Examiner, Art Unit 2151