DETAILED ACTION
This communication is responsive to Amendment filed 06/02/2026.
Claims 1 and 3-8 are pending in this application. In the Amendment, claims 1 and 7-8 are amended. This action is made Final.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant's arguments filed 06/02/2026 have been fully considered but they are not persuasive.
Applicant argued Sirangimoorthy discloses that these intermediate documents are sequential stages of text extraction generated internally by a single "document converter 142". Sirangimoorthy does not teach routing these specific files back and forth between a distinct buyer-side computer and a web service provider.
Per A) The Examiner respectfully disagrees in response to applicant's arguments against the references individually, one cannot show nonobviousness by attacking references individually where the rejections are based on combinations of references. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981); In re Merck & Co., 800 F.2d 1091, 231 USPQ 375 (Fed. Cir. 1986). Jayaraman teaches the routing of files back and forth. Sirangimoorthy is combined to teach the further files which the process of generating and sending may be done partially locally or remotely (Sirangimoorthy, para.67, instructions may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server).
Applicant argued Sirangimoorthy does not teach that these intermediate documents comprise the specifically claimed contents, namely neural network classifications for segments of unstructured data, concatenated modified text (text extracted from the segments of unstructured data concatenated with the classifications), and extracted key terms from the concatenated modified text.
Per B) The Examiner respectfully disagrees in response to applicant's arguments against the references individually, one cannot show nonobviousness by attacking references individually where the rejections are based on combinations of references. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981); In re Merck & Co., 800 F.2d 1091, 231 USPQ 375 (Fed. Cir. 1986). Liang teaches a text analysis method that includes concatenating, by the processor of the buyer-side computer, classifications with text to generate modified text for each of the plurality of segments (Liang, para.23, 137-139, 166, 215, entity 514 may modify text tag).
Applicant argued that selecting an ontology based on keywords does not teach or suggest selecting an extraction model from a plurality of extraction models based on a neural network classification of a segment.
Per C) The Examiner respectfully disagrees as Sirangimoorthy teaches selecting the extraction model based on the classification of each of the plurality of segments. The documents are each associated with a domain and based on the classified domain a domain-specific ontology is used for extraction (Sirangimoorthy, para.26, 35, 39, 58, document analyzer extracts based on domain-specific ontology determined by key words).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1 and 5-8 are rejected under 35 U.S.C. 103 as being unpatentable over Jayaraman et al. (“Jayaraman”, US 2021/0256097) in view of Priestas et al. (“Priestas”, US 2019/0005012) in view of Liang et al. (“Liang” US 2018/0060302) and further in view of Sirangimoorthy et al. (“Sirangimoorthy”, US 2021/0248153).
As per claim 1, Jayaraman teaches a computer-implemented method of transforming an unstructured set of data to a structured set of data (Jayaraman, para.13, extract data into structured format), wherein the unstructured set of data comprises a common file format document (Jayaraman, para.14, 33, input file formats i.e. DOC/PDF/RTF/TXT); the method comprising:
receiving, by at least one processor of a web service provider (Jayaraman, para.15-16, 20, document content extraction platform 100 communicatively connected via Internet), a plurality of segments of the unstructured set of data, wherein at least a portion of the plurality of segments are clauses of a document that contain text (Jayaraman, para.12, 22-23, document preprocessing of contract document that includes clauses and identifying relevant text blocks);
generating, by the at least one processor of the web service provider, a JSON (Java Script Object Notation) file that includes locations of the segments within the common file format document (Jayaraman, para.19-20, 25, 30, 36, row level identifiers; extracted entities located and classified in JSON format);
forwarding, by the at least one processor of the web service provider to a processor of a buyer-side computer, the JSON file (Jayaraman, para.19-20, interface 104 used to transfer data from storage 112 to device 102 where data storage 112 is separately connected to platform 100; para.30, structured records viewable by user interface 104);
adding, by the processor of the buyer-side computer, bookmarks to the common file format document (Jayaraman, para.20, 22-23, 31, 36-38, detected text blocks and table structures tagged with unique identifier to identify relevant content; mark relevant text);
forwarding, by the processor of the buyer-side computer to the at least one processor of the web service provider, the text of each segment extracted from the common file format document (Jayaraman, para.20, 25, 31, 33-34, 38, feedback updates record);
classifying, by at least one processor of the web service provider, each of the plurality segments by applying each segment to a neural network classification model (Jayaraman, para.17-18, 20, 25-26, 30, 36-38, entity classified into categories by machine learning model i.e. R-CNN; designated text blocks);
generating, by the at least one processor of the web service provider, a JSON file that includes the classifications for the plurality of segments and forwarding the JSON file to the processor of the buyer-side computer (Jayaraman, para.20, 38, text blocks with relevant entities tagged and passed on for further processing);
extracting key terms from the text of the at least some of the plurality of segments by the at least one processor of the web service provider using an extraction model (Jayaraman, para.17, 20, 30, 38, extracting key/value pairs corresponding to entities), and
generating, by the at least one processor of the web service provider, a JSON file that includes the extracted key terms for the at least some of the plurality of segments and sending the JSON file to the processor of the buyer-side computer (Jayaraman, para.19-20, 30, 38, structured stored records store extracted key/value pairs corresponding to entities).
However, Jayaraman does not explicitly teach the bookmarks indicating the segments and extracting, by the processor of the buyer-side computer, the text of the segment from the common file format document using the bookmarks. Priestas teaches a document extraction method that includes bookmarking to indicate the segments of a document (Priestas, para.17, 25-29, 39-41, markup tags) and extracting, by the processor of the buyer-side computer, the text of the segment from the common file format document using the bookmark (Priestas, para.40, 48, 53, text extractor 208 extracts text from markup file). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include Priestas’ teaching with Jayaraman’s method in order to extract data based on tagged sections.
Furthermore, the method of Jayaraman and Priestas does not explicitly teach concatenating, by the processor of the buyer-side computer, the classifications with the text to generate modified text for each of the plurality of segments. Liang teaches a text analysis method that includes concatenating, by the processor of the buyer-side computer, classifications with text to generate modified text for each of the plurality of segments (Liang, para.23, 137-139, 166, 215, entity 514 may modify text tag). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include Liang’s teaching with the method of Jayaraman and Priestas in order to provide user feedback.
Additionally, the method of Jayaraman, Priestas and Liang does not teach generating a first further JSON file and forwarding the first further JSON file to the processor of the buyer-side computer; generating, by the processor of the buyer-side computer, a second further JSON file comprising the modified text for at least some of the plurality of segments, and sending the second further JSON file to the at least one processor of the web service provider; the extraction model selected from a plurality of extraction models based on the classification of each of the plurality of segments; generating, a third further JSON file and sending the third further JSON file to the processor of the buyer-side computer; and generating the structured set of data, by the processor of the buyer-side computer using the modified text of the classified segments and the extracted key terms in the third further JSON file.
Sirangimoorthy teaches a method of generating structured documents that includes generating a first further JSON file and forwarding the first further JSON file to the processor of the buyer-side computer (Sirangimoorthy, para.27-33, 38, 57-58, document converter 142 generates first/second/third intermediate documents which are structured by the structured document generator 144); generating, by the processor of the buyer-side computer, a second further JSON file comprising the modified text for at least some of the plurality of segments (Sirangimoorthy, para.21, 28-31, second intermediate document), and sending the second further JSON file to the at least one processor of the web service provider (Sirangimoorthy, para.20, 23, 27-33, 36, 38, 40, 57, intermediate plaintext i.e. JSON document with location of extracted text; document converter 142 forwards intermediate document; Fig.1, para. 20, 67 remote application server 130); the extraction model selected from a plurality of extraction models based on the classification of each of the plurality of segments (Sirangimoorthy, para.26, 35, 39, 58, document analyzer extracts based on domain-specific ontology determined by key words); generating, a third further JSON file and sending the third further JSON file to the processor of the buyer-side computer (Sirangimoorthy, para.27-33, 38, 57-58, document converter 142 generates first/second/third intermediate documents which are structured by the structured document generator 144); and generating the structured set of data, by the processor of the buyer-side computer using the modified text of the classified segments and the extracted key terms in the third further JSON file (Sirangimoorthy, para.21, 28-33, 38, 41, 58, structured document generator uses intermediate documents). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include Sirangimoorthy’s teaching with the method of Jayaraman, Priestas and Liang in order to extract data based on context and modify structured records.
As per claim 5, the method of Jayaraman, Priestas, Liang and Sirangimoorthy teaches the method of claim 1, comprising performing one or more of the following, by the at least one processor of the web service provider: annotating the segments; clustering a plurality of structured sets of data; generating a sematic meaning for the segments; scoring the segments and/or the structured set of data; querying the structured set of data; normalizing the data of the structured set of data; navigating the segments using the logical grouping of segments having the same classification (Sirangimoorthy, para.25, 45, navigating segments).
As per claim 6, the method of Jayaraman, Priestas, Liang and Sirangimoorthy teaches the method according to claim 1, wherein the unstructured set of data and the structured set of data correspond to a contract document (Jayaraman, para.22, contract).
Claims 7-8 are similar in scope to claim 1, and are therefore rejected under similar rationale.
Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over Jayaraman et al. (“Jayaraman”, US 2021/0256097), Priestas et al. (“Priestas”, US 2019/0005012), Liang et al. (“Liang” US 2018/0060302) and Sirangimoorthy et al. (“Sirangimoorthy”, US 2021/0248153) in view of Hurd et al. (“Hurd”, US 2021/0082062).
As per claim 3, the method of Jayaraman, Priestas, Liang and Sirangimoorthy teaches the method of claim 1, wherein the extraction model selected from the plurality of extraction models is dependent on the classification (Sirangimoorthy, p.26, 39, document analyzer extracts based on domain determined by key words), however does not teach wherein the classification model outputs, by the at least one processor of the web service provider, a confidence score for the classification and wherein the extraction model selected from the plurality of extraction models is dependent on the confidence score. Hurd teaches a method of summarizing unstructured documents wherein the classification model outputs, by the at least one processor of the web service provider, a confidence score for the classification and wherein the extraction model selected from the plurality of extraction models is dependent on the confidence score (Hurd, para.19, 35-37, 49-50, confidence level). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include Hurd’s teaching with the method of Jayaraman, Priestas, Liang and Sirangimoorthy in order to determine the accuracy of extraction.
Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over Jayaraman et al. (“Jayaraman”, US 2021/0256097), Priestas et al. (“Priestas”, US 2019/0005012), Liang et al. (“Liang” US 2018/0060302) and Sirangimoorthy et al. (“Sirangimoorthy”, US 2021/0248153) in view of Fujimoto et al. (“Fujimoto”, US 2023/0010202).
As per claim 4, the method of Jayaraman, Priestas, Liang and Sirangimoorthy teaches the method according to claim 1, wherein the set of extraction models comprises an extraction model corresponding to each of a plurality of classifications (Sirangimoorthy, p.26, 39, document analyzer extracts based on domain determined by key words), however does not teach wherein a generic extraction model is selected, by the at least one processor of the web service provider, when the classification of the segment does not correspond to one of said plurality of classifications. Fujimoto teaches a method of extracting data wherein a generic extraction model is selected dependent on the data classification (Fujimoto, para.126-127, 164-165, 187, 202, generic feature extractor/handwriting recognition model). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include Fujimoto’s teaching with the method of Jayaraman, Priestas, Liang and Sirangimoorthy in order to extract newly acquired data.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Inquiries
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SAJEDA MUHEBBULLAH whose telephone number is (571)272-4065. The examiner can normally be reached Mon-Tue/Thur-Fri 10am-8pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, William L Bashore can be reached on 571-272-4088. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/S.M./
Sajeda MuhebbullahExaminer, Art Unit 2174
/WILLIAM L BASHORE/ Supervisory Patent Examiner, Art Unit 2174