DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Specification
The use of the terms “PYTHON”, “JAVA” AND “JSON”, which is a trade name or a mark used in commerce, has been noted in this application. The term should be accompanied by the generic terminology; furthermore the term should be capitalized wherever it appears or, where appropriate, include a proper symbol indicating use in commerce such as ™, SM , or ® following the term.
Although the use of trade names and marks used in commerce (i.e., trademarks, service marks, certification marks, and collective marks) are permissible in patent applications, the proprietary nature of the marks should be respected and every effort made to prevent their use in any manner which might adversely affect their validity as commercial marks.
Claim Objections
Claims 1-4 are objected to because of the following informalities: The use of the terms “PYTHON”, “JAVA” AND “JSON”, which is a trade name or a mark used in commerce, has been noted in this application. The term should be accompanied by the generic terminology; furthermore the term should be capitalized wherever it appears or, where appropriate, include a proper symbol indicating use in commerce such as ™, SM , or ® following the term . Appropriate correction is required.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Agarwal et al. (US Patent Application Publication No. 2024/0005640) in view of Sirvastava et al. (US Patent No. 11,087.081).
Regarding claim 1, Agarwal discloses a method of training a document parsing artificial intelligence (AI) system, the method comprising [see abstract, para. 0003; a device that receives an instruction to generate a document to be used as a training instance for a first machine learning model, the instruction including an element configuration, a document class configuration, a format configuration, an augmentation configuration, and data bias and fairness. The device can receive an element from an interface based at least in part on the element configuration, the element can simulate a real-world image, real-world text, or real-world machine-readable visual code. The device can generate metadata describe a layout for the element on the document based on the document class configuration]:
configuring, by a processing device, a PYTHON data structure (word processing format document) for generating a simulated document for training the document parsing AI system, wherein the simulated document comprises a list of characters and associated characteristics [see para. 0006-0015; to generate metadata describing a layout for the element on the document based at least in part on the document class configuration to generate the document by arranging the element on the document based at least in part on the metadata, wherein the document is generated in a format based on the format configuration and including receiving an element from an interface based at least in part on the element configuration, the element can be an image, a text, or machine-readable visual code simulating a real-world image, real-world text or real-world machine-readable visual code; which corresponds to generating a simulated document for training the document parsing AI system];
configuring, by the processing device, a JAVA data structure for generating a non-simulated document for training the document parsing AI system [see para. 0041-0044; Extracting information from imaged documents and processing the information in a reliable and accurate manner involves solving multiple sub-problems such as: using optical character recognition techniques to extract document contents, key-value extraction, layout parsing, named entity recognition, entity-linking, table parsing and extraction, document image quality assessment, document image classification; which corresponds to generate a non-simulated document for training the document parsing AI system];
receiving, by the processing device, a set of one or more parameters for training the document parsing AI system [see para. 0043, 0051; a synthetic document generation pipeline for training Al models. A unified training data-generation framework is described that is capable of creating a large volume of training data with labels and annotations in an automated manner. The framework enables a user to have control over the type of training data that is generated. For example, the training data can include synthetically generated documents, and the framework enables a user to control the content (e.g., text, images, handwritten text, background images, different fonts, different languages, etc.) of the documents, and the format/layouts of the documents; which corresponds to using AI to process receiving documents];
generating, by the processing device, via the PYTHON data structure and based on the set of one or more parameters, a simulated document comprising a list of one or more characters associated with one or more respective characteristics [see para. 0063-0064; the input interface generate elements that are appropriate for the given document class. Therefore, the control instructions include metadata that includes element parameters that are configured for the document class to guide the input interface, the input interface further provide labels describing the elements of a configuration layer of a synthetic document generation pipeline. The configuration layer includes three modules: a task-based configurator, an augmentation-based configurator, and a data bias and fairness configurator. Each module receive inputs establishing parameters for a task and convert the inputs into control instructions for the pipeline; which corresponds to generating a simulated document for training the document parsing AI system];
parsing, by the processing device, a received document, with the trained document parsing AI system to determine one or more characteristics associated with textual data written to the received document; and generating, by the processing device, an output of the parsed received document [see para. 0041-0044, 0081 and figures 5-6; key value extraction (KVE), named entity recognition (NER), optical character recognition (OCR), visual question and answering (VQA), layout parsing, entity learning (EL), table parsing and extraction, digital image correlation (DIC), document image quality assessment (DIQA), and machine-readable zone extraction (MRZE). The pipeline configured to generate training instances for each of these tasks and others, the pipeline allows for the configuration of input elements, layout format, and augmentation. Therefore, regardless of the machine learning task, the pipeline configured to generate a training instance suitable for training an ML model to perform the task; which corresponds to generating documents in a specific format, producing training data for a document parsing AI, training the AI on the generated document]; however, Agarwal fails to explicitly teach generating, by the processing device, via the JAVA data structure, a non-parsed JSON file comprising a description for a non-simulated document based on the set of one or more parameters.
Srivastava discloses generating, by the processing device, via the JAVA data structure, a non-parsed JSON file comprising a description for a non-simulated document based on the set of one or more parameters [see col. 4, lines 41-53 and figure 3; controller analyzes the configuration to a set of rules to determine which types of element templates used to generate the synthetic documents and generates a configuration that indicates the types of element templates and weights for the different types of element templates based on the configuration a JSON (JavaScript Object Notation) file]; reading, by the processing device, the non-parsed JSON file; generating, by the processing device, based on the reading of the non-parsed JSON file, a word-processing format file comprising [see col. 2, lines 62-67; provide a configuration-driven approach that addresses these problems with conventional methods. The synthetic document generation system takes a configuration (e.g., a JSON (JavaScript Object Notation) file) specifying which elements must be present in a form or document (key-value pairs, tables, checkboxes, text, etc.). The configuration may also specify whether or not the layout of the elements is structured (number of rows and columns), and which styles the elements. In addition, a weight can be specified for each element to attain a weighted probabilistic distribution of element]; a first set of one or more objects, each object of the first set being associated with a respective object type, wherein each object type in the first set corresponds to a specific and repeatable manner in which associated text of that object is placed in the non-simulated document [see col. 2, lines 30-40; generating configuration-controlled synthetic documents for training machine learning models such as neural networks. A document analysis service or system may analyze real-world documents such as forms, receipts, and dense text documents using machine learning models (e.g., neural networks) to generate digital and semantic information for the documents. Machine learning models are trained and tested using ground truth data. For a machine learning model used by a document analysis system, the ground truth data describes the various elements which make up a given document];
generating, by the processing device, a parsed JSON file for the simulated document comprising a second set of one or more objects, each object in the second set being associated with a respective object type, wherein each object type corresponds to a specific and repeatable manner in which associated text of that object is placed in the simulated document [see col. 4, lines 30-40; Configuration also specify whether or not the layout of the elements is to be structured (e.g., number of rows and columns), and which styles the elements should adhere to. In addition, a weight may be specified for each element to attain a weighted probabilistic distribution of elements in the synthetic documents. For text documents (or for text portions of form documents), the configuration used to specify which types of text elements are needed (short words, long words, hyphenated words, punctuation symbols, numeric symbols, etc.), configuration stored to a configuration data store on one or more storage devices. Configuration a JSON (JavaScript Object Notation) file. However, other methods used to specify a configuration]; training, by the processing device, the document parsing AI system based on the generated word-processing format file and on the parsed JSON file for the simulated document [see col. 7, lines 1-10 and figures 2A-2B; the configuration file may be passed to multiple instances of a synthetic document generator; each instance configured to perform element, a markup language document (e.g., an HTML document) is generated from the configuration file (e.g., a JSON file) using element templates from a repository as specified in the configuration file. The element templates in the markup language document populated with example content. Content, size, style, and location of the element templates in the markup language documents randomized to provide diversity in the synthetic documents, a synthetic document and an annotation document are generated from the markup language document].
It would have been obvious to one of an ordinary skill in the art, having the teachings of Agarwal and Srivastava before the affective filing date of the claimed invention to modify, synthetic document generation pipeline for training AI of Agarwal to include synthetic document generator, as taught by Srivastava.
One would have been motivated to make such a combination in order to implement synthetic document when adapting a J-SON configured synthetic document generator for training a document parsing AI.
Regarding claim 2, Agarwal discloses wherein the training comprises training the document parsing AI system with training data derived from the generated word-processing format file associated with non-simulated document and confirming an accuracy [see para. 0003-0006; a synthetic document generation pipeline for training artificial intelligence models. A method including a device that receives an instruction to generate a document to be used as a training instance for a first machine learning model, the instruction including an element configuration, a document class configuration, a format configuration, an augmentation configuration].
Srisvastava discloses the training based on the parsed JSON file for the simulated document [see col. 7, lines 1-10 and figures 2A-2B; the configuration file may be passed to multiple instances of a synthetic document generator; each instance configured to perform element, a markup language document (e.g., an HTML document) is generated from the configuration file (e.g., a JSON file) using element templates from a repository as specified in the configuration file. The element templates in the markup language document populated with example content. Content, size, style, and location of the element templates in the markup language documents randomized to provide diversity in the synthetic documents, a synthetic document and an annotation document are generated from the markup language document].
One would have been motivated to make such a combination in order to implement synthetic document when adapting a J-SON configured synthetic document generator for training a document parsing AI.
Regarding claim 3, Srivastava discloses wherein confirming the accuracy comprises determining whether the AI system can identify data associated with the generated word-processing format file and a measure of correlation with parsed information in the parsed JSON file for the simulated document [see col. 2, lines 20-37 and figure 10; A document analysis service or system may analyze real-world documents such as forms, receipts, and dense text documents using machine learning models (e.g., neural networks) to generate digital and semantic information for the documents. Machine learning models are trained and tested using ground truth data. For a machine learning model used by a document analysis system, the ground truth data describes the various elements which make up a given document. For example, form documents may contain key-value pairs, tables, text, headers, footers, and so on, text documents may contain headers, footers, columns or blocks of text, and so on. Conventionally, to train a machine learning model for a document analysis system, a large number of real-world documents are analyzed and annotated through human effort to provide a sufficiently large set of ground truth data].
Regarding claim 4, Srivastava discloses wherein determining whether to process the set with a JAVA data structure or a PYTHON data structure comprising a random determination [see col. 7, lines 40-50; Configuration file, for example, be a JSON file “componentWeights” lists components (or element template types) that are to be included in synthetic documents (“keyValuePair” and “checkbox”), and gives weights for the components. “keyValueStyles” indicates two styles of keyValuePair (i.e., two styles of element templates) that are to be used (“XformKeyAtTop” and “XformKeyAtBottom”), and gives weights for the styles. “Xform” refer to a particular type of form from which the element templates were extracted].
Regarding claim 5, Srivastava discloses wherein the first set of one or more objects correspond to the second set of one or more objects [see col. 9, lines 20-35 and figures 6A-7F; the rendered markup language documents (e.g. annotated HTML documents) are parsed to generate annotation documents for respective synthetic documents, the annotation documents be JSON (JavaScript Object Notation) files. However, other methods used to specify annotation documents. Each annotation document includes information describing a respective synthetic document. For example, an annotation document include information describing the element template type, location, size, style, and content of the elements in the respective synthetic document, and also include information indicating associations and relationships between elements in the synthetic document (for example, which words are associated with a text element, which words are in a line, which words are in a key, which words are in a value, etc.)].
Regarding claim 6, Srivastava discloses wherein the first set of one or more objects are randomly generated having one or more random strings [see col. 10, lines 21-39 and figure 7A; annotation document includes a list of words that appear in the respective synthetic document. For each word, document specifies a location and dimensions of a bounding box for the word in the synthetic document (<X,Y,W,H>), a bounding box identifier for the word, and the content and transcription type of the word. Word 1 and word 2 are of type text, and the content of these words are strings].
Regarding claim 7, Srivastava discloses wherein the second set of one or more objects are randomly generated having one or more random strings [see col. 2, lines 15-30 and figures 5a-5D; Machine learning models are trained and tested using ground truth data. For a machine learning model used by a document analysis system, the ground truth data describes the various elements which make up a given document. For example, form documents may contain key-value pairs, tables, text, headers, footers, and so on. As another example, text documents may contain headers, footers, columns or blocks of text].
Regarding claim 8, Srivastava discloses wherein the first set of one or more objects comprises a KeyValuePair object [see col. 4, lines 17-35; a configuration extraction process to generate a configuration for generating synthetic documents based on the real-world documents to train a machine learning model (e.g., a neural network). Configuration extraction a manual process, an automated process, or a combination of manual and automated steps. Configuration specify which elements should be present in the synthetic documents (key-value pairs, tables, text, etc.). Configuration also specify whether or not the layout of the elements is to be structured (e.g., number of rows and columns). In addition, a weight specified for each element to attain a weighted probabilistic distribution of elements in the synthetic documents].
Regarding claim 9, Srivastava discloses wherein an object type associated with the KeyValuePair object comprises one of the following: "right_offset," "left_under," "right_offset_list," or "left_under_list."; see col. 3, lines 40-55 and figures 3-7A; Configuration file, for example, be a JSON file. In this example “componentWeights” lists components (or element template types) that are to be included in synthetic documents (“keyValuePair” and “checkbox”), and gives weights for the components. “keyValueStyles” indicates two styles of keyValuePair (i.e., two styles of element templates) that are to be used (“XformKeyAtTop” and “XformKeyAtBottom”), and gives weights for the styles. “Xform” refer to a particular type of form from which the element templates were extracted. The configuration file indicate other information regarding the layout and style of the synthetic document, such as indicating that the form is to be (or not to be) an Xform-like form, and indicating the number of sections in the form].
Regarding claim 10, Srivastava discloses wherein the second set of one or more objects comprises a KeyValuePair object; [see figures 3-7A].
Regarding claim 11, Srivastava discloses wherein an object type associated with the KeyValuePair object comprises one of the following: "right_offset," "left_under," "right_offset_list," or "left_under_list."; see col. 10, lines 21-45 and figure 7A; Annotation document includes a list of key-value elements that appear in the respective synthetic document. A first key-value element (key value element 1) is a single key-single value element. A second key-value element (key value element 2) is a single key-multiple value element. For each key-value element, document specifies a key and one or more values. For the key, document specifies a location and dimensions of a bounding box for the key in the synthetic document (<X,Y,W,H>), a bounding box identifier for the key, and one or more word boxes (word bounding box identifiers) that are in the key. For each value, document specifies a location and dimensions of a bounding box for the value in the synthetic document (<X,Y,W,H>), a bounding box identifier for the value, and one or more word boxes (word bounding box identifiers) that are in the value. Document also specifies a list of child elements (as bounding box identifiers) of the value, as well as the type of each child element (for example, a word). Document also specifies the bounding box location and dimensions, bounding box identifier, and style (e.g., single key-single value, single key-multiple value, etc.) of a container for the key-value element].
Regarding claim 12, Agarwal discloses wherein an object type associated with the first set of one or more objects comprises a table format characteristic [see para. 0075; augmentation processes include selective node & edge pruning, node shuffling, altering node and edge characteristics by changing word sizes in the annotation, or re-weighting the edges. The graph augmentation extensible and can add additional modules for graph augmentation as needed for different use cases. For instance, documents can include one or more graphical elements like barcodes, QR codes, holograms, stamps, logos, tables, infographics, charts, diagrams, etc. An interface of the input interface can be scaled and configured to include multiple graphical elements while generating documents. The graphical element passed through the graph augmentation to create similar documents that are structurally different].
Regarding claim 13, Agarwal discloses wherein an object type associated with the second set of one or more objects comprises a table format characteristic [see para. 0084; The table interface used to generate table elements, such as rows, columns, and cells for inclusion in a document. The table interface further generate spatial and textual information. The spatial information can include a label describing a spatial layout of the table elements. The textual information can include a content included in one of the table elements].
Regarding claim 14, Agarwal discloses wherein the one or more parameters are hard-coded configuration parameters [see figure 6].
Regarding claim 15, Agarwal discloses wherein the one or more parameters comprise locational data [see para. 0121].
Regarding claim 16, Agarwal discloses wherein the one or more parameters comprise font data [see para. 0086].
Regarding claim 17, Agarwal discloses trained document parsing AI system [see para. 0051].
Regarding claim 18, Agarwal discloses wherein the training comprises training the document parsing AI system with training data derived from a plurality of documents [see para. 0051; The configuration layer provides the user the flexibility to customize the training instance. Certain artificial intelligence (AI) models are configured to receive and process certain types of data. For example, certain AI models can receive and process image data, while other models can receive and process textual data. The configuration layer can generate control instructions such that the data generator receives elements and labels that are appropriate for a target AI model].
Regarding claim 19, Agarwal discloses wherein the one or more parameters comprise alignment information [see para. 0063-0064 and figure 3; The configuration layer also transmit the control instructions to the input interface, such that the input interface instructed to generate the proper elements that are configured for the document class. The input interface receive the control instructions from the configuration layer and generate one or more elements pursuant to the control instructions. The one or more elements can be configured for the document class. In other words, the input interface generate elements that are appropriate for the given document class. Therefore, the control instructions include metadata that includes element parameters that are configured for the document class to guide the input interface, the input interface can further provide labels describing the elements of a configuration layer of a synthetic document generation pipeline. The configuration layer includes three modules: a task-based configurator, an augmentation-based configurator, and a data bias & fairness configurator. Each module can receive inputs establishing parameters for a task and convert the inputs into control instructions for the pipeline].
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure (See PTO-892).
A reference to specific paragraphs, columns, pages, or figures in a cited prior art reference is not limited to preferred embodiments or any specific examples. It is well settled that a prior art reference, in its entirety, must be considered for all that it expressly teaches and fairly suggests to one having ordinary skill in the art. Stated differently, a prior art disclosure reading on a limitation of Applicant's claim cannot be ignored on the ground that other embodiments disclosed were instead cited. Therefore, the Examiner's citation to a specific portion of a single prior art reference is not intended to exclusively dictate, but rather, to demonstrate an exemplary disclosure commensurate with the specific limitations being addressed. In re Heck, 699 F.2d 1331, 1332-33,216 USPQ 1038, 1039 (Fed. Cir. 1983) (quoting In re Lemelson, 397 F.2d 1006,1009, 158 USPQ 275, 277 (CCPA 1968)). In re: Upsher-Smith Labs. v. Pamlab, LLC, 412 F.3d 1319, 1323, 75 USPQ2d 1213, 1215 (Fed. Cir. 2005); In re Fritch, 972 F.2d 1260, 1264, 23 USPQ2d 1780, 1782 (Fed. Cir. 1992); Merck & Co. v. Biocraft Labs., Inc., 874 F.2d 804, 807, 10 USPQ2d 1843, 1846 (Fed. Cir. 1989); In re Fracalossi, 681 F.2d 792,794 n.1,215 USPQ 569, 570 n.1 (CCPA 1982); In re Lamberti, 545 F.2d 747, 750, 192 USPQ 278, 280 (CCPA 1976); In re Bozek, 416 F.2d 1385, 1390, 163 USPQ 545, 549 (CCPA 1969).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CAO H NGUYEN whose telephone number is (571)272-4053. The examiner can normally be reached on Mon-Fri 9am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kieu Vu can be reached on 571-272-4057. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/CAO H NGUYEN/ Primary Examiner, Art Unit 2171