DETAILED ACTION
1. This communication is in response to the Application No. 18/433,565 filed on February 6, 2024 in which Claims 1-20 are presented for examination.
Notice of Pre-AIA or AIA Status
2. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
3. The information disclosure statement submitted on 02/06/2024 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 112
4. The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
5. Claims 6 and 16 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. The term “maximum benefit” in Claims 6 and 16 is a relative term which renders the claim indefinite. The term “maximum benefit” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. Applicant’s specification does not recite the term, and the instant claim language merely recites identifying the alternative valid sequence which provides a “maximum benefit” without providing any threshold/requisite/degree to which the alternative sequences may provide any such “benefit”, let alone a “maximum benefit”. This renders the claims indefinite.
Claim Rejections - 35 USC § 103
6. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
7. Claims 1-8, 11-18, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Pito et al. (hereinafter Pito) (US PG-PUB 20210319177), in view of McNeill et al. (hereinafter McNeill) (US PG-PUB 20230153335).
Regarding Claim 1, Pito teaches a method, in a data processing system, for extracting a structure from an unstructured electronic document (Pito, Abstract, “The present disclosure is directed towards systems and methods for extracting structure and headers from a body of text.”, thus, a method, in a data processing system, for extracting a structure from an unstructured electronic document (See Par. [0080] which explicitly discloses an unstructured electronic document)), the method comprising:
executing a data preparation operation to convert the unstructured electronic document into a plain text document (Pito, Par. [0006], “The present inventions may fall into the domain of logical structure analysis and take as input text blocks in reading order that are annotated with layout and formatting information and produces a hierarchy of sections and/or list items in the form of a tree. The present inventions deal, therefore, not only with scanned documents but natively electronic documents that do not have structure annotations”, thus, a data preparation operation (logical structure analysis) is executed to convert the unstructured electronic document into a plain text document);
executing candidate structure identification rules (See introduction of McNeill reference below) on the plain text document to identify candidate structural elements (Pito, Par. [0087], “In one exemplary approach, we define a document D=(c1, . . . , cl) as an ordered list of l text chunks ci in reading order of the document. Each text chunk is endowed with a set of features fci such as font size, emphasis and justification. The features of each text chunk are assumed to be corrupted by noise, e.g. a chunk that is actually underlined may instead be reported as italic. Some exemplary raw text features that may be used in this work are listed in Table 1.”, thus, candidate structural elements (features) within the plain text document are identified – these structural elements are depicted by Table 1);
executing candidate structural element semantic consistency rules on the candidate structural elements to remove candidate structural elements that do not match expected semantic ordering of structural elements, to thereby generate a first group of structural elements (Pito, Par. [0079], “In one exemplary embodiment, text of a document is input in reading order and produces a tree-structured hierarchy of the document's text where each node of the tree represents a logical section of text. All potential sections are identified from top level sections with headings to bulleted list items. The present innovations are not restricted to a finite nesting depth. See FIG. 1(b) for example output. In addition, our algorithm may be configured to identify and remove “boilerplate” items like page headers, page footers, page numbers, etc.”, therefore, candidate structural element semantic consistency rules (algorithm) is executed on the candidate structural elements (features/potential sections) to remove candidate structural elements (boilerplate items like page headers, page footers, page numbers, etc.) that do not match expected semantic ordering (based on comparing the headers – see Par. [0021-0024]) to thereby generate a first group of structural elements);
generating a plurality of second groupings of candidate structural elements based on the first grouping of structural elements, wherein each second grouping comprises a different combination of structural elements than other second groupings in the plurality of second groupings (Pito, Par. [0100], “The proposed solution begins by identifying section headings H from the text of a document. A crossing partition of H is derived in polynomial time and is used to identify boilerplate sequences. This is because boilerplate sequences will normally cross/bisect other sequences, i.e. consider a sequence of page numbers. Once boilerplate sequences are removed, the remaining sequences are “shattered” to create a noncrossing partition of H whose coherence can be monotonically increased using search-based optimization techniques. A tree can be produced directly from the resulting noncrossing partition.” & Par. [0151], “With boilerplate removed, the final sequencing and partitioning of the remaining header or title sequences may begin. For the remainder of the proposed solution it may be assumed that boilerplate sequences have been removed from S.”, therefore, a plurality of second groupings of candidate structural elements may be generated based on the first grouping of structural elements (first grouping with boilerplate sequences removed) where each second grouping comprises a different combination of structural elements than other second groupings (noncrossing partition of data chunks));
merging the candidate structural elements of the second groupings to generate a hierarchical structure for the unstructured electronic document (Pito, Par. [0078], “One approach may be to identify and extract possible titles or headers from a document and then to group them into sequences that belong together based on their “appearance” in the document. These titles or headers can then be merged into a hierarchy or structure, ones that have certain properties can be eliminated as boilerplate, and we can search the remaining sequences for the most “clause like”. In a simple case, each title or heading in a “best” sequence is then used to delimit each clause while in a more complex case, sequences can be merged to produce a full hierarchy of titles, e.g. sections, clauses, sub-clauses, etc.”, therefore, the candidate structural elements of the second groupings may be merged to generate a hierarchical structure for the unstructured electronic document); and
generating a structured electronic document, corresponding to the unstructured electronic document, based on the hierarchical structure (Pito, Par. [0101], “The steps of an exemplary solution according to one embodiment are shown generally in FIG. 8. First, an input document is received and the headings H of the document are found 10, for example, by classifying each text chunk. Then, an initial partition S is generated 12 that approximates a solution to (2) using, for example, a polynomial time algorithm. Then, boilerplate sequences are identified and removed from S 14. Then, the remaining sequences are “shattered” 16 to form a set of noncrossing sequences P. Then, the remaining sequences are merged 18, for example, to incrementally improve the coherence G(P) by selectively merging its elements while ensuring P remains noncrossing. Then, the document's structure is constructed 20, for example, directly from the final partition.”, therefore, a structured electronic document is generated, corresponding to the unstructured electronic document, based on the hierarchical structure – see Figure 8).
Although Pito broadly overviews the use of domain specific rules in related prior art references (See Pito Par. [0007]) and Pito additionally discloses identifying candidate structural elements, as outlined in the rejection above, Pito does not explicitly disclose executing candidate structure identification rules on the plain text document to identify candidate structural elements;
However, McNeill teaches executing candidate structure identification rules on the plain text document to identify candidate structural elements (McNeill, Par. [0051], “The document parsing model(s) 114 designate the functional regions 134, assign category labels 140 to the functional regions 134, or both, based on a probabilistic analysis of the pixel data associated with the electronic document 124. In some implementations, the document parsing model(s) 114 may also apply one or more rules or heuristics to assign the category labels 140. For example, when the text 138 of a functional region 134 includes one or more special characters, the document parsing model(s) 114 may assign a particular category label 140 to the functional region 134 (or may perform operations to indicate an increased probability that the functional region 134 is associated with the particular category label 140). To illustrate, when the first character of each line of the text 138 of a functional region 134 includes a bullet point character, the document parsing model(s) 114 determine a high probability that the functional region 134 corresponds to a list.”, thus, candidate structure identification rules (one or more rules or heuristics) are executed on the plain text document to identify candidate structural elements (e.g., semantic regions which correlate to chapters, headings, paragraphs, lists, etc.));
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method for extracting a structure from an unstructured electronic document, as disclosed by Pito to include executing candidate structure identification rules on the plain text document to identify candidate structural elements, as disclosed by McNeill. One of ordinary skill in the art would have been motivated to make this modification to enable improved precision for identifying predictable patterns in the electronic document, with low computational overhead (McNeill, Par. [0051], “To illustrate, when the first character of each line of the text 138 of a functional region 134 includes a bullet point character, the document parsing model(s) 114 determine a high probability that the functional region 134 corresponds to a list. The high probability can be determined by assigning a default probability value (e.g., 1) or by weighting output of the probabilistic analysis of the document parsing model(s) 114 to increase the probability associated with the list category label. In some implementations, a rule can also, or in the alternative, be used to decrease the probability that a particular category label is assigned to a functional region 134. To illustrate, a rule may indicate that text 138 with a large font size (e.g., greater than an average font size for the electronic document), a bold font, and a centered alignment has a low probability of being assigned a footnote category label.”)
Regarding Claim 2, Pito in view of McNeill teaches the method of claim 1, wherein the candidate structural elements comprise one or more of candidate chapters or candidate sub-chapters (McNeill, Par. [0181], “Although FIGS. 17A-17B illustrate example of particular types (e.g., root, section heading, paragraph, page footer, and table) of semantic regions, the electronic document(s) 124 can include different or fewer types of semantic regions. Examples of other types of semantic regions include a chapter, a heading, a section, a subsection, a column, a page header, a figure, a caption, an image, etc.”, thus, the candidate structural elements comprise one or more of candidate chapters. It is worth noting that Pito does not explicitly disclose the elements described in their disclosure as “chapters” but instead refers to them as “headings” which may include different text sections such as “introduction”, “related work”, “problem definition”, etc. as supported by Pito Par. [0085] and [0094] – thus, McNeill is relied upon for the explicit disclosure of the term “chapter”).
The reasons of obviousness have been noted in the rejection of Claim 1 above and applicable herein.
Regarding Claim 3, Pito in view of McNeill teaches the method of claim 2, wherein generating a plurality of second groupings of candidate structural elements based on the first grouping of structural elements, comprises generating a second grouping for each possible combination of chapters and sub-chapters in the plain text document (Pito, Par. [0101], “The steps of an exemplary solution according to one embodiment are shown generally in FIG. 8. First, an input document is received and the headings H of the document are found 10, for example, by classifying each text chunk. Then, an initial partition S is generated 12 that approximates a solution to (2) using, for example, a polynomial time algorithm. Then, boilerplate sequences are identified and removed from S 14. Then, the remaining sequences are “shattered” 16 to form a set of noncrossing sequences P.”, thus, the plurality of second groupings is generated based on the first grouping of structural elements, comprising generating a second grouping for each possible combination of sequences (through “shattering” which is a partitioning mechanism which realizes every possible combination of elements of the set/sequence). Pito does not explicitly disclose the elements described in their disclosure as “chapters” but instead refers to them as “headings” which may include different text sections such as “introduction”, “related work”, “problem definition”, etc. as supported by Pito Par. [0085] and [0094]. Although Pito does not explicitly disclose the elements as “chapters/sub-chapters”, this is explicitly taught by McNeill as shown above in the rejection of Claims 1 & 2 above and also supported by McNeill Par. [0181])
The reasons of obviousness have been noted in the rejection of Claim 1 above and applicable herein.
Regarding Claim 4, Pito in view of McNeill teaches the method of claim 3, wherein the different second groupings are different based on which chapters or sub-chapters are maintained and which chapters or sub-chapters are eliminated from the combination of chapters and sub-chapters in the plain text document (Pito, Par. [0159], “Some valid sequences in S may, however, have been cut when the set was shattered. In this stage, we merge sequences in P that increase global coherence while maintaining the set as noncrossing. The candidates for merging are those sequences which lie between edges in P.”, therefore, the second groupings may be different based on which elements are maintained and/or eliminated from the combination of elements (noncrossing partition/maintaining the set as noncrossing) in the document. Pito does not explicitly disclose the elements described in their disclosure as “chapters” but instead refers to them as “headings” which may include different text sections such as “introduction”, “related work”, “problem definition”, etc. as supported by Pito Par. [0085] and [0094]. Although Pito does not explicitly disclose the elements as “chapters/sub-chapters”, this is explicitly taught by McNeill as shown above in the rejection of Claims 1 & 2 above and also supported by McNeill Par. [0181])
The reasons of obviousness have been noted in the rejection of Claim 1 above and applicable herein.
Regarding Claim 5, Pito in view of McNeill teaches the method of claim 2, wherein generating the first grouping of structural elements comprises executing candidate structural element semantic consistency rules specifying conditions of expected sequences of chapters and sub-chapters to identify one or more of chapters or sub-chapters that have invalid sequences which do not follow the expected sequences, and wherein generating a plurality of second groupings comprises identifying valid sequences of chapters and sub-chapters based on the structural element semantic rules (Pito, Par. [0078], “One approach may be to identify and extract possible titles or headers from a document and then to group them into sequences that belong together based on their “appearance” in the document. These titles or headers can then be merged into a hierarchy or structure, ones that have certain properties can be eliminated as boilerplate, and we can search the remaining sequences for the most “clause like”. In a simple case, each title or heading in a “best” sequence is then used to delimit each clause while in a more complex case, sequences can be merged to produce a full hierarchy of titles, e.g. sections, clauses, sub-clauses, etc”, therefore, generating the first grouping comprises executing consistency rules (algorithm) specifying conditions of expected sequences of elements to identify one or more elements that have invalid sequences that do not follow expected sequences, where generating a plurality of second groupings comprises identifying valid sequences of elements based on the rules – this is further supported by Pito Par. [0025-0029] which mentions how the elements may be grouped based on similarities and matching non-adjacent/adjacent groups. Pito does not explicitly disclose the elements described in their disclosure as “chapters” but instead refers to them as “headings” which may include different text sections such as “introduction”, “related work”, “problem definition”, etc. as supported by Pito Par. [0085] and [0094]. Although Pito does not explicitly disclose the elements as “chapters/sub-chapters”, this is explicitly taught by McNeill as shown above in the rejection of Claims 1 & 2 above and also supported by McNeill Par. [0181])
The reasons of obviousness have been noted in the rejection of Claim 1 above and applicable herein.
Regarding Claim 6, Pito in view of McNeill teaches the method of claim 5, wherein generating the first group of candidate structural elements comprises, for valid sequences that are alternative to each other, evaluating the alternative valid sequences to identify which alternative provides a maximum benefit for maintaining the alternative based on a function of a number of chapters, sub-chapters, and leaf chapters in each alternative, and wherein alternatives that do not provide a maximum benefit are not maintained in the first group of candidate structural elements (Pito, Par. [0111], “The edges of G′ are thus enhanced with a score indicating the similarity of each pair of headings. A matching of the nodes in G′ that maximizes these weights would represent a partition P of H that maximizes the pairwise coherency between the elements of each sequence in P. In addition, since we expect each header to participate in a sequence, we add the additional constraint that all nodes should be matched. This is an instance of a maximum weight perfect matching problem which can be cast as a linear sum assignment problem (LSAP) for which there exist polynomial time solutions.”, therefore, generating the group of elements comprises evaluating alternative valid sequences to identify which alternative provides a maximum benefit (See 35 U.S.C. 112(b) rejection above regarding this term – as such, Examiner interprets this “maximum benefit” to be analogous to maximizing pairwise coherency and application of linear sum assignment problem (LSAP)) for maintaining the alternative based on a function of a number of elements and where the alternatives that do not provide a maximum benefit are not maintained in the first group – See also Par. [0159] which mentions that sequences are merged that increase global coherence while maintaining the set as noncrossing. Pito does not explicitly disclose the elements described in their disclosure as “chapters” but instead refers to them as “headings” which may include different text sections such as “introduction”, “related work”, “problem definition”, etc. as supported by Pito Par. [0085] and [0094]. Although Pito does not explicitly disclose the elements as “chapters/sub-chapters”, this is explicitly taught by McNeill as shown above in the rejection of Claims 1 & 2 above and also supported by McNeill Par. [0181])
The reasons of obviousness have been noted in the rejection of Claim 1 above and applicable herein.
Regarding Claim 7, Pito in view of McNeill teaches the method of claim 1, wherein executing structural element identification rules comprises performing a line-by-line execution of the candidate structure identification rules on each line of the plain text document to determine whether the line comprises a pattern of content matching a pattern specified in one or more of the candidate structure identification rules (McNeill, Par. [0137], “FIG. 12A includes a diagram 1200 illustrating an example of the page 1122 of the electronic document 124 with various graphical regions identified. In a particular aspect, the page 1122 corresponds to a graphical region (GR) 1220. In the diagram 1200, each graphical region within the GR 1220 is denoted by a dashed line indicating a boundary of the graphical region. For example, in the diagram 1200, the graphical region 1220 includes a plurality of graphical sub-regions, such as a graphical region 1222 (e.g., a line of text), a graphical region 1224 (e.g., a text box), a graphical region 1226 (e.g., a text box), and a graphical region 1228 (e.g., a line of text).”, thus, the structural element identification rule may comprise performing line-by-line execution of the rules (See McNeill Par. [0051] for the explicit recitation of the rules) on each line to determine whether the line comprises a pattern of content specified by the rules. Alternatively, Pito Par. [0153-0154] similarly discloses such line-by-line pattern matching)
The reasons of obviousness have been noted in the rejection of Claim 1 above and applicable herein.
Regarding Claim 8, Pito in view of McNeill teaches the method of claim 1, wherein executing the data preparation operation comprises:
splitting the unstructured electronic document into blocks of content (Pito, Par. [0087], “In one exemplary approach, we define a document D=(c1, . . . , cl) as an ordered list of l text chunks ci in reading order of the document. Each text chunk is endowed with a set of features fci such as font size, emphasis and justification. The features of each text chunk are assumed to be corrupted by noise, e.g. a chunk that is actually underlined may instead be reported as italic” & Par. [0101], “The steps of an exemplary solution according to one embodiment are shown generally in FIG. 8. First, an input document is received and the headings H of the document are found 10, for example, by classifying each text chunk. Then, an initial partition S is generated 12 that approximates a solution to (2) using, for example, a polynomial time algorithm.”, thus, the unstructured electronic document may be split into blocks of content (text chunks));
removing noise from the blocks of content to generate denoised blocks of content (Pito, Par. [0100], “The proposed solution begins by identifying section headings H from the text of a document. A crossing partition of H is derived in polynomial time and is used to identify boilerplate sequences. This is because boilerplate sequences will normally cross/bisect other sequences, i.e. consider a sequence of page numbers. Once boilerplate sequences are removed, the remaining sequences are “shattered” to create a noncrossing partition of H whose coherence can be monotonically increased using search-based optimization techniques.”, therefore, noise (boilerplate sequences which may include page headers, page footers, page numbers as supported by instant dependent claim 9) is removed from the blocks of content (text chunks) to generate denoised blocks of content);
converting the denoised blocks of content to plain text pages of content (Pito, Par. [0102], “In another exemplary embodiment, the steps of an exemplary solution may be simplified to finding the headings 10, identifying and removing boilerplate 14 and constructing final partitions and sequences of headings 20.”, thus, the denoised blocks of content (removed boilerplate) are converted to plain text pages of content through the construction of final partitions and sequences of headings); and
combining the plain text pages of content to generate the plain text document (Pito, Par. [0101], “Then, the remaining sequences are merged 18, for example, to incrementally improve the coherence G(P) by selectively merging its elements while ensuring P remains noncrossing. Then, the document's structure is constructed 20, for example, directly from the final partition.”, therefore, the plain text pages of content may be combined/merged to generate the structured document).
Regarding Claim 11, Pito in view of McNeill teaches a computer program product comprising a computer readable storage medium having a computer readable program stored therein, wherein the computer readable program, when executed on a computing device (Pito, Par. [0181], “In software implementations, computer software (e.g., programs or other instructions) and/or data is stored on a machine readable medium as part of a computer program product, and is loaded into a computer system or other device or machine via a removable storage drive, hard drive, or communications interface. Computer programs (also called computer control logic or computer readable program code) are stored in a main and/or secondary memory, and executed by one or more processors (controllers, or the like) to cause the one or more processors to perform the functions of the disclosure as described herein.”, thus, a computer program product comprising a computer readable storage medium having a program stored thereon to be executed by a computing device is disclosed), causes the computing device to: […]
The rest of the claim language in Claim 11 recites substantially the same limitations as Claim 1, in the form of a computer program product, therefore it is rejected under the same rationale.
The reasons of obviousness have been noted in the rejection of Claim 1 above and applicable herein.
Claim 12 recites substantially the same limitations as Claim 2 in the form of a computer program product, therefore it is rejected under the same rationale.
Claim 13 recites substantially the same limitations as Claim 3 in the form of a computer program product, therefore it is rejected under the same rationale.
Claim 14 recites substantially the same limitations as Claim 4 in the form of a computer program product, therefore it is rejected under the same rationale.
Claim 15 recites substantially the same limitations as Claim 5 in the form of a computer program product, therefore it is rejected under the same rationale.
Claim 16 recites substantially the same limitations as Claim 6 in the form of a computer program product, therefore it is rejected under the same rationale.
Claim 17 recites substantially the same limitations as Claim 7 in the form of a computer program product, therefore it is rejected under the same rationale.
Claim 18 recites substantially the same limitations as Claim 8 in the form of a computer program product, therefore it is rejected under the same rationale.
Regarding Claim 20, Pito in view of McNeill teaches an apparatus comprising: at least one processor; and at least one memory coupled to the at least one processor, wherein the at least one memory comprises instructions which, when executed by the at least one processor (Pito, Par. [0181], “In software implementations, computer software (e.g., programs or other instructions) and/or data is stored on a machine readable medium as part of a computer program product, and is loaded into a computer system or other device or machine via a removable storage drive, hard drive, or communications interface. Computer programs (also called computer control logic or computer readable program code) are stored in a main and/or secondary memory, and executed by one or more processors (controllers, or the like) to cause the one or more processors to perform the functions of the disclosure as described herein. “, thus, an apparatus comprising at least one processor and at least one memory coupled to the at least one processor comprising instructions to be executed by the processor is disclosed), cause the at least one processor to: […]
The rest of the claim language in Claim 20 recites substantially the same limitations as Claim 1, in the form of an apparatus, therefore it is rejected under the same rationale.
The reasons of obviousness have been noted in the rejection of Claim 1 above and applicable herein.
8. Claim 9-10 and 19 is rejected under 35 U.S.C. 103 as being unpatentable over Pito et al. (hereinafter Pito) (US PG-PUB 20210319177), in view of McNeill et al. (hereinafter McNeill) (US PG-PUB 20230153335), further in view of Tonkin et al. (hereinafter Tonkin) (US PG-PUB 20190311003).
Regarding Claim 9, Pito in view of McNeill teaches the method of claim 8, wherein:
the noise comprises at least one of page headers, page footers, author information, page numbers, or data of the block of content that is not directed specifically to topics/subjects of the content (Pito, Par. [0079], “In addition, our algorithm may be configured to identify and remove “boilerplate” items like page headers, page footers, page numbers, etc.”, thus, the noise (boilerplate items to be removed) comprises at least one of page headers, page footers, or page numbers),
the blocks of content are pages of the unstructured electronic document (Pito, Par. [0072], “In another example, on another page of an exemplary document, shown in FIG. 11, there is another title “STANDARD LEASE PROVISIONS STANDARD LEASEPROVISIONS” which is likely at the “same level” as “BASIC LEASE PROVISIONS” from two pages prior (i.e. FIG. 10). In addition, we see other items like “1. TERM” which likely begins a clause and sub-items that begin with text like “(a)” or “(b)”. This page also shows how simplistic approaches to identifying clauses will fail. For example, if one decided to look for “numbered items” as clause headings it would be easy to mis-identify the page number “3” as the third clause in this list of clauses”, thus, the blocks of content may correlate to pages of the unstructured electronic document), and
Although Pito discloses removing noise in the form of “boilerplate” items like page headers, page footers, page numbers, etc. as supported by Pito Par. [0079], Pito in view of McNeill does not explicitly disclose removing noise from the pages of the unstructured electronic document comprises executing a trained machine learning computer model, specifically trained to identify portions of the pages that are indicative of noise data, on the data of each page of the unstructured electronic document.
However, Tonkin teaches removing noise from the pages of the unstructured electronic document comprises executing a trained machine learning computer model, specifically trained to identify portions of the pages that are indicative of noise data, on the data of each page of the unstructured electronic document (Tonkin, Par. [0227], “Machine learning techniques can be applied to each of these processes. The application of Deep Learning techniques allows the use of multiple contexts to be evaluated simultaneously and the detection and minimisation of noise words. The difference between Deep learning as normally applied to NLP is that a specific deep learning context is available in terms of the various context specific templates being used (DTT, Concept Maps (CM) and Thought Bubbles (TB)).”, therefore, noise may be removed from the pages of the unstructured document by executing a trained machine learning model which is trained to identify and minimize noise words on pages of an electronic document – this is further supported by the application of Stanford’s DeepDive in Par. [0063-0064]).
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of claim 8, as disclosed by Pito in view of McNeill to include removing noise from the pages of the unstructured electronic document comprises executing a trained machine learning computer model, specifically trained to identify portions of the pages that are indicative of noise data, on the data of each page of the unstructured electronic document, as disclosed by Tonkin. One of ordinary skill in the art would have been motivated to make this modification to enable the application of machine learning for noise removal, which may allow for efficient and simultaneous detection and minimization of noise words (Tonkin, Par. [0227], “Machine learning techniques can be applied to each of these processes. The application of Deep Learning techniques allows the use of multiple contexts to be evaluated simultaneously and the detection and minimisation of noise words. The difference between Deep learning as normally applied to NLP is that a specific deep learning context is available in terms of the various context specific templates being used (DTT, Concept Maps (CM) and Thought Bubbles (TB)).”).
Regarding Claim 10, Pito in view of McNeill teaches the method of claim 1.
Pito in view of McNeill does not explicitly disclose:
generating a knowledge base based on the structured electronic document corresponding to the unstructured electronic document; and
training a machine learning computer model based on the knowledge base.
However, Tonkin teaches:
generating a knowledge base based on the structured electronic document corresponding to the unstructured electronic document (Tonkin, Abstract, “A system for categorising and referencing a document using an electronic processing device, wherein: the electronic o processing device reviews the content of the document to identify structures within the document; wherein the identified structures are referenced against a library of structures stored in a database; wherein the document is categorised according to the conformance of the identified structures with those of the stored library of structures; and wherein the categorised structure is added to the stored library.”, therefore, a knowledge base (see Par. [0001] for explicit recitation) based on the structured electronic document corresponding to the unstructured electronic document is disclosed); and
training a machine learning computer model based on the knowledge base (Tonkin, Par. [0465-0466], “Extract RDF components Within a Penn tree, we identify and derive subjects, predicates and objects. Using AI techniques and prepared training data, we perform a classification to extract the concepts/entities for subjects and objects.”, therefore, a machine learning model is trained based on the knowledge base).
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of claim 1, as disclosed by Pito in view of McNeill to include generating a knowledge base based on the structured electronic document corresponding to the unstructured electronic document and training a machine learning computer model based on the knowledge base, as disclosed by Tonkin. One of ordinary skill in the art would have been motivated to make this modification to enable the use of a semantically accurate knowledge base and trained machine learning model, which may provide improved cognitive processing with contextual understanding and adaptation through iterative learning (Tonkin, Par. [0098-0099], “Most techniques for extracting knowledge are based on keyword search and match techniques. As such they are really information extraction techniques, containing little or no knowledge and hence unable to accurately infer any knowledge based facts. The technique of the present invention is based on building a knowledge model based utilising ontologies. It depends heavily on Natural Language Processing (NLP) to determine the semantics and hence the meaning. Without the ability to convert the document to semantics, cognitive processing is not possible.”).
Claim 19 recites substantially the same limitations as Claim 9 in the form of a computer program product, therefore it is rejected under the same rationale.
Conclusion
9. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Devika S Maharaj whose telephone number is (571)272-0829. The examiner can normally be reached Monday - Thursday 8:30am - 5:30pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571)270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DEVIKA S MAHARAJ/Examiner, Art Unit 2123