DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1-20 are pending.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 05/06/2024 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Contractor et al. (Contractor), US Patent Application Publication No. US 2019/0012405 A1, and further in view of Li et al. (Li), US Patent Application Publication No. US 2021/0382944 A1.
Claims 1-20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Contractor et al. (Contractor), US Patent Application Publication No. US 2019/0012405 A1.
As to independent claim 1, Contractor discloses a document management system (paragraph [0005]: provide a method, system and computer-readable storage medium for generating a knowledge graph) comprising:
at least one memory configured to a store a program (paragraph [0019]: the system includes a system memory; paragraph [0022]: the system memory may include at least one program product having a set of program modules that are configured to carry out the functions); and
at least one processor communicatively connected to the at least one memory and configured to execute the stored program to (paragraph [0019]: the system includes at least one processor or processing unit and a bus that couples system memory to processor):
receive a collection of documents (Figure 3 and paragraphs [0024], [0025]: receiving set of texts 100, 110, 120 (documents));
generate a first categorization of contents of each document of the collection
of documents using a first natural language processing (NLP) (paragraphs [0005], [0016], [0024] and [0025]: a document 1 is parsed by a natural language system/processing (NLP) with a rhetorical structure component, and the document is categorized into a plurality of rhetorical structure zones, wherein categorizing plurality of portions of the document as one of i) an introduction section and ii) a theory section, according to a Rhetorical Structure Theory (“RST”) scheme);
extract a plurality of terms from each document of the collection of documents
using the first categorization as input to a second natural language processing (paragraphs [0016], [0024] and [0025]: the system will identify a glossary of terms based on a pre-defined set of terms or by using automated operations, such as key-phrase extraction on a block of text to extract the pre-supplied terms, and where rhetorical structure zones will enable the natural language system to determine if a relationship exists between or among the pre-supplied terms);
generate a knowledge graph for each document of the collection of documents,
each knowledge graph having a plurality of nodes corresponding to the extracted plurality of terms from each document, the knowledge graphs for each document being linked to each other by common terms to form a collection of knowledge graphs (paragraphs [0016] and [0025]: the glossary of terms will in turn be used by a graph generating component of a system to create a knowledge graph that reviews the hierarchy of data, and the terms or subjects that underlie the knowledge graph, wherein the knowledge graph will have nodes and links associated with the terms and the relationship of those terms, based on the glossary of terms; paragraph [0036]: the node graph generation component can link nodes of each of the respective graphs by identifying or computing the lowest common ancestor or ancestors of each graph and linking them at that point);
receive a new document (Figure 6 and paragraph [0041]: the NLP component (second natural language processing) receives at least one text input (i.e., electronic text data), such as a text block (i.e., block of electronic text));
extract terms from the new document using the first categorization as input to
the second language model (Figure 6 and paragraph [0041]: the RST characterization component of the NLP component creates an RST characterization scheme that divides the text input into an introduction zone, where an item or term is introduced, and a theory zone where the item or term is elaborated on, the RST characterization component will characterize, line-by-line, each line of the text input into either an introduction zone or a theory zone); and
generate a new knowledge graph for the new document, the new knowledge
graph having a plurality of nodes corresponding to the extracted terms from the new
document, the new knowledge graph being linked to the knowledge graphs in the collection of knowledge graphs using common terms (Figure 6 and paragraph [0041]: an exemplary RST scheme with at least one text block is shown in Figure 3. The RST characterization scheme identifies a glossary or introduction set of terms from the introduction zones, as having an RST relationship, which can then be used by the node graph creation component to create a node knowledge graph, wherein the NLP component can receive a pre-supplied set of terms supplied by a user and located in memory or provided by the NLP using automated operations, such as key-phrase extraction, on a block of text; paragraph [0042]: the node graph generation components creates plurality of nodes from the glossary of terms or introduction terms provided by the RST characterization scheme, and the node graph generation component determines the lowest common ancestor of the two node graphs, the node generation component links the two graphs at the lowest common ancestor).
Contractor discloses the document is categorized into a plurality of rhetorical structure zones by using a natural language processing NLP (paragraph [0005]). In addition, Contractor discloses in paragraph [0004] that modern NLP algorithms are based on machine learning, especially statistical machine learning. Thus, one of ordinary skill in the art would interpret that the natural language processing is considered as a language model.
To support Examiner’s interpretation, Li discloses a technique that produces a graph data structure based on at least partially unstructured information dispersed over web documents, wherein the technique involves applying a machine-trained model to a set of documents to identify topics in the documents (Abstract). Li further discloses the topic-detecting system can implement the multi-class classifier using any type of neural network, such a convolutional neural network (CNN), a transformer-based neural network, etc., or any combination thereof (paragraphs [0080], [0082]).
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the method and apparatus for generating a knowledge graph of Contractor to include a neural network that incorporates transformer-based technology to detect topics in documents, as taught by Li for the purpose of producing a graph data structure.
As to dependent claim 2, Contractor and Li disclose wherein the at least one processor is further configured to execute the stored program to:
train the first language model using the collection of documents as input and a
predefined categorization of the contents of each document of the collection of documents as ground truth for training the first language model (Contractor, paragraphs [0016], [0024]-[0025]; Li, Abstract and paragraphs [0080], [0082]);
generate the first categorization of the contents of each document of the collection of documents using the first language model (Contractor, paragraphs [0016], [0024]-[0025]; Li, Abstract and paragraphs [0080], [0082]); and
train the second language model using the first categorization of the contents of each document of the collection of documents and the contents of each document of the collection of documents as input and a predefined plurality of terms for each document of the collection of documents as ground truth for training the second language model (Contractor, paragraph [0016]; Li, Abstract and paragraphs [0080], [0082]).
As to dependent claim 3, Contractor and Li disclose wherein the second language model is a same language model as the first language model (Contractor, Figure 2, paragraph [0024); Li, paragraphs [0080, [0082]).
As to dependent claim 4, Contractor and Li disclose wherein the at least one processor is further configured to execute the stored program to:
receive a user defined categorization of the contents of each document in the
collection of documents (Contractor, paragraph [0027]; paragraphs [0080, [0082]); and
train the second language model using the first categorization of the contents of each document of the collection of documents, the user-defined categorization of the contents of each document in the collection of documents, and the contents of each document of the collection of documents as input and the predefined plurality of terms for each document of the collection of documents as the ground truth for training the second language model (Contractor, paragraphs [0026], [0027]; paragraphs [0080, [0082]).
As to dependent claim 5, Contractor and Li disclose wherein the first categorization defines the plurality of nodes for the knowledge graph (Contractor, Abstract).
As to dependent claim 6, Contractor and Li disclose wherein, in a case where the new document includes additional content that cannot be categorized using the first categorization, the at least one processor is further configured to execute the stored program to:
generate a second categorization of contents of the new document using the first
language model (Contractor, paragraph [0032]; Li, paragraphs [0080], [0082]);
extract the terms from the new document using the second categorization as input to the second language model (Contractor, Figures 3, 6 and paragraph [0041]; Li, paragraphs [0080], [0082]); and
generate the new knowledge graph for the new document (Contractor, paragraph [0042]).
As to dependent claim 7, Contractor and Li disclose wherein the at least one processor is further configured to execute the stored program to:
receive a query including one or more nodes of the plurality of nodes of the collection of knowledge graphs (Li, paragraphs [0049], [0064], [0066]), [0077]);
receive a search value for the query (Li, paragraph [0066]);
identify a subset of knowledge graphs, of the collection of knowledge graphs, that
include the one or more nodes corresponding to the query and have term values
corresponding to the search value (Li, paragraphs [0064], [0066]), [0077]); and
output the subset of knowledge graphs in response to the query (Li, paragraph [0049]).
As to dependent claim 8, Contractor and Li disclose wherein the first language model is a transformer neural network model (Li, paragraphs [0080], [0082]).
As to dependent claim 9, Contractor and Li disclose wherein the second language model is a transformer neural network model (Li, paragraphs [0080], [0082]).
As to dependent claim 10, Contractor and Li disclose wherein at least one the first language model and the second language model further includes a conventional neural network model in addition to the transformer neural network model (Li, paragraphs [0080], [0082]).
Claims 11-19 are method claims that contain similar limitations of claims 1-10. Therefore, claims 11-19 are rejected under the same rationale.
Claim 20 is medium claim that contains similar limitations of claim 1. Therefore, claim 20 is rejected under the same rationale.
Conclusion
Any inquiry concerning this communication should be directed to CHAU T NGUYEN at telephone number (571)272-4092. The examiner can normally be reached on M-F from 8am to 5pm (PT).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) Form at https://www.uspto.gov/patents/uspto-automated-interview-request-air-form.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula, can be reached at telephone number 5712724128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from Patent Center and the Private Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from Patent Center or Private PAIR. Status information for unpublished applications is available through Patent Center and Private PAIR for authorized users only. Should you have questions about access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free).
/CHAU T NGUYEN/Primary Examiner, Art Unit 2145