DETAILED ACTION
Notice of Pre-AIA or AIA Status
Claims 1-20 are pending in this application. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Specification
The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1, 10, and 16 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Oh et al. (US PGPub US 2018/0253426 A1, hereby referred to as “Oh”).
Consider Claims 1, 10 and 16.
Oh teaches:
1. An artificial intelligence device comprising: / 10. A computerized method comprising: / 16. A non-transitory computer-readable medium configured to store instructions (Oh: abstract, Described herein are systems and methods that efficiently search for documents related to chemical structures of interest to a user. In certain embodiments, text data and chemical structure data provided in a user query are simultaneously searched with a text-based search method to efficiently produce search results. Subsequent structure-based searching on the results of the text-based search produces precise results for a particular user query. This approach increases the speed of the structure-based search by reducing the amount of data the structure-based search searches over. Additionally described herein are systems and methods for indexing document data in order to facilitate this efficient searching.)
1. a memory configured to store a structural formula image; / 16. that when executed by one or more processors, cause the one or more processors to perform operations comprising: (Oh: [0119] The computing device 1000 includes a processor 1002, a memory 1004, a storage device 1006, a high-speed interface 1008 connecting to the memory 1004 and multiple high-speed expansion ports 1010, and a low-speed interface 1012 connecting to a low-speed expansion port 1014 and the storage device 1006. Each of the processor 1002, the memory 1004, the storage device 1006, the high-speed interface 1008, the high-speed expansion ports 1010, and the low-speed interface 1012, are interconnected using various busses, and may be mounted on a common motherboard or in other manners as appropriate. The processor 1002 can process instructions for execution within the computing device 1000, including instructions stored in the memory 1004 or on the storage device 1006 to display graphical information for a GUI on an external input/output device, such as a display 1016 coupled to the high-speed interface 1008. In other implementations, multiple processors and/or multiple buses may be used, as appropriate, along with multiple memories and types of memory. Also, multiple computing devices may be connected, with each device providing portions of the necessary operations (e.g., as a server bank, a group of blade servers, or a multi-processor system). [0120] The memory 1004 stores information within the computing device 1000. In some implementations, the memory 1004 is a volatile memory unit or units. In some implementations, the memory 1004 is a non-volatile memory unit or units. The memory 1004 may also be another form of computer-readable medium, such as a magnetic or optical disk. [0121]-[0125] The processor 1052 can execute instructions within the mobile computing device 1050, including instructions stored in the memory 1064. The processor 1052 may be implemented as a chipset of chips that include separate and multiple analog and digital processors. The processor 1052 may provide, for example, for coordination of the other components of the mobile computing device 1050, such as control of user interfaces, applications run by the mobile computing device 1050, and wireless communication by the mobile computing device 1050.)
1. and a processor configured to obtain information on a plurality of atomic regions from the structural formula image, / 10. obtaining information on a plurality of atomic regions from a structural formula image; / 16. obtaining information on a plurality of atomic regions from a structural formula image; (Oh: [0125] The processor 1052 can execute instructions within the mobile computing device 1050, including instructions stored in the memory 1064. The processor 1052 may be implemented as a chipset of chips that include separate and multiple analog and digital processors. The processor 1052 may provide, for example, for coordination of the other components of the mobile computing device 1050, such as control of user interfaces, applications run by the mobile computing device 1050, and wireless communication by the mobile computing device 1050. [0035] In another aspect, the present invention is directed to a method for searching a set of indexed documents comprising chemical information using sequential searches, the method comprising the steps of: (a) receiving, by a processor of a computing device, a user query comprising user-input chemical structure data and text data; (b) querying, using a text-based search method, by the processor, a database comprising document data corresponding to the set of indexed documents, wherein querying comprises correlating at least a portion of the user-input chemical structure data with the document data (e.g., by augmenting or converting the chemical structure data prior to correlating with the document data) and at least a portion of the text data of the user query with the document data to generate filtered document data; (c) searching, using a structure-based search method, subsequent to the querying step, by the processor, the filtered document data, wherein searching comprises correlating at least a portion of user-input chemical structure data with relevant filtered chemical structure data in the filtered document data to generate one or more search results; and (d) outputting, by the processor, (e.g., rendering for display, or outputting to another processor for rendering for display) the one or more search results. In certain embodiments, the method comprises the step of: converting, by the processor, the chemical structure data to one or more strings.)
1. obtain information on bonding relationships between a plurality of atoms based on the information on the plurality of atomic regions, / 10. obtaining information on bonding relationships between a plurality of atoms based on the information on the plurality of atomic regions; / 16. obtaining information on bonding relationships between a plurality of atoms based on the information on the plurality of atomic regions; (Oh: [0074] Constituent element: As used herein, the phrase “constituent element” refers to a portion of a chemical structure. A constituent element may be a bond, an atom, a fragment, a functional group, a heteroatom, a moiety or any combination thereof that forms in whole or in part a chemical structure. A constituent element may be used to identify, describe, and/or classify a chemical structure. A constituent element may be used as a search term when querying for documents related to a chemical structure that comprises the constituent element. [0087] Documents are indexed in a format such that they can be fully searched with text-based searching methods. Document data corresponding to the documents to be indexed may be loaded (e.g., uploaded) to a service such as ChemSearch. FIG. 1 shows an exemplary hierarchy of data structures that correspond to a document. Document data 100 comprise chemical structure data 110, text data 130, and metadata 140. Chemical structure data 110 corresponds to chemical structure information, such as a chemical structure representation. Chemical structure data may be stored in any number of standard formats (e.g., a simplified molecular input line entry specification (SMILES) or SMILES arbitrary target specification (SMARTS) based string or as formatted binary data). Chemical structure data 110 comprises bit-screening data 150 and connection data 160. Bit-screening data 150 correspond to one or more constituent elements of a chemical structure. Connection data 160 correspond to one or more connections (e.g., interactions, bonds) between a plurality of the one or more constituent elements. In certain embodiments, chemical structure data (e.g., bit-screening data and connection data) are stored as strings or converted to strings such that all document data used for searching is searchable with a text search engine. Text data 130 corresponds to descriptive information about the chemical and/or its structure. For example, text data may describe properties of a chemical (e.g., its structure) and/or it may describe processes, reactions, or formulations/mixtures involving the chemical. In certain embodiments, document data may include metadata that can be used to identify the document and its contents. For example, a document's metadata may include a unique ID and bucket ID. The metadata may be persisted to allow the document to be referenced in a database.)
1.generate an adjacency matrix based on the information on the plurality of atomic regions and the information on the bonding relationships between the plurality of atoms, / 10. generating an adjacency matrix based on the information on the plurality of atomic regions and the information on the bonding relationships between the plurality of atoms; / 16. generating an adjacency matrix based on the information on the plurality of atomic regions and the information on the bonding relationships between the plurality of atoms; (Oh: [0004] In order to reproduce chemical structures in a document, a range of standard formats are used to efficiently store the chemical structure data. One type of format uses connection tables, adjacency matrices, or similar data structures to relate atoms and bonds as edges and nodes. Another type of format uses linear string notations based on depth first or breadth first traversal. The use of standardized data formats for storing chemical structure data enables algorithmic searching of the data. Furthermore, chemical structure data in standard formats can be indexed with a document in a database. [0088] Document data 100 has been augmented during indexing to comprise string tags 120 (as depicted by the dashed line connecting the two in FIG. 1). String tags are a sequence of characters that provide an alphanumeric text-based string for identifying, classifying, and/or describing chemical structures corresponding to chemical structure data in document data. In certain embodiments, string tags are generated using bit-screening data by performing an atom-by-atom or similar structure-based search on the bit-screening data therein to identify constituent elements corresponding to the bit-screening data and populating the string tags with strings in a predefined list or array that identify, classify, and/or describe the constituent elements. In certain embodiments, string tags are populated using an array that comprises the strings and corresponding reference bit screening data that is compared to the bit screening data in document data. The predefined list may be manually created by storing strings for common constituent elements in chemical structures and associations to reference bit-screening data that correspond to those common constituent elements. Thus, the reference bit-screening data associated with the pre-defined strings can be matched, using the structure-based search, to bit-screening data in document data in order to generate string tags that are populated with appropriate descriptive strings from the pre-defined list for constituent elements corresponding to the bit-screening data in the document data. String tags may also be generated using appropriate ad hoc structure-based methods that populate the string tags with appropriate descriptive strings. String tags may be associated with chemical structure data or directly with directly with the document data that comprises the chemical structure data. Referring again to FIG. 1, string tags 120 are associated with document data 100, but not directly associated with chemical structure data 110.)
1.and generate a string format corresponding to the structural formula image based on the adjacency matrix. / 10. and generating a string format corresponding to the structural formula image based on the adjacency matrix./ 16. and generating a string format corresponding to the structural formula image based on the adjacency matrix. (Oh: [0090] FIG. 2 is a block diagram of an exemplary method for indexing documents comprising chemical structure information. Indexing method 200 is used to augment document data by generating one or more string tags from chemical structure data in the document data. In step 210, document data comprising chemical structure data is received by a processor of a computing device. In step 220, bit-screening data and connection data in the chemical structure data is identified or extracted. In step 230, the bit-screening data identified or extracted in step 220 is used to generate a string tag. In step 240, the string tag generated in step 230 is associated with the document data directly. In step 250, the string tag is outputted. The string tag outputted in step 250 is stored with the document data for later searching. In some embodiments, document data is augmented to comprise a string tag. In some embodiments, a string tag is stored separate from document data. When a string tag is stored separate from document data, the document data may be augmented to comprise the association of the string tag to the document data such that the string tag is searchable when the document data is being searched. [0091]-[0095])
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent may not be obtained though the invention is not identically disclosed or described as set forth in section 102 of this title, if the differences between the subject matter sought to be patented and the prior art are such that the subject matter as a whole would have been obvious at the time the invention was made to a person having ordinary skill in the art to which said subject matter pertains. Patentability shall not be negatived by the manner in which the invention was made.
Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Oh et al. (US PGPub US 2018/0253426 A1, hereby referred to as “Oh”), in view of Chen et al. (US PGPub US 2023/0052865), hereby referred to as “Chen”.
Consider Claims 1, 10 and 16.
Oh teaches:
1. An artificial intelligence device comprising: / 10. A computerized method comprising: / 16. A non-transitory computer-readable medium configured to store instructions (Oh: abstract, Described herein are systems and methods that efficiently search for documents related to chemical structures of interest to a user. In certain embodiments, text data and chemical structure data provided in a user query are simultaneously searched with a text-based search method to efficiently produce search results. Subsequent structure-based searching on the results of the text-based search produces precise results for a particular user query. This approach increases the speed of the structure-based search by reducing the amount of data the structure-based search searches over. Additionally described herein are systems and methods for indexing document data in order to facilitate this efficient searching.)
1. a memory configured to store a structural formula image; / 16. that when executed by one or more processors, cause the one or more processors to perform operations comprising: (Oh: [0119] The computing device 1000 includes a processor 1002, a memory 1004, a storage device 1006, a high-speed interface 1008 connecting to the memory 1004 and multiple high-speed expansion ports 1010, and a low-speed interface 1012 connecting to a low-speed expansion port 1014 and the storage device 1006. Each of the processor 1002, the memory 1004, the storage device 1006, the high-speed interface 1008, the high-speed expansion ports 1010, and the low-speed interface 1012, are interconnected using various busses, and may be mounted on a common motherboard or in other manners as appropriate. The processor 1002 can process instructions for execution within the computing device 1000, including instructions stored in the memory 1004 or on the storage device 1006 to display graphical information for a GUI on an external input/output device, such as a display 1016 coupled to the high-speed interface 1008. In other implementations, multiple processors and/or multiple buses may be used, as appropriate, along with multiple memories and types of memory. Also, multiple computing devices may be connected, with each device providing portions of the necessary operations (e.g., as a server bank, a group of blade servers, or a multi-processor system). [0120] The memory 1004 stores information within the computing device 1000. In some implementations, the memory 1004 is a volatile memory unit or units. In some implementations, the memory 1004 is a non-volatile memory unit or units. The memory 1004 may also be another form of computer-readable medium, such as a magnetic or optical disk. [0121]-[0125] The processor 1052 can execute instructions within the mobile computing device 1050, including instructions stored in the memory 1064. The processor 1052 may be implemented as a chipset of chips that include separate and multiple analog and digital processors. The processor 1052 may provide, for example, for coordination of the other components of the mobile computing device 1050, such as control of user interfaces, applications run by the mobile computing device 1050, and wireless communication by the mobile computing device 1050.)
1. and a processor configured to obtain information on a plurality of atomic regions from the structural formula image, / 10. obtaining information on a plurality of atomic regions from a structural formula image; / 16. obtaining information on a plurality of atomic regions from a structural formula image; (Oh: [0125] The processor 1052 can execute instructions within the mobile computing device 1050, including instructions stored in the memory 1064. The processor 1052 may be implemented as a chipset of chips that include separate and multiple analog and digital processors. The processor 1052 may provide, for example, for coordination of the other components of the mobile computing device 1050, such as control of user interfaces, applications run by the mobile computing device 1050, and wireless communication by the mobile computing device 1050. [0035] In another aspect, the present invention is directed to a method for searching a set of indexed documents comprising chemical information using sequential searches, the method comprising the steps of: (a) receiving, by a processor of a computing device, a user query comprising user-input chemical structure data and text data; (b) querying, using a text-based search method, by the processor, a database comprising document data corresponding to the set of indexed documents, wherein querying comprises correlating at least a portion of the user-input chemical structure data with the document data (e.g., by augmenting or converting the chemical structure data prior to correlating with the document data) and at least a portion of the text data of the user query with the document data to generate filtered document data; (c) searching, using a structure-based search method, subsequent to the querying step, by the processor, the filtered document data, wherein searching comprises correlating at least a portion of user-input chemical structure data with relevant filtered chemical structure data in the filtered document data to generate one or more search results; and (d) outputting, by the processor, (e.g., rendering for display, or outputting to another processor for rendering for display) the one or more search results. In certain embodiments, the method comprises the step of: converting, by the processor, the chemical structure data to one or more strings.)
1. obtain information on bonding relationships between a plurality of atoms based on the information on the plurality of atomic regions, / 10. obtaining information on bonding relationships between a plurality of atoms based on the information on the plurality of atomic regions; / 16. obtaining information on bonding relationships between a plurality of atoms based on the information on the plurality of atomic regions; (Oh: [0074] Constituent element: As used herein, the phrase “constituent element” refers to a portion of a chemical structure. A constituent element may be a bond, an atom, a fragment, a functional group, a heteroatom, a moiety or any combination thereof that forms in whole or in part a chemical structure. A constituent element may be used to identify, describe, and/or classify a chemical structure. A constituent element may be used as a search term when querying for documents related to a chemical structure that comprises the constituent element. [0087] Documents are indexed in a format such that they can be fully searched with text-based searching methods. Document data corresponding to the documents to be indexed may be loaded (e.g., uploaded) to a service such as ChemSearch. FIG. 1 shows an exemplary hierarchy of data structures that correspond to a document. Document data 100 comprise chemical structure data 110, text data 130, and metadata 140. Chemical structure data 110 corresponds to chemical structure information, such as a chemical structure representation. Chemical structure data may be stored in any number of standard formats (e.g., a simplified molecular input line entry specification (SMILES) or SMILES arbitrary target specification (SMARTS) based string or as formatted binary data). Chemical structure data 110 comprises bit-screening data 150 and connection data 160. Bit-screening data 150 correspond to one or more constituent elements of a chemical structure. Connection data 160 correspond to one or more connections (e.g., interactions, bonds) between a plurality of the one or more constituent elements. In certain embodiments, chemical structure data (e.g., bit-screening data and connection data) are stored as strings or converted to strings such that all document data used for searching is searchable with a text search engine. Text data 130 corresponds to descriptive information about the chemical and/or its structure. For example, text data may describe properties of a chemical (e.g., its structure) and/or it may describe processes, reactions, or formulations/mixtures involving the chemical. In certain embodiments, document data may include metadata that can be used to identify the document and its contents. For example, a document's metadata may include a unique ID and bucket ID. The metadata may be persisted to allow the document to be referenced in a database.)
1.generate an adjacency matrix based on the information on the plurality of atomic regions and the information on the bonding relationships between the plurality of atoms, / 10. generating an adjacency matrix based on the information on the plurality of atomic regions and the information on the bonding relationships between the plurality of atoms; / 16. generating an adjacency matrix based on the information on the plurality of atomic regions and the information on the bonding relationships between the plurality of atoms; (Oh: [0004] In order to reproduce chemical structures in a document, a range of standard formats are used to efficiently store the chemical structure data. One type of format uses connection tables, adjacency matrices, or similar data structures to relate atoms and bonds as edges and nodes. Another type of format uses linear string notations based on depth first or breadth first traversal. The use of standardized data formats for storing chemical structure data enables algorithmic searching of the data. Furthermore, chemical structure data in standard formats can be indexed with a document in a database. [0088] Document data 100 has been augmented during indexing to comprise string tags 120 (as depicted by the dashed line connecting the two in FIG. 1). String tags are a sequence of characters that provide an alphanumeric text-based string for identifying, classifying, and/or describing chemical structures corresponding to chemical structure data in document data. In certain embodiments, string tags are generated using bit-screening data by performing an atom-by-atom or similar structure-based search on the bit-screening data therein to identify constituent elements corresponding to the bit-screening data and populating the string tags with strings in a predefined list or array that identify, classify, and/or describe the constituent elements. In certain embodiments, string tags are populated using an array that comprises the strings and corresponding reference bit screening data that is compared to the bit screening data in document data. The predefined list may be manually created by storing strings for common constituent elements in chemical structures and associations to reference bit-screening data that correspond to those common constituent elements. Thus, the reference bit-screening data associated with the pre-defined strings can be matched, using the structure-based search, to bit-screening data in document data in order to generate string tags that are populated with appropriate descriptive strings from the pre-defined list for constituent elements corresponding to the bit-screening data in the document data. String tags may also be generated using appropriate ad hoc structure-based methods that populate the string tags with appropriate descriptive strings. String tags may be associated with chemical structure data or directly with directly with the document data that comprises the chemical structure data. Referring again to FIG. 1, string tags 120 are associated with document data 100, but not directly associated with chemical structure data 110.)
1.and generate a string format corresponding to the structural formula image based on the adjacency matrix. / 10. and generating a string format corresponding to the structural formula image based on the adjacency matrix./ 16. and generating a string format corresponding to the structural formula image based on the adjacency matrix. (Oh: [0090] FIG. 2 is a block diagram of an exemplary method for indexing documents comprising chemical structure information. Indexing method 200 is used to augment document data by generating one or more string tags from chemical structure data in the document data. In step 210, document data comprising chemical structure data is received by a processor of a computing device. In step 220, bit-screening data and connection data in the chemical structure data is identified or extracted. In step 230, the bit-screening data identified or extracted in step 220 is used to generate a string tag. In step 240, the string tag generated in step 230 is associated with the document data directly. In step 250, the string tag is outputted. The string tag outputted in step 250 is stored with the document data for later searching. In some embodiments, document data is augmented to comprise a string tag. In some embodiments, a string tag is stored separate from document data. When a string tag is stored separate from document data, the document data may be augmented to comprise the association of the string tag to the document data such that the string tag is searchable when the document data is being searched. [0091]-[0095])
Even if Oh does not specifically teach limitations from claims 8, 14 and 20 for: “having the information on bonding relationships between the plurality of atoms as edges based on the information on the plurality of atomic regions”
Chen teaches:
1. An artificial intelligence device comprising: / 10. A computerized method comprising: / 16. A non-transitory computer-readable medium configured to store instructions (Chen: abstract, The present invention is a molecular graph representation learning method based on contrastive learning, the method comprising: obtaining a molecular fingerprint representation of each molecule, and calculating a similarity between each two molecular fingerprints; collecting a full amount of chemical functional group information, and matching a corresponding functional group for each atom in the molecule; using a heterogeneous graph to model a molecular graph; using a RGCN in the structure-aware molecular encoder to encode the representation of each atom in the molecule and the representation of the functional group to which the atom belongs, and mapping the molecule to a feature space through an aggregation function to obtain a structure-aware feature representation; according to the fingerprint similarity between molecules, selecting positive and negative samples, and carrying out a comparative learning in the feature space; obtaining the structure-aware molecular encoder by using the contrastive learning method for training on a large-sample molecular dataset, and applying the structure-aware molecular encoder to a prediction task of downstream molecular attributes. The present invention helps to capture more abundant molecular structure information and solve the problem on molecular property prediction.)
1. a memory configured to store a structural formula image; / 16. that when executed by one or more processors, cause the one or more processors to perform operations comprising: (Chen: [0004] In recent years, graph representation learning using a Graph Neural Network (GNN) has received extensive attention. The graph neural network usually updates a hidden state of a node by a weighted sum of neighborhood states. By passing information between nodes, the graph neural network is able to capture information from its neighborhood. [0005] Molecular graphs are a type of natural graph data with rich structural information. At present, many studies have used deep learning methods to encode molecules to accelerate drug development and molecular identification. To represent the molecules in a vector space, traditional molecular fingerprints attempt to encode the molecules as fixed-length binary vectors, with each bit on the molecular fingerprint corresponding to a molecular fragment.)
1. and a processor configured to obtain information on a plurality of atomic regions from the structural formula image, / 10. obtaining information on a plurality of atomic regions from a structural formula image; / 16. obtaining information on a plurality of atomic regions from a structural formula image; (Chen: [0033]-[0038], [0035] Compared with the prior art, the present invention has the following beneficial effects: [0036] 1. Different from the existing supervised pre-training method, the present invention uses the self-supervised contrastive learning method to train the structure-aware molecular encoder. Supervised learning has the problem of insufficient labeled data, and the model obtained through label training often only involve specific knowledge, which is far less rich than the structural information of the data itself. Therefore, using the self-supervised contrastive learning method to construct labels based on the structure or characteristics of the molecular graph data itself for molecular graph representation learning is helpful to capture richer molecular structural information, and it is easier to obtain discriminative high-level features.)
1. obtain information on bonding relationships between a plurality of atoms based on the information on the plurality of atomic regions, / 10. obtaining information on bonding relationships between a plurality of atoms based on the information on the plurality of atomic regions; / 16. obtaining information on bonding relationships between a plurality of atoms based on the information on the plurality of atomic regions; (Chen: [0019] The molecular fingerprint is selected as one of Morgan fingerprint, Molecular ACCess System (MACCs) fingerprint and topological fingerprint. The Morgan fingerprint sets a radius from a specific atom to count the number of partial molecular structures within the radius to form the molecular fingerprint. The MACCs fingerprint pre-specifies the partial molecular structures of 166 molecules, and when the molecular structure is contained, the corresponding position is recorded as 1, otherwise it is recorded as 0. The topological fingerprint does not need to pre-specify the partial molecular structures, but calculates all molecular paths between the minimum number of bonds and the maximum number of bonds, and hashes each subgraph to generate an ID of each bit, and then generates the molecular fingerprint. [0023] In step (3), modeling the molecular graph with heterogeneous graph is beneficial to characterize the different attributes of each node and edge. [0024]-[0029], [0037] 2. In the present invention, modeling the molecular graph with the heterogeneous graph is beneficial to characterize the different attributes of each atom and bond. [0038] 3. Different from the existing molecular graph representation learning methods lacking prior knowledge in the field of chemistry, the present invention proposes to use the structure-aware graph neural network to learn the molecular representation, and directly encode the functional group information that plays a decisive role in molecular properties into the feature representation of the graph.)
1.generate an adjacency matrix based on the information on the plurality of atomic regions and the information on the bonding relationships between the plurality of atoms, / 10. generating an adjacency matrix based on the information on the plurality of atomic regions and the information on the bonding relationships between the plurality of atoms; / 16. generating an adjacency matrix based on the information on the plurality of atomic regions and the information on the bonding relationships between the plurality of atoms; (Chen: [0023] In step (3), modeling the molecular graph with heterogeneous graph is beneficial to characterize the different attributes of each node and edge. [0024] The specific process of step (4) is: [0025] taking the heterogeneous graph with initialized node features and functional group features as an input of the structure-aware molecular encoder, transferring information by the relational graph convolutional network (RGCN) in the structure-aware molecular encoder through calculating and aggregating information for different types of edges, and integrating the information aggregated by different edges for different types of nodes; [0026] after obtaining the feature representation of each atom and the functional group that the atom belongs to, then aggregating the features of the nodes and the functional groups to obtain the structure-aware feature representation of the molecule. [0027] A formula for the information transfer of the relational graph convolutional network (RGCN) is as follows:
PNG
media_image1.png
38
336
media_image1.png
Greyscale
[0028] wherein, R is a set of all edges, Ni r is all neighbor nodes which are adjacent to the node i and are of edge type r, ci,r is a parameter that can be learned, Wr l is a weight matrix of the current layer l, hi l is a feature vector of the current layer l to the current node i; the feature of each neighbor node is multiplied by a weight corresponding to the edge type, and then is multiplied by a learnable parameter, and then summed, and finally, the information transferred by a self-loop edge is added and the activation function σ is passed, which is used as an output of the layer and an input of a next layer.)
1.and generate a string format corresponding to the structural formula image based on the adjacency matrix. / 10. and generating a string format corresponding to the structural formula image based on the adjacency matrix./ 16. and generating a string format corresponding to the structural formula image based on the adjacency matrix. (Chen: [0029] In step (5), when selecting the positive and negative samples, one molecule of which similarity with a target molecule is greater than a certain threshold is selected as the positive sample, K molecules of which each similarity is less than a certain threshold are selected as the negative samples; a feature representation corresponding to the target molecule is denoted as q, a feature representation of the positive sample is denoted as k0, and the feature representations of K negative samples are denoted as k1, . . . , kK. [0030] After obtaining the feature representations of each target molecule and the positive and negative samples thereof, a loss is calculated by using a loss function, and the parameters of the structure-aware molecular encoder are updated through a back-propagation algorithm, which causes the model to recognize the target molecule and the positive samples as similar instances and distinguish the target molecule and the positive samples from dissimilar samples.)
8/14/20. wherein the generating of the adjacency matrix comprises generating the adjacency matrix having the plurality of atoms as vertices and having the information on bonding relationships between the plurality of atoms as edges based on the information on the plurality of atomic regions (Chen: [0046] Modeling the target molecule and its corresponding positive and negative samples by using a heterogeneous graph, which aims to characterize the different attributes of each node and edge. Inputting the sample data of the molecules into a structure-aware molecular encoder shown in FIG. 2 , and obtaining the feature representations corresponding to the target sample and the positive and negative samples. Denoting a feature representation corresponding to the target molecule as q, denoting a feature representation of the positive sample as k0, and denoting the feature representations of K negative samples as k1, . . . , kK. [0050]-[0051], [0050] As shown in FIG. 2 , it is a schematic diagram of a structure-aware graph neural network provided by an embodiment of the present invention. Modeling the molecules by using the heterogeneous graph with initialized node features and functional group features, and characterizing the different attributes of each node and edge. Taking the heterogeneous graph as an input of the structure-aware molecular encoder, and then calculating and aggregating information for different types of edges by utilizing RGCN, and integrating the information aggregated by different edges for different types of nodes to transfer information. The RGCN takes into account the type of edge, and in order to transfer the features of the nodes in a previous layer to a next layer, the RGCN adds a special self-loop edge for each node.)
It would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to modify Oh’s method and system for indexing molecular structures with the structural representations and text-based search method for indexing data of Chen. The determination of obviousness is predicated upon the following findings: One skilled in the art would have been motivated to modify Oh in order to incorporate a more efficient indexing method to facilitate searching as proposed by Chen. Furthermore, the prior art collectively includes each element claimed (though not all in the same reference), and one of ordinary skill in the art could have combined the elements in the manner explained above using known engineering design, interface and/or programming techniques, without changing a “fundamental” operating principle of Oh, while the teaching of Chen continues to perform the same function as originally taught prior to being combined, in order to produce the repeatable and predictable result of indexing data structures for more efficient searching. It is for at least the aforementioned reasons that the examiner has reached a conclusion of obviousness with respect to the claim in question.
Consider Claim 2.
The combination of Oh and Chen teaches:
2. The artificial intelligence device of claim 1, wherein the processor is configured to input the structural formula image to an atomic region recognition model and obtain the information on the plurality of atomic regions output from the atomic region recognition model. (Oh: [0077] In some embodiments, a first data structure is stored on a first computer readable medium, a second data structure is stored on a second computer readable medium, and the association between the first data structure and second data structure is stored on the first computer readable medium. In some embodiments, a first data structure is stored on a first computer readable medium, a second data structure is stored on a second computer readable medium, and the association between the first data structure and second data structure is stored on the second computer readable medium. [0078], [0088] Document data 100 has been augmented during indexing to comprise string tags 120 (as depicted by the dashed line connecting the two in FIG. 1). String tags are a sequence of characters that provide an alphanumeric text-based string for identifying, classifying, and/or describing chemical structures corresponding to chemical structure data in document data. In certain embodiments, string tags are generated using bit-screening data by performing an atom-by-atom or similar structure-based search on the bit-screening data therein to identify constituent elements corresponding to the bit-screening data and populating the string tags with strings in a predefined list or array that identify, classify, and/or describe the constituent elements. In certain embodiments, string tags are populated using an array that comprises the strings and corresponding reference bit screening data that is compared to the bit screening data in document data. The predefined list may be manually created by storing strings for common constituent elements in chemical structures and associations to reference bit-screening data that correspond to those common constituent elements. Thus, the reference bit-screening data associated with the pre-defined strings can be matched, using the structure-based search, to bit-screening data in document data in order to generate string tags that are populated with appropriate descriptive strings from the pre-defined list for constituent elements corresponding to the bit-screening data in the document data. String tags may also be generated using appropriate ad hoc structure-based methods that populate the string tags with appropriate descriptive strings. String tags may be associated with chemical structure data or directly with directly with the document data that comprises the chemical structure data. Referring again to FIG. 1, string tags 120 are associated with document data 100, but not directly associated with chemical structure data 110.Chen: [0010] A molecular graph representation learning method based on contrastive learning, wherein, the method comprises the following steps: [0011] (1) obtaining a molecular fingerprint representation of each molecule, and calculating a similarity between each two molecular fingerprints; [0012] (2) collecting a full amount of chemical functional group information, and matching a corresponding functional group for each atom in the molecule; wherein, when an atom belongs to a plurality of functional groups, a functional group containing a larger number of atoms is preferentially matched; [0013] (3) using a heterogeneous graph to model a molecular graph, wherein the heterogeneous graph is a graph containing different types of nodes and edges, different atoms correspond to different node types, and different bonds correspond to different edge types)
Consider Claim 3.
The combination of Oh and Chen teaches:
3. The artificial intelligence device of claim 2, wherein the information on the plurality of atomic regions includes information on positions of the plurality of atomic regions and information on the plurality of atoms. (Oh: [0077] In some embodiments, a first data structure is stored on a first computer readable medium, a second data structure is stored on a second computer readable medium, and the association between the first data structure and second data structure is stored on the first computer readable medium. In some embodiments, a first data structure is stored on a first computer readable medium, a second data structure is stored on a second computer readable medium, and the association between the first data structure and second data structure is stored on the second computer readable medium. [0078], [0088] Document data 100 has been augmented during indexing to comprise string tags 120 (as depicted by the dashed line connecting the two in FIG. 1). String tags are a sequence of characters that provide an alphanumeric text-based string for identifying, classifying, and/or describing chemical structures corresponding to chemical structure data in document data. In certain embodiments, string tags are generated using bit-screening data by performing an atom-by-atom or similar structure-based search on the bit-screening data therein to identify constituent elements corresponding to the bit-screening data and populating the string tags with strings in a predefined list or array that identify, classify, and/or describe the constituent elements. In certain embodiments, string tags are populated using an array that comprises the strings and corresponding reference bit screening data that is compared to the bit screening data in document data. The predefined list may be manually created by storing strings for common constituent elements in chemical structures and associations to reference bit-screening data that correspond to those common constituent elements. Thus, the reference bit-screening data associated with the pre-defined strings can be matched, using the structure-based search, to bit-screening data in document data in order to generate string tags that are populated with appropriate descriptive strings from the pre-defined list for constituent elements corresponding to the bit-screening data in the document data. String tags may also be generated using appropriate ad hoc structure-based methods that populate the string tags with appropriate descriptive strings. String tags may be associated with chemical structure data or directly with directly with the document data that comprises the chemical structure data. Referring again to FIG. 1, string tags 120 are associated with document data 100, but not directly associated with chemical structure data 110.Chen: [0010] A molecular graph representation learning method based on contrastive learning, wherein, the method comprises the following steps: [0011] (1) obtaining a molecular fingerprint representation of each molecule, and calculating a similarity between each two molecular fingerprints; [0012] (2) collecting a full amount of chemical functional group information, and matching a corresponding functional group for each atom in the molecule; wherein, when an atom belongs to a plurality of functional groups, a functional group containing a larger number of atoms is preferentially matched; [0013] (3) using a heterogeneous graph to model a molecular graph, wherein the heterogeneous graph is a graph containing different types of nodes and edges, different atoms correspond to different node types, and different bonds correspond to different edge types)
Consider Claim 4.
The combination of Oh and Chen teaches:
4. The artificial intelligence device of claim 3, wherein the processor is configured to obtain a bonding image between a first atom and a second atom based on information on a first atomic region and information on a second atomic region from the information on the plurality of atomic regions, and obtain information on a bonding relationship between the first atom and the second atom based on the bonding image. (Oh: [0077] In some embodiments, a first data structure is stored on a first computer readable medium, a second data structure is stored on a second computer readable medium, and the association between the first data structure and second data structure is stored on the first computer readable medium. In some embodiments, a first data structure is stored on a first computer readable medium, a second data structure is stored on a second computer readable medium, and the association between the first data structure and second data structure is stored on the second computer readable medium. [0078], [0088] Document data 100 has been augmented during indexing to comprise string tags 120 (as depicted by the dashed line connecting the two in FIG. 1). String tags are a sequence of characters that provide an alphanumeric text-based string for identifying, classifying, and/or describing chemical structures corresponding to chemical structure data in document data. In certain embodiments, string tags are generated using bit-screening data by performing an atom-by-atom or similar structure-based search on the bit-screening data therein to identify constituent elements corresponding to the bit-screening data and populating the string tags with strings in a predefined list or array that identify, classify, and/or describe the constituent elements. In certain embodiments, string tags are populated using an array that comprises the strings and corresponding reference bit screening data that is compared to the bit screening data in document data. The predefined list may be manually created by storing strings for common constituent elements in chemical structures and associations to reference bit-screening data that correspond to those common constituent elements. Thus, the reference bit-screening data associated with the pre-defined strings can be matched, using the structure-based search, to bit-screening data in document data in order to generate string tags that are populated with appropriate descriptive strings from the pre-defined list for constituent elements corresponding to the bit-screening data in the document data. String tags may also be generated using appropriate ad hoc structure-based methods that populate the string tags with appropriate descriptive strings. String tags may be associated with chemical structure data or directly with directly with the document data that comprises the chemical structure data. Referring again to FIG. 1, string tags 120 are associated with document data 100, but not directly associated with chemical structure data 110.Chen: [0010] A molecular graph representation learning method based on contrastive learning, wherein, the method comprises the following steps: [0011] (1) obtaining a molecular fingerprint representation of each molecule, and calculating a similarity between each two molecular fingerprints; [0012] (2) collecting a full amount of chemical functional group information, and matching a corresponding functional group for each atom in the molecule; wherein, when an atom belongs to a plurality of functional groups, a functional group containing a larger number of atoms is preferentially matched; [0013] (3) using a heterogeneous graph to model a molecular graph, wherein the heterogeneous graph is a graph containing different types of nodes and edges, different atoms correspond to different node types, and different bonds correspond to different edge types)
Consider Claim 5.
The combination of Oh and Chen teaches:
5. The artificial intelligence device of claim 4, wherein the processor is configured to input the bonding image to a bonding relationship recognition model, and obtain the information on the bonding relationship output from the bonding relationship recognition model. (Oh:[0054] In another aspect, the present invention is directed to a system for indexing a document to facilitate chemical structure searching, the system comprising: a processor; and a non-transitory computer readable medium having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to: receive, by a processor of a computing device, document data corresponding to the document, wherein the document data comprise chemical structure data corresponding to a chemical structure; identify or extract, by the processor, bit-screening data and connection data in the chemical structure data, wherein the bit-screening data correspond to one or more constituent elements of the chemical structure, and the connection data correspond to connections (e.g., interactions, bonds) between the one or more constituent elements; generate, by the processor, a string tag based on at least a portion of the identified bit-screening data, the string tag comprising an alphanumeric value for describing the chemical structure that corresponds to the chemical structure data (e.g., for use in querying for documents comprising the chemical structure data); associate, by the processor, the string tag with the chemical structure data or the document data; and output, by the processor, the string tag (e.g., for storage on a non-transitory computer readable medium). In certain embodiments, the instructions, when executed by the processor, cause the processor to: convert, by the processor, the bit-screening data and the connection data to one or more strings. In certain embodiments, the string tag comprises natural language text. Chen: [0018] In step (1), the SMILES representation of the molecules is transformed into the molecular fingerprint by Rdkit which is a powerful tool for cheminformatics. According to different calculation methods, different kinds of molecular fingerprints of the same molecule can be obtained. [0019] The molecular fingerprint is selected as one of Morgan fingerprint, Molecular ACCess System (MACCs) fingerprint and topological fingerprint. The Morgan fingerprint sets a radius from a specific atom to count the number of partial molecular structures within the radius to form the molecular fingerprint. The MACCs fingerprint pre-specifies the partial molecular structures of 166 molecules, and when the molecular structure is contained, the corresponding position is recorded as 1, otherwise it is recorded as 0. The topological fingerprint does not need to pre-specify the partial molecular structures, but calculates all molecular paths between the minimum number of bonds and the maximum number of bonds, and hashes each subgraph to generate an ID of each bit, and then generates the molecular fingerprint.[0020] The evaluation method often used in the calculation of similarity between compound molecules is a Tanimoto coefficient. The similarity between two molecular fingerprints is calculated using the Tanimoto coefficient, and the formula is as follows:
PNG
media_image2.png
39
109
media_image2.png
Greyscale
)
Consider Claim 6.
The combination of Oh and Chen teaches:
6. The artificial intelligence device of claim 4, wherein the processor is configured to obtain information on a position of a center point of the first atomic region and a position of a center point of the second atomic region based on the information on the first atomic region and the information on the second atomic region. (Oh:[0054] In another aspect, the present invention is directed to a system for indexing a document to facilitate chemical structure searching, the system comprising: a processor; and a non-transitory computer readable medium having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to: receive, by a processor of a computing device, document data corresponding to the document, wherein the document data comprise chemical structure data corresponding to a chemical structure; identify or extract, by the processor, bit-screening data and connection data in the chemical structure data, wherein the bit-screening data correspond to one or more constituent elements of the chemical structure, and the connection data correspond to connections (e.g., interactions, bonds) between the one or more constituent elements; generate, by the processor, a string tag based on at least a portion of the identified bit-screening data, the string tag comprising an alphanumeric value for describing the chemical structure that corresponds to the chemical structure data (e.g., for use in querying for documents comprising the chemical structure data); associate, by the processor, the string tag with the chemical structure data or the document data; and output, by the processor, the string tag (e.g., for storage on a non-transitory computer readable medium). In certain embodiments, the instructions, when executed by the processor, cause the processor to: convert, by the processor, the bit-screening data and the connection data to one or more strings. In certain embodiments, the string tag comprises natural language text. Chen: [0018] In step (1), the SMILES representation of the molecules is transformed into the molecular fingerprint by Rdkit which is a powerful tool for cheminformatics. According to different calculation methods, different kinds of molecular fingerprints of the same molecule can be obtained. [0019] The molecular fingerprint is selected as one of Morgan fingerprint, Molecular ACCess System (MACCs) fingerprint and topological fingerprint. The Morgan fingerprint sets a radius from a specific atom to count the number of partial molecular structures within the radius to form the molecular fingerprint. The MACCs fingerprint pre-specifies the partial molecular structures of 166 molecules, and when the molecular structure is contained, the corresponding position is recorded as 1, otherwise it is recorded as 0. The topological fingerprint does not need to pre-specify the partial molecular structures, but calculates all molecular paths between the minimum number of bonds and the maximum number of bonds, and hashes each subgraph to generate an ID of each bit, and then generates the molecular fingerprint.[0020] The evaluation method often used in the calculation of similarity between compound molecules is a Tanimoto coefficient. The similarity between two molecular fingerprints is calculated using the Tanimoto coefficient, and the formula is as follows:
PNG
media_image2.png
39
109
media_image2.png
Greyscale
)
Consider Claim 7.
The combination of Oh and Chen teaches:
7. The artificial intelligence device of claim 4, wherein the processor is configured to select the second atomic region located within a predetermined distance from the first atomic region. (Oh: [0088] Document data 100 has been augmented during indexing to comprise string tags 120 (as depicted by the dashed line connecting the two in FIG. 1). String tags are a sequence of characters that provide an alphanumeric text-based string for identifying, classifying, and/or describing chemical structures corresponding to chemical structure data in document data. In certain embodiments, string tags are generated using bit-screening data by performing an atom-by-atom or similar structure-based search on the bit-screening data therein to identify constituent elements corresponding to the bit-screening data and populating the string tags with strings in a predefined list or array that identify, classify, and/or describe the constituent elements. In certain embodiments, string tags are populated using an array that comprises the strings and corresponding reference bit screening data that is compared to the bit screening data in document data. The predefined list may be manually created by storing strings for common constituent elements in chemical structures and associations to reference bit-screening data that correspond to those common constituent elements. Thus, the reference bit-screening data associated with the pre-defined strings can be matched, using the structure-based search, to bit-screening data in document data in order to generate string tags that are populated with appropriate descriptive strings from the pre-defined list for constituent elements corresponding to the bit-screening data in the document data. String tags may also be generated using appropriate ad hoc structure-based methods that populate the string tags with appropriate descriptive strings. String tags may be associated with chemical structure data or directly with directly with the document data that comprises the chemical structure data. Referring again to FIG. 1, string tags 120 are associated with document data 100, but not directly associated with chemical structure data 110.)
Consider Claims 8, 14 and 20.
The combination of Oh and Chen teaches:
8. The artificial intelligence device of claim 1, wherein the processor is configured to generate the adjacency matrix having the plurality of atoms as vertices and having the information on bonding relationships between the plurality of atoms as edges based on the information on the plurality of atomic regions. / 14. The computerized method of claim 10, wherein the generating of the adjacency matrix comprises generating the adjacency matrix having the plurality of atoms as vertices and having the information on bonding relationships between the plurality of atoms as edges based on the information on the plurality of atomic regions. / 20. The non-transitory computer-readable medium of claim 16, wherein the generating of the adjacency matrix comprises generating the adjacency matrix having the plurality of atoms as vertices and having the information on bonding relationships between the plurality of atoms as edges based on the information on the plurality of atomic regions. (Oh: [0004] In order to reproduce chemical structures in a document, a range of standard formats are used to efficiently store the chemical structure data. One type of format uses connection tables, adjacency matrices, or similar data structures to relate atoms and bonds as edges and nodes. Another type of format uses linear string notations based on depth first or breadth first traversal. The use of standardized data formats for storing chemical structure data enables algorithmic searching of the data. Furthermore, chemical structure data in standard formats can be indexed with a document in a database. [0088] Document data 100 has been augmented during indexing to comprise string tags 120 (as depicted by the dashed line connecting the two in FIG. 1). String tags are a sequence of characters that provide an alphanumeric text-based string for identifying, classifying, and/or describing chemical structures corresponding to chemical structure data in document data. In certain embodiments, string tags are generated using bit-screening data by performing an atom-by-atom or similar structure-based search on the bit-screening data therein to identify constituent elements corresponding to the bit-screening data and populating the string tags with strings in a predefined list or array that identify, classify, and/or describe the constituent elements. In certain embodiments, string tags are populated using an array that comprises the strings and corresponding reference bit screening data that is compared to the bit screening data in document data. The predefined list may be manually created by storing strings for common constituent elements in chemical structures and associations to reference bit-screening data that correspond to those common constituent elements. Thus, the reference bit-screening data associated with the pre-defined strings can be matched, using the structure-based search, to bit-screening data in document data in order to generate string tags that are populated with appropriate descriptive strings from the pre-defined list for constituent elements corresponding to the bit-screening data in the document data. String tags may also be generated using appropriate ad hoc structure-based methods that populate the string tags with appropriate descriptive strings. String tags may be associated with chemical structure data or directly with directly with the document data that comprises the chemical structure data. Referring again to FIG. 1, string tags 120 are associated with document data 100, but not directly associated with chemical structure data 110. Chen: [0046] Modeling the target molecule and its corresponding positive and negative samples by using a heterogeneous graph, which aims to characterize the different attributes of each node and edge. Inputting the sample data of the molecules into a structure-aware molecular encoder shown in FIG. 2 , and obtaining the feature representations corresponding to the target sample and the positive and negative samples. Denoting a feature representation corresponding to the target molecule as q, denoting a feature representation of the positive sample as k0, and denoting the feature representations of K negative samples as k1, . . . , kK. [0050]-[0051], [0050] As shown in FIG. 2 , it is a schematic diagram of a structure-aware graph neural network provided by an embodiment of the present invention.)
Consider Claim 11.
The combination of Oh and Chen teaches:
11. The computerized method of claim 10, wherein the obtaining of the information on the plurality of atomic regions comprises: inputting the structural formula image to an atomic region recognition model; and obtaining the information on the plurality of atomic regions output from the atomic region recognition model, wherein the information on the plurality of atomic regions includes information on positions of the plurality of atomic regions and information on the plurality of atoms. (Oh:[0054] In another aspect, the present invention is directed to a system for indexing a document to facilitate chemical structure searching, the system comprising: a processor; and a non-transitory computer readable medium having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to: receive, by a processor of a computing device, document data corresponding to the document, wherein the document data comprise chemical structure data corresponding to a chemical structure; identify or extract, by the processor, bit-screening data and connection data in the chemical structure data, wherein the bit-screening data correspond to one or more constituent elements of the chemical structure, and the connection data correspond to connections (e.g., interactions, bonds) between the one or more constituent elements; generate, by the processor, a string tag based on at least a portion of the identified bit-screening data, the string tag comprising an alphanumeric value for describing the chemical structure that corresponds to the chemical structure data (e.g., for use in querying for documents comprising the chemical structure data); associate, by the processor, the string tag with the chemical structure data or the document data; and output, by the processor, the string tag (e.g., for storage on a non-transitory computer readable medium). In certain embodiments, the instructions, when executed by the processor, cause the processor to: convert, by the processor, the bit-screening data and the connection data to one or more strings. In certain embodiments, the string tag comprises natural language text. Chen: [0018] In step (1), the SMILES representation of the molecules is transformed into the molecular fingerprint by Rdkit which is a powerful tool for cheminformatics. According to different calculation methods, different kinds of molecular fingerprints of the same molecule can be obtained. [0019] The molecular fingerprint is selected as one of Morgan fingerprint, Molecular ACCess System (MACCs) fingerprint and topological fingerprint. The Morgan fingerprint sets a radius from a specific atom to count the number of partial molecular structures within the radius to form the molecular fingerprint. The MACCs fingerprint pre-specifies the partial molecular structures of 166 molecules, and when the molecular structure is contained, the corresponding position is recorded as 1, otherwise it is recorded as 0. The topological fingerprint does not need to pre-specify the partial molecular structures, but calculates all molecular paths between the minimum number of bonds and the maximum number of bonds, and hashes each subgraph to generate an ID of each bit, and then generates the molecular fingerprint.[0020] The evaluation method often used in the calculation of similarity between compound molecules is a Tanimoto coefficient. The similarity between two molecular fingerprints is calculated using the Tanimoto coefficient, and the formula is as follows:
PNG
media_image2.png
39
109
media_image2.png
Greyscale
)
Consider Claim 12.
The combination of Oh and Chen teaches:
12. The computerized method of claim 11, wherein the obtaining of the information on the bonding relationships comprises: obtaining a bonding image between a first atom and a second atom based on information on a first atomic region and information on a second atomic region from the information on the plurality of atomic regions; and obtaining information on a bonding relationship between the first atom and the second atom based on the bonding image, wherein the obtaining of the information on the bonding relationship comprises: inputting the bonding image to a bonding relationship recognition model; and obtaining the information on the bonding relationship output from the bonding relationship recognition model. (Oh: [0004] In order to reproduce chemical structures in a document, a range of standard formats are used to efficiently store the chemical structure data. One type of format uses connection tables, adjacency matrices, or similar data structures to relate atoms and bonds as edges and nodes. Another type of format uses linear string notations based on depth first or breadth first traversal. The use of standardized data formats for storing chemical structure data enables algorithmic searching of the data. Furthermore, chemical structure data in standard formats can be indexed with a document in a database. [0088] Document data 100 has been augmented during indexing to comprise string tags 120 (as depicted by the dashed line connecting the two in FIG. 1). String tags are a sequence of characters that provide an alphanumeric text-based string for identifying, classifying, and/or describing chemical structures corresponding to chemical structure data in document data. In certain embodiments, string tags are generated using bit-screening data by performing an atom-by-atom or similar structure-based search on the bit-screening data therein to identify constituent elements corresponding to the bit-screening data and populating the string tags with strings in a predefined list or array that identify, classify, and/or describe the constituent elements. In certain embodiments, string tags are populated using an array that comprises the strings and corresponding reference bit screening data that is compared to the bit screening data in document data. The predefined list may be manually created by storing strings for common constituent elements in chemical structures and associations to reference bit-screening data that correspond to those common constituent elements. Thus, the reference bit-screening data associated with the pre-defined strings can be matched, using the structure-based search, to bit-screening data in document data in order to generate string tags that are populated with appropriate descriptive strings from the pre-defined list for constituent elements corresponding to the bit-screening data in the document data. String tags may also be generated using appropriate ad hoc structure-based methods that populate the string tags with appropriate descriptive strings. String tags may be associated with chemical structure data or directly with directly with the document data that comprises the chemical structure data. Referring again to FIG. 1, string tags 120 are associated with document data 100, but not directly associated with chemical structure data 110. Chen: [0046] Modeling the target molecule and its corresponding positive and negative samples by using a heterogeneous graph, which aims to characterize the different attributes of each node and edge. Inputting the sample data of the molecules into a structure-aware molecular encoder shown in FIG. 2 , and obtaining the feature representations corresponding to the target sample and the positive and negative samples. Denoting a feature representation corresponding to the target molecule as q, denoting a feature representation of the positive sample as k0, and denoting the feature representations of K negative samples as k1, . . . , kK. [0050]-[0051], [0050] As shown in FIG. 2 , it is a schematic diagram of a structure-aware graph neural network provided by an embodiment of the present invention.)
Consider Claim 13.
The combination of Oh and Chen teaches:
13. The computerized method of claim 12, further comprising: obtaining information on a position of a center point of the first atomic region and a position of a center point the second atomic region based on the information on the first atomic region and the information on the second atomic region; and selecting the second atomic region located within a predetermined distance from the first atomic region. (Oh: [0088] Document data 100 has been augmented during indexing to comprise string tags 120 (as depicted by the dashed line connecting the two in FIG. 1). String tags are a sequence of characters that provide an alphanumeric text-based string for identifying, classifying, and/or describing chemical structures corresponding to chemical structure data in document data. In certain embodiments, string tags are generated using bit-screening data by performing an atom-by-atom or similar structure-based search on the bit-screening data therein to identify constituent elements corresponding to the bit-screening data and populating the string tags with strings in a predefined list or array that identify, classify, and/or describe the constituent elements. In certain embodiments, string tags are populated using an array that comprises the strings and corresponding reference bit screening data that is compared to the bit screening data in document data. The predefined list may be manually created by storing strings for common constituent elements in chemical structures and associations to reference bit-screening data that correspond to those common constituent elements. Thus, the reference bit-screening data associated with the pre-defined strings can be matched, using the structure-based search, to bit-screening data in document data in order to generate string tags that are populated with appropriate descriptive strings from the pre-defined list for constituent elements corresponding to the bit-screening data in the document data. String tags may also be generated using appropriate ad hoc structure-based methods that populate the string tags with appropriate descriptive strings. String tags may be associated with chemical structure data or directly with directly with the document data that comprises the chemical structure data. Referring again to FIG. 1, string tags 120 are associated with document data 100, but not directly associated with chemical structure data 110.)
Consider Claim 9 and 15.
The combination of Oh and Chen teaches:
9. The artificial intelligence device of claim 1, wherein the processor is configured to convert the structural formula image into the string format including a Simplified Molecular Input Line Entry System (SMILE) based on the generated adjacency matrix. / 15. The computerized method of claim 10, wherein the generating of the string format comprises converting the structural formula image into the string format including a Simplified Molecular Input Line Entry System (SMILE) based on the generated adjacency matrix. (Oh: [0074] Constituent element: As used herein, the phrase “constituent element” refers to a portion of a chemical structure. A constituent element may be a bond, an atom, a fragment, a functional group, a heteroatom, a moiety or any combination thereof that forms in whole or in part a chemical structure. A constituent element may be used to identify, describe, and/or classify a chemical structure. A constituent element may be used as a search term when querying for documents related to a chemical structure that comprises the constituent element. [0087] Documents are indexed in a format such that they can be fully searched with text-based searching methods. Document data corresponding to the documents to be indexed may be loaded (e.g., uploaded) to a service such as ChemSearch. FIG. 1 shows an exemplary hierarchy of data structures that correspond to a document. Document data 100 comprise chemical structure data 110, text data 130, and metadata 140. Chemical structure data 110 corresponds to chemical structure information, such as a chemical structure representation. Chemical structure data may be stored in any number of standard formats (e.g., a simplified molecular input line entry specification (SMILES) or SMILES arbitrary target specification (SMARTS) based string or as formatted binary data). Chemical structure data 110 comprises bit-screening data 150 and connection data 160. Bit-screening data 150 correspond to one or more constituent elements of a chemical structure. Connection data 160 correspond to one or more connections (e.g., interactions, bonds) between a plurality of the one or more constituent elements. In certain embodiments, chemical structure data (e.g., bit-screening data and connection data) are stored as strings or converted to strings such that all document data used for searching is searchable with a text search engine. Text data 130 corresponds to descriptive information about the chemical and/or its structure. For example, text data may describe properties of a chemical (e.g., its structure) and/or it may describe processes, reactions, or formulations/mixtures involving the chemical. In certain embodiments, document data may include metadata that can be used to identify the document and its contents. For example, a document's metadata may include a unique ID and bucket ID. The metadata may be persisted to allow the document to be referenced in a database. Chen: [0044] As shown in FIG. 1 , a molecular graph representation learning method based on contrastive learning, wherein, the method comprises the following steps: [0045] firstly, transforming the SMILES representation of the molecules into the molecular fingerprint by Rdkit which is a powerful tool for cheminformatics. For each molecule, after calculating the fingerprint similarities between it and all other molecules using a Tanimoto coefficient, selecting one molecule of which the similarity with the molecule is greater than a certain threshold as a positive sample, and selecting K molecules of which the similarities are less than a certain threshold as negative samples.)
Consider Claim 17.
The combination of Oh and Chen teaches:
17. The non-transitory computer-readable medium of claim 16, wherein the obtaining of the information on the plurality of atomic regions comprises: inputting the structural formula image to an atomic region recognition model; and obtaining the information on the plurality of atomic regions output from the atomic region recognition model, wherein the information on the plurality of atomic regions includes information on positions of the plurality of atomic regions and information on the plurality of atoms. (Oh: [0077] In some embodiments, a first data structure is stored on a first computer readable medium, a second data structure is stored on a second computer readable medium, and the association between the first data structure and second data structure is stored on the first computer readable medium. In some embodiments, a first data structure is stored on a first computer readable medium, a second data structure is stored on a second computer readable medium, and the association between the first data structure and second data structure is stored on the second computer readable medium. [0078], [0088] Document data 100 has been augmented during indexing to comprise string tags 120 (as depicted by the dashed line connecting the two in FIG. 1). String tags are a sequence of characters that provide an alphanumeric text-based string for identifying, classifying, and/or describing chemical structures corresponding to chemical structure data in document data. In certain embodiments, string tags are generated using bit-screening data by performing an atom-by-atom or similar structure-based search on the bit-screening data therein to identify constituent elements corresponding to the bit-screening data and populating the string tags with strings in a predefined list or array that identify, classify, and/or describe the constituent elements. In certain embodiments, string tags are populated using an array that comprises the strings and corresponding reference bit screening data that is compared to the bit screening data in document data. The predefined list may be manually created by storing strings for common constituent elements in chemical structures and associations to reference bit-screening data that correspond to those common constituent elements. Thus, the reference bit-screening data associated with the pre-defined strings can be matched, using the structure-based search, to bit-screening data in document data in order to generate string tags that are populated with appropriate descriptive strings from the pre-defined list for constituent elements corresponding to the bit-screening data in the document data. String tags may also be generated using appropriate ad hoc structure-based methods that populate the string tags with appropriate descriptive strings. String tags may be associated with chemical structure data or directly with directly with the document data that comprises the chemical structure data. Referring again to FIG. 1, string tags 120 are associated with document data 100, but not directly associated with chemical structure data 110.Chen: [0010] A molecular graph representation learning method based on contrastive learning, wherein, the method comprises the following steps: [0011] (1) obtaining a molecular fingerprint representation of each molecule, and calculating a similarity between each two molecular fingerprints; [0012] (2) collecting a full amount of chemical functional group information, and matching a corresponding functional group for each atom in the molecule; wherein, when an atom belongs to a plurality of functional groups, a functional group containing a larger number of atoms is preferentially matched; [0013] (3) using a heterogeneous graph to model a molecular graph, wherein the heterogeneous graph is a graph containing different types of nodes and edges, different atoms correspond to different node types, and different bonds correspond to different edge types)
Consider Claim 18.
The combination of Oh and Chen teaches:
18. The non-transitory computer-readable medium of claim 17, wherein the obtaining of the information on the bonding relationships comprises: obtaining a bonding image between a first atom and a second atom based on information on a first atomic region and information on a second atomic region from the information on the plurality of atomic regions; and obtaining information on a bonding relationship between the first atom and the second atom based on the bonding image, wherein the obtaining of the information on the bonding relationship comprises: inputting the bonding image to a bonding relationship recognition model; and obtaining the information on the bonding relationship output from the bonding relationship recognition model. (Oh: [0077] In some embodiments, a first data structure is stored on a first computer readable medium, a second data structure is stored on a second computer readable medium, and the association between the first data structure and second data structure is stored on the first computer readable medium. In some embodiments, a first data structure is stored on a first computer readable medium, a second data structure is stored on a second computer readable medium, and the association between the first data structure and second data structure is stored on the second computer readable medium. [0078], [0088] Document data 100 has been augmented during indexing to comprise string tags 120 (as depicted by the dashed line connecting the two in FIG. 1). String tags are a sequence of characters that provide an alphanumeric text-based string for identifying, classifying, and/or describing chemical structures corresponding to chemical structure data in document data. In certain embodiments, string tags are generated using bit-screening data by performing an atom-by-atom or similar structure-based search on the bit-screening data therein to identify constituent elements corresponding to the bit-screening data and populating the string tags with strings in a predefined list or array that identify, classify, and/or describe the constituent elements. In certain embodiments, string tags are populated using an array that comprises the strings and corresponding reference bit screening data that is compared to the bit screening data in document data. The predefined list may be manually created by storing strings for common constituent elements in chemical structures and associations to reference bit-screening data that correspond to those common constituent elements. Thus, the reference bit-screening data associated with the pre-defined strings can be matched, using the structure-based search, to bit-screening data in document data in order to generate string tags that are populated with appropriate descriptive strings from the pre-defined list for constituent elements corresponding to the bit-screening data in the document data. String tags may also be generated using appropriate ad hoc structure-based methods that populate the string tags with appropriate descriptive strings. String tags may be associated with chemical structure data or directly with directly with the document data that comprises the chemical structure data. Referring again to FIG. 1, string tags 120 are associated with document data 100, but not directly associated with chemical structure data 110.Chen: [0010] A molecular graph representation learning method based on contrastive learning, wherein, the method comprises the following steps: [0011] (1) obtaining a molecular fingerprint representation of each molecule, and calculating a similarity between each two molecular fingerprints; [0012] (2) collecting a full amount of chemical functional group information, and matching a corresponding functional group for each atom in the molecule; wherein, when an atom belongs to a plurality of functional groups, a functional group containing a larger number of atoms is preferentially matched; [0013] (3) using a heterogeneous graph to model a molecular graph, wherein the heterogeneous graph is a graph containing different types of nodes and edges, different atoms correspond to different node types, and different bonds correspond to different edge types)
Consider Claim 19.
The combination of Oh and Chen teaches:
19. The non-transitory computer-readable medium of claim 18, further comprising: obtaining information on a position of a center point of the first atomic region and a position of a center point the second atomic region based on the information on the first atomic region and the information on the second atomic region; and selecting the second atomic region located within a predetermined distance from the first atomic region. (Oh: [0004] In order to reproduce chemical structures in a document, a range of standard formats are used to efficiently store the chemical structure data. One type of format uses connection tables, adjacency matrices, or similar data structures to relate atoms and bonds as edges and nodes. Another type of format uses linear string notations based on depth first or breadth first traversal. The use of standardized data formats for storing chemical structure data enables algorithmic searching of the data. Furthermore, chemical structure data in standard formats can be indexed with a document in a database. [0088] Document data 100 has been augmented during indexing to comprise string tags 120 (as depicted by the dashed line connecting the two in FIG. 1). String tags are a sequence of characters that provide an alphanumeric text-based string for identifying, classifying, and/or describing chemical structures corresponding to chemical structure data in document data. In certain embodiments, string tags are generated using bit-screening data by performing an atom-by-atom or similar structure-based search on the bit-screening data therein to identify constituent elements corresponding to the bit-screening data and populating the string tags with strings in a predefined list or array that identify, classify, and/or describe the constituent elements. In certain embodiments, string tags are populated using an array that comprises the strings and corresponding reference bit screening data that is compared to the bit screening data in document data. The predefined list may be manually created by storing strings for common constituent elements in chemical structures and associations to reference bit-screening data that correspond to those common constituent elements. Thus, the reference bit-screening data associated with the pre-defined strings can be matched, using the structure-based search, to bit-screening data in document data in order to generate string tags that are populated with appropriate descriptive strings from the pre-defined list for constituent elements corresponding to the bit-screening data in the document data. String tags may also be generated using appropriate ad hoc structure-based methods that populate the string tags with appropriate descriptive strings. String tags may be associated with chemical structure data or directly with directly with the document data that comprises the chemical structure data. Referring again to FIG. 1, string tags 120 are associated with document data 100, but not directly associated with chemical structure data 110. Chen: [0046] Modeling the target molecule and its corresponding positive and negative samples by using a heterogeneous graph, which aims to characterize the different attributes of each node and edge. Inputting the sample data of the molecules into a structure-aware molecular encoder shown in FIG. 2 , and obtaining the feature representations corresponding to the target sample and the positive and negative samples. Denoting a feature representation corresponding to the target molecule as q, denoting a feature representation of the positive sample as k0, and denoting the feature representations of K negative samples as k1, . . . , kK. [0050]-[0051], [0050] , Figure 2)
Conclusion
The prior art made of record in form PTO-892 and not relied upon is considered pertinent to applicant's disclosure.
PNG
media_image3.png
175
905
media_image3.png
Greyscale
Any inquiry concerning this communication or earlier communications from the examiner should be directed to TAHMINA ANSARI whose telephone number is 571-270-3379. The examiner can normally be reached on IFP Flex - Monday through Friday 9 to 5.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, O’NEAL MISTRY can be reached on 313-446-4912. The fax phone numbers for the organization where this application or proceeding is assigned are 571-273-8300 for regular communications and 571-273-8300 for After Final communications. TC 2600’s customer service number is 571-272-2600.
Any inquiry of a general nature or relating to the status of this application or proceeding should be directed to the receptionist whose telephone number is 571-272-2600.
2674
/Tahmina Ansari/
August 18, 2026
/TAHMINA N ANSARI/Primary Examiner, Art Unit 2674