DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on February 19th, 2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Specification
The disclosure is objected to because the Brief Description of Drawings provides inconsistent descriptions of the figures:
At paragraph [0007], Figure 3 is described as “an illustration of a process flow for managing virtual spaces,” while paragraph [0073] describes the figure as “an illustration of a process flow for identifying positions of uncommon characters,” and Figure 3 does not appear to show management of virtual spaces.
At paragraph [0008], Figure 4 is described as “a flowchart of a process for creating virtual spaces,” while paragraph [0086] describes the figure as “an illustration of a process flow for identifying positions of uncommon characters,” and Figure 4 does not appear to show creation of virtual spaces.
Paragraph [0009] does not describe Figure 5.
Paragraph [0010] does not describe Figure 6.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claims recite a mental process that can be performed in the human mind or with the aid of pen and paper. This judicial exception is not integrated into a practical application because a computer is invoked merely as a tool to execute an abstract idea. The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because an abstract idea is merely applied on a generic computer without any element that would otherwise preclude performance of the abstrac.
Regarding claim 1, the claim recites “A computer implemented method for transforming data, the computer implemented method comprising:receiving, by a processor set, a number of data pairs, wherein each data pair in the number of data pairs comprises an input data and an output data that is semantically equivalent to the input data;creating, by the processor set, a program graph for each data pair in the number of data pairs, wherein nodes in each program graph represent positions for characters from each data pair;identifying, by the processor set, a number of paths between nodes in each program graph, wherein each path in the number of paths represents a sequence of characters in a data pair from the number of data pairs;identifying, by the processor set, a number of common paths from the number of paths based on common characters between the input data and the output data for each data pair;identifying, by the processor set, a set of nodes in the program graphs based on the number of common paths, wherein the set of nodes represent positions of unmatched characters between input data and output data in the number of data pairs; andgenerating, by the processor set, a prompt for a large language model based on the number of data pairs and the set of nodes that represent positions of unmatched characters between input data and output data in the number of data pairs.”
The limitations of “receiving… a number of data pairs,” “creating… a program graph for each data pair,” “identifying… a number of paths between nodes,” “identifying… a number of common paths from the number of paths,” “identifying… a set of nodes in the program graphs,” and “generating… a prompt for a large language model” as drafted cover mental activities which can be performed in the mind or with the aid of pen and paper. Taken individually, or as a whole, these limitations describe acts which are equivalent to human mental work of comparing sentences.
Each of the limitations preceding “generating… a prompt for a large language model” could additionally be interpreted as well-understood and routine methods in the art of computer science (see included reference Shifrin). String comparison and graph data structures are commonplace in sorting and matching algorithms that would be known to any person having ordinary skill in the art. Under this interpretation, the mental process of creating a prompt for a language model is not incorporated into a practical application by the inclusion of well-understood string processing techniques.
Each of the limitations preceding “generating… a prompt for a large language model” could additionally be interpreted as insignificant extra-solution activity, as mere pre-processing performed outside of the inventive method of semantic transformation. Under this interpretation, the mental process of creating a prompt for a large language model is not made patent eligible by the inclusion of sorting or filtering of input data.
The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the steps of the claimed invention can be performed mentally, and no additional features in the claims would preclude them from being performed as such. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claim 2, the claim depends from claim 1, and thus recites the limitations of claim 1, “further comprising:identifying, by the processor set, uncommon characters between input data and output data in the number of data pairs based on the set of nodes; andidentifying, by the processor set, semantic transformations between input data and output data in the number of data pairs based on the uncommon characters.”
Taken individually, or as a whole with claim 1, these limitations describe acts which are equivalent to human mental work of comparing sentences. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claim 3, the claim depends from claim 1, and thus recites the limitations of claim 1, “wherein the set of nodes are identified based on possible paths between a number of source nodes in the program graph and a number of target nodes in the program graph, wherein the number of source nodes represent positions for first characters in the output data from the number of data pairs and the number of target nodes represent positions for last characters in the output data from the number of data pairs.”
Taken individually, or as a whole with claim 1, these limitations describe acts which are equivalent to human mental work of comparing sentences. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claim 4, the claim depends from claim 1, and thus recites the limitations of claim 1, “wherein each path in the number of paths is identified by matching common characters between the input data and the output data in each data pair from the number of data pairs.”
Taken individually, or as a whole with claim 1, these limitations describe acts which are equivalent to human mental work of comparing sentences. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claim 5, the claim depends from claim 1, and thus recites the limitations of claim 1, “further comprising:inputting, by the processor set, a new input data to the large language model;performing, by the processor set using the large language model, semantic transformation to a number of characters in the new input data based on the prompt; andoutputting, by the processor set using the large language model, a new output data, wherein the new output data comprises semantically transformed characters.”
Taken individually, or as a whole with claim 1, these limitations describe acts which are equivalent to human mental work of comparing and rewriting sentences. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claim 6, the claim depends from claim 1, and thus recites the limitations of claim 1, “wherein each path from the number of paths is identified by a different program instruction.”
Taken individually, or as a whole with claim 1, these limitations describe acts which are equivalent to organizing human mental work of comparing sentences. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claim 7, the claim depends from claim 1, and thus recites the limitations of claim 1, “wherein the prompt comprises an index constructed based on the number of data pairs and position of matched characters between input data and output data in the number of data pairs.”
Taken individually, or as a whole with claim 1, these limitations describe acts which are equivalent to organizing human mental work of comparing sentences. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible.
Regarding claims 8-14, system claims 8-14 and method claims 1-7 are related as a method and system of using the same, with each system element’s function corresponding to the method step. Accordingly, claims 8-14 are similarly rejected under the same rationale as applied to claims 1-7.
Regarding claims 15-20, computer-readable medium claims 15-20 and method claims 1-6 are related as method and computer-readable medium for performing the same, with each computer-readable medium element’s function corresponding to the method step. Accordingly, claims 15-20 are similarly rejected under the same rationale as applied to claims 1-6.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 3-5, 7-8, 10-12, 14-15 and 17-20 are rejected under 35 U.S.C. 103 as being unpatentable over China invention application 11900740 to Zhao et al. (hereinafter, “Zhao”) in view of MIT Advanced Data Structures Lecture 10 by Demaine et al. (hereinafter, “Demaine”).
Regarding claims 1, 8 and 15, Zhao teaches a method, system and computer readable medium comprising: receiving, by a processor set, a number of data pairs, wherein each data pair in the number of data pairs comprises an input data and an output data that is semantically equivalent to the input data (page 3, content of the invention, "S1: obtaining source data and target data, said source data and target data are respectively from patient information and medicine data;");
identifying, by the processor set, a number of common paths from the number of paths based on common characters between the input data and the output data for each data pair (page 3, content of the invention, "S3: using the active learning method, selecting a subset of the marked source data as the candidate source data, primarily filtering the source data for the large language model learning; Specifically, the method includes: according to the uncertainty of each record in the source data and the relevance with the target data, selecting the record from the source data as the candidate;");
identifying, by the processor set, a set of nodes in the program graphs based on the number of common paths, wherein the set of nodes represent positions of unmatched characters between input data and output data in the number of data pairs (page 3, content of the invention, "S4: changing the demonstration selected range from the source data to the candidate source data; combining the similarity of structure and semanteme to select more valuable demonstration;"); and
generating, by the processor set, a prompt for a large language model based on the number of data pairs and the set of nodes that represent positions of unmatched characters between input data and output data in the number of data pairs (page 3, content of the invention, "S5: injecting the domain information of each entity pair into the pre-defined format, sending the prompt to the large language model for processing, the large language model returns the result of the special entity pair according to the received prompt;").
Zhao does not disclose the particular structure of the subject entity paris, and thus Demaine is introduced. Demaine teaches data structures for text processing including creating, by the processor set, a program graph for each data pair in the number of data pairs, wherein nodes in each program graph represent positions for characters from each data pair (page 3, section 4.1 Suffix Trees - Description, "Description 1. A trie is a tree in which each node has children labeled by letter in [an alphabet] Σ." See also Figure 1.);
identifying, by the processor set, a number of paths between nodes in each program graph, wherein each path in the number of paths represents a sequence of characters in a data pair from the number of data pairs (page 5, section 4.2 Suffix Trees - Solving String Matching, "The idea is that we start at the root of S and index through P, using the current letter of P to decide which branch to take from the current node of S. When edges in S are labeled with multiple letters, we index through the edge and through P, matching the edge labels to the current letter of P. We continue in this fashion, taking branches whenever we have indexed to the end of the current edge label.").
Zhao and Demaine are considered analogous because they are each concerned with data comparison. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have replaced the entity pairs of Zhao with the suffix trees of Demaine, given that the substitution of one known element for another yields predictable results.
Regarding claims 3, 10 and 17, Demaine further teaches the computer-implemented method of claims 1, 8 and 15 wherein the set of nodes are identified based on possible paths between a number of source nodes in the program graph and a number of target nodes in the program graph, wherein the number of source nodes represent positions for first characters in the output data from the number of data pairs and the number of target nodes represent positions for last characters in the output data from the number of data pairs (page 6, "We note that the algorithm returns a subtree (since it returns a pointer to a node), and that the nodes of this returned subtree correspond to all of the occurrences of P in T. To see this, imagine that P occurs at indices i1, i2, …ik.").
Regarding claims 4, 11 and 18, Demaine further teaches the computer-implemented method of claims 1, 8 and 15 wherein each path in the number of paths is identified by matching common characters between the input data and the output data in each data pair from the number of data pairs (page 6, "We note that the algorithm returns a subtree (since it returns a pointer to a node), and that the nodes of this returned subtree correspond to all of the occurrences of P in T. To see this, imagine that P occurs at indices i1, i2, …ik.").
Regarding claims 5, 12 and 19, Zhao further teaches inputting, by the processor set, a new input data to the large language model (page 6, "In the active candidate source data generation module, the present application uses the concept of active learning to select a subset of the tagged source data as the candidate source data, which subset will be used in the subsequent context presentation selection module.");
performing, by the processor set using the large language model, semantic transformation to a number of characters in the new input data based on the prompt (page 6, "In the demonstration selection module in the context, the application changes the demonstration selection range from the source data to the candidate source data. The application redefines the structural similarity between entity pairs to cope with the challenges brought by the heterogeneity of data structures in different fields."); and
outputting, by the processor set using the large language model, a new output data, wherein the new output data comprises semantically transformed characters (page 6, "The prompt is sent to the large language model for processing, the large language model returns the result of the specific entity pair according to the received prompt.").
Regarding claims 7 and 14, Zhao and Demaine may be combined to teach a computer-implemented method wherein the prompt comprises an index constructed based on the number of data pairs and position of matched characters between input data and output data in the number of data pairs (Zhao page 3, content of the invention, "Finally, the application uses the selected source entity pair and its tag information, target instance and cross-domain information to form a prompt of a large language model, and then obtains a prediction result," and Demaine page 6, "We note that the algorithm returns a subtree (since it returns a pointer to a node), and that the nodes of this returned subtree correspond to all of the occurrences of P in T. To see this, imagine that P occurs at indices i1, i2, …ik.").
Zhao and Demaine are considered analogous because they are each concerned with data comparison. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have included the subtree and indices taught by Demaine in the prompting as taught by Zhao for the purpose of improving language model accuracy.
Claims 2, 6, 9, 13, 16 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Zhao and Demaine as applied to claims 1, 8 and 15 above, and further in view of "A Software Package for the Manipulation and Analysis of Character Strings" by Gerald Shifrin (hereinafter, "Shifrin").
Regarding claims 2, 9 and 16, Zhao teaches a computer-implemented method including identifying, by the processor set, semantic transformations between input data and output data in the number of data pairs based on the uncommon characters (page 6, "The candidate source data generation can be viewed as a preliminary filtering of the source data, in order to select records with higher values for large language model learning. In order to achieve this object, the present application comprehensively takes into account the inherent uncertainty of each record in the source data and its association with the target data, thereby selecting useful records from the source data as candidates").
The combination of Zhao and Demaine considers but does not explicitly teach “identifying, by the processor set, uncommon characters between input data and output data in the number of data pairs based on the set of nodes,” and thus, Shifrin is introduced. Shifrin teaches identifying, by the processor set, uncommon characters between input data and output data in the number of data pairs based on the set of nodes (file page 35, document page 32, BREAK, "Search a character string or string segment for a character different from one or more of a specified set of characters, and indicate the string position in which it was found, if any.").
Zhao, Demaine and Shifrin are considered analogous because they are each concerned with data comparison. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified the combination of Zhao and Demaine with the teachings of Shifrin for the purpose of analyzing dissimilar strings. Given that all the claimed elements were known in the prior art, one skilled in the art could have combined the elements by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Regarding claims 6, 13 and 20, the combination of Zhao and Demaine does not teach “each path from the number of paths is identified by a different program instruction,” however, Shifrin teaches a computer-implemented method wherein each path from the number of paths is identified by a different program instruction (file pages 89-90, document pages 86-87, SPAN, "To search a string or string segment for the first appearance of any one of a set of characters -nptr=SPAN(string [ , sptr [ , slen] ] , char [ , clen] )… where
string (input; character) is the CHARACTER type string to be searched.
char (input; alphanumeric or character) is either an alphanumeric variable, array, or string in quotes; or a CHARACTER type string. char contains one or more characters to search string for.").
Zhao, Demaine and Shifrin are considered analogous because they are each concerned with data comparison. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified the combination of Zhao and Demaine with the teachings of Shifrin for the purpose of analyzing dissimilar strings. Given that all the claimed elements were known in the prior art, one skilled in the art could have combined the elements by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
U.S. Patent 11,928,126 to Guttula et al.
U.S. Patent 7,945,525 to Ananthanarayanan et al.
China invention application 119149786 to Qiao.
WIPO publication 2013/106989 to Ling et al.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SEAN T SMITH whose telephone number is (571)272-6643. The examiner can normally be reached Monday - Friday 8:00am - 5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, PIERRE-LOUIS DESIR can be reached at (571) 272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SEAN THOMAS SMITH/Examiner, Art Unit 2659
/PIERRE LOUIS DESIR/Supervisory Patent Examiner, Art Unit 2659