Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
Status of Claims
The present Office Action is pursuant to Applicant’s communication on 08-21-2024; current application filed on 08-21-2024; Continuation-in-part of application No. 18/417,695, filed on Jan. 19, 2024, which is a continuation-in-part of application No. 17/732,322, filed on Apr. 28, 2022; Provisional application No. 63/180,919, filed on Apr. 28, 2021.
Examiner’s Note
The rejections below group claims that may not be identical, but whose language and scope are so substantively similar as to lend themselves to grouping, in the interests of clarity and conciseness.
Information Disclosure Statement
The information disclosure statements (IDS) filed on 02-27-2025, 04-23-2025, 08-01-2025, 10-28-2025, 01-27-2026, 04-22-2026, have been acknowledged. The submissions are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception without significantly more.
Claim 1 recites:
A machine learning system for automatically extracting information from medical records, comprising: a memory storing a plurality of medical records; and a processor in communication with the memory, the processor programmed to perform the steps of: retrieving the plurality of medical records from the memory; retrieving at least one document having pages of text from the plurality of medical records; processing the pages of text to clean data in the pages of text; and processing the pages of text to extract medical service data from the text using at least one of a regular expression algorithm or a trained machine learning model.
Analysis of Independent Claims (Claims 1 and 11)
Step 1: Statutory Category Claims 1 and 11 are directed to a "machine learning system" and a "machine learning method," respectively, which fall within the statutory categories of machine and process under 35 U.S.C. § 101.
Step 2A Prong One: Abstract Idea The claims recite an abstract idea falling under the judicial exception of organizing human activity and mental processes. The core limitations involve retrieving medical records, cleaning text data, and extracting service/date information using regex or machine learning models. These steps essentially describe data gathering, processing, and organization that can be performed mentally or with pen and paper.
Specification discloses the system's data handling and extraction capabilities.
[00117]: "The data cleaning process 1084 begins in step 1086, wherein the system removes any e-mail addresses or links that exist in the pages of text..."
[00119]: "date of service extraction process 1098 occurs... processes the text pages to identify a date of medical service using a suitable pattern matching algorithm, such as a regular expression ("regex")..." [00120]: "a second date of service extraction process 1112 occurs, which extracts a date of medical service from the text pages using a pre-trained machine learning model."
Step 2A Prong Two: Integration into a Practical Application The additional elements (e.g., "a memory storing a plurality of medical records," "a processor in communication with the memory") recite generic computer components performing their basic, routine functions. Merely implementing the abstract idea on a generic computer or using conventional machine learning/regex algorithms does not integrate the abstract idea into a practical application. The claims do not improve the functioning of the computer itself, nor do they apply the abstract idea in a particular technological field beyond general data processing.
Step 2B: Inventive Concept (Significantly More) The claims lack an inventive concept. Using a generic processor to retrieve documents, clean text by removing specific characters/words, and extract dates via conventional regex or pre-trained ML models amounts to no more than routine, conventional computer activity. There is no unconventional arrangement of elements or technical improvement that would render the claims patent-eligible.
Analysis of Dependent Claims (Claims 2-6 and 12-16)
These claims further limit the data cleaning steps by specifying the removal of emails, non-English words, punctuation, stop words, small-length words, and extra spaces/lowercase "the".
Step 1: Directed to statutory categories (system/method). Step 2A Prong One: The additional limitations recite further mental processes/organizing human activity. Filtering text by removing specific character types or word lengths is a conventional data-cleaning technique that can be performed manually. Specification discloses these exact cleaning steps.
[00117]: "Next, in step 1088, the system removes all non-English words and punctuation marks from the text pages. Then, in step 1090, the system removes any stop words from the pages of text. In step 1092, the system removes any small-length words from the text pages (e.g., words having 2 or fewer letters). Finally, in step 1094, the system removes extra spaces and lower-case "the" letters from the text pages."
Step 2A Prong Two: The limitations are implemented using generic computer processing. They do not integrate the abstract idea into a practical application because they merely automate routine text-filtering rules.
Step 2B: No inventive concept is added. Automating standard text-cleaning heuristics (removing punctuation, stop words, or short strings) is well-understood, routine, and conventional in data processing.
Analysis of Dependent Claims (Claims 7-8 and 17-18)
These claims limit the extraction step to searching for keywords around dates or using a trained ML model to extract all dates on a page.
Step 1: Directed to statutory categories (system/method).
Step 2A Prong One: The limitations recite mental processes/organizing human activity. Searching surrounding text for keywords to locate dates, or using a pre-trained model to identify dates, is essentially information retrieval and pattern recognition.
Specification discloses the keyword search and ML extraction methods.
[00119]: "Specifically, in step 1100, the system searches surrounding words in the text pages using a few key words (which could be pre-programmed into the system). Then, in step 1102, if the system identifies a date in the surrounding words, the date is extracted by the system."
[00120]: "In parallel with the date of service extraction process 1098, a second date of service extraction process 1112 occurs, which extracts a date of medical service from the text pages using a pre-trained machine learning model."
Step 2A Prong Two: The claims merely apply conventional regex and ML techniques to generic document text. They do not solve a technology-specific problem or improve computer functionality; they simply automate manual date-finding tasks.
Step 2B: No inventive concept is present. Using regex for pattern matching and pre-trained models for entity extraction are routine, conventional tools in natural language processing and data mining.
Analysis of Dependent Claims (Claims 9-10 and 19-20)
These claims limit the system/method to classifying pages as "start," "end," or "other," and bundling pages with the same date.
Step 1: Directed to statutory categories (system/method).
Step 2A Prong One: The limitations recite organizing human activity. Classifying document pages into categories and grouping/bundling them based on shared attributes (like dates) is a conventional data organization technique.
Specification discloses the classification and bundling steps.
[00121]: "Start and end page classification process 1104 processes the text pages to identify the starting and ending pages for a particular medical event. Specifically, in step 1106, the system labels the data using three classes, namely, "start page," "end page," and "other.""
[00122]: "Specifically, in step 1120, the system uses the start and end labels identified by the start and end page classifier model in process 1104 to bundle the text pages... Then, in step 1124, the system assigns all the pages in the same bundle the same date with a maximum count."
Step 2A Prong Two: The classification and bundling steps are performed by generic computer components. They do not integrate the abstract idea into a practical application because they merely automate manual document sorting and grouping procedures without any technical improvement to the underlying computing system.
Step 2B: No inventive concept is added. Categorizing pages and bundling them based on extracted metadata (dates) is a routine, conventional data management practice that does not amount to significantly more than the abstract idea itself.
Conclusion
Claims 1-20 are directed to an abstract idea of organizing human activity and mental processes (data cleaning, date extraction, page classification, and bundling). The additional elements recite generic computer components performing routine, conventional functions. Under the Alice/Mayo framework, the claims fail at Step 2A Prong Two (no practical application) and Step 2B (no inventive concept), rendering them ineligible under 35 U.S.C. § 101.
Thus, taken alone, the additional elements do not amount to significantly more than the abstract idea identified above. Furthermore, looking at the limitations as an ordered combination adds nothing that is not already present when looking at the elements taken individually, and there is no indication that the combination of elements improves the functioning of a computer or improves any other technology, and their collective functions merely provide conventional computer implementation.
Therefore, whether taken individually or as an ordered combination, claim(s) 1-20 is/are nonetheless rejected under 35 U.S.C. § 101 as being directed to non-statutory subject matter.
Double Patenting
The non-statutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A non-statutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on non-statutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP §§ 706.02(l)(1) - 706.02(l)(3) for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/process/file/efs/guidance/eTD-info-I.jsp.
Claims 1-20 of continuation application 18/811,440 are rejected on the ground of non-statutory double patenting as being unpatentable over claims 1-14 of parent patent US 12,665,087 B2. Although the claims at issue are not identical, they are not patentably distinct from each other because the pending continuation-in-part claims encompass the same core data processing and extraction methodologies taught in the parent patent.
I. Analysis of Independent Claims
Continuation Claim 1 vs. Parent Claim 1 Continuation Claim 1 recites a machine learning system comprising: "processing the pages of text to clean data in the pages of text; and processing the pages of text to extract medical service data from the text using at least one of a regular expression algorithm or a trained machine learning model." Parent Claim 1 discloses substantially overlapping limitations, specifically: "processing the pages of text to clean data in the pages of text” and "processing the pages of text to extract a date of service from the text using a pattern matching algorithm executed by the processor, the pattern matching algorithm searching for a plurality of surrounding words in the pages of text using at least one key word and identifying and extracting a date from the plurality of surrounding words”. The pending claim's broader recitation of extracting "medical service data" via regex or ML is encompassed by the parent patent's teachings on cleaning text and extracting service dates via algorithmic pattern matching. Furthermore, Parent Claim 1 includes additional classifier and bundling steps, making the pending claim a broader variation of the same invention.
Continuation Claim 11 vs. Parent Claim 8 Continuation Claim 11 recites a machine learning method comprising: "processing the pages of text to clean data in the pages of text; and processing the pages of text to extract medical service data from the text using at least one of a regular expression algorithm or a trained machine learning model." Parent Claim 8 discloses identical foundational steps: "processing the pages of text to clean data in the pages of text" and "processing the pages of text to extract a date of service from the text using a pattern matching algorithm executed by the processor, the pattern matching algorithm searching for a plurality of surrounding words in the pages of text using at least one key word and identifying and extracting a date from the plurality of surrounding words”. The method claim in the continuation is not patentably distinct from the parent's method claim, as both rely on the same sequence of retrieving records, cleaning text, and extracting service data/dates using algorithmic or machine learning techniques.
II. Analysis of Dependent Claims
Continuation Claims 2-6 vs. Parent Claims 2-6 (System) The dependent claims in the continuation application recite identical data-cleaning limitations to those in the parent patent:
Continuation Claim 2: "removing e-mail addresses and links from the pages of text." matches Parent Claim 2: "removing e-mail addresses and links from the pages of text."
Continuation Claim 3: "removing non-English words and punctuation from the pages of text." matches Parent Claim 3: "removing non-English words and punctuation from the pages of text."
Continuation Claim 4: "removing stop words from the pages of text." matches Parent Claim 4: "removing stop words from the pages of text."
Continuation Claim 5: "removing small-length words from the pages of text." matches Parent Claim 5: "removing small-length words from the pages of text."
Continuation Claim 6: "removing extra spaces and lower-case 'the' letters from the pages of text." matches Parent Claim 6: "removing extra spaces and lower-case 'the' letters from the pages of text."
Continuation Claims 7-8 vs. Parent Claims 7 & 14 (Extraction Variations)
Continuation Claim 7 recites: "searching for key words within surrounding words to find a date in the surrounding words and extracting the date." This is directly disclosed in Parent Claim 1/8 as: "the pattern matching algorithm searching for a plurality of surrounding words in the pages of text using at least one key word and identifying and extracting a date from the plurality of surrounding words”.
Continuation Claim 8 recites: "extracting all date in the page using the trained machine learning model." This overlaps with Parent Claim 7/14 which discloses: "extracting all dates in the page." Additionally, the parent specification confirms ML-based extraction is within the scope of the invention: "a second date of service extraction process 1112 occurs, which extracts a date of medical service from the text pages using a pre-trained machine learning model”.
Continuation Claims 9-10 vs. Parent Claim 1/8 (Classifier & Bundling)
Continuation Claim 9 recites: "identifying each page as one of a start page, and end page, or another page." This is explicitly taught in the parent patent: "processing the pages of text using a classifier model to identify a type of page for each page of text, the classifier model identifying and labeling each page with a label comprising one of a start page label, and end page label, and an other page label”.
Continuation Claim 10 recites: "bundling a group of the pages of text and assigning the same date to each page of the group." This is directly disclosed in the parent patent: "bundling the pages of text into a bundle using the labels such that bundled pages are considered to correspond to a visit by a patient to a medical provider" and "assigning a date having a maximum count to each page of the bundle.”.
Continuation Claims 12-16 vs. Parent Claims 9-13 (Method) The method-dependent claims in the continuation mirror the system-dependent claims and are identically rejected over their parent counterparts:
Continuation Claim 12 matches Parent Claim 9: "removing e-mail addresses and links from the pages of text."
Continuation Claim 13 matches Parent Claim 10: "removing non-English words and punctuation from the pages of text."
Continuation Claim 14 matches Parent Claim 11: "removing stop words from the pages of text."
Continuation Claim 15 matches Parent Claim 12: "removing small-length words from the pages of text."
Continuation Claim 16 matches Parent Claim 13: "removing extra spaces and lower-case 'the' letters from the pages of text."
Continuation Claims 17-20 vs. Parent Claims 8, 14 & Specification (Method Extraction/Classification)
Continuation Claim 17 mirrors Claim 7 and is rejected over Parent Claim 8's pattern matching limitation: "searching for a plurality of surrounding words in the pages of text using at least one key word and identifying and extracting a date”.
Continuation Claim 18 mirrors Claim 8 and is rejected over Parent Claim 14: "extracting all dates in the page."
Continuation Claim 19 mirrors Claim 9 and is rejected over Parent Claim 8's classifier limitation: "identifying and labeling each page with a label comprising one of a start page label, and end page label, and an other page label”.
Continuation Claim 20 mirrors Claim 10 and is rejected over Parent Claim 8's bundling limitation: "bundling the pages of text into a bundle using the labels" and "assigning a date having a maximum count to each page of the bundle.”.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless -
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1, 7, 8, 9, 10, 11, 17, 18, 19, and 20 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Lucas (US 2020/0176098).
Claim 1 - Complete Limitation Listing and Mapping
1[a]: "A machine learning system for automatically extracting information from medical records, comprising:"
1[b]: "a memory storing a plurality of medical records;"
1[c]: "a processor in communication with the memory, the processor programmed to perform the steps of:"
1[d]: "retrieving the plurality of medical records from the memory;"
1[e]: "retrieving at least one document having pages of text from the plurality of medical records;"
1[f]: "processing the pages of text to clean data in the pages of text; and"
1[g]: "processing the pages of text to extract medical service data from the text using at least one of a regular expression algorithm or a trained machine learning model."
Limitation-by-Limitation Mapping
1[a] - "A machine learning system for automatically extracting information from medical records":
Lucas discloses a system that "identifies information in clinical documents or other records" using "a combination of text extraction techniques, text cleaning techniques, natural language processing techniques, machine learning algorithms, and medical concept (Entity) identification, normalization, and structuring techniques."
[0024]: "In one aspect, a system is disclosed that identifies information in clinical documents or other records. The system may use a combination of text extraction techniques, text cleaning techniques, natural language processing techniques, machine learning algorithms, and medical concept (Entity) identification, normalization, and structuring techniques."
1[b] - "a memory storing a plurality of medical records":
Lucas discloses a data storage device and a Clinical Data Vault that stores patient documents.
[0264]: "The example computer system 900 includes a processing device 902, a main memory 904… a static memory 906… and a data storage device 918, which communicate with each other via a bus 930." [0247]: "the system receives documents, for example, from a Clinical Data Vault 815, new documents from a Document Pipeline 805, or corrected documents via the Workbench 810."
1[c] - "a processor in communication with the memory, the processor programmed to perform the steps of":
Lucas discloses a processing device in communication with memory via a bus, programmed to execute instructions. [0264]: "The example computer system 900 includes a processing device 902, a main memory 904… and a data storage device 918, which communicate with each other via a bus 930." [0265]: "[P]rocessing device 902 is configured to execute instructions 922 for performing the operations and steps discussed herein."
1[d] - "retrieving the plurality of medical records from the memory":
Lucas discloses retrieving patient documents from the Clinical Data Vault.
[0247]: "the system receives documents, for example, from a Clinical Data Vault 815, new documents from a Document Pipeline 805, or corrected documents via the Workbench 810 (introduced below), uploads documents, and posts them to a server that coordinates a number of tasks and manages the intake of documents for the intake pipeline described in FIG. 1."
[0260]: "When a Workbench user loads a patient, Workbench 810 may pull the patient record from the Document Pipeline 805."
1[e] - "retrieving at least one document having pages of text from the plurality of medical records":
Lucas discloses receiving a clinical document that may include machine-readable text or an image file, with pages. [0097]: "The intake pipeline 110 receives a clinical document that may include machine readable text or that may be received as an image file."
[0216]: "a document named Progress Note 01_01_01 may be presumed to have a date of Jan. 1, 2001. Other concept candidates from the document may be referenced to validate the date/time… a page number may be identified, for example, in a document that has 5 pages by referencing the page number by performing an OCR of text at the bottom of the page."
1[f] - "processing the pages of text to clean data in the pages of text":
Lucas discloses a pre-processing stage that performs text cleaning and error detection.
[0097]: "the document may be submitted to a pre-processor stage 120 that performs text cleaning and error detection (i.e., format conversion, resolution conversion, batch sizing, text cleaning, etc.)."
[0100]: "it may be necessary to further improve upon the quality of the OCR output by performing a text cleaning algorithm on the OCR output. Text cleaning may be implemented by a category of NLP models designed for Language Modeling."
1[g] - "processing the pages of text to extract medical service data from the text using at least one of a regular expression algorithm or a trained machine learning model":
Lucas discloses both regular expression-based extraction and trained machine learning model-based extraction.
[0179]: "Regular Expressions: The system may identify anchor strings regularly occurring in text that identify where key health information may reside. For example, the system may recognize that 'DOB' is a string to search for dates of birth and 'Pathological Diagnosis' may be a header to a section that provides concepts for linking to a pathological diagnosis."
[0024]: "The system may use a combination of text extraction techniques, text cleaning techniques, natural language processing techniques, machine learning algorithms, and medical concept (Entity) identification, normalization, and structuring techniques."
[0095]: "pipeline stage 120 for pre-processing may include OCR and text cleaning, stage 130 for parsing may include NLP algorithms for sentence splitting and candidate extraction, stage 140 for dictionary lookups may include entity linking, stage 150 for normalization may include entity normalization, stage 160 for structuring may include entity structuring."
Claim 11 - Complete Limitation Listing and Mapping
11[a]: "A machine learning method for automatically extracting information from medical records, comprising:"
11[b]: "retrieving the plurality of medical records from the memory;"
11[c]: "retrieving at least one document having pages of text from the plurality of medical records;"
11[d]: "processing the pages of text to clean data in the pages of text; and"
11[e]: "processing the pages of text to extract medical service data from the text using at least one of a regular expression algorithm or a trained machine learning model."
11[a]-[e] are mapped identically to 1[a]-[g] above. Lucas discloses the method in
[0007]: "a method includes the steps of determining a first concept from a text of a medical record from an electronic health record system, the first concept relating to a patient, identifying a match to the first concept in a first list of concepts… referencing the first concept with an entity in a database of related concepts, identifying a match to a second concept in a second list of concepts… and providing the second concept as an identifier of the patient's medical record." The method steps of retrieving, cleaning, and extracting are disclosed in [0095], [0097], [0100], and [0179] as cited above.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 2, 3, 4, 5, 6, 12, 13, 14, 15, and 16 are rejected under 35 U.S.C. § 103 as being unpatentable over Lucas (US 2020/0176098) in view of Munzert1.
Claim 2 - Complete Limitation Listing and Mapping
2[a]: "The system of Claim 1, wherein the step of processing the pages of text to clean the data in the pages of text comprises removing e-mail addresses and links from the pages of text."
2[a] - "removing e-mail addresses and links from the pages of text":
Lucas discloses text cleaning as a pre-processing step but does not explicitly disclose removing e-mail addresses and links.
[0100]: "it may be necessary to further improve upon the quality of the OCR output by performing a text cleaning algorithm on the OCR output. Text cleaning may be implemented by a category of NLP models designed for Language Modeling."
Munzert discloses data cleansing operations that include removing specific character patterns from text. [Section 10.2.3.1]: "In order to take care of some of these errors, one typically runs several data preparation operations. Furthermore, the data preparation addresses some of the concerns that are leveled against (semi-)automated text classification… one might consider removing numbers and period characters from the texts without losing much information. This can either be done on the raw textual data or while setting up the term-document matrix."
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the text cleaning step of Lucas to include the removal of e-mail addresses and hyperlinks, as these are standard non-clinical artifacts in medical documents that would interfere with the extraction of medical service data. Munzert teaches that data preparation operations remove extraneous characters and patterns from text to improve downstream processing. [Section 10.2.3.1]: "one might consider removing numbers and period characters from the texts without losing much information." The motivation to remove e-mail addresses and links is analogous to removing non-informative characters, as both are extraneous to the medical content being extracted.
Claim 12 - Complete Limitation Listing and Mapping
12[a]: "The method of Claim 11, wherein the step of processing the pages of text to clean the data in the pages of text comprises removing e-mail addresses and links from the pages of text."
12[a] is mapped identically to 2[a] above. The same Lucas and Munzert combination applies.
Claim 3 - Complete Limitation Listing and Mapping
3[a]: "The system of Claim 2, wherein the step of processing the pages of text to clean the data in the pages of text comprises removing non-English words and punctuation from the pages of text."
3[a] - "removing non-English words and punctuation from the pages of text":
Lucas discloses text cleaning as part of the pre-processing pipeline.
[0100]: "it may be necessary to further improve upon the quality of the OCR output by performing a text cleaning algorithm on the OCR output. Text cleaning may be implemented by a category of NLP models designed for Language Modeling."
Munzert explicitly discloses removing punctuation characters as a data cleansing operation.
[Section 10.2.3.1]: "one might consider removing numbers and period characters from the texts without losing much information." The heading of the section is "Removing punctuation characters." Specifically, Munzert explicitly discloses “Dropping non-english texts” as part of “data preparation [Section 17.3.1].”
It would have been obvious to one of ordinary skill in the art to combine Lucas's text cleaning pipeline with Munzert's teaching of removing punctuation characters, as both are directed to improving the quality of text data prior to extraction of meaningful information. The removal of non-English words is a standard NLP text cleaning/data preparation operation that would be obvious in view of Lucas's disclosure of using "a category of NLP models designed for Language Modeling" [0100] to clean text, combined with Munzert's teaching of removing non-informative characters.
Claim 13 - Complete Limitation Listing and Mapping
13[a]: "The method of Claim 12, wherein the step of processing the pages of text to clean the data in the pages of text comprises removing non-English words and punctuation from the pages of text."
13[a] is mapped identically to 3[a] above. The same Lucas and Munzert combination applies.
Claim 4 - Complete Limitation Listing and Mapping
4[a]: "The system of Claim 3, wherein the step of processing the pages of text to clean the data in the pages of text comprises removing stop words from the pages of text."
4[a] - "removing stop words from the pages of text":
Lucas discloses text cleaning using NLP language models. [0100]: "Text cleaning may be implemented by a category of NLP models designed for Language Modeling. Additionally, machine learning algorithms and deep learning algorithms may be utilized to further improve upon the OCR results."
Munzert discloses data preparation operations that remove non-informative elements from text. [Section 10.2.3.1]: "In order to take care of some of these errors, one typically runs several data preparation operations… one might consider removing numbers and period characters from the texts without losing much information."
It would have been obvious to one of ordinary skill in the art to include stop-word removal as part of the text cleaning step, as stop-word removal is a well-known and standard NLP text preprocessing technique. Lucas's disclosure of using "NLP models designed for Language Modeling" [0100] for text cleaning inherently encompasses standard NLP preprocessing operations including stop-word removal, and Munzert's teaching of running "several data preparation operations" [Section 10.2.3.1] supports the combination.
Claim 14 - Complete Limitation Listing and Mapping
14[a]: "The method of Claim 13, wherein the step of processing the pages of text to clean the data in the pages of text comprises removing stop words from the pages of text."
14[a] is mapped identically to 4[a] above. The same Lucas and Munzert combination applies.
Claim 5 - Complete Limitation Listing and Mapping
5[a]: "The system of Claim 4, wherein the step of processing the pages of text to clean the data in the pages of text comprises removing small-length words from the pages of text."
5[a] - "removing small-length words from the pages of text":
Lucas discloses text cleaning as a pre-processing step.
[0100]: "it may be necessary to further improve upon the quality of the OCR output by performing a text cleaning algorithm on the OCR output."
Munzert discloses removing specific character patterns and short elements from text.
[Section 10.2.3.1]: "one might consider removing numbers and period characters from the texts without losing much information."
It would have been obvious to one of ordinary skill in the art to remove small-length words (e.g., words of one or two characters) as part of text cleaning, as this is a standard text preprocessing operation. The combination of Lucas's text cleaning pipeline [0100] with Munzert's teaching of removing non-informative elements [Section 10.2.3.1] makes the removal of small-length words obvious.
Claim 15 - Complete Limitation Listing and Mapping
15[a]: "The method of Claim 14, wherein the step of processing the pages of text to clean the data in the pages of text comprises removing small-length words from the pages of text."
15[a] is mapped identically to 5[a] above. The same Lucas and Munzert combination applies.
Claim 6 - Complete Limitation Listing and Mapping
6[a]: "The system of Claim 5, wherein the step of processing the pages of text to clean the data in the pages of text comprises removing extra spaces and lower-case 'the' letters from the pages of text."
6[a] - "removing extra spaces and lower-case 'the' letters from the pages of text":
Lucas discloses text cleaning including error correction. [0097]: "the document may be submitted to a pre-processor stage 120 that performs text cleaning and error detection (i.e., format conversion, resolution conversion, batch sizing, text cleaning, etc.)." [0100]: "it may be necessary to further improve upon the quality of the OCR output by performing a text cleaning algorithm on the OCR output."
Munzert discloses removing specific characters and patterns from text. [Section 10.2.3.1]: "one might consider removing numbers and period characters from the texts without losing much information. This can either be done on the raw textual data or while setting up the term-document matrix."
It would have been obvious to one of ordinary skill in the art to remove extra spaces and the lower-case word "the" (a common stop word) as part of text cleaning. Lucas's text cleaning and error detection [0097], combined with Munzert's teaching of removing specific characters and patterns [Section 10.2.3.1], makes this combination obvious.
Claim 16 - Complete Limitation Listing and Mapping
16[a]: "The method of Claim 15, wherein the step of processing the pages of text to clean the data in the pages of text comprises removing extra spaces and lower-case 'the' letters from the pages of text."
16[a] is mapped identically to 6[a] above. The same Lucas and Munzert combination applies.
Claim 7 - Complete Limitation Listing and Mapping
7[a]: "The system of Claim 1, wherein the step of processing the pages of text to extract the medical data from the text comprises searching for key words within surrounding words to find a date in the surrounding words and extracting the date."
7[a] - "searching for key words within surrounding words to find a date in the surrounding words and extracting the date":
Lucas explicitly discloses using regular expressions to search for key words (anchor strings) to find dates in surrounding text. [0179]: "Regular Expressions: The system may identify anchor strings regularly occurring in text that identify where key health information may reside. For example, the system may recognize that 'DOB' is a string to search for dates of birth and 'Pathological Diagnosis' may be a header to a section that provides concepts for linking to a pathological diagnosis."
Lucas further discloses extracting dates from surrounding context. [0216]: "a document named Progress Note 01_01_01 may be presumed to have a date of Jan. 1, 2001. Other concept candidates from the document may be referenced to validate the date/time or select the date absent any other validating/corroborating information. For example, the time 11:35 am may have been provided as a concept candidate spatially near the 'Tylenol 50 mg' concept candidate. The Relational Extraction MLA may then identify 11:35 am as the time the medication was administered based on the concept candidate time being the next concept candidate in the list, a spatial proximity of the concept candidate, a new application of NLP to the OCRed text string, or any combination thereof."
Claim 17 - Complete Limitation Listing and Mapping
17[a]: "The method of Claim 11, wherein the step of processing the pages of text to extract the medical service data from the text comprises searching for key words within surrounding words to find a date in the surrounding words and extracting the date."
17[a] is mapped identically to 7[a] above. Lucas discloses the same method in [0179] and [0216] as cited above.
Claim 8 - Complete Limitation Listing and Mapping
8[a]: "The system of Claim 1, wherein the step of processing the pages of text to extract the medical data the text comprises extracting all date in the page using the trained machine learning model."
8[a] - "extracting all date in the page using the trained machine learning model":
Lucas discloses using a trained machine learning model to extract dates and other structured data from medical record pages. [0024]: "The system may use a combination of text extraction techniques, text cleaning techniques, natural language processing techniques, machine learning algorithms, and medical concept (Entity) identification, normalization, and structuring techniques." [0095]: "pipeline stage 120 for pre-processing may include OCR and text cleaning, stage 130 for parsing may include NLP algorithms for sentence splitting and candidate extraction, stage 140 for dictionary lookups may include entity linking, stage 150 for normalization may include entity normalization, stage 160 for structuring may include entity structuring, and stage 170 for post-processing may include structuring the data and formatting it into a universal EMR or institution based EMR format."
Lucas further discloses that the trained model extracts dates from the text. [0216]: "a document named Progress Note 01_01_01 may be presumed to have a date of Jan. 1, 2001. Other concept candidates from the document may be referenced to validate the date/time or select the date absent any other validating/corroborating information." [0069]: "Dosage & Dosage Units: The dosage (i.e., 50 mg) associated with the medication mentioned… normalizing the dosage and dosage units by separating value 50 into the dosage field and string 'mg' or by selecting a known value entry for the milligram units within a list may be preferable."
Claim 18 - Complete Limitation Listing and Mapping
18[a]: "The method of Claim 11, wherein the step of processing the pages of text to extract the medical service data from the text comprises extracting all date in the page using the trained machine learning model."
18[a] is mapped identically to 8[a] above. Lucas discloses the same method in [0024], [0095], and [0216] as cited above.
Claim 9 - Complete Limitation Listing and Mapping
9[a]: "The system of Claim 1, further comprising processing the pages of text using a classifier model to identify the type of page for each page of text comprises identifying each page as one of a start page, and end page, or another page."
9[a] - "processing the pages of text using a classifier model to identify the type of page for each page of text… identifying each page as one of a start page, and end page, or another page":
Lucas discloses a document classification model that identifies the type of document/page. [0178]: "Document classification: The system may generate an image or text classification model to: determine whether a given document belongs to one of the templates that may be extracted from, assign the document an identifier for linking the document to the template, use the identifier to look up the classification model optimized for the document, and classify the document."
Lucas further discloses classifying pages by their content type.
[0106]: "An exemplary report featuring Sections 1-4 as described in FIG. 7 may be processed by an MLA or DLNN to identify Section 1 as a header which lists patient demographics such as name and date of birth, Section 2 as a listing of genetic variants which are linked to Section 3, Section 3 as a corresponding sequencing result, and Section 4 as a multi-page table summarizing conclusions made from the sequencing results."
Lucas discloses identifying pages by their position and content, which corresponds to identifying a page as a start page (header/first page), end page (conclusions/last page), or another page (intermediate content).
[0116]: "Identifying regions of interest, features within the region of interest, or relationships between regions may be performed from the OCR text itself or processed from the image itself prior to OCR. For example, identifying a region of interest may be performed by identifying a border (e.g., black box) that encapsulates some segment of text."
Claim 19 - Complete Limitation Listing and Mapping
19[a]: "The method of Claim 11, further comprising identifying each page as one of a start page, and end page, or another page."
19[a] is mapped identically to 9[a] above. Lucas discloses the same method in [0178] and [0106] as cited above.
Claim 10 - Complete Limitation Listing and Mapping
10[a]: "The system of Claim 1, further comprising bundling a group of the pages of text and assigning the same date to each page of the group."
10[a] - "bundling a group of the pages of text and assigning the same date to each page of the group":
Lucas discloses assigning a single date to a multi-page document (i.e., bundling pages and assigning the same date).
[0216]: "a document named Progress Note 01_01_01 may be presumed to have a date of Jan. 1, 2001. Other concept candidates from the document may be referenced to validate the date/time or select the date absent any other validating/corroborating information."
Lucas further discloses that a document with multiple pages shares a single date.
[0070]: "Document & Page: The document and page where the text is found (i.e., Progress Note 01_01_01.pdf and page 3)."
[0077]: "Document & Page: The document and page where the text is found (i.e., Progress Note 01_01_01.pdf and page 5)."
[0083]: "Document & Page: The document and page where the text is found (i.e., Progress Note 01_01_01.pdf and page 6)."
In each of these examples, the same document "Progress Note 01_01_01.pdf" (dated Jan. 1, 2001) spans multiple pages (pages 3, 5, and 6), and the same date is assigned to all pages of that document. This constitutes bundling a group of pages and assigning the same date to each page of the group.
Claim 20 – Complete Limitation Listing and Mapping
20[a]: "The method of Claim 11, further comprising bundling a group of the pages of text and assigning the same date to each page of the group."
20[a] is mapped identically to 10[a] above. Lucas discloses the same method in [0216], [0070], [0077], and [0083] as cited above.
Conclusion
The prior art made of record2 and NOT relied upon is considered pertinent to applicant's disclosure:
Mossin (US 2019/0034591): A system for predicting and summarizing medical events from electronic health records includes a computer memory storing aggregated electronic health records from a multitude of patients of diverse age, health conditions, and demographics including medications, laboratory values, diagnoses, vital signs, and medical notes. The aggregated electronic health records are converted into a single standardized data structure format and ordered arrangement per patient, e.g., into a chronological order. A computer (or computer system) executes one or more deep learning models trained on the aggregated health records to predict one or more future clinical events and summarize pertinent past medical events related to the predicted events on an input electronic health record of a patient having the standardized data structure format and ordered into a chronological order. An electronic device configured with a healthcare provider-facing interface displays the predicted one or more future clinical events and the pertinent past medical events of the patient.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHAEL EZEWOKO whose telephone number is 571 272 7850. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Fonya Long can be reached on 571 270 5096. The fax phone number for the organization where this application or proceeding is assigned is 571-273-7850.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MICHAEL I EZEWOKO/Primary Examiner, Art Unit 3682
1 See Form 892
2Please see Form 892 for complete listing