Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
This communication is responsive to Amendment, filed 05/11/2026.
Claims 1-21 are pending in this application. Claim 21 was added. This action is made Final.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Claims 1-3, 5-10, 12-17, 19-21 are rejected under 35 U.S.C. 103 as being unpatentable over Xu et al. (US Pat No. 10,929,420), in view of BARNES et al. (US Pub No. 2022/0044812).
As to claims 1, 8, 15, Xu teaches a method for processing and distributing text data from a dataset, the method comprising:
receiving raw data and converting the raw data into a set of text chunks (i.e. The pre-processing unit 404 may split the text of the medical text report 403 into n sentences (sentence1, sentence2 . . . sentencen). The pre-processing unit 404 may segment each sentence into words (word11 . . . word1s). The NLP unit 406 may acquire each sentence (sentence1, sentence2 . . . sentencen) including each segmented word (word11 . . . word1s) thereof, from the pre-processing unit 404, col. 14, line 39 to col. 15, lines 16; one or more medical text reports may be acquired from a database or pre-existing medical text reports. For example, the database may be an Electronic Medical Record (EMR) database, an Electronic Health Record (HER) database, a Radiology Information System (RIS) database, and/or another form of database, col. 4, lines 35-40; a new sentence may be acquired from voice-to-text software, col. 4, lines 41-61);
determining a set of classifications for the raw data, wherein classification of the set of classification correspond to different data stores of a set of data stores (i.e. the text classification model may classify which of a plurality of predefined labels a given natural language statement corresponds, col. 6, line 59 to col. 7, lines 14; the text classification model may determine a first label for a given natural language statement and a second label for the given natural language statement, col. 7, lines 15-37; The database may include a predetermined field structure populated with the label and the associated natural language data, which may be structured hierarchically, for example such that the labels are in a first level of the hierarchical structure and the associated natural language data are in a second, lower, level of the hierarchical structure. Further, the database may be structured to provide that data derived from a given natural language statement is stored hierarchically, col. 12, line 63 to col. 13, line 18; using a neural network, and the text classification model may map the vector representing the sentence onto one or more labels. The text classification model may be based on or implemented using a deep learning network, col.6, lines 28-58);
augmenting a text chunk in the set of text chunks with metadata, the augmenting comprising (i.e. using the text classification model, for each natural language statement, one or more classifications of the natural language statement with respect to a medical finding; wherein, for each natural language statement, the natural language data includes the one or more classifications, col. 2, lines 38-45; The deep learning network 408 determines, based on the determined one or more vectors, and using the text classification model, one or more labels 410 associated with the input sentence ... The label 410 may then be used to generate a report 412 including the label 410 in association with natural language data including the natural language statement, col. 14, line 39 to col. 15, line 16):
extracting retrieval metadata from the text chunk (i.e. using a computer implemented text analysis process, the medical text report to determine, for each natural language statement, one or more labels for the natural language statement, col. 4, lines 62-65; The data derived from a given natural language statement may itself be structured in a predetermined format ... a predetermined field structure may be populated with data derived from the natural language statement ... the first heading or section or field of the data derived from the natural language statement may be “nodules” and under this heading or section or field may be the classification data related to nodules, col. 11, lines 29-50);
determining and assigning, with a machine learning model configured to classify text information, a classification in the set of classifications for the text chunk (i.e. a structured report structured such that the determined one or more labels are each presented in association with the natural language data to which the label corresponds, col. 2, lines 24-28; The computer implemented text analysis process of act 104 of the method includes determining, for each word of each natural language statement of the acquired medical text report, and using word embeddings, a vector representing the word, col. 5, lines 6-10; The word embeddings may be learned using machine learning techniques, col. 5, lines 19-23; The text classification model training process may include machine learning techniques, col. 7, lines 48-49); and
sequencing the text chunk (i.e. acquiring a first of the natural language statements of the medical text report when the first natural language statement has been produced and before a second of the natural language statements of the medical text report has been produced. Optionally, the method includes training the text analysis process, col. 2, lines 46-51; acquiring the medical text report may include acquiring a first natural language statement of the medical text report when the first natural language statement has been produced and before a second natural language statement of the medical text report has been produced, col. 5, lines 46-61);
generating a windowed chunk by appending context to the augmented chunk (i.e. the determined label 208, specifically “Pulmonary arteries,” in association with the natural language statement 210 to which the label corresponds ... The label 206 is presented in association with the natural language statement 210. Specifically, the label 206 immediately precedes the natural language statement 210 ... The label 206 is presented in association with the natural language statement 210. Specifically, the label 206 immediately precedes the natural language statement 210, col. 12, lines 4-24);
generating a chunk embedding with the windowed chunk and the extracted retrieval metadata (i.e. the structured report data may be structured such that the one or more labels are in a first level of the hierarchical structure and the corresponding natural language data are in a second level of the hierarchical structure, the second level being lower than the first level in the hierarchical structure; col. 9, lines 29-54; the word embeddings may be learned using a skip-gram model implemented on a neural network. The skip-gram model may learn word embeddings for words given the local usage context of the words, where the context is defined by a window of neighbouring words. This window is a configurable parameter of the model, col. 5, lines 37-63; the label 216 is presented on a line that immediately precedes the data 218 derived from the natural language statement 210, and the label 214 is presented on a line that immediately precedes the label 216. The labels 214, 216 act as headings for the data 218 derived from the natural language statement 210, col. 12, lines 25-44); and
distributing the chunk embedding by (i.e. the structured report data is structured such that the one or more labels are in a first level of a hierarchical structure and the corresponding natural language data are in a second level of the hierarchical structure, the second level being lower than the first level in the hierarchical structure. Optionally, the method includes performing a computer implemented searching process for said natural language data stored in the structured database based at least in part on said labels stored in the structured database, col. 2, lines 29-33):
determining a data store in the set of data stores corresponding to the assigned classification (i.e. storing the generated structured report data in a structured database such that the determined one or more labels are each stored in association with the natural language data to which the label corresponds, col. 2, lines 19-23; the generated structured report data may include a table including the label and natural language data to which the label corresponds in a common row of the table, col. 8, line 61 to col. 9, line 4; the first example structured report 206 illustrated in FIG. 2. The structured report data of the first example report 206 is generated from the medical text report 202 ... The structured report data of the second example structured report 212 includes determined labels 214, 216 (a total of eight are shown) each in association with data 218 derived from the natural language statement to which the label corresponds, col. 9, lines 5-28); and
providing the chunk embedding to the determined data store (i.e. the generated structured report data may include the natural language data corresponding to a first natural language statement under a first heading or section or field including the label associated with the first natural language statement, and may include the natural language data corresponding to a second natural language statement under a second heading or section or field including the label associated with the second natural language statement, col. 9, lines 29-54; The structured report data may be transmitted ... The generated structured report data may be stored, for example in a suitable database, for example as a suitably structured text file or another format, col. 11, lines 51-58).
Xu implicitly teaches the term "converting" (i.e. receiving raw data and converting the raw data into a set of text chunks) (i.e. The pre-processing unit 404 may split the text of the medical text report 403 into n sentences (sentence1, sentence2 . . . sentencen). The pre-processing unit 404 may segment each sentence into words (word11 . . . word1s). The NLP unit 406 may acquire each sentence (sentence1, sentence2 . . . sentencen) including each segmented word (word11 . . . word1s) thereof, from the pre-processing unit 404, col. 14, line 39 to col. 15, lines 16; one or more medical text reports may be acquired from a database or pre-existing medical text reports. For example, the database may be an Electronic Medical Record (EMR) database, an Electronic Health Record (HER) database, a Radiology Information System (RIS) database, and/or another form of database, col. 4, lines 35-40; a
new sentence may be acquired from voice-to-text software, col. 4, lines 41-61).
Xu does not clearly state this term.
BARNES teaches this term (i.e. retrieving patient data of a patient ... The patient data can include raw structured, [0005]; automated extraction of patient data ... conversion of the extracted patient data into structured patient data records, such as a cancer registry, which can substantially speed up the generation of structured patient data records, [0023]; generating structured patient data records ... Each patient data record 110 may include a plurality of sections or tables including a patient biography information section 112, a tumor information section 114, a treatment information section 116, a biomarkers section 118, ... The structured medical data of patient data records 110 can be provided to, for example, different medical applications including, for example, a clinical decision application, a care evaluation application, a research application, regional/national cancer registries, accreditation boards, etc. In some examples, patient data records 110 can include a cancer registry, [0025]).
BARNES further teaches:
generating a chunk embedding with the windowed chunk and the extracted retrieval metadata (i.e. In some examples, natural language processor 304 and data normalization module 306 can operate together in various ways to handle the extracted data. For example, the natural language processor 304 and data normalization module 306 can operate in parallel to handle different sets of extracted data. In one example, data normalization module 306 can be assigned to handle shorter text strings, numerical values, etc., for which data normalization rules can define a reference numerical range or a set of standardized text data candidates. Natural language processor 304 can be assigned to handle more complex text strings, which may require some forms of contextual and semantic analyses to determine the intended meaning of the text strings for the output. Data normalization module 306 and natural language processor 304 can also operate in a serial fashion on the same set of extracted data. For example, data normalization module 306 can perform pre-processing on the extracted data to correct typos and/or out-of-range values. Natural language processor 304 can then process the pre-processed data to generate an output associated with data elements in patient data records 110, [0058]);
distributing the chunk embedding (i.e. patient data abstraction module 202 can automatically populate different fields of patient data records 110 using the processed data, or assist an abstractor in populating the fields of patient data records 110. For example, in one operation mode, patient data abstraction module 202 can automatically populate, via server 122, different fields of patient data records 110 of database 120 based on pre-determined mapping between the pre-defined data representations and the fields of patient data records 110, [0033], and Fig. 1B)).
It would have been obvious to one of ordinary skill of the art having the teaching of Xu, BARNES before the effective filing date of the claimed invention to modify the system of Xu to include the limitations as taught by BARNES. One of ordinary skill in the art would be motivated to make this combination in order to perform automated extraction of patient data from electronic medical records and conversion into a structured patient data record, in view of BARNES ([0027]), as doing so would give the added benefit of populating various fields of a structured patient data record based on the structured data, as taught by BARNES ([0027]).
As to claims 2, 9, 16, Xu teaches training the machine learning model with ground truth data (i.e. wherein the ground-truth for each natural language statement includes the label, col. 2, line 65 to col. 3, line 2; The ground-truth for each natural language statement used in the training process may include the label from the structured report ... The training process may therefore use the headings or names of the sections of the reports under which a given natural language statement as the ground truth label for that statement, col. 8, lines 16-33).
As to claims 3, 10, 17, Xu teaches ordering the augmented text chunk (i.e. for each natural language statement: determining, for each of a plurality of predefined labels, an association parameter indicating a degree to which the natural language statement is associated with the pre-defined label; wherein the determining the one or more labels associated with the natural language statement is based on the determined association parameter, col. 2, lines 34-37; the plurality of predefined labels may be “Pulmonary arteries,” “Lungs and Airways” and “Pleura,” and the model may determine that the natural language statement “There are multiple subsegmental pulmonary emboli throughout the right lung” has a larger association parameter for the label “Pulmonary arteries” than for the labels “Lungs and Airways” or “Pleura,” and hence the text classification model may determine that the label accordingly. In some examples, the text classification algorithm may be arranged such that any one of the plurality of predefined labels that has an association parameter with a given natural language statement higher than a predefined threshold may be determined as a label for that natural language statement, col. 6, line 59 to col. 7, line 14).
As to claims 5, 12. 19, Xu teaches the determined data store comprises a repository (i.e. Optionally, the structured report data is structured such that the one or more labels are in a first level of a hierarchical structure and the corresponding natural language data are in a second level of the hierarchical structure, the second level being lower than the first level in the hierarchical structure. Optionally, the method includes performing a computer implemented searching process for said natural language data stored in the structured database based at least in part on said labels stored in the structured database, col. 2, lines 29-33).
As to claims 6, 13, 20, BARNES teaches the determined data store comprises a second machine learning model (i.e. The learning system can include, for example, a rule-based extraction system, a machine learning (ML) model (which may include a deep learning neural network or other machine learning models), a natural language processor (NLP), etc., which can extract data elements from the unstructured patient data and determine their data categories based on a trained language extraction model, such as language extraction model 312 of FIG. 3B. Some of the data elements can also be mapped to pre-defined data representations (e.g., codes, fields, etc.) to form structured data, based on data table 330 of FIG. 3C. Moreover, as part of a normalization process, the learning system can also detect and correct data errors in the extracted data elements, and convert the extracted data elements to standardized data formats, [0057]).
As to claims 7, 14, Xu teaches the raw data comprises medical record data (i.e. In some examples, one or more medical text reports may be acquired from a database or pre-existing medical text reports. For example, the database may be an Electronic Medical Record (EMR) database, an Electronic Health Record (HER) database, a Radiology Information System (RIS) database, and/or another form of database, col. 4, lines 35-40).
As per claim 21, Xu teaches each classification of the set of classification corresponds to different data store of the set of data stores (i.e. the text classification model may classify which of a plurality of predefined labels a given natural language statement corresponds, col. 6, line 59 to col. 7, lines 14; the text classification model may determine a first label for a given natural language statement and a second label for the given natural language statement, col. 7, lines 15-37; The database may include a predetermined field structure populated with the label and the associated natural language data, which may be structured hierarchically, for example such that the labels are in a first level of the hierarchical structure and the associated natural language data are in a second, lower, level of the hierarchical structure. Further, the database may be structured to provide that data derived from a given natural language statement is stored hierarchically, col. 12, line 63 to col. 13, line 18; using a neural network, and the text classification model may map the vector representing the sentence onto one or more labels. The text classification model may be based on or implemented using a deep learning network, col.6, lines 28-58).
Claims 4, 11, 18 are rejected under 35 U.S.C. 103 as being unpatentable over Xu et al. (US Pat No. 10,929,420), in view of BARNES et al. (US Pub No. 2022/0044812), as applied to claims above, and further in view of MOK et al. (US Pub No. 2010/0094658).
As to claims 4, 11, 18, BARNES teaches a general data store, and based on a determination that the text chunk is not assigned to a classification in the set of classifications, assigning the text chunk to the general data store (i.e. in one operation mode, patient data abstraction module 202 can automatically populate, via server 122, different fields of patient data records 110 of database 120 based on pre-determined mapping between the pre-defined data representations and the fields of patient data records 110. Moreover, in a different operation mode, patient data abstraction module 202 may allow manual extraction as a backup option when, for example, AI-assisted clinical extraction tool outputs a low confidence level for the output, which may indicate that raw patients data 210 include data that are inconsistent with the training data set. In some examples, patient data abstraction module 202 may adopt a hybrid approach by allowing a human abstractor to populate certain data element fields, via a display interface 206 and server 122, while using the AI-assisted clinical extraction tool to populate other data element fields, [0033]).
Xu, BARNES do not seem to specifically teach "the general data store".
MOK teaches this limitation (i.e. A Document log sorted by medical category
provides an inventory of pages in a patient's compiled records, organized by sections made up of commonly used medical sub-categories such as medications & allergies; hospital summaries; labs & cultures; etc. (see FIG. 15). Within each sub-category, pages are sorted in reversed chronological order. When a document cannot be placed in a specific medical sub-category, it is placed in a sub-category labeled “Other: Unclassified” and placed at the end of the log, [0095]; A document log sorted by specialization organizes the pages based on the specialty of the doctor or provider who wrote or created each page. This log provides an inventory of pages in a patient's compiled records, organized by name of the specialty of the physician(s) who authored the pages. The ability to sort charts by specialty helps patients bring information that is most relevant to their doctors, especially when they see a specialist about a particular condition. Pages authored by a location or an organization (such as a health clinic or laboratory) can be difficult to classify into specialties and are left for the patients or their doctors to categorize by relevant specialty. A document which cannot be categorized by specialty of author is included in a category labeled, “Unknown Specialty” and placed at the end of the log, [0096]).
It would have been obvious to one of ordinary skill of the art having the teaching of Xu, BARNES, MOK before the effective filing date of the claimed invention to modify the system of Xu, BARNES to include the limitations as taught by MOK. One of ordinary skill in the art would be motivated to make this combination in order to label the uncategorized document as "Other: Unclassified" in view of MOK ([0095]), as doing so would give the added benefit of providing a comprehensive, organized, and accessible medical history of the patient, and allowing easy retrieval and meaningful display of
relevant information, as taught by MOK ([0096]).
Response to Arguments
Applicant's arguments with respect to claims 1-21 have been considered but are moot in view of the new ground(s) of rejection.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MIRANDA LE whose telephone number is (571)272-4112. The examiner can normally be reached M-F 7AM-5PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kavita Stanley can be reached on 571-272-8352. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MIRANDA LE/ Primary Examiner, Art Unit 2153