Prosecution Insights
Last updated: August 14, 2026
Application No. 19/334,767

SYSTEM AND METHOD FOR AUTOMATED HEALTH RECORD GENERATION

Non-Final OA §101§103
Filed
Sep 19, 2025
Priority
Sep 19, 2024 — provisional 63/696,639
Examiner
MACCAGNO, PIERRE L
Art Unit
3687
Tech Center
3600 — Transportation & Electronic Commerce
Assignee
Acuity Holding Corp.
OA Round
1 (Non-Final)
23%
Grant Probability
At Risk
1-2
OA Rounds
2y 2m
Est. Remaining
54%
With Interview

Examiner Intelligence

Grants only 23% of cases
23%
Career Allowance Rate
32 granted / 139 resolved
-29.0% vs TC avg
Strong +32% interview lift
Without
With
+31.5%
Interview Lift
resolved cases with interview
Typical timeline
3y 1m
Avg Prosecution
21 currently pending
Career history
182
Total Applications
across all art units

Statute-Specific Performance

§101
47.2%
+7.2% vs TC avg
§103
35.3%
-4.7% vs TC avg
§102
9.2%
-30.8% vs TC avg
§112
7.5%
-32.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 139 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Status of Claims This action is a non-final rejection Claims 1-20 are pending Claims 1-20 are rejected under 35 USC § 101 Claim 1- 20 are rejected under 35 USC § 103 Priority Acknowledgement is made of Applicant’s claim for a foreign priority date of 9-19-2024 Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are not patent eligible because the claimed invention is directed to an abstract idea without significantly more. Analysis First, claims are directed to one or more of the following statutory categories: a process, a machine, a manufacture, and a composition of matter. Regarding claims 1-20 the claims recite an abstract idea of “generating a patient health record based on vocalized expression”. Independent Claims 1, 12, 18 are rejected under 35 U.S.C 101 based on the following analysis. -Step 1 (Does the claim fall within a statutory category? YES): claims 1, 12, 18 recite respectively a method, a processor and a non-transitory computer-readable medium to generate a patient health record based on vocalized expression. -Step 2A Prong One (Does the claim fall within at least one of the groupings of abstract ideas?: YES): The claimed invention: receiving an audio signal indicative of the vocalized expression; generating a first transcript using a first transcription model based on the audio signal; generating a second transcript using a second transcription model based on the audio signal, independently of the first transcription model; determining a third transcript based on the first and second transcripts by using a ... model; using the third transcript and a ... model to determine data suitable for generating the patient record; generate the patient record belonging to the grouping of mental processes under concepts performed in the human mind (including an observation, evaluation, judgement, opinion) as it recites “generating a patient health record based on vocalized expression”. Alternatively, the selected abstract idea belongs to the grouping of certain methods of organizing human activity under managing personal behavior or relationships or interactions between people as it recites “generating a patient health record based on vocalized expression” (refer to MPP 2106.04(a)(2)). Accordingly this claim recites an abstract idea. -Step 2A Prong Two (Are there additional elements in the claim that imposes a meaningful limit on the abstract idea? NO). Claims 1, 12, 18 recite: automatically generating a patient record, via a user-oriented record-keeping application, based on vocalized expression; using a large language model machine learning model; operating the record-keeping application using a desktop automation tool based on the data. Claim 1 recites: automatically generating a patient record, via a user-oriented record-keeping application, based on vocalized expression. Claim 12 recites: A system for automatically generating a patient record, via a user-oriented record-keeping application, based on vocalized expression, comprising: a processor; and computer-readable memory coupled to the processor and storing processor- executable instructions that, when executed, configure the processor. Claim 18 recites: A non-transitory computer-readable medium having stored thereon machine interpretable instructions which, when executed by a processor, cause the processor to perform a computer-implemented method of automatically generating a patient record, via a user-oriented record-keeping application, based on vocalized expression. Amounting to additional elements that are recited at a high-level of generality such that it amounts to no more than mere instructions to implement an abstract idea on a computer, or merely use a computer as a tool to implement the abstract idea. (refer to MPEP 2106.05(f)). Accordingly, the claim as a whole does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. -Step 2B (Does the additional elements of the claim provide an inventive concept?: NO. As discussed previously with respect to Step 2A Prong Two: Claims 1, 12, 18 recite: automatically generating a patient record, via a user-oriented record-keeping application, based on vocalized expression; using a large language model machine learning model; operating the record-keeping application using a desktop automation tool based on the data. Claim 1 recites: automatically generating a patient record, via a user-oriented record-keeping application, based on vocalized expression. Claim 12 recites: A system for automatically generating a patient record, via a user-oriented record-keeping application, based on vocalized expression, comprising: a processor; and computer-readable memory coupled to the processor and storing processor- executable instructions that, when executed, configure the processor. Claim 18 recites: A non-transitory computer-readable medium having stored thereon machine interpretable instructions which, when executed by a processor, cause the processor to perform a computer-implemented method of automatically generating a patient record, via a user-oriented record-keeping application, based on vocalized expression. Amount to additional elements that are recited at a high-level of generality such that it amounts to no more than mere instructions to implement an abstract idea on a computer, or merely use a computer as a tool to implement the abstract idea. (refer to MPEP 2106.05(f)) Accordingly, even in combination the additional elements of the claim do not provide an inventive concept (significantly more than the abstract idea) and hence the claim is ineligible. Dependent Claims: Step 2A Prong One: The following dependent claims recite additional limitations that further define the abstract idea of “generating a patient health record based on vocalized expression”: claims 2-11, 13-17, 19-20. Step 2A Prong Two (Are there additional elements in the claim that imposes a meaningful limit on the abstract idea? NO). The following dependent claims 2-3, 5-7, 9-11, 13-14, 16-17, 19 recite mere instructions to implement an abstract idea on a computer, or merely use a computer as a tool to implement the abstract idea. (refer to MPEP 2106.05(f)). Accordingly, the claims as a whole do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. Step 2B (Does the additional elements of the claim provide an inventive concept?: NO). As discussed previously with respect to Step 2A Prong Two, the following dependent claims 2-3, 5-7, 9-11, 13-14, 16-17, 19recite mere instructions to implement an abstract idea on a computer, or merely use a computer as a tool to implement the abstract idea. (refer to MPEP 2106.05(f)). Accordingly, the claim does not provide an inventive concept (significantly more than the abstract idea) and hence the claim is ineligible. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or non-obviousness. Claims 1, 7-8, 10-12, 18-20 are rejected under 35 U.S.C. 103 as being un-patentable by Xu et.al. (US 20250029612 A1) hereinafter “Xu” in view of Kim et.al. (US 20260073123 A1) hereinafter “Kim”, in further view of Grillo et.al. (US 20250063140 A1) hereinafter “Grillo”. Regarding claims 1, 12, 18 Xu teaches: receiving an audio signal indicative of the vocalized expression; (See at least [0011] via: “... In some applications of automatic speech recognition, such as in the healthcare domain, medical summaries of doctor-patient conversations may be generated that are transcribed using automatic speech recognition for recorded audio of clinical visits...”; in addition see at least [0051] via: “...As indicated at 610, audio data may be received for generating a transcription...”) generating a second transcript using a second transcription model based on the audio signal, independently of the first transcription model; (See at least [0052] via: “.... As indicated at 620, a first version of a transcript may be generated for speech in a portion of the audio data...”; in addition see at least [claim 1] via: “...the automatic speech recognition system is configured to: receive audio data for generating a transcription; generate a first version of a transcript for speech in a portion of the audio data according to a machine learning model trained to recognize the speech in the portion of the audio data..”) operating the record-keeping application using a desktop automation tool based on the data to generate the patient record. (See at least [0040] via: “....Once the summary is generated, the summarization task processing engine 232 may provide the generated summary to an output interface. The output interface may notify the customer of the completed job request. In some embodiments, the output interface may provide a notification of a completed job to the output API. In some embodiments, the output API may be implemented to provide the summary for upload to an electronic health record (EHR) or may push the summary out to an electronic health record (EHR), in response to a notification of a completed job..”) However, Xu is silent the following limitation that is taught by Kim: generating a first transcript using a first transcription model based on the audio signal; (See at least [0018] via: “....collecting transcript texts generated during a meeting to generate data chunks for every configured unit; generating summarized meeting minutes, when the data chunks are generated, by summarizing the transcript texts included in the data chunks using a large language model (LLM); and summarizing, when the data chunk is added, transcript text included in the added data chunk using the large language model, and updating the summarized meeting minutes to reflect the transcript text...”) It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Xu to incorporate the teachings of Kim. Those in the art would have recognized that Xu’s teaching regarding receiving audio data for generating a transcription, followed by the generation of a transcript for speech in a portion of the audio data based on a machine learning model trained to recognize the speech in the portion of the audio data, and whereby transcript summaries may be created using a special class of machine learning models based on large language models (LLM) could be modified to include Kim’s teaching regarding receiving audio data for generating a transcription, followed by the generation of a different transcript for speech in a different portion of the audio data based on a different machine learning model trained to recognize the speech in the different portion of the audio data, and whereby transcript summaries may be created using a different special class of machine learning models based on different large language models (LLM). This combination would be beneficial in separating an audio interaction between two people by obtaining different transcript versions of an audio interaction. As an example an audio interaction between doctor and patient, with one transcript obtained based on the doctor’s utterances and the second separate transcript obtained based on the patient utterances. The need based on the fact that the doctor and patient might speak different languages that may require separate LLMs. However, Xu and Kim are silent the following limitation that is taught by Grillo: determining a third transcript based on the first and second transcripts by using a large language model; (See at least [0165] via: “.... first portion of the first transcript comprising subject matter associated with the smart topic and a second portion of the second transcript comprising subject matter associated with the smart topic and providing the first portion of the first transcript and the second portion of the second transcript to the large language model to generate the combined smart topic output....”) It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Xu and Kim to incorporate the teachings of Grillo. Those in the art would have recognized that Xu’s teaching regarding receiving audio data for generating a transcription, followed by the generation of a transcript for speech in a portion of the audio data and Kim’a teaching regarding receiving audio data for generating a transcription, followed by the generation of a different transcript for speech in a different portion of the audio, could be modified to include Grillo’s teaching regarding the combination of both transcripts into one summary transcript by means of a LLM. This combination would be beneficial in grouping two or more different transcript sections into one transcript to obtain the complete audio transcription between two or more users. using the third transcript and a machine learning model to determine data suitable for generating the patient record; (See at least [0131] via: “.... the smart topic generation system 102 provides transcripts to a context transformer engine 1206 to identify related subject matter or subject matter related to a smart topic in the transcripts. Specifically, the smart topic generation system 102 provides first transcript 1202 and second transcript 1204 to the context transformer engine 1206 to generate portions 1210, including a first transcript portion and a second transcript portion (from the same or different transcripts). For example, the smart topic generation system 102 receives text input from a client device that comprises a prompt (e.g., for a large language model) indicating an objective (e.g., an action, task, or goal) for a combined smart topic output 1214. The smart topic generation system 102 utilizes the context transformer engine 206 to process the objective and to identify portions of first transcript 1202 and second transcript 1204 that relate to the objective and process (e.g., break downs or separate) first transcript 1202 and second transcript 1204 into to portions 1210 that relate to the objective It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Xu and Kim to incorporate the teachings of Grillo. Those in the art would have recognized that Xu’s teaching regarding receiving audio data for generating a transcription, followed by the generation of a transcript for speech in a portion of the audio data and Kim’a teaching regarding receiving audio data for generating a transcription, followed by the generation of a different transcript for speech in a different portion of the audio, could be modified to include Grillo’s teaching regarding identifying related subject matter or subject matter related to a smart topic in the combined transcripts. This combination would be beneficial in grouping two or more different transcript sections and extracting into one transcript subject matter related to a smart topic such as suitability of the transcript information to a patient. further comprising obtaining an overall score for each of the candidate principal investigators based on two or more factors, wherein the overall score includes a weighted sum of the two or more factors. (See at least [0035] via: “...the step of determining the score for each of the merged data entries based on the rankings for each of the one or more ranking factors for each of the merged data entries further comprising the steps of:..”; in addition see at least [0036] via: “.. receiving a weight for each of the one or more ranking factors; in addition see at least [0037] via: “... applying the weight for each of the one or more ranking factors to each of the one or more rankings before determining the score for each of the merged data entries...”; in addition see at least [0041] via: “...wherein each of the plurality of markers corresponds to an entity of each of the merged data entries, and wherein the entity is a principal investigator or a clinical trial site..”) It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Xu and Kim to incorporate the teachings of Grillo. Those in the art would have recognized that Xu’s teaching regarding receiving audio data for generating a transcription, followed by the generation of a transcript for speech in a portion of the audio data and Kim’a teaching regarding receiving audio data for generating a transcription, followed by the generation of a different transcript for speech in a different portion of the audio, could be modified to include Grillo’s teaching regarding applying the weight for each of the one or more ranking factors to each of the one or more rankings before determining the score for each of the merged data entries. This combination would be beneficial in as an example determining the score of a principal investigator. Regarding claims 7, 19 Xu, Kim and Grillo teach the invention as claimed and detailed above with respect to claims 1 and 18 respectively. Xu also teaches: wherein the machine learning model is a second large language model independent of the first large language model (See at least [claim 1] via: “...the automatic speech recognition system is configured to: receive audio data for generating a transcription; generate a first version of a transcript for speech in a portion of the audio data according to a machine learning model trained to recognize the speech in the portion of the audio data....; in addition see at least [0011] via: “... medical summaries of doctor-patient conversations may be generated that are transcribed using automatic speech recognition for recorded audio of clinical visits. The summaries capture a patient's reason for visit, history of illness as well as the doctor's assessment and plan for the patient. The summaries may be created using a special class of machine learning models, generative large language models (LLM) that are tuned to follow natural language instructions describing any task. This class of LLMs (e.g., InstructGPT) are typically trained on massive general-purpose text corpora and on a variety of tasks, including summarization...”) However, Xu is silent the following limitation that is taught by Kim: wherein the large language model is a first large language model (See at least [0018] via: “....collecting transcript texts generated during a meeting to generate data chunks for every configured unit; generating summarized meeting minutes, when the data chunks are generated, by summarizing the transcript texts included in the data chunks using a large language model (LLM); and summarizing, when the data chunk is added, transcript text included in the added data chunk using the large language model, and updating the summarized meeting minutes to reflect the transcript text...”) It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Xu to incorporate the teachings of Kim. Those in the art would have recognized that Xu’s teaching regarding receiving audio data for generating a transcription, followed by the generation of a transcript for speech in a portion of the audio data based on a machine learning model trained to recognize the speech in the portion of the audio data, and whereby transcript summaries may be created using a special class of machine learning models based on large language models (LLM) could be modified to include Kim’s teaching regarding receiving audio data for generating a transcription, followed by the generation of a different transcript for speech in a different portion of the audio data based on a different machine learning model trained to recognize the speech in the different portion of the audio data, and whereby transcript summaries may be created using a different special class of machine learning models based on different large language models (LLM). This combination would be beneficial in separating an audio interaction between two people by obtaining different transcript versions of an audio interaction. As an example an audio interaction between doctor and patient, with one transcript obtained based on the doctor’s utterances and the second separate transcript obtained based on the patient utterances. The need based on the fact that the doctor and patient might speak different languages that may require separate LLMs. Regarding claims 8, 20 Xu, Kim and Grillo teach the invention as claimed and detailed above with respect to claims 1 and 18 respectively. Xu also teaches: generating the second transcript using the second transcription model based on the audio signal, independently of the first transcription model, includes generating a second candidate transcript using the second transcription model based on the audio signal, (See at least [0011] via: “... In some applications of automatic speech recognition, such as in the healthcare domain, medical summaries of doctor-patient conversations may be generated that are transcribed using automatic speech recognition for recorded audio of clinical visits...”; in addition see at least [0051] via: “...As indicated at 610, audio data may be received for generating a transcription...”). receiving a second user input responsive to the second candidate transcript (See at least [0024] via: “... guided transcript generation 130 may use the section type to modify (or re-generate first transcription version) to provide final transcript 104...”) modifying the second candidate transcript based on the second user input to generate the second transcript. (See at least [0024] via: “... guided transcript generation 130 may use the section type to modify (or re-generate first transcription version) to provide final transcript 104...”) However, Xu is silent the following limitation that is taught by Kim: generating the first transcript using the first transcription model based on the audio signal includes generating a first candidate transcript using the first transcription model based on the audio signal, (See at least [0018] via: “....collecting transcript texts generated during a meeting to generate data chunks for every configured unit; generating summarized meeting minutes, when the data chunks are generated, by summarizing the transcript texts included in the data chunks using a large language model (LLM); and summarizing, when the data chunk is added, transcript text included in the added data chunk using the large language model, and updating the summarized meeting minutes to reflect the transcript text...”) receiving a first user input responsive to the first candidate transcript, (See at least [0117] via: “....a request for modification of transcript text may be input..”) modifying the first candidate transcript based on the first user input to generate the first transcript; (See at least [0016] via: “....the generative AI-based meeting minutes summarization method according to an embodiment of the disclosure may further include, when a request for modification of the transcript text is input, modifying the transcript text in the data chunk to generate a modified data chunk, and regenerating the summary data by summarizing the modified data chunk using the large language model...”) It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Xu to incorporate the teachings of Kim. Those in the art would have recognized that Xu’s teaching regarding receiving audio data for generating a transcription, followed by the generation of a transcript for speech in a portion of the audio data based on a machine learning model trained to recognize the speech in the portion of the audio data, and whereby transcript summaries may be created using a special class of machine learning models based on large language models (LLM) could be modified to include Kim’s teaching regarding receiving audio data for generating a transcription, followed by the generation of a different transcript for speech in a different portion of the audio data based on a different machine learning model trained to recognize the speech in the different portion of the audio data, and whereby transcript summaries may be created using a different special class of machine learning models based on different large language models (LLM). This combination would be beneficial in separating an audio interaction between two people by obtaining different transcript versions of an audio interaction. As an example an audio interaction between doctor and patient, with one transcript obtained based on the doctor’s utterances and the second separate transcript obtained based on the patient utterances. The need based on the fact that the doctor and patient might speak different languages that may require separate LLMs. Regarding claim 10 Xu, Kim and Grillo teach the invention as claimed and detailed above with respect to claims 1&8. However Xu is silent the following claim that is taught by Kim: generating a first candidate transcript using the first transcription model based on the audio signal, (See at least [0018] via: “....collecting transcript texts generated during a meeting to generate data chunks for every configured unit; generating summarized meeting minutes, when the data chunks are generated, by summarizing the transcript texts included in the data chunks using a large language model (LLM); and summarizing, when the data chunk is added, transcript text included in the added data chunk using the large language model, and updating the summarized meeting minutes to reflect the transcript text...”) receiving a first user input responsive to the first candidate transcript, (See at least [0117] via: “....a request for modification of transcript text may be input..”) modifying the first candidate transcript based on the first user input to generate the first transcript. (See at least [0016] via: “....the generative AI-based meeting minutes summarization method according to an embodiment of the disclosure may further include, when a request for modification of the transcript text is input, modifying the transcript text in the data chunk to generate a modified data chunk, and regenerating the summary data by summarizing the modified data chunk using the large language model...”) It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Xu to incorporate the teachings of Kim. Those in the art would have recognized that Xu’s teaching regarding receiving audio data for generating a transcription, followed by the generation of a transcript for speech in a portion of the audio data based on a machine learning model trained to recognize the speech in the portion of the audio data, and whereby transcript summaries may be created using a special class of machine learning models based on large language models (LLM) could be modified to include Kim’s teaching regarding receiving audio data for generating a transcription, followed by the generation of a different transcript for speech in a different portion of the audio data based on a different machine learning model trained to recognize the speech in the different portion of the audio data, and whereby transcript summaries may be created using a different special class of machine learning models based on different large language models (LLM). This combination would be beneficial in separating an audio interaction between two people by obtaining different transcript versions of an audio interaction. As an example an audio interaction between doctor and patient, with one transcript obtained based on the doctor’s utterances and the second separate transcript obtained based on the patient utterances. The need based on the fact that the doctor and patient might speak different languages that may require separate LLMs. Regarding claim 11 Xu, Kim and Grillo teach the invention as claimed and detailed above with respect to claim 1. Xu also teaches: wherein operating the record-keeping application using a desktop automation tool based on the data to generate the patient record includes storing the patient record in a database. (See at least [0027] via: “....FIG. 2 illustrates an example provider network that may implement a medical audio summarization service that implements guiding transcript generation using detected section types as part of automatic speech recognition,. ... The provider network ... may include numerous data centers hosting various resource pools, such as collections of physical and/or virtualized computer servers, storage devices, networking equipment and the like .. needed to implement and distribute the infrastructure and services offered by the provider network 200. For example, the provider network 200 may implement various computing resources or services, such as ...network-based services 290 (which may include a virtual compute service and various other types of storage, .. data cataloging, data ingestion (e.g., ETL),..”) Claims 2-4, 9, 13-15 are rejected under 35 U.S.C. 103 as being un-patentable by Xu in view of Kim, in view of Grillo, in further view of Halyal et.al. (US 20250200098 A1) hereinafter “Halyal” Regarding claims 2, 13 Xu, Kim and Grillo teach the invention as claimed and detailed above with respect to claims 1 and 12 respectively. Xu, Kim and Grillo are silent the following claim taught by Halyal: wherein determining the third transcript based on the first and second transcripts by using a large language model includes determining the third transcript based on the first and second transcripts by using a plurality of independent large language models. (See at least [0078] via: “....Flow mining refers to the process of analyzing conversation transcripts and extracting patterns (or other data from which patterns can be ascertained) that represent the likely flow(s) of the conversation. It should be appreciated that a particular mined flow can be visually represented as a state machine, where the state machine includes a set of states at which a conversation could be in and its context within the respective states. A mined flow may provide an aggregated view of the most common pattern of states present in a set of conversations and their most common sequence of events. For example, for conversations containing a booking flight intent, the most likely first slot may be the destination/arrival city, which may then become the first state in the booking flight flow. In some embodiments, the flows are mined using a multi-step approach and leveraging one or more large language models (LLMs) as described herein...”) It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Xu, Kim and Grillo to incorporate the teachings of Halyal. Those in the art would have recognized that Xu’s teaching regarding receiving audio data for generating a transcription, followed by the generation of a transcript for speech in a portion of the audio data; Kim’a teaching regarding receiving audio data for generating a transcription, followed by the generation of a different transcript for speech in a different portion of the audio, and Grillo’s teaching regarding combining two or more transcripts could be modified to include Halyal’s teaching regarding the use of one or more LLMs used to mine transcripts. This combination would be beneficial in “analyzing conversation transcripts and extracting patterns (or other data from which patterns can be ascertained) that represent the likely flow(s) of the conversation” Halyal [0078]. Regarding claims 3, 14 Xu, Kim and Grillo teach the invention as claimed and detailed above with respect to claims 1&2 and 12&13 respectively. Xu, Kim and Grillo are silent the following claim taught by Halyal: generating a plurality of candidate transcripts using the plurality of independent large language models, each of the plurality of candidate transcripts being uniquely associated with a corresponding one of the plurality of independent large language models; and generating the third transcript based on the plurality of candidate transcripts. (See at least [0078] via: “....Flow mining refers to the process of analyzing conversation transcripts and extracting patterns (or other data from which patterns can be ascertained) that represent the likely flow(s) of the conversation. It should be appreciated that a particular mined flow can be visually represented as a state machine, where the state machine includes a set of states at which a conversation could be in and its context within the respective states. A mined flow may provide an aggregated view of the most common pattern of states present in a set of conversations and their most common sequence of events. For example, for conversations containing a booking flight intent, the most likely first slot may be the destination/arrival city, which may then become the first state in the booking_flight flow. In some embodiments, the flows are mined using a multi-step approach and leveraging one or more large language models (LLMs) as described herein...”) It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Xu, Kim and Grillo to incorporate the teachings of Halyal. Those in the art would have recognized that Xu’s teaching regarding receiving audio data for generating a transcription, followed by the generation of a transcript for speech in a portion of the audio data; Kim’a teaching regarding receiving audio data for generating a transcription, followed by the generation of a different transcript for speech in a different portion of the audio, and Grillo’s teaching regarding combining two or more transcripts could be modified to include Halyal’s teaching regarding the use of one or more LLMs used to mine transcripts. This combination would be beneficial in “analyzing conversation transcripts and extracting patterns (or other data from which patterns can be ascertained) that represent the likely flow(s) of the conversation” Halyal [0078]. Regarding claims 4, 15 Xu, Kim and Grillo teach the invention as claimed and detailed above with respect to claims 1&2&3 and 12&13&14 respectively. Xu, Kim and Grillo are silent the following claim taught by Halyal: receiving a user input responsive to the plurality of candidate transcripts, and modifying the plurality of candidate transcripts based on the user input to generate the third transcript. (See at least [0012] via: “.... clustering the plurality of transcripts into the plurality of intent categories may include modifying the consolidated intents based on user feedback...”) It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Xu, Kim and Grillo to incorporate the teachings of Halyal. Those in the art would have recognized that Xu’s teaching regarding receiving audio data for generating a transcription, followed by the generation of a transcript for speech in a portion of the audio data; Kim’a teaching regarding receiving audio data for generating a transcription, followed by the generation of a different transcript for speech in a different portion of the audio, and Grillo’s teaching regarding combining two or more transcripts could be modified to include Halyal’s teaching regarding the use of one or more LLMs used to mine transcripts. This combination would be beneficial in “analyzing conversation transcripts and extracting patterns (or other data from which patterns can be ascertained) that represent the likely flow(s) of the conversation” Halyal [0078]. Regarding claim 9 Xu, Kim and Grillo teach the invention as claimed and detailed above with respect to claims 1&8. Xu, Kim and Grillo are silent the following claim taught by Halyal: updating the first machine learning model based on the first user input. (See at least [0071] via: “... the analytics module 250 also may generate, update, train, and modify predictors or models (e.g., machine learning models) based on collected data, such as, for example, customer data, agent data, and interaction data...”) It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Xu, Kim and Grillo to incorporate the teachings of Halyal. Those in the art would have recognized that Xu’s teaching regarding receiving audio data for generating a transcription, followed by the generation of a transcript for speech in a portion of the audio data; Kim’a teaching regarding receiving audio data for generating a transcription, followed by the generation of a different transcript for speech in a different portion of the audio, and Grillo’s teaching regarding combining two or more transcripts could be modified to include Halyal’s teaching regarding the use of one or more LLMs used to mine transcripts. This combination would be beneficial in “analyzing conversation transcripts and extracting patterns (or other data from which patterns can be ascertained) that represent the likely flow(s) of the conversation” Halyal [0078] Claims 5-6, 16-17 are rejected under 35 U.S.C. 103 as being un-patentable by Xu in view of Kim, in view of Grillo, in view of Halyal, in further view of Chalana et.al. (US 20220215052 A1) hereinafter “Chalana” Regarding claims 5, 16 Xu, Kim and Grillo teach the invention as claimed and detailed above with respect to claims 1 and 12 respectively and Xu, Kim, Grillo and Halyal teach the invention as claimed and detailed above with respect to claims 2&3&4 and 13&14&15 respectively. Xu, Kim, Grillo and Halyal are silent the following claim taught by Chalana: updating the supervised learning model based on the user input (See at least [0063] via: “...The supervised neural network may be trained on a training dataset training dataset comprising a plurality of audio-visual media files and or the transcript, wherein the plurality of audio-visual media and the transcript are annotated with a human feedback regarding summary status...) It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Xu, Kim, Grillo and Halyal to incorporate the teachings of Chalana. Those in the art would have recognized that Xu’s teaching regarding receiving audio data for generating a transcription, followed by the generation of a transcript for speech in a portion of the audio data; Kim’a teaching regarding receiving audio data for generating a transcription, followed by the generation of a different transcript for speech in a different portion of the audio; Grillo’s teaching regarding combining two or more transcripts and Halyal’s teaching regarding the use of one or more LLMs could be modified to include Chalana’s’s teaching regarding the use of a supervised model to summarize transcripts which by the use of labeled data for training results in summaries with reduced errors and improved accuracy. Regarding claims 6, 17 Xu, Kim and Grillo teach the invention as claimed and detailed above with respect to claims 1 and 12 respectively and Xu, Kim, Grillo and Halyal teach the invention as claimed and detailed above with respect to claims 2&3 and 13&14 respectively. Xu, Kim, Grillo and Halyal are silent the following claim taught by Chalana: wherein generating the third transcript based on the plurality of candidate transcripts includes using a supervised learning model to process the plurality of candidate transcripts to generate the third transcript (See at least [0063] via: “...At block 600, audio-visual media synopsis module 400 may call or trigger execution of, for example, transcription summarization module 600. Transcription summarization module 600 may, for example, identify meaning clusters in sentences in input transcript 335 and input video 305 records, referred to herein as “sentence meaning clusters”, wherein meaning within each sentence meaning cluster is similar and is dissimilar between sentence meaning clusters. Identification of sentence meaning clusters may be performed by one or more neural networks in transcription neural network. The one or more neural networks in transcription neural network may comprise .. a supervised neural network. .. The supervised neural network may be trained on a training dataset training dataset comprising a plurality of audio-visual media files and or the transcript, wherein the plurality of audio-visual media and the transcript are annotated with a human feedback regarding summary status and not summary status...”) It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the invention to have modified Xu, Kim, Grillo and Halyal to incorporate the teachings of Chalana. Those in the art would have recognized that Xu’s teaching regarding receiving audio data for generating a transcription, followed by the generation of a transcript for speech in a portion of the audio data; Kim’a teaching regarding receiving audio data for generating a transcription, followed by the generation of a different transcript for speech in a different portion of the audio; Grillo’s teaching regarding combining two or more transcripts and Halyal’s teaching regarding the use of one or more LLMs could be modified to include Chalana’s’s teaching regarding the use of a supervised model to summarize transcripts which by the use of labeled data for training results in summaries with reduced errors and improved accuracy. Prior Art Made of Record The prior art made of record and not relied upon is considered pertinent to Applicant's disclosure, and is listed in the attached form PTO-892 (Notice of References Cited). Unless expressly noted otherwise by the Examiner, all documents listed on form PTO-892 are cited in their entirety. Peng (US 20260066066 A1)- Generating Clinical Documentation Using Large Language Models And Artificial Intelligence - teaches: Systems and methods generate clinical documentation using large language models and artificial intelligence (AI). A template management module is provided to create customizable templates. A processing unit can receive input data from various sources and use AI to generate transcripts, summarize sessions, and produce clinical documentation such as clinical notes. The processing unit may also generate Current Procedural Terminology (CPT) and diagnosis codes, generate after-visit summaries, and generate referral letters. The AI may be trained on past clinical notes and can adapt to the clinician's style over time, with a feedback loop for continuous improvement. Additional features include cohort-based training, real-time language translation, predictive text, and analytics for documentation trends. The system supports customization of note length, style, and keywords, as well as integration with external medical databases and patient portals. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to PIERRE L MACCAGNO whose telephone number is (571)270-5408. The examiner can normally be reached M-F 8:00 to 5:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Mamon Obeid can be reached at (571)270-1813. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /PIERRE L MACCAGNO/Examiner, Art Unit 3687 /MAMON OBEID/Supervisory Patent Examiner, Art Unit 3687
Read full office action

Prosecution Timeline

Sep 19, 2025
Application Filed
Jul 22, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705585
VIRTUAL SUBACCOUNTS
5y 0m to grant Granted Aug 11, 2026
Patent 12639767
SYSTEMS AND METHODS FOR CLAIM PROCESSING
3y 7m to grant Granted May 26, 2026
Patent 12580057
SYSTEMS AND METHODS FOR NATURAL LANGUAGE PROCESSING-BASED CLASSIFICATION OF ELECTRONIC MEDICAL RECORDS
1y 7m to grant Granted Mar 17, 2026
Patent 12423674
SECURE QR CODE TRANSACTIONS
6y 1m to grant Granted Sep 23, 2025
Patent 12263019
APPARATUS AND A METHOD FOR THE GENERATION OF A PLURALITY OF PERSONAL TARGETS
1y 11m to grant Granted Apr 01, 2025
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
23%
Grant Probability
54%
With Interview (+31.5%)
3y 1m (~2y 2m remaining)
Median Time to Grant
Low
PTA Risk
Based on 139 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month