Prosecution Insights
Last updated: September 17, 2026
Application No. 19/042,936

Systems and Methods for Digital Transcript Creation Using Automated Speech Recognition

Non-Final OA §101§103§112
Filed
Jan 31, 2025
Priority
Sep 13, 2018 — provisional 62/730,700 +2 more
Examiner
SMITH, SEAN THOMAS
Art Unit
Tech Center
Assignee
Magna Legal Services LLC
OA Round
1 (Non-Final)
71%
Grant Probability
Favorable
1-2
OA Rounds
1y 1m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 71% — above average
71%
Career Allowance Rate
12 granted / 17 resolved
+10.6% vs TC avg
Strong +32% interview lift
Without
With
+31.9%
Interview Lift
resolved cases with interview
Typical timeline
2y 9m
Avg Prosecution
21 currently pending
Career history
48
Total Applications
across all art units

Statute-Specific Performance

§101
26.0%
-14.0% vs TC avg
§103
52.7%
+12.7% vs TC avg
§102
12.5%
-27.5% vs TC avg
§112
7.5%
-32.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 17 resolved cases

Office Action

§101 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Applicant’s claim for the benefit of a prior-filed application under 35 U.S.C. 119(e) or under 35 U.S.C. 120, 121, 365(c), or 386(c) is acknowledged. Accordingly, the current application is afforded the benefit of the earlier filing date of application 17/465,509, filed September 1st, 2021, having priority to provisional application 62/730,700, filed September 13th, 2018. Information Disclosure Statement The information disclosure statements (IDS) submitted on January 31st, 2025 and May 20th, 2025 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. Claims 32-34 and 38-39 are rejected under 35 U.S.C. 112(b) as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor regards as the invention. Claim 32 recites the limitation "the generating the real-time transcript of the audio stream includes using the selected formatting template; and the displaying includes displaying the real-time transcript in the selected format and the visual annotation." There is insufficient antecedent basis for “the real-time transcript” in the claim. Claim 33 depends from claim 32, and thus also fails to particularly point out or distinctly claim the subject matter recited. Claim 34 recites the limitation "editing a portion of the transcript earlier in time than a current time of the audio stream while generating the real-time transcript." There is insufficient antecedent basis for “the real-time transcript” in the claim. Claim 38 recites the limitation "synchronizing the audio stream and the real-time transcript using the first timestamp and the second timestamp." There is insufficient antecedent basis for the first and second timestamps in the claim. Claim 39 depends from claim 38, and thus also fails to particularly point out or distinctly claim the subject matter recited. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 21-42 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claims recite a mental process that can be performed in the human mind or with the aid of pen and paper. This judicial exception is not integrated into a practical application because a computer is invoked merely as a tool to execute an abstract idea. The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because an abstract idea is merely applied on a generic computer without any element that would otherwise preclude performance of the abstract idea as a mental process. Regarding claim 21, the claim recites “A method of synchronizing an audio playback and a transcript, comprising:receiving an audio stream, wherein the audio stream includes a first timestamp for each word of the audio stream;generating a real-time transcript of the audio stream, wherein the real-time transcript includes a second timestamp for each word in the real-time transcript;synchronizing the audio stream and the real-time transcript using the first timestamp and the second timestamp; andplaying back the audio stream and displaying a visual annotation in the real-time transcript of a corresponding word in the audio stream.” The limitations of “receiving an audio stream,” “generating a real-time transcript,” “synchronizing…” and “playing back the audio stream…” as drafted cover mental activities which can be performed in the mind or with the aid of pen and paper. Taken individually, or as a whole, these limitations describe acts which are equivalent to human mental work of court reporting, as described in paragraph [0002] and Figure 2. The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the steps of the claimed invention can be performed mentally, and no additional features in the claims would preclude them from being performed as such. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 22, the claim depends from claim 21, and thus recites the limitations of claim 21, “wherein the audio stream timestamps are associated with a transcript participant.” Taken individually, or as a whole with claim 21, these limitations describe acts which are equivalent to human mental work of court reporting. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 23, the claim depends from claim 21, and thus recites the limitations of claim 21, “wherein displaying the visual annotation in the transcript includes adding emphasis to the corresponding word in the audio recording.” Taken individually, or as a whole with claim 21, these limitations describe acts which are equivalent to human mental work of court reporting with further organization of human activity, in that a human actor can indicate – for example, with a finger – a word in a transcript corresponding to a point in audio data. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 24, the claim depends from claim 23, and thus recites the limitations of claims 21 and 23, “wherein adding emphasis to the corresponding word includes any one or more of bolding, underlining, or using color in either the text or the background of the corresponding word.” Taken individually, or as a whole with the preceding claims, these limitations describe acts which are equivalent to human mental work of court reporting with predefined formatting methods. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 25, the claim depends from claim 21, and thus recites the limitations of claim 21, “further comprising:selecting a formatting template for the transcript; wherein:the generating the real-time transcript of the audio stream includes using the selected formatting template; andthe displaying includes displaying the real-time transcript in the selected format and the visual annotation.” Taken individually, or as a whole with claim 21, these limitations describe acts which are equivalent to human mental work of court reporting with predefined formatting methods. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 26, the claim depends from claim 25, and thus recites the limitations of claim21 and 25, “wherein the selected formatting template includes settings for any one or more of indentation, page margins, page headers, page footers, font type, font size, line spacing, or line numbering.” Taken individually, or as a whole with the preceding claims, these limitations describe acts which are equivalent to human mental work of court reporting with predefined formatting methods. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 27, the claim depends from claim 25, and thus recites the limitations of claims 21 and 25, “wherein the real-time transcript of the audio stream includes data associating a transcript participant with one or more portions of the transcript.” Taken individually, or as a whole with the preceding claims, these limitations describe acts which are equivalent to human mental work of court reporting. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 28, the claim depends from claim 21, and thus recites the limitations of claim 21, “further comprising: editing a portion of the transcript earlier in time than a current time of the audio stream while generating the real-time transcript.” Taken individually, or as a whole with claim 21, these limitations describe acts which are equivalent to human mental work of court reporting. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 29, the claim recites “A method of synchronizing an audio playback and a transcript, comprising:obtaining an audio recording, wherein the audio recording includes a first timestamp for each word of the audio recording;generating a transcript of the audio recording, wherein the transcript includes a second timestamp for each word in the transcript;synchronizing the audio recording and the transcript using the first timestamp and the second timestamp; andplaying back the audio recording and displaying a visual annotation in the transcript of a corresponding word in the audio recording.” The limitations of “obtaining an audio recording,” “generating a transcript,” “synchronizing…” and “playing back the audio recording…” as drafted cover mental activities which can be performed in the mind or with the aid of pen and paper. Taken individually, or as a whole, these limitations describe acts which are equivalent to human mental work of post-hoc court reporting. The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the steps of the claimed invention can be performed mentally, and no additional features in the claims would preclude them from being performed as such. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 30, the claim depends from claim 29, and thus recites the limitations of claim 29, “wherein displaying the visual annotation in the transcript includes adding emphasis to the corresponding word in the audio recording.” Taken individually, or as a whole with claim 29, these limitations describe acts which are equivalent to human mental work of court reporting and organization of human activity. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 31, the claim depends from claim 30, and thus recites the limitations of claims 29 and 30, “wherein adding emphasis to the corresponding word includes any one or more of bolding, underlining, or highlighting.” Taken individually, or as a whole with the preceding claims, these limitations describe acts which are equivalent to human mental work of court reporting with predefined formatting methods. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 32, the claim depends from claim 29, and thus recites the limitations of claim 29, “further comprising:selecting a formatting template for the transcript; wherein:the generating the real-time transcript of the audio stream includes using the selected formatting template; andthe displaying includes displaying the real-time transcript in the selected format and the visual annotation.” Taken individually, or as a whole with claim 29, these limitations describe acts which are equivalent to human mental work of court reporting with predefined formatting methods. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 33, the claim depends from claim 32, and thus recites the limitations of claims 29 and 32, “wherein the selected formatting template includes settings for any one or more of indentation, page margins, page headers, page footers, font type, font size, line spacing, or line numbering.” Taken individually, or as a whole with the preceding claims, these limitations describe acts which are equivalent to human mental work of court reporting with predefined formatting methods. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 34, the claim depends from claim 29, and thus recites the limitations of claim 29, “further comprising: editing a portion of the transcript earlier in time than a current time of the audio stream while generating the real-time transcript.” Taken individually, or as a whole with claim 29, these limitations describe acts which are equivalent to human mental work of court reporting. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 35, the claim recites “A method for real-time formatting of a transcript of an audio stream, comprising:selecting a formatting template for the transcript;receiving the audio stream;generating a real-time transcript of the audio stream using the selected formatting template; anddisplaying the real-time transcript in the selected format.” The limitations of “selecting a formatting template,” “receiving the audio stream,” “generating a real-time transcript,” and “displaying the real-time transcript” as drafted cover mental activities which can be performed in the mind or with the aid of pen and paper. Taken individually, or as a whole, these limitations describe acts which are equivalent to human mental work of transcribing audio, or filling out a form. The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the steps of the claimed invention can be performed mentally, and no additional features in the claims would preclude them from being performed as such. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 36, the claim depends from claim 35, and thus recites the limitations of claim 35, “wherein the selected formatting template includes settings for any one or more of indentation, page margins, page headers, page footers, font type, font size, line spacing, or line numbering.” Taken individually, or as a whole with claim 35, these limitations describe acts which are equivalent to human mental work of audio transcription. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 37, the claim depends from claim 35, and thus recites the limitations of claim 35, “wherein: the audio stream includes a first timestamp for each word of the audio stream;the real-time transcript includes a second timestamp for each word in the real-time transcript.” Taken individually, or as a whole with claim 35, these limitations describe acts which are equivalent to human mental work of audio transcription. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 38, the claim depends from claim 35, and thus recites the limitations of claim 35, “further comprising: synchronizing the audio stream and the real-time transcript using the first timestamp and the second timestamp.” Taken individually, or as a whole with claim 35, these limitations describe acts which are equivalent to human mental work of audio transcription. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 39, the claim depends from claim 38, and thus recites the limitations of claims 35 and 38, “further comprising: playing the audio stream while displaying the real-time transcript in the selected format.” Taken individually, or as a whole with the preceding claims, these limitations describe acts which are equivalent to human mental work of audio transcription. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 40, the claim depends from claim 39, and thus recites the limitations of claims 35 and 38-39, “further comprising: displaying a visual annotation in the real-time transcript corresponding to a word being played in the audio stream.” Taken individually, or as a whole with the preceding claims, these limitations describe acts which are equivalent to human mental work of audio transcription. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 41, the claim depends from claim 35, and thus recites the limitations of claim 35, “further comprising: editing a portion of the transcript earlier in time than a current time of the audio stream while generating the real-time transcript.” Taken individually, or as a whole with claim 35, these limitations describe acts which are equivalent to human mental work of audio transcription. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 42, the claim depends from claim 35, and thus recites the limitations of claim 35, “wherein the audio stream includes data associating a transcript participant with one or more portions of the real-time transcript.” Taken individually, or as a whole with claim 35, these limitations describe acts which are equivalent to human mental work of audio transcription. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 21-22, 25-29, 32-39 and 41-42 are rejected under 35 U.S.C. 103 as being unpatentable over U.S. Patent 10,360,915 to Taple et al. (hereinafter, "Taple") in view of U.S. Patent 10,013,407 to Grimm (hereinafter, "Grimm"). Regarding claim 21, Taple teaches a method of synchronizing an audio playback and a transcript, comprising: receiving an audio stream, wherein the audio stream includes a first timestamp for each word of the audio stream (column 7, lines 23 through 36, "As the deposition proceeds, audio storage module 230 receives an output signal from microphone(s) 105, and stores one or more audio recordings representing what was said at the deposition in memory… In some examples, audio storage module stores audio recordings with a plurality of timestamps that identify when a particular recording was made," and column 8, line 66, "Transcript generator 240 may review timestamps or other information contained in stored audio, and piece together a transcript reflecting sequentially the content of what was said, and by whom, during the deposition proceeding."); generating a real-time transcript of the audio stream, wherein the real-time transcript includes a second timestamp for each word in the real-time transcript (column 9, line 23, "In some examples, by sequentially generating transcript portions in real time, transcript generator 240 can quickly generate a transcript of the deposition that is available to the deposition participants immediately upon conclusion of the deposition proceeding."); synchronizing the audio stream and the real-time transcript using the first timestamp and the second timestamp (column 16, line 40, "Regardless of how it is accomplished (all audio from a deposition, in one embodiment) whether by being captured in a single file, or by capturing and synchronizing multiple files, acquired across multiple audio detection devices (e.g., microphones), once these files are obtained, the system 200 may utilize them to create a transcript that accurately captures and orders speech event into a transcript, which in preferred embodiments is rendered by attributing speech events to an identified speaker."). Taple does not explicitly teach playback of synchronized audio and transcripts, and thus, Grimm is introduced. Grimm teaches a computerized method for generating processed files of deposition testimony transcripts including playing back the audio stream and displaying a visual annotation in the real-time transcript of a corresponding word in the audio stream (column 6, line 8, "As shown in FIG. 3, the textual transcript 100 of FIG. 1 may be synchronized with the video 200 of FIG. 2, such that a designation of a page and line number may be used to designate either or both of corresponding content of the textual transcript and the video. One skilled in the art will readily recognize that a computer-implemented program may be used to automatically recognize audio of the video that corresponds to text of the transcript. Such a computer-implemented program may automatically synchronize text of the textual transcript 100 with time stamps of the video 200, and automatically generate closed captions for the video 200 based on the corresponding text of the textual transcript 100."). Taple and Grimm are considered analogous because they are each concerned with transcribing legal proceedings. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified Taple with the playback of Grimm for the purpose of providing a more consumable work product. Given that all the claimed elements were known in the prior art, one skilled in the art could have combined the elements by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Regarding claim 22, Taple teaches a method for recording sworn testimony wherein the audio stream timestamps are associated with a transcript participant (column 8, line 14, "Speaker identification module 232 further operates to identify in audio recordings stored by audio storage module 230, a speaker source for each word or phrase."). Regarding claim 25, Taple does not teach a recording method that uses formatting templates, however, Grimm’s method includes selecting a formatting template for the transcript (column 10, line 18, "FIG. 8 illustrates a graphical user interface 800 adapted to receive user selections for generating a processed textual transcript in accordance with the present disclosure. For example, a folder selection control 802 may be provided whereby a user may select a folder, stored on a non-transitory computer-readable medium, of files to be processed. Also, a set of markup controls 804 may be provided that permit a user to select a hue of highlight or line for use in marking, by the one or more computer processors, contents of the textual transcript corresponding to two or more pre-defined categories, such as a plaintiff affirmative category, a defense counter category, a defense affirmative category, a plaintiff counter category, and/or custom user categories."); wherein: the generating the real-time transcript of the audio stream includes using the selected formatting template (column 11, line 19, "At block 1006, the one or more computer processors may markup text of the textual transcript according to the user selections, as described above."); and the displaying includes displaying the real-time transcript in the selected format and the visual annotation (column 11, line 44, "At block 1010, the one or more computer processors may generate the marked up textual transcript. Generating the marked up textual transcript at block 1010 may include recording the marked up transcript in a computer-readable medium, and automatically assigning a name according to a naming convention that indicates the processed designations and/or, textual transcript to which it relates. Alternatively or additionally, a memory location, such as a folder, may be automatically selected to indicate the processed designations and/or textual transcript to which it relates. It is envisioned, for example, that the marked up textual transcript may be saved in a portable document format (PDF). Alternatively or additionally, generating the marked up textual transcript at block 1010 may include printing the marked up textual transcript, generating an electronic display of the marked up textual transcript, or otherwise rendering the marked up textual transcript."). Taple and Grimm are considered analogous because they are each concerned with transcribing legal proceedings. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified Taple with the template controls of Grimm for the purpose of standardizing Taple’s transcripts. Given that all the claimed elements were known in the prior art, one skilled in the art could have combined the elements by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Regarding claim 26, Grimm’s method further teaches a formatting template includes settings for any one or more of indentation, page margins, page headers, page footers, font type, font size, line spacing, or line numbering (column 10, line 31, "It is envisioned that text boxes may be provided for the user to specify the categories to be used with each hue, and that radio buttons may be provided to allow the user to select a highlight type of markup or a line type of markup for each category. Additional controls 806, such as a checkbox and drop down menu, may be provided to allow the user to select to remove non-designated text, and to select a margin to be employed between lines."). Taple and Grimm are considered analogous because they are each concerned with transcribing legal proceedings. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified Taple with the template controls of Grimm for the purpose of standardizing Taple’s transcripts. Given that all the claimed elements were known in the prior art, one skilled in the art could have combined the elements by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Regarding claim 27, Taple teaches a testimony recording method wherein the real-time transcript of the audio stream includes data associating a transcript participant with one or more portions of the transcript (column 9, line 15, "According to these examples, as audio storage module 230 receive and stores audio data from microphone(s) 105, STT module 234 converts the stored audio data into a text representation, and speaker identification module 232 associates a deposition participant to each converted text representation."). Regarding claim 28, Taple teaches a testimony recording method that includes editing a portion of the transcript earlier in time than a current time of the audio stream while generating the real-time transcript (column 15, line 59, "While the deposition is still proceeding, system 200 may provide users with the option to edit text to reflect what was said by a user, in the instance of errors."). Regarding claim 29, Taple teaches a method for synchronizing audio and transcripts comprising: obtaining an audio recording, wherein the audio recording includes a first timestamp for each word of the audio recording (column 3, line 47, "As also shown in FIG. 1, system 100 includes an audio translation engine 107. Audio translation engine 107 receives (directly or indirectly) from microphone 105 digital or other data reflecting audio recordings of oral statements and other audible sounds made by deposer 103A and deponent 103B in the course of a deposition proceeding," and column 8, line 66, "Transcript generator 240 may review timestamps or other information contained in stored audio, and piece together a transcript reflecting sequentially the content of what was said, and by whom, during the deposition proceeding."); generating a transcript of the audio recording, wherein the transcript includes a second timestamp for each word in the transcript (column 3, line 52, "Audio translation engine 107 stores, for example in temporary memory such as Random Access Memory (RAM), or long term storage such as a magnetic hard disk or other long-term storage device (or, in other embodiments, otherwise accesses electronically) the received data reflecting audio recordings, and processes the data to generate a transcript 113 reflecting the orally communicated content of the deposition proceeding."; synchronizing the audio recording and the transcript using the first timestamp and the second timestamp (column 16, line 40, "Regardless of how it is accomplished (all audio from a deposition, in one embodiment) whether by being captured in a single file, or by capturing and synchronizing multiple files, acquired across multiple audio detection devices (e.g., microphones), once these files are obtained, the system 200 may utilize them to create a transcript that accurately captures and orders speech event into a transcript, which in preferred embodiments is rendered by attributing speech events to an identified speaker."). Taple does not explicitly teach playback of synchronized audio and transcripts, however, Grimm’s method teaches playing back the audio recording and displaying a visual annotation in the transcript of a corresponding word in the audio recording (column 6, line 8, "As shown in FIG. 3, the textual transcript 100 of FIG. 1 may be synchronized with the video 200 of FIG. 2, such that a designation of a page and line number may be used to designate either or both of corresponding content of the textual transcript and the video. One skilled in the art will readily recognize that a computer-implemented program may be used to automatically recognize audio of the video that corresponds to text of the transcript. Such a computer-implemented program may automatically synchronize text of the textual transcript 100 with time stamps of the video 200, and automatically generate closed captions for the video 200 based on the corresponding text of the textual transcript 100."). Taple and Grimm are considered analogous because they are each concerned with transcribing legal proceedings. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified Taple with the playback of Grimm for the purpose of providing a more consumable work product. Given that all the claimed elements were known in the prior art, one skilled in the art could have combined the elements by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Regarding claim 32, Taple does not teach a recording method that uses formatting templates, however Grimm’s method includes selecting a formatting template for the transcript (column 10, line 18, "FIG. 8 illustrates a graphical user interface 800 adapted to receive user selections for generating a processed textual transcript in accordance with the present disclosure. For example, a folder selection control 802 may be provided whereby a user may select a folder, stored on a non-transitory computer-readable medium, of files to be processed. Also, a set of markup controls 804 may be provided that permit a user to select a hue of highlight or line for use in marking, by the one or more computer processors, contents of the textual transcript corresponding to two or more pre-defined categories, such as a plaintiff affirmative category, a defense counter category, a defense affirmative category, a plaintiff counter category, and/or custom user categories."); wherein: the generating the real-time transcript of the audio stream includes using the selected formatting template (column 11, line 19, "At block 1006, the one or more computer processors may markup text of the textual transcript according to the user selections, as described above."); and the displaying includes displaying the real-time transcript in the selected format and the visual annotation (column 11, line 44, "At block 1010, the one or more computer processors may generate the marked up textual transcript. Generating the marked up textual transcript at block 1010 may include recording the marked up transcript in a computer-readable medium, and automatically assigning a name according to a naming convention that indicates the processed designations and/or, textual transcript to which it relates. Alternatively or additionally, a memory location, such as a folder, may be automatically selected to indicate the processed designations and/or textual transcript to which it relates. It is envisioned, for example, that the marked up textual transcript may be saved in a portable document format (PDF). Alternatively or additionally, generating the marked up textual transcript at block 1010 may include printing the marked up textual transcript, generating an electronic display of the marked up textual transcript, or otherwise rendering the marked up textual transcript."). Taple and Grimm are considered analogous because they are each concerned with transcribing legal proceedings. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified Taple with the template controls of Grimm for the purpose of standardizing Taple’s transcripts. Given that all the claimed elements were known in the prior art, one skilled in the art could have combined the elements by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Regarding claim 33, Grimm further teaches the selected formatting template includes settings for any one or more of indentation, page margins, page headers, page footers, font type, font size, line spacing, or line numbering (column 10, line 31, "It is envisioned that text boxes may be provided for the user to specify the categories to be used with each hue, and that radio buttons may be provided to allow the user to select a highlight type of markup or a line type of markup for each category. Additional controls 806, such as a checkbox and drop down menu, may be provided to allow the user to select to remove non-designated text, and to select a margin to be employed between lines."). Taple and Grimm are considered analogous because they are each concerned with transcribing legal proceedings. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified Taple with the template controls of Grimm for the purpose of standardizing Taple’s transcripts. Given that all the claimed elements were known in the prior art, one skilled in the art could have combined the elements by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Regarding claim 34, Taple teaches a testimony recording method that includes editing a portion of the transcript earlier in time than a current time of the audio stream while generating the real-time transcript (column 15, line 59, "While the deposition is still proceeding, system 200 may provide users with the option to edit text to reflect what was said by a user, in the instance of errors."). Regarding claim 35, Taple teaches a testimony recording method that includes receiving the audio stream (column 9, line 14, "In some examples, transcript generator 240 may generate portions of a transcript in real-time during a deposition proceeding. According to these examples, as audio storage module 230 receive and stores audio data from microphone(s) 105, STT module 234 converts the stored audio data into a text representation, and speaker identification module 232 associates a deposition participant to each converted text representation. At the same time transcript generator 240 sequentially generates transcript portions as the deposition proceeding takes place. In some examples, by sequentially generating transcript portions in real time, transcript generator 240 can quickly generate a transcript of the deposition that is available to the deposition participants immediately upon conclusion of the deposition proceeding."). Taple does not explicitly teach a recording method that uses formatting templates, however, Grimm’s method teaches selecting a formatting template for the transcript (column 10, line 18, "FIG. 8 illustrates a graphical user interface 800 adapted to receive user selections for generating a processed textual transcript in accordance with the present disclosure. For example, a folder selection control 802 may be provided whereby a user may select a folder, stored on a non-transitory computer-readable medium, of files to be processed. Also, a set of markup controls 804 may be provided that permit a user to select a hue of highlight or line for use in marking, by the one or more computer processors, contents of the textual transcript corresponding to two or more pre-defined categories, such as a plaintiff affirmative category, a defense counter category, a defense affirmative category, a plaintiff counter category, and/or custom user categories."); generating a real-time transcript of the audio stream using the selected formatting template (column 11, line 19, "At block 1006, the one or more computer processors may markup text of the textual transcript according to the user selections, as described above."); and displaying the real-time transcript in the selected format (column 11, line 44, "At block 1010, the one or more computer processors may generate the marked up textual transcript. Generating the marked up textual transcript at block 1010 may include recording the marked up transcript in a computer-readable medium, and automatically assigning a name according to a naming convention that indicates the processed designations and/or, textual transcript to which it relates. Alternatively or additionally, a memory location, such as a folder, may be automatically selected to indicate the processed designations and/or textual transcript to which it relates. It is envisioned, for example, that the marked up textual transcript may be saved in a portable document format (PDF). Alternatively or additionally, generating the marked up textual transcript at block 1010 may include printing the marked up textual transcript, generating an electronic display of the marked up textual transcript, or otherwise rendering the marked up textual transcript."). Taple and Grimm are considered analogous because they are each concerned with transcribing legal proceedings. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified Taple with the template controls of Grimm for the purpose of standardizing Taple’s transcripts. Given that all the claimed elements were known in the prior art, one skilled in the art could have combined the elements by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Regarding claim 36, Grimm’s method further teaches the selected formatting template includes settings for any one or more of indentation, page margins, page headers, page footers, font type, font size, line spacing, or line numbering (column 10, line 31, "It is envisioned that text boxes may be provided for the user to specify the categories to be used with each hue, and that radio buttons may be provided to allow the user to select a highlight type of markup or a line type of markup for each category. Additional controls 806, such as a checkbox and drop down menu, may be provided to allow the user to select to remove non-designated text, and to select a margin to be employed between lines."). Taple and Grimm are considered analogous because they are each concerned with transcribing legal proceedings. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified Taple with the template controls of Grimm for the purpose of standardizing Taple’s transcripts. Given that all the claimed elements were known in the prior art, one skilled in the art could have combined the elements by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Regarding claim 37, Taple teaches a testimony recording method wherein the audio stream includes a first timestamp for each word of the audio stream (column 8, line 66, "Transcript generator 240 may review timestamps or other information contained in stored audio, and piece together a transcript reflecting sequentially the content of what was said, and by whom, during the deposition proceeding."); the real-time transcript includes a second timestamp for each word in the real-time transcript (column 16, line 63, "To ensure that no inappropriate or inaccurate post-deposition changes are made to the transcript, in some embodiments, system 200 preserves an audio recording of the deposition and a time stamp applied to both the audio recording and a time stamp to the translation, so there is no doubt of what was said if there is a difference of opinion among the participants."). Regarding claim 38, Taple teaches a testimony recording method that includes synchronizing the audio stream and the real-time transcript using the first timestamp and the second timestamp (column 16, line 40, "Regardless of how it is accomplished (all audio from a deposition, in one embodiment) whether by being captured in a single file, or by capturing and synchronizing multiple files, acquired across multiple audio detection devices (e.g., microphones), once these files are obtained, the system 200 may utilize them to create a transcript that accurately captures and orders speech event into a transcript, which in preferred embodiments is rendered by attributing speech events to an identified speaker."). Regarding claim 39, Taple does not explicitly teach playback of synchronized audio and transcripts, nor the use of formatting templates, however, Grimm’s method includes playing the audio stream while displaying the real-time transcript in the selected format (column 6, line 8, "As shown in FIG. 3, the textual transcript 100 of FIG. 1 may be synchronized with the video 200 of FIG. 2, such that a designation of a page and line number may be used to designate either or both of corresponding content of the textual transcript and the video. One skilled in the art will readily recognize that a computer-implemented program may be used to automatically recognize audio of the video that corresponds to text of the transcript. Such a computer-implemented program may automatically synchronize text of the textual transcript 100 with time stamps of the video 200, and automatically generate closed captions for the video 200 based on the corresponding text of the textual transcript 100."). Regarding claim 41, Taple teaches a testimony recording method that includes editing a portion of the transcript earlier in time than a current time of the audio stream while generating the real-time transcript (column 15, line 59, "While the deposition is still proceeding, system 200 may provide users with the option to edit text to reflect what was said by a user, in the instance of errors."). Regarding claim 42, Taple teaches a testimony recording method wherein the audio stream includes data associating a transcript participant with one or more portions of the real-time transcript (column 6, line 19, "In some examples, speaker identification module 232 may be generally configured to utilize identification of a microphone or microphones that captured audio to identify which deposition participant is associated with recorded audio segments, but may utilize processing to identify speaker(s) based on stored user profiles as a fail-safe," and column 9, line 15, "According to these examples, as audio storage module 230 receive and stores audio data from microphone(s) 105, STT module 234 converts the stored audio data into a text representation, and speaker identification module 232 associates a deposition participant to each converted text representation."). Claims 23-24, 30-31 and 40 are rejected under 35 U.S.C. 103 as being unpatentable over Taple and Grimm as applied to claims 21, 29 and 35 above, and further in view of "Video Closed Captions Formatting" as offered by Capital Captions (hereinafter, "Capital Captions"). Regarding claim 23, the combination of Tpale and Grimm does not explicitly teach visual annotations for audio transcription, and thus, Capital Captions is introduced. Capital Captions offers a tool for creating video captions wherein displaying the visual annotation in the transcript includes adding emphasis to the corresponding word in the audio recording (Closed Caption Formats, "Highlight captions draw focus to significant words or punchlines using bolding effects or color pops in conjunction with dialogue,"). Taple, Grimm and Capital Captions are considered analogous because they are each concerned with audio transcriptions. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified Taple and Grimm with the emphasis techniques of Capital Captions for the purpose of improving readability of synchronized transcripts. Given that all the claimed elements were known in the prior art, one skilled in the art could have combined the elements by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Regarding claim 24, Capital Captions offers a tool for creating video captions wherein adding emphasis to the corresponding word includes any one or more of bolding, underlining, or using color in either the text or the background of the corresponding word (Closed Caption Formats, "Highlight captions draw focus to significant words or punchlines using bolding effects or color pops in conjunction with dialogue,"). Taple, Grimm and Capital Captions are considered analogous because they are each concerned with audio transcriptions. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified Taple and Grimm with the emphasis techniques of Capital Captions for the purpose of improving readability of synchronized transcripts. Given that all the claimed elements were known in the prior art, one skilled in the art could have combined the elements by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Regarding claim 30, the combination of Taple and Grimm does not explicitly teach visual annotations for audio transcriptions, however, Capital Captions offers a tool for creating video captions wherein displaying the visual annotation in the transcript includes adding emphasis to the corresponding word in the audio recording (Closed Caption Formats, "Highlight captions draw focus to significant words or punchlines using bolding effects or color pops in conjunction with dialogue,"). Taple, Grimm and Capital Captions are considered analogous because they are each concerned with audio transcriptions. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified Taple and Grimm with the emphasis techniques of Capital Captions for the purpose of improving readability of synchronized transcripts. Given that all the claimed elements were known in the prior art, one skilled in the art could have combined the elements by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Regarding claim 31, Capital Captions offers a tool for creating video captions wherein adding emphasis to the corresponding word includes any one or more of bolding, underlining, or highlighting (Closed Caption Formats, "Highlight captions draw focus to significant words or punchlines using bolding effects or color pops in conjunction with dialogue,"). Taple, Grimm and Capital Captions are considered analogous because they are each concerned with audio transcriptions. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified Taple and Grimm with the emphasis techniques of Capital Captions for the purpose of improving readability of synchronized transcripts. Given that all the claimed elements were known in the prior art, one skilled in the art could have combined the elements by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Regarding claim 40, the combination of Taple and Grimm does not explicitly teach visual annotations for audio transcriptions, however, Capital Captions offers a tool for creating video captions that includes displaying a visual annotation in the real-time transcript corresponding to a word being played in the audio stream (Closed Caption Formats, "Highlight captions draw focus to significant words or punchlines using bolding effects or color pops in conjunction with dialogue,"). Taple, Grimm and Capital Captions are considered analogous because they are each concerned with audio transcriptions. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified Taple and Grimm with the emphasis techniques of Capital Captions for the purpose of improving readability of synchronized transcripts. Given that all the claimed elements were known in the prior art, one skilled in the art could have combined the elements by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: U.S. Patent 7,881,930 to Faisman et al. teaches speech-to-text system that learns speaker characteristics. U.S. Patent 10,304,458 to Woo teaches a system and method for video transcription with speaker identification. U.S. Patent Application Publication 2007/0260457 to Bennet et al. teaches a court reporter transcription system with transcription synchronization. U.S. Patent Application Publication 2009/0037171 to McFarland et al. teaches a voice transcription system with speaker identification. U.S. Patent Application Publication 2016/0133251 to Kadirkamanathan et al. teaches an audio transcription system transcript synchronization. U.S. Patent Application Publication 2018/0174587 to Bermundo teaches formatted transcript generation. U.S. Patent Application Publication 2019/0132265 to Nowak-Przygodzki et al. teaches a conferencing assistant system that provides speech-to-text services for participants. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SEAN T SMITH whose telephone number is (571)272-6643. The examiner can normally be reached Monday - Friday 8:00am - 5:00pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, PIERRE-LOUIS DESIR can be reached at (571) 272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SEAN THOMAS SMITH/Examiner, Art Unit 2659 /PIERRE LOUIS DESIR/Supervisory Patent Examiner, Art Unit 2659
Read full office action

Prosecution Timeline

Jan 31, 2025
Application Filed
Aug 12, 2026
Non-Final Rejection mailed — §101, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737549
METHOD AND APPARATUS FOR SENTIMENT ANALYSIS, ELECTRONIC DEVICE AND COMPUTER-READABLE STORAGE MEDIUM
2y 3m to grant Granted Sep 15, 2026
Patent 12688358
SYSTEMS AND METHODS FOR DYNAMICALLY PROVIDING A CORRECT PRONUNCIATION FOR A USER NAME BASED ON USER LOCATION
2y 6m to grant Granted Jul 21, 2026
Patent 12626056
GENERATING NATURAL LANGUAGE MODEL INSIGHTS FOR DATA CHARTS USING LIGHT LANGUAGE MODELS DISTILLED FROM LARGE LANGUAGE MODELS
2y 10m to grant Granted May 12, 2026
Patent 12602540
LEVERAGING A LARGE LANGUAGE MODEL ENCODER TO EVALUATE PREDICTIVE MODELS
2y 3m to grant Granted Apr 14, 2026
Patent 12530534
SYSTEM AND METHOD FOR GENERATING STRUCTURED SEMANTIC ANNOTATIONS FROM UNSTRUCTURED DOCUMENT
3y 0m to grant Granted Jan 20, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
71%
Grant Probability
99%
With Interview (+31.9%)
2y 9m (~1y 1m remaining)
Median Time to Grant
Low
PTA Risk
Based on 17 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month